Hernán Astudillo

dblp:69/3668 · also Hernán Astudillo-Rojas · DBLP profile ↗
← Back
26ranked-venue papers in the field
0as first author
14since 2021 · last 2025
0000-0002-6487-5813ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 24Database Systems & Data Management · 1Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Enhancing Software Requirements Education Through Active Learning: A Pilot Study on Role-Playing and Real-World Simulations
abstract
Software engineering education (SEE) requires a careful balance between theoretical concepts and practical application, and this is especially true of teaching software requirements. However, traditional teaching methods, such as lectures, often fail to engage students and foster the practical skills necessary for real world software development. We present a pilot study on the application of active learning methodologies in a software engineering course involving 21 midlevel undergraduate informatics students. The approach covers the entire requirements engineering process, from stakeholder identification to effort estimation techniques, to make learning more experiential and applied. We split each class into three segments: (i) theoretical instruction; (ii) collaborative group activities, including role playing to simulate real world dynamics and encourage students to apply concepts in context; and (iii) feedback and guided reflection. The course capstone was a requirements elicitation activity conducted in an actual real world scenario, aimed at consolidating and demonstrating the skills developed throughout the course. Initial results, as shown by ungraded assessment of the elicitation activity and by a student survey, indicate that this structured active learning approach allowed us to achieve the intended learning outcomes and was also well received by the students. Our piloted approach shows promise to provide a replicable SEE method that enhances student engagement, promotes deeper learning, and strengthens the development of practical skills in requirements engineering.
Mauricio Hidalgo, Laura M. Castro, Hernán Astudillo, Fernando Montoya, Manuel Alejandro Goyo
CLEI3
2025 Framework to Support Students in Defining DevOps Technology Stacks
abstract
DevOps is an approach to automated software development and deployment that combines development and operations, with the goal of improving collaboration between teams to optimize the software lifecycle processes, from planning and development to testing, deployment, and monitoring. Despite being an innovative approach, DevOps presents a significant challenge for students in understanding and implementing software development projects. This challenge includes understanding the problem to the solution abstraction that contemplates the design, implementation, and automation of the software development process. This study proposes and evaluates a methodological framework to support students in defining DevOps-oriented technology stacks. The framework combines software architecture and software engineering practices that collectively provide a learning approach based on design decision-making and selection of technologies, frameworks, and tools. We evaluated the framework in two iterations of a capstone course, using a case study that considered the implementation of a DevOps stack on (i) a pre-existing system and (ii) a system from scratch. Results show that students who implemented a DevOps-oriented stack on systems developed from scratch were successful, but those who implemented it on a pre-existing system confronted challenges in configuration management and system flexibility. The proposed framework facilitates the pedagogical experience of implementing DevOps in software development projects, thus rendering it beneficial for students.
Gastón Marquez, Hernán Astudillo
CLEI2
2025 Researchers want to Explain but Practitioners just want to Observe
abstract
Context: The Explainability term is increasingly used in informatics literature, especially in relation to Artificial Intelligence, but it is unclear whether and how it is used in the Software Architecture (SA) community. Objective: This article describes the design and results of a study of the presence and evolution of explainability-related concepts in SA discourse, in both academic and practitioner literature. Method: Topic analysis was done using Latent Dirichlet Allocation (LDA) on postings/articles tagged as "software architecture," between 2008 and 2024, on several representative practitioner sources (StackOverflow, InfoQ and DZone) and a researcher source (IEEE Xplore). Results: A search for "explainability," "observability" and other Explainability-related terms (as identified in previous work) yielded over 48,000 online primary documents. Besides descriptive results on relative use of terms, a key finding is that academic literature focuses on Traceability, but practitioner discussions focus on Observability. Conclusions: The terminology gap between academia and industry on Explainability-related terms matches other reported academia-industry vocabulary differences. This line of research may help to identify alternative concepts in use for a term and alternative terms in use for a concept, and align practitioner-research exchanges on Explainability-related concepts.
Fernando Muraca, María Florencia Pollo Cattaneo, Hernán Astudillo
CLEI3
2025 Cloud-Based Customer Segmentation to Enhance Experience and Performance in Commercial Retail Spaces
abstract
In today’s rapidly accelerating digital transformation landscape, shopping malls face significant challenges stemming from changes in consumer behavior, the rise of e-commerce, and the impact of the COVID-19 pandemic on in-person visits. This situation has led to a decline in foot traffic, shorter dwell times, and increased difficulty in making timely, data-driven decisions.To address this problem, we propose a Big Data platform focused on dynamic customer segmentation, aimed at enhancing the visitor experience and optimizing commercial management within physical retail spaces. The solution leverages Amazon Web Services (AWS) cloud technologies and integrates multiple data sources, including Wi-Fi networks, captive portals, and sensors, to identify real-time behavioral patterns.The platform was developed using the agile Scrum methodology across twelve sprints, enabling an iterative implementation validated in development, testing, and production environments. Deployment was carried out in two shopping malls from Chile and one from Colombia, utilizing real customer data captured through the mall’s technological infrastructure. However, the specific names of the malls cannot be disclosed due to confidentiality agreements and company policies. To enable the public release of the models and datasets, all identifying information has been anonymized.Validation was conducted using key performance indicators such as average monthly visits, average dwell time, and Net Promoter Score (NPS). Short-term post-deployment measurements showed increases of 5%, 30%, and 1 point, respectively, compared to the immediate pre-deployment baseline. However, when comparing long-term trends between 2019 and 2021, average dwell time showed an 8.98% decrease, likely reflecting post-pandemic shifts in consumer behavior rather than a limitation of the implemented system. We conclude that scalable platforms and data-driven segmentation can effectively address the challenges of the commercial real estate sector, paving the way for predictive models and artificial intelligence for advanced personalization.
David Ruete, Paolo Caviedes-Saavedra, Patricio Lagos-Gatica, Carla Taramasco, Hernán Astudillo, Jean Paul Maidana González
CLEI5
2024 SETMAT: A Representation Instrument for Software Engineering Teaching Practices
abstract
Software engineering (SE) teaching includes different strategies such as project-based learning, collaborative learning, simulations, and case methods, among others, seeking to balance theory and practice and encourage teamwork. However, although some repositories of teaching practices are identified, each SE teacher incorporates different strategies/practices either barely documented or subjectively represented. The latter hinders the transfer of knowledge among the software engineering education community. Therefore, in this paper we propose SETMAT (Software Engineering Teaching Methods and Theory), a descriptive theory of SE education offering a common conceptual framework for representing software engineering teaching practices, and we also describe a pilot with teachers from Colombia, Chile, and Mexico where they represented and socialized their SE teaching practices in SETMAT. This pilot proves the usage of SETMAT for representing teaching practices easies comparison, composition, and transferring such practices. This poses an input of interest for SE education.
María Clara Gómez-Alvarez, Carlos Mario Zapata Jaramillo, Hernán Astudillo
CLEI3
2024 Mapping Kolb's Learning Style to Roles in Software Development Team
abstract
The effective transfer and acquisition of necessary knowledge, methods, and attitudes pose significant challenges for Software Engineering Education. Furthermore, training in software development skills and knowledge currently lacks a clear set of techniques to link learning styles and preferences with development team roles. This paper characterizes the learning style of four traditional roles in software development (Analyst, Architect, Developer, and Project Manager) using Kolb's Learning Styles Inventory. Kolb's Learning Styles Test was administered to 110 software development practitioners (15 analysts, 18 architects, 50 developers, and 27 project managers). The test results show that, with some differences, architects and analysts have the Deciding style, while developers and project managers exhibit the Thinking style. Finally, in alignment with Kolb's learning strengths and challenges, this work provides a set of teaching strategies for each role based on their inferred learning styles.
Mauricio Hidalgo, Kattia Rodriguez, Hernán Astudillo
CLEI3
2024 Security Mechanisms Used in Systems Based on Zero Trust Architecture: A Systematic Mapping
abstract
Zero Trust Architecture (ZTA) is a novel security approach for building secure systems. ZTA-based systems are built with specific security mechanisms to enforce their basic tenets, for example, explicit verification and least privilege. Although existing security mechanisms have been useful in building ZTA-based systems, the current literature does not provide clear guidance on which security mechanisms should be used by developers of these systems. This article describes the design and results of a systematic mapping study to identify the security mechanisms used in the building of ZTA-based systems. The review yielded 290 articles, of which 30 primary studies were selected. Key findings are: (i) 24 different security mechanisms were reported; (ii) 37 % of them are classified into access control techniques to implement ZTA least priveleges tenet; (iii) ABAC and AIM are the most used mechanisms; (iv) over half of security mechanisms (69 %) focus on resisting attacks (instead of detecting or recovering); and (v) experimentation is a predominant empirical strategy within ZTA security research. The identification of these security mechanisms will enable developers of ZTA-based systems to effectively address the security challenges associated with implementing ZTA tenets.
Carlos Manzano, Gastón Marquez, Hernán Astudillo
CLEI3
2024 Causal Organizational Mining in Software Engineering: Evaluating Improvement Strategies in Development Team Dynamics
abstract
In the context of software engineering, analyzing causality in the dynamics of development teams is essential for optimizing performance. Causality involves understanding how certain factors (organizational structures, individual interactions) directly influence the team's performance, allowing for the identification and implementation of effective improvements in development processes. Identifying and understanding the underlying causes of heterogeneous effects in team dynamics is essential for improving collaboration and productivity. The lack of specific methodologies to explore these causal relationships in complex organizational settings limits the ability of project leaders to implement effective changes. This article describes a method that uses causal inference in organizational mining to assess software development teams. Data are analyzed to identify interactions and causal factors affecting role dynamics, and causal inference techniques are employed to evaluate the effects of improvement actions. The effectiveness of this approach was confirmed with a pilot software development team at a Chilean payment processing company. By employing causal organizational analysis methods, managers were able to select more focused strategies, based on a deep and detailed understanding of the underlying causal dynamics. This work contributes to the field of software engineering by introducing a structured and causal approach to analyze team dynamics and providing project managers with tools to address the underlying factors that hinder team effectiveness.
Fernando Montoya, Cristian C. Beltran-Hernandez, Hernán Astudillo
CLEI3
2024 Is Explainability in the Radar of Software Architects? A Rapid Multivocal Review
abstract
With growing system complexity, explainability is being heard more frequently in software development ecosystems. However, it is unclear whether this situation also exists among software architects in industry. This paper reports the design, execution, and findings of a rapid multivocal review exploring perspectives on explainability within current software architecture discourse. Analysis of seven well-known practitioner and academic sources finds that “explainability” is used in discussions about documenting trade-offs, surfacing assumptions, and framing architecture maps, all of which resonate with motivations for comprehensible, adaptable design. These insights may allow the research community to build further bridges with practitioners for sustaining collective system understanding.
Fernando Muraca, María Florencia Pollo Cattaneo, Hernán Astudillo
CLEI3
2024 On the Variability of Microservice Decompositions: A Data-Driven Analysis
abstract
The problem of migrating monolithic applications to microservices has become popular both in industry and academia, particularly when using automated tools to assist developers in the decomposition. While a variety of tools and techniques have been proposed, deciding which is the most appropriate decomposition for a given monolith is challenging because the selected technique can return alternative decompositions depending on how the parameters of that technique are configured. This issue has not received enough attention in the literature, and therefore, developers have to resort to their intuition or use the default parameters reported by the authors of the technique. To investigate this problem further, in this work we perform a study of the parameters and variability of the MicroMiner approach, assessing its parameter sensitivity when dealing with two monolithic applications from the literature. Based on a systematic, data-driven analysis of the landscape of possible decompositions, our results show that, depending on the monolithic application provided as input, certain parameters of MicroMiner have more or less importance on the characteristics of the generated decompositions. These findings provide initial guidelines for developers to configure MicroMiner, as well as other approaches, in order to obtain microservice decompositions with relatively low variability.
Ana C. Martínez Saucedo, Jorge Andrés Díaz Pace, Hernán Astudillo, Guillermo Rodríguez 0002
CLEI3
2023 Extending the SEMAT Kernel to Represent and Assess Software Architecture Evaluations
abstract
Software architecture evaluation (SAE) is a key area in software architecture design. Some of its key challenges are describing and assessing the architecture itself, the architectural decisions, the business or mission goals, and the quality attributes; and further, the adoption itself of SAE practices. The lack of a standard representation for SAE endeavors hampers its adoption by development teams, its (semi-)automated support by tool providers, and its normative assessment by process specialists. In this paper we introduce SAEMET (Software Architecture Evaluation MEthod and Theory), an extension of the Essence standard proposed by SEMAT and adopted by OMG; The Essence kernel defines “things” (called Alphas) any software engineering endeavor should include, and provides an extensible representation to be used for assessing an endeavor progression. SAEMET includes five sub-alphas (Quality Attributes, Business Goals, Architecture Description, Architecture Decision, and Evaluation Adoption), and provides a complete description for each one, their progression levels, and the relationships among them. Our approach is useful for representing an already published architecture review, conducted using DCAR (Decision-Centric Architecture Review method), and assessing its suitability for actionable support of adoption, automated support, and normative assessment. Ongoing empirical evaluation of SAEMET is underway, and early results indicate it is usable and useful for guiding and auditing SAE endeavor, as well as planning courses to train teams for adopting SAE.
Pablo Cruz, Hernán Astudillo, Carlos Mario Zapata Jaramillo
CLEI2
2023 Attention Mechanisms in Process Mining: A Systematic Literature Review
abstract
Process Mining (PM) focuses on monitoring and optimizing long-running business processes by examining their execution event logs (usually complex and heterogeneous) to obtain insights and enable data-driven decisions. Several Machine Learning (ML) techniques have been recently proposed to exploit these logs as learning datasets and enable examination of past events and predict future ones, but their black-box nature makes hard for human analysts to interpret their results and recognize the key parts of input data. Attention mechanisms (AM) is an ML technique that does address these shortcomings, but it has been little used for PM. This article describes the design, results and findings of a systematic literature review of attention mechanisms for PM. We addressed three research questions: (a) for which applications are AM used? (b) which kinds of AM are used? and (c) how are AM combined with other ML techniques? An initial search yield 73 papers, and inclusion/exclusion criteria left sixteen, published between 2017 and 2023. Key finding are that: (1) the most common application is sequence prediction, (2) most studies combine global and item-wise attention, added as layers after an encoder generates the continuous representation, and (3) emerging research topics include anomaly detection and data representation. This study shows that attention mechanisms can help process analysts to get some sense of interpretability, and showcases the bright potential for process mining of attention mechanisms, which paradoxically have received little attention themselves.
Gonzalo Rivera Lazo, Hernán Astudillo, Ricardo Ñanculef
CLEI2
2023 Causal Graph: Interpretation of Causal Relationships in Temporary Deviations of Business Processes
abstract
Process deviations can be difficult and costly to identify. Therefore, it's imperative for organizations to detect temporary deviations during execution and understand their causal relationships. This enables decision-makers to implement targeted corrective actions. This article presents a method to construct a causal graph for the analysis of process variants, combining techniques from process mining, unsupervised machine learning, and causal discovery and inference. This graph is not susceptible to Simpson's paradox, where aggregating the feature space can lead to incorrect interpretations of causal effects. The technique has been initially validated with a well-known event log, namely a loan application process taken from the BPI Challenge 2017 and containing 16,299 records. This pilot run successfully identified the causal variables and their direction. Wider use of this approach will allow organizations to interpret and estimate the causal effect of an action plan on process variants with temporary deviations.
Fernando Montoya, Hernán Astudillo
CLEI2
2023 To Security and Beyond: On The Impacts of Microservice Security Smells and Refactorings
abstract
Microservices gained momentum in enterprise IT, as they enable building cloud-native applications. At the same time, they come with new security challenges, including security smells, viz., symptoms of bad (though often unintentional) design decisions that might affect application security. This study aims to explore the impacts of microservice security smells- and of the refactorings known to mitigate their effects-beyond security. In particular, we systematically elicit possible impacts of smells and refactorings on applications' maintainability, performance efficiency, and adherence to microservices' key design principles. We then validate the elicited impacts by means of an online survey targeting experienced practitioners and researchers. Our main contributions include 35 validated impacts, and a discussion of the survey results geared towards analyzing the (mis)alignment between practitioners and researchers.
Francisco Ponce 0001, Jacopo Soldani, Carla Taramasco, Hernán Astudillo, Antonio Brogi
CLEI4
2019 Assessing Architectural Patterns Trade-offs using Moment-based Pattern Taxonomies
abstract
Large software systems are designed to satisfy or accommodate many requirements; architectural patterns are a well-known technique to reuse design knowledge. However, requested quality attributes (QA) may be inconsistent at times; e.g., high security typically hampers performance and scalability. Thus, a key concern of systems architects is understanding trade-offs among alternative solutions; e.g., a pattern may favor performance at the expense of scalability or security, another may privilege scalability, and yet another may push security. This article argues that the usual organization of individual patterns in topic-related pattern languages is not too helpful to identify trade-offs, and proposes to borrow a taxonomic principle of architectural tactics, organizing the patterns for each QA into “moments”. This enables architects to use simple tradeoff highlighting techniques to understand trade-offs in complex systems. The approach was used in the systematic design of a SCADA-to-ERP secure bridge, where moment-oriented pattern taxonomies for availability, confidentiality, and performance were used. This approach offers the promise of enabling the trade-offenabled, pattern-driven design of large systems by supporting the systematic exploration of trade-offs among patterns for specific QA's.
Cristian Orellana, Mónica M. Villegas, Hernán Astudillo
CLEI3
2019 Security Mechanisms Used in Microservices-Based Systems: A Systematic Mapping
abstract
Microservices is an architectural style that conceives systems as a modular, costumer, independent and scalable suite of services; it offers several advantages but its growing popularity has given rise to security challenges. Building secure systems is greatly helped by deploying existing security mechanisms, but current literature does not guide developers about which mechanisms are actually used by developers of microservices-based systems. This article describes the design and results of a systematic mapping study to identify the security mechanisms used in microservices-based systems described in the literature. The study yielded 321 articles, of which 26 are primary studies. Key findings are that (i) the studies mention 18 security mechanisms; (ii) the most mentioned security mechanisms are authentication, authorization and credentials; and (iii) almost 2/3 of security mechanisms focus on stopping or mitigating attacks, but none on recovering from them. Additionally, it emerges that experiments and case studies are the most used empirical strategies in microservices security research. The clear identification of most-used security solutions will facilitate the reuse of existing architectural knowledge to address security problems in microservices-based systems.
Anelis Pereira Vale, Gastón Marquez, Hernán Astudillo, Eduardo B. Fernández
CLEI3
2017 ATAM-RPG: A role-playing game to teach architecture trade-off analysis method (ATAM)
abstract
Teaching software architecture to undergraduate students is particularly hard because they typically have no experience with medium or large systems with competing stakeholders. A particularly hard case is ATAM (Architecture Trade-off Analysis Method), which allows the evaluation of architectural designs and quality attributes by competing stakeholders. This article describes ATAM-RPG, a role-playing game to support the teaching of ATAM by simulating stakeholder's interaction and trade-offs. The initial ATAM-RPG case incorporates the architecture, scenarios and design trade-offs of the Chilean national tsunami alert system (SNAM). The approach was tested by deploying the SNAM case in undergraduate courses; initial results show that ATAM-RPG was well-evaluated regarding trade-off description and understanding (and especially utility trees). Students also recognized the importance of exercising technically-based negotiation skills. We conclude that role playing games can be fruitfully used for software architecture education.
Claudia Hidalgo Montenegro, Hernán Astudillo, María Clara Gómez-Alvarez
CLEI2
2017 The Benefit of Thinking Small in Big Data
Kurt Englmeier, Hernán Astudillo
DATA2
2016 Identifying potential suppliers for competitive bidding using Latent Semantic Analysis
abstract
Chilecompra is the Chilean governmental agency in charge of Chile's Public Procurement System. One of its main current problems is that around 54% (supply indicator) of the public procurements that are submitted get less than three bids, which is the minimum required so that the procurement process may be free of technical, legal and ethical questionings. This article presents a method based on Natural Language Processing and Semantic Analysis in order to automatically identify prospective bidding candidates from a historical application dataset with a software tool which implements this method. The results are proven by using a dataset of procurement data and historical applications, biddings and adjudications of the suppliers. We compared the number of suppliers that applied for a tender and the number of prospective bidders identified by this method, and the domain procurement experts proved that the identified bidders were relevant to the new procurement.
Victor Aravena-Diaz, Ricardo Gacitúa, Hernán Astudillo, José Emilio Labra Gayo
CLEI3
2016 An approach for software knowledge sharing based on architectural decisions
abstract
In Software Architecture (SA), design decisions are considered valuable knowledge, and Architectural Knowledge (AK) is key to increase software quality. Several proposals address implementation of Knowledge Management for Software Architecture, but very few address sharing and distribution of AK. We propose that AK be shared as concrete usage scenarios, elicited from design meetings; this proposal builds on Knowledge Management sharing insights and on DVIA (Design Verbal Interventions Analysis), which identifies design decisions by mapping verbal interventions in recorded design meetings. Specifically, we show an scenario of decision impact analysis for introducing a new functional requirement. The approach was initially validated in a case study with undergraduate practitioners video-recorded while team-designing a command-and-control center for the pan-Andean spatial project. Initial qualitative results suggest that AK reuse is actually helped by sharing specific design scenarios.
Gilberto Pedraza-Garcia, Hernán Astudillo, Darío Correal
CLEI2
2015 Comparing scalability of message queue system: ZeroMQ vs RabbitMQ
abstract
Modern web apps handle huge and increasing numbers of users and operations. A rise of event-driven architecture approach and message queue systems provide a new alternative to face this scenario evidencing the lack of quantitative measurements comparing performance and scalability between specific message queue products. This article proposes a prototype architecture applied in ZeroMQ and RabbitMQ, used for measure the impact of (1) the number of messages over performance, and (2) the numbers of consuming nodes over scalability. The results show that for both criteria, the degradation threshold of ZeroMQ is higher than RabbitMQ, thus more scalable and faster.
Nicolas Estrada, Hernán Astudillo
CLEI2
2014 Trust-based improved recommendation of IT-related Web resources
abstract
The explosive growth of the Web makes increasingly harder to identify relevant resources among large result sets yield by search engines. Researcher looking for specialized documents have little guidance beyond the documents' own metadata. This article describes the use of trust-based computing to improve ranking and relevance of results by combining topic-specific trust among community members and individual's resources evaluation. The approach has been prototyped with Antu, built over the IT-specific faceted keyword-based search tool Yarquen. An experimental study was conducted with informatics graduate students, and found improved relevance of the suggested documents. This result suggests that trust-based approaches have huge potential to improve recommendations in specialized communities.
Pablo Cruz, Oscar Cornejo 0001, Hernán Astudillo
CLEI3
2014 Analysis of design meetings for understanding software architecture decisions
Gilberto Pedraza-Garcia, Hernán Astudillo, Darío Correal
CLEI2
2014 Time-Based Hesitant Fuzzy Information Aggregation Approach for Decision-Making Problems
abstract
Hesitant fuzzy sets have been proposed as an extension of fuzzy sets to address situations in which decision makers exhibit variations in their alternatives' assessment values. However, in real-world problems, the decision-making process has to be accomplished under situations where these assessment values may also drastically change over time. In this paper, we propose a prioritized aggregation operator to combine a time sequence of hesitant fuzzy information, where the time-based hesitancy due to changing environment is mitigated. The proposed method is applied to the service selection problem in service-based systems, where software architects must select as a group the service that has the best combination of features based on their historical assessments. We claim that the time-based hesitant fuzzy information aggregation method addresses the hesitancy at intra- and interexpert levels obtaining more robust decisions.
Romina Torres, Rodrigo Salas 0001, Hernán Astudillo
Int. J. Intell. Syst.3
2013 Ontology and semantic wiki for an Intangible Cultural Heritage inventory
abstract
The increasing globalization has made the preservation of the Intangible Cultural Heritage (ICH) an urgent need, and the UNESCO's states parties have compromised to make collaborative inventories of ICH. Many traditional inventories become obsolete quickly because they present rigid data models and/or because data adquisition from scarce specialists are required. This article proposes to introduce a participative inventory as a semantic wiki, combining the familiarity of the audience with the free text, the expressive power of ontologies, and the benefits of wikis for the controled social enrichment. The scope has been validated in a pilot inventory with public access, that combine an Ontowiki instance, a new ontology based on UNESCO's clasification, and the data of an existent Chilean ICH inventory. Preliminary results indicate that the semantic-based catalog is much more flexible than existing traditional catalogs. Using these collaborative enrichment will enable active involvement to citizenship in recording their immaterial heritage.
Renzo Stanley, Hernán Astudillo
CLEI2
2008 No mining, no meaning: relating documents across repositories with ontology-driven information extraction
abstract
Far from eliminating documents as some expected, the Internet has lead to a proliferation of digital documents, without a centralized control or indexing. Thus, identifying relevant documents becomes simultaneously more important and much harder, since what users require may be dispersed across many documents and many repositories. This paper describes Ontologic Anchoring, a technique to relate documents in domain ontologies, using named entity recognition (a natural-language processing approach) and semantic annotation to relate individual documents to elements in ontologies. This approach allows document retrieval using domain-level inferences, and integration of repositories with heterogeneous media, languages and structure. Ontological anchoring is a two-way street: ontologies allow semantic indexing of documents, and simultaneously new documents enrich ontologies. The approach is illustrated with an initial deployment for heritage documents in Spanish.
Víctor Codocedo, Hernán Astudillo
ACM Symposium on Document Engineering2