Ciara Breathnach

dblp:169/4761 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0002-4065-0660ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2023 Curating History Datasets and Training Materials as OER: An Experience
abstract
Evidence shows that practice-based learning is beneficial to students’ understanding of threshold concepts in all disciplines. Teaching activities that draw on research data provide students with real-world examples of how they might apply this new knowledge, and this reinforces their understanding. While research data from projects in disciplines such as computer science or economics is widely published, it is only recently that humanities scholars, in particular historians, have started considering publishing their research data in digital format. Using a case study of a 5-year funded collaborative project between historians and computer scientists, this paper discusses how the research data was created and applied in a classroom context, providing students with a ’real-world’ experience of working as a historian. It shows how the project data developed from a basic transcription of historical records into a fully enriched open dataset that can be used by teachers from a range of disciplines including history, historical geography, demography, computer science and medicine/medical humanities. It concludes that lessons planned around such open research data comprise a valuable educational resource to teachers and students.
Ciara Breathnach, Rachel Murphy, Alexander Schieweck, Enda O'Shea, Stuart Clancy, Tiziana Margaria
COMPSAC1
2022 CensusIRL: Historical census data preparation with MDD support
abstract
Census returns are a critical source of information for governments globally. They underpin a wide spectrum of public planning including health, housing, work and education. Historically, census forms have captured names, places, dates, age, occupation, family structure, and religion. In more recent times, sexual orientation and ethnicity, queries that can be intrusive to vulnerable communities, have been added to the criteria, and for such reasons data security is of paramount importance. Most governments restrict access to individual census returns, presenting the data in aggregate report format. The Irish government is particularly strict, enforcing a statutory closure period of 100 years. An exception was made for the Irish 1911 census which were digitised and released for free online consultation in 2009 [1]. They are an excellent source for genealogists and historians alike but exist as separate digital siloes. This project uses an eXtreme Model-Driven Development (XMDD) environment to create linkages between both datasets. It will discuss the development process of the CensusIrl application and the process used in developing the matching algorithm used. We will discuss the census records and the data cleansing process used in creating the initial proof of concept application. We detail the different approaches to the development life-cycle of the application and describe the different utilises used in the sanitation of data points in the records and the match-making process.
Adam J. Doherty, Rachel Murphy, Alexander Schieweck, Stuart Clancy, Ciara Breathnach, Tiziana Margaria
IEEE Big Data5
2022 Evolution of the Historian Data Entry Application: Supporting Transcribathons in the Digital Humanities through MDD
abstract
Death and Burial Data: Ireland 1864–1922 (DBDIrl), is a digital humanities project, which uses historical civil registration of death as its primary dataset. The overarching aim of this project is to provide enriched and clean historical Irish data for analysis, in a eXtreme Model-Driven Development (XMDD) fashion. This paper discusses how e-learning environments were used to enrich these partially indexed data in an online, hybrid and blended learning group instruction format over four years. It describes how the DBDIrl data entry application, called Historian Dime App (HDA), evolved over a number of iterations to create a more user friendly interface, in an interdisciplinary collaboration of historians and computer scientists enabled by the XMDD approach. It discusses how the development process of HDA benefitted successive cohorts of history students engaged in a curricular Practice-based learning (PBL) project that follows a transcribathon model as defined by the Folger Library11https://folgerpedia.folger.edu/Transcribathon, We adapted the model for postgraduate teaching and learning in the humanities and took a reflexive approach to student/user feedback to evolve the HDA over four versions. It resulted in enhanced features, higher rates of user satisfaction, and a more responsive data curation and storage mechanism. This effort achieved our original aim of obtaining clean and accurate outputs from the students' project work.
Alexander Schieweck, Rachel Murphy, Rafflesia Khan, Ciara Breathnach, Tiziana Margaria
COMPSAC4
2021 Transcribathons as Practice-Based Learning for Historians and Computer Scientists
abstract
This paper discusses the collaboration of higher education computer scientists, data scientists and historians to design and adapt a set of resources and digital assets that can be used by students and local communities to transcribe historical data. We show that an agile software design and development approach to co-creating such tools has great pedagogical merit both for the end users and the creators of the tools themselves. We note that effective interdisciplinary collaboration requires pooling resources and mutual respect for domain expertise. As part of the project took place during the SARs-Cov-2 pandemic, we also pivoted to a completely online environment. This means that we have a proven model for classroom, online and blended formats alike that we intend to use in the future also for "citizen scientist" events. In this way, the Open Education Resources can be used in a more accessible and inclusive way.
Ciara Breathnach, Rachel Murphy, Tiziana Margaria
COMPSAC1
2020 Towards Automatic Data Cleansing and Classification of Valid Historical Data An Incremental Approach Based on MDD
abstract
The project Death and Burial Data: Ireland 1864-1922 (DBDIrl) examines the relationship between historical death registration data and burial data to explore the history of power in Ireland from 1864 to 1922. Its core Big Data arises from historical records from a variety of heterogeneous sources, some aspects are pre-digitized and machine readable. A huge data set (over 4 million records in each source) and its slow manual enrichment (ca 7,000 records processed so far) pose issues of quality, scalability, and creates the need for a quality assurance technology that is accessible to non-programmers. An important goal for the researcher community is to produce a reusable, high-level quality assurance tool for the ingested data that is domain specific (historic data), highly portable across data sources, thus independent of storage technology.This paper outlines the step-wise design of the finer granular digital format, aimed for storage and digital archiving, and the design and test of two generations of the techniques, used in the first two data ingestion and cleaning phases.The first small scale phase was exploratory, based on metadata enrichment transcription to Excel, and conducted in parallel with the design of the final digital format and the discovery of all the domain-specific rules and constraints for the syntax and semantic validity of individual entries. Excel embedded quality checks or database-specific techniques are not adequate due to the technology independence requirement. This first phase produced a Java parser with an embedded data cleaning and evaluation classifier, continuously improved and refined as insights grew. The next, larger scale phase uses a bespoke Historian Web Application that embeds the Java validator from the parser, as well as a new Boolean classifier for valid and complete data assurance built using a Model-Driven Development technique that we also describe. This solution enforces property constraints directly at data capture time, removing the need for additional parsing and cleaning stages. The new classifier is built in an easy to use graphical technology, and the ADD-Lib tool it uses is a modern low-code development environment that auto-generates code in a large number of programming languages. It thus meets the technology independence requirement and historians are now able to produce new classifiers themselves without being able to program. We aim to infuse the project with computational and archival thinking in order to produce a robust data set that is FAIR compliant (Free Accessible Inter-operable and Re-useable).
Enda O'Shea, Rafflesia Khan, Ciara Breathnach, Tiziana Margaria
IEEE BigData3