EDBT 2026 Demo / reviewers in the wild / expert
Greg Jansen
dblp:170/0283 · also Gregory Jansen, Gregory N. Jansen
· DBLP profile ↗
7ranked-venue papers in the field
3as first author
2since 2021 · last 2024
0000-0001-6591-6595ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 7 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Sifting US Census Records with Computer Vision and Machine LearningabstractThis paper shares the culmination of my work to computationally enhance researcher access to U. S. Census records, by targeting their personal transcription labor on those document pages that are most likely to contain relevant information. Much research on the United States population over time concerns demographic groups that may be identified, for example, through the race column on census population schedules, which are the handwritten forms on which census takers would record household information. This project was created to support the researcher efforts of Dr. Richard Marciano and the study of the community impact of the forced relocation of Japanese American households during the second world war. In particular through a detailed comparison between the Japanese American households and people recorded in 1940 and in 1950 Sacramento California. While the census forms have a different layout in each decade, the general design is tabular with rows and columns that may be used to visually segment the document. This paper, and the code notebooks that are published along with it, demonstrate a computer vision technique for segmenting population schedules to extract the individual cell images from their race column. Then the individual cell images are cleaned up and fed into two different neural network models, for identifying the handwritten race code within them. Finally, we created a user interface that allows a researcher to perform a visual review of uncertain results from the above process and thereby create a reliable dataset containing only those population schedule pages that pertain to their research. The Python code notebooks that were used to perform this analysis and the review process are linked within the paper and are freely available for reuse under a Creative Commons share-alike license. Greg Jansen |
IEEE Big Data | 1 |
| 2021 | A Framework for Unlocking and Linking WWII Japanese American Incarceration Biographical DataabstractEntity Resolution (ER) is increasingly being used to identify and link names across archival collections. We describe a framework for unlocking and linking biographical data from WWII Japanese American Incarceration Camps using Entity Resolution and other computational approaches. We demonstrate the construction of social graphs that link people, places, and events and which support further scholarship and reveal hidden stories in historical events, especially given contested archival sources. Finally, we show the power of computational analysis to recreate event networks and represent movement of people using maps. This type of modeling is captured through interactive Jupyter Notebooks that integrate these various elements and document our interpretation of Japanese American experiences and events at the Tule Lake concentration camp. Lencia Beltran, Emily Ping O'Brien, Greg Jansen, Richard Marciano |
IEEE BigData | 3 |
| 2019 | Digital Legacies on Paper: Reading Punchcards with Computer VisionabstractWe describe the development of a computer visionbased workflow for normalizing images of the legacy punchcard data format (IBM 029 - 80 column punchcard standard) and then reading the encoded data. We show the role of a newly developed Punchcard Extractor Tool within the Brown Dog service API. We also point to our showcase of these same computer vision techniques in a Jupyter notebook system. Greg Jansen |
IEEE BigData | 1 |
| 2019 | Using Data Partitions and Stateless Servers to Scale Up Fedora RepositoriesabstractWe describe the development and testing of the next-generation Trellis Linked Data Platform with Memento versioning support. In addition to highlighting several features that set this system apart from others, we elaborate on the extensive testing and compatibility work that was done in order to align this system with the Fedora 5.0 specification. We draw attention to the performance and scaling features provided by the Trellis Linked Data Platform in general and by the Cassandra database back end. We review the profound impact that such a system can have on demanding, next generation use cases, such as crowd sourcing, machine learning, and direct file access by desktop applications. Greg Jansen, Aaron Coburn, Adam Soroka, Richard Marciano |
IEEE BigData | 1 |
| 2018 | A Case Study in Creating Transparency in Using Cultural Big Data: The Legacy of Slavery ProjectabstractThe Maryland State Archives (MSA) and the Digital Curation Innovation Center (DCIC) of the University of Maryland's iSchool are collaborating on a digital project that utilizes digital strategies and technologies to create an in-depth understanding of the African-American experience in Maryland during the era of slavery. Utilizing crowdsourcing for transcription, data cleaning and transformation techniques, and data visualization strategies, the joint project team is creating new avenues for understanding the complex web of relationships that undergirded the institution of slavery. iSchool students, full participants on the project team, are learning digital curation and other technical skills while gaining insights into the multiple uses of how cultural Big Data can penetrate the past and illuminate the present. Ryan Cox, Sohan Shah, William Frederick, Tammie Nelson, Will Thomas, Greg Jansen, Noah Dibert, Michael Kurtz, Richard Marciano |
IEEE BigData | 6 |
| 2015 | Mixed-initiative social media analytics at the World Bank: Observations of citizen sentiment in Twitter data to explore "trust" of political actors and state institutions and its relationship to social protestabstractThis paper discusses a project that studied the relationship between citizen trust and social protest using visual analysis of approximately 11 million sentiment classified Tweets from the period of the 2014 Brazilian World Cup. The results of the study reveal that the 2014 World Cup protests in Brazil sprang from a wide range of grievances coupled with a relative sense of deprivation compared with emergent comparative `standards'. This sense of grievance gave rise to sentiments that activated online protest that may have led to other forms of social protest, such as demonstrations. The paper describes an innovative approach to big data analytics-mixed initiative social media analytics - and discusses the potential of using big data in social science research of this kind, as well as some of the open methodological, technical and ethical issues still to be addressed. Nadya A. Calderón, Brian D. Fisher, Jeff Hemsley, Billy Ceskavich, Greg Jansen, Richard Marciano, Victoria L. Lemieux |
IEEE BigData | 5 |
| 2015 | Brown Dog: Leveraging everything towards autocurationabstractWe present Brown Dog, two highly extensible services that aim to leverage any existing pieces of code, libraries, services, or standalone software (past or present) towards providing users with a simple to use and programmable means of automated aid in the curation and indexing of distributed collections of uncurated and/or unstructured data. Data collections such as these encompassing large varieties of data, in addition to large amounts of data, pose a significant challenge within modern day "Big Data" efforts. The two services, the Data Access Proxy (DAP) and the Data Tilling Service (DTS), focusing on format conversions and content based analysis/extraction respectively, wrap relevant conversion and extraction operations within arbitrary software, manages their deployment in an elastic manner, and manages job execution from behind a deliberately compact REST API. We describe both the motivation and need/scientific drivers for such services, the constituent components that allow for arbitrary software/code to be used and managed, and lastly an evaluation of the systems capabilities and scalability. Smruti Padhy, Greg Jansen, Jay Alameda, Edgar F. Black, Liana Diesendruck, Mike Dietze, Praveen Kumar 0002, Rob Kooper, Jong Lee, Richard Marciano, Luigi Marini, Dave Mattson, Barbara S. Minsker, Christopher M. Navarro, Marcus Slavenas, William C. Sullivan, Jason Votava, Inna Zharnitsky, Kenton McHenry |
IEEE BigData | 2 |