Arumoy Shome

dblp:251/3802 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0002-3778-5150ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 Towards Understanding Machine Learning Testing in Practise
abstract
Visualisations drive all aspects of the Machine Learning (ML) Development Cycle but remain a vastly untapped resource by the research community. ML testing is a highly interactive and cognitive process which demands a human-in-the-loop approach. Besides writing tests for the code base, bulk of the evaluation requires application of domain expertise to generate and interpret visualisations. To gain a deeper insight into the process of testing ML systems, we propose to study visualisations of ML pipelines by mining Jupyter notebooks. We propose a two prong approach in conducting the analysis. First, gather general insights and trends using a qualitative study of a smaller sample of notebooks. And then use the knowledge gained from the qualitative study to design an empirical study using a larger sample of notebooks. Computational notebooks provide a rich source of information in three formats—text, code and images. We hope to utilise existing work in image analysis and Natural Language Processing for text and code, to analyse the information present in notebooks. We hope to gain a new perspective into program comprehension and debugging in the context of ML testing.
Arumoy Shome, Luis Cruz 0002, Arie van Deursen
CAIN1
2022 Data smells in public datasets
abstract
The adoption of Artificial Intelligence (AI) in high-stakes domains such as healthcare, wildlife preservation, autonomous driving and criminal justice system calls for a data-centric approach to AI. Data scientists spend the majority of their time studying and wrangling the data, yet tools to aid them with data analysis are lacking. This study identifies the recurrent data quality issues in public datasets. Analogous to code smells, we introduce a novel catalogue of data smells that can be used to indicate early signs of problems or technical debt in machine learning systems. To understand the prevalence of data quality issues in datasets, we analyse 25 public datasets and identify 14 data smells.
Arumoy Shome, Luis Cruz 0002, Arie van Deursen
CAIN1
2019 ACE: Art, Color and Emotion
abstract
We present ACE, the Art, Color and Emotion browser. ACE is a data driven web based platform for exploring the visual sentiment and emotion in artistic paintings over time. To that end, we train our own visual artistic sentiment extraction model by leveraging the artworks from the OmniArt dataset. With our model we are able to estimate the overall sentiment dominating in groups of artworks belonging to a specific time interval. To make the results interactive and explorable we designed an intuitive interface with a carefully considered shape, color and element placement enforcing a top-down interaction scheme. Moreover, we perform extensive control on resource utilisation to provide the smoothest possible user experience and quality of service while using ACE.
Gjorgji Strezoski, Arumoy Shome, Riccardo Bianchi, Shruti Rao, Marcel Worring
ACM Multimedia2