Anj Simmons

dblp:232/1933 · also Andrew J. Simmons, Angie Simmons · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-8402-2853ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Green Runner: A Tool for Efficient Deep Learning Component Selection
abstract
For software that relies on machine-learned functionality, model selection is key to finding the right model for the task with desired performance characteristics. Evaluating a model requires developers to i) select from many models (e.g. the Hugging face model repository), ii) select evaluation metrics and training strategy, and iii) tailor trade-offs based on the problem domain. However, current evaluation approaches are either ad-hoc resulting in sub-optimal model selection or brute force leading to wasted compute. In this work, we present GreenRunner, a novel tool to automatically select and evaluate models based on the application scenario provided in natural language. We leverage the reasoning capabilities of large language models to propose a training strategy and extract desired trade-offs from a problem description. GreenRunner features a resource-efficient experimentation engine that integrates constraints and trade-offs based on the problem into the model selection process. Our preliminary evaluation demonstrates that GreenRunner is both efficient and accurate compared to ad-hoc evaluations and brute force. This work presents an important step toward energy-efficient tools to help reduce the environmental impact caused by the growing demand for software with machine-learned functionality. Our tool is available at Figshare GreenRunner.
Jai Kannan, Scott Barnett, Anj Simmons, Taylan Selvi, Luis Cruz 0002
CAIN3
2024 Comparative analysis of real issues in open-source machine learning projects
abstract
Abstract Context In the last decade of data-driven decision-making, Machine Learning (ML) systems reign supreme. Because of the different characteristics between ML and traditional Software Engineering systems, we do not know to what extent the issue-reporting needs are different, and to what extent these differences impact the issue resolution process. Objective We aim to compare the differences between ML and non-ML issues in open-source applied AI projects in terms of resolution time and size of fix. This research aims to enhance the predictability of maintenance tasks by providing valuable insights for issue reporting and task scheduling activities. Method We collect issue reports from Github repositories of open-source ML projects using an automatic approach, filter them using ML keywords and libraries, manually categorize them using an adapted deep learning bug taxonomy, and compare resolution time and fix size for ML and non-ML issues in a controlled sample. Result 147 ML issues and 147 non-ML issues are collected for analysis. We found that ML issues take more time to resolve than non-ML issues, the median difference is 14 days. There is no significant difference in terms of size of fix between ML and non-ML issues. No significant differences are found between different ML issue categories in terms of resolution time and size of fix. Conclusion Our study provided evidence that the life cycle for ML issues is stretched, and thus further work is required to identify the reason. The results also highlighted the need for future work to design custom tooling to support faster resolution of ML issues.
Tuan Dung Lai, Anj Simmons, Scott Barnett, Jean-Guy Schneider, Rajesh Vasa
Empir. Softw. Eng.2
2021 Towards a taxonomy for annotation of data science experiment repositories
abstract
Data scientists, like software engineers, use search engines, code repositories, tutorials, and question and answer sites for finding code snippets. The objective of this study is to understand what information can be extracted from data science experiment repositories for quicker availability of relevant information when data scientists search for information. In this paper, we investigated a set of notebooks to identify recurring data science techniques for efficient information retrieval and easy adaptation from online solutions to support their search during experimentation. From the manual annotation of 57 natural language processing notebooks, a taxonomy on 106 data science techniques was developed, grouped by data science workflow stages. The preliminary evaluation shows that our constructed taxonomy is relevant to retrieve information that data scientists are searching for. Future work will continue to investigate the creation of a context aware code snippet engine designed for data scientists.
Shangeetha Sivasothy, Scott Barnett, Niroshinie Fernando, Rajesh Vasa, Roopak Sinha, Anj Simmons
SCAM6
2020 Visual Languages for Supporting Big Data Analytics Development
abstract
We present BiDaML (Big Data Analytics Modeling Languages), an integrated suite of visual languages and supporting tool to help end-users with the engineering of big data analytics solutions. BiDaML, our visual notations suite, comprises six diagrammatic notations: brainstorming diagram, process diagram, technique diagrams, data diagrams, output diagrams and deployment diagram. BiDaML tool provides a platform for efficiently producing BiDaML visual models and facilitating their design, creation, code generation and integration with other tools. To demonstrate the utility of BiDaML, we illustrate our approach with a realworld example of traffic data analysis. We evaluate BiDaML using two types of evaluations, the physics of notations and a cognitive walkthrough with several target end-users e.g. data scientists and software engineers.
Hourieh Khalajzadeh, Anj Simmons, Mohamed Almorsy, John C. Grundy, John G. Hosking, Qiang He 0001
ENASE2
2020 A large-scale comparative analysis of Coding Standard conformance in Open-Source Data Science projects
abstract
Background: Meeting the growing industry demand for Data Science requires cross-disciplinary teams that can translate machine learning research into production-ready code. Software engineering teams value adherence to coding standards as an indication of code readability, maintainability, and developer expertise. However, there are no large-scale empirical studies of coding standards focused specifically on Data Science projects. Aims: This study investigates the extent to which Data Science projects follow code standards. In particular, which standards are followed, which are ignored, and how does this differ to traditional software projects? Method: We compare a corpus of 1048 Open-Source Data Science projects to a reference group of 1099 non-Data Science projects with a similar level of quality and maturity. Results: Data Science projects suffer from a significantly higher rate of functions that use an excessive numbers of parameters and local variables. Data Science projects also follow different variable naming conventions to non-Data Science projects. Conclusions: The differences indicate that Data Science codebases are distinct from traditional software codebases and do not follow traditional software engineering conventions. Our conjecture is that this may be because traditional software engineering conventions are inappropriate in the context of Data Science projects.
Anj Simmons, Scott Barnett, Jessica Rivera-Villicana, Akshat Bajaj, Rajesh Vasa
ESEM1
2020 End-User-Oriented Tool Support for Modeling Data Analytics Requirements
abstract
Big data and analytics are increasingly used in different domains to gain insights and to improve decision-making. Developing big data analytics solutions is a complex task involving multidisciplinary teams and users - with no data science and programming background - to professional data scientists and software engineers. Different stakeholders work with a variety of data types, tasks and concepts in different languages from high- level domain concepts to low level programming languages and technical concepts. In order to advance the level of abstraction beyond low-level data analysis technical details, we demonstrate our BiDaML tool. BiDaML brings all stakeholders around one tool to specify, model and document their big data applications using a novel set of domain-specific visual languages (DSVLs).
Hourieh Khalajzadeh, Anj Simmons, Mohamed Almorsy, John C. Grundy, John G. Hosking, Qiang He 0001
VL/HCC2
2017 Spatio-Temporal Reference Frames as Geographic Objects
abstract
It is often desirable to analyse trajectory data in local coordinates relative to a reference location. Similarly, temporal data also needs to be transformed to be relative to an event. Together, temporal and spatial contextualisation permits comparative analysis of similar trajectories taken across multiple reference locations. To the GIS professional, the procedures to establish a reference frame at a location and reproject the data into local coordinates are well known, albeit tedious. However, GIS tools are now often used by subject matter experts who may not have the deep knowledge of coordinate frames and projections required to use these techniques effectively.
Anj Simmons, Rajesh Vasa
SIGSPATIAL/GIS1
2015 Hub Map: A new approach for visualizing traffic data sets with multi-attribute link data
abstract
Visualizing road traffic datasets involves representing junctions, their links, and the attributes of those links. Current traffic visualization techniques are not sufficient for professional traffic engineers, as they are limited in the number of attributes that can be represented. This paper proposes a new approach to visualize multiple attributes on graph edges without compromising their visibility. In particular, we introduce a parameterized connector symbol that increases the number of attributes that can be displayed on graph edges. We demonstrate that our approach can significantly increase the number of traffic parameters that can be displayed compared to existing traffic visualizations.
Anj Simmons, Iman Avazpour, Hai Le Vu 0001, Rajesh Vasa
VL/HCC1