Abdussalam Alawini

dblp:147/2931 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
17since 2021 · last 2024
0000-0003-1106-7190ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 GA2Graph: A Data-Driven Approach to Visualizing and Analyzing Collaborative Learning
abstract
This innovative practice full paper describes an approach to visualizing and analyzing collaborative learning. Collaborative learning is vital for enriching computer science education and developing key skills. However, educators often lack tools for effectively tracking and analyzing student engagement in group work. Our study introduces GA2Graph, a data-driven system that uses Ne04j to visualize and analyze online collaborative patterns in classroom activities. It graphically represents students and their interactions, helping educators understand group dynamics. We tested this system in a large-enrollment database course, revealing that despite emphasis on assigned roles like those in POGIL, students frequently disregarded them and adopted more individualistic or varied collaborative strategies.
Abdussalam Alawini, Isa Hajara-Yasmin, Jiabao Xu, Yutong Zhang 0011, Zhijun Zhao
FIE1
2024 Clustering Entity Relationship Diagrams: Enhancing Feedback Quality and Grading Consistency in Large Database Courses
abstract
This innovative practice full paper introduces a tool for clustering Entity Relationship Diagrams (ERDs) and explores its application in large classes. ERDs are fundamental for database design in courses related to databases, data science, and software engineering. However, processing ERD homework submissions in large classes poses significant challenges due to the variety of design decisions made by students, leading to numerous diagram variations. This paper presents an ERD clustering tool designed to group similar ERD submissions, aiding instructors and teaching assistants in identifying popular solutions and common mistakes. The tool employs advanced object detection, OCR, and clustering technologies. We evaluated the tool using four datasets from two large public U.S. universities, with submissions ranging from 130 to 430 diagrams. Various clustering methodologies were assessed, highlighting the importance of incorporating ERD structure into the clustering process. Our findings indicate that the tool successfully generated adequate clusters, and that aiming for 10 clusters is appropriate regardless of the dataset size. The generated clusters included common approaches and mistakes, proving helpful for providing feedback and simplifying the grading process.
Sohum Thadani, Andrey Shor, Soohong Ahn, Abdussalam Alawini, Hisham Benotman
FIE5
2024 Optimizing SQL Learning: Identifying Prime Study Times Using Time-Data Analysis
abstract
This research full paper explores the optimal times for studying and learning Structured Query Language (SQL), a critical skill in managing relational databases across various domains, to improve learning, problem-solving, and academic performance. By identifying prime study times, productivity and retention can be enhanced, particularly by scheduling breaks when cognitive function wanes. This study analyzes over 129,000 SQL submissions from students in a Fall 2022 Database Systems course at the University of Illinois Urbana-Champaign, examining correlations between time of day and answer accuracy (correct, syntax error, or semantic error). Time series analysis was done to analyze the data collected over a set of intervals, and null-hypothesis significance testing was utilized to calculate the p-value, determining the statistical significance of the collected data for drawing confident conclusions.
Sophia Yang, Colin Li, Abdussalam Alawini
FIE3
2024 Curriculum Analysis for Data Systems Education
abstract
The field of data systems has seen quick advances due to the popularization of data science, machine learning, and real-time analytics. In industry contexts, system features such as recommendation systems, chatbots and reverse image search require efficient infrastructure and data management solutions. Due to recent advances, it remains unclear (i) which topics are recommended to be included in data systems studies in higher education, (ii) which topics are a part of data systems courses and how they are taught, and (iii) which data-related skills are valued for roles such as software developers, data engineers, and data scientists. This working group aims to answer these points to explain the state of data systems education today and to uncover knowledge gaps and possible discrepancies between recommendations, course implementations, and industry needs. We expect the results to be applicable in tailoring various data systems courses to better cater to the needs of industry, and for teachers to share best practices.
Daphne Miedema, Toni Taipalus, Vangel V. Ajanovski, Abdussalam Alawini, Martin Goodfellow, Michael Liut, Svetlana Peltsverger, Tiffany Young
ITiCSE (2)4
2024 Exploring Computing Students' Sense of Belonging Before and After a Collaborative Learning Course
abstract
Prior work has found that women tend to report lower sense of belonging compared to men in STEM and computing contexts, which may discourage women's persistence. Collaborative learning has been shown to improve students' sense of belonging in some STEM and computing courses relative to traditional lecturing; however, these studies tend to focus on a single course or the first implementation of such pedagogical changes. Our study explores whether these trends generalize by measuring students' sense of belonging across three non-introductory computing courses that have consistently used collaborative learning activities over three semesters. We ask the following research question: Is collaborative learning generally associated with an increased sense of belonging, especially for women? We found that while there were variations across courses, students' reported sense of belonging improved in all courses. Notably, women's reported sense of belonging improved 15% whereas men's reported sense of belonging improved by 11%. Our findings complement prior studies by providing evidence of a relationship between increased sense of belonging and collaborative learning, and suggest students' sense of belonging is malleable beyond the first year. These findings challenge critiques of past studies as being isolated to single courses or conducted only immediately after an effort to change a course, suggesting pedagogical changes may hold promise in improving students' affective outcomes.
Morgan M. Fong, Shan Huang 0008, Abdussalam Alawini, Mariana Silva, Geoffrey L. Herman
SIGCSE (1)3
2023 Beyond Courses: Towards Supporting Goal-Oriented Learning in MOOC Platforms
abstract
Existing Massive Open Online Course (MOOC) platforms limit their course offering to degree programs and certifications. The course design of existing MOOCs adopts a one-size-fits-all pedagogy where all learners watch the same video lectures and complete the same assignments regardless of their learning goals or background. However, research shows that MOOC learners may have none-traditional learning goals, such as exploring a course or learning specific concepts within a course. The design of current MOOC platforms lacks support to learners who want to utilize MOOCs as modularized resources instead of regular courses. Learners need to manually locate the educational materials that address their learning goals which can be time-consuming and sometimes confusing. To support the diverse needs of MOOC learners, we introduce a customized learning approach, CustomLearn, which facilitates goal-oriented learning in MOOC platforms. Our approach of customizing learning proposes four learning goals: Exploring (explore a subject), Conforming (learn popular topics in a subject), Mastering (being expert in a subject), and Personalizing (creating a personalized study plan). These learning goals use subjects to deliver customized learning plans that span multiple courses. Thus, the customized learning approach organizes similar courses to make the educational materials easily accessible by learners. To assess our proposed customized learning approach, we developed an early prototype of CustomLearn and used it to conduct a user study with MOOC learners to evaluate the viability of CustomLearn and learners' acceptability of heterogeneous study plans. Our evaluation indicated that CustomLearn is a viable approach, evident by participants' high perceptions of the usefulness of CustomLearn and their willingness to use and recommend it if offered by MOOC platforms. Additionally, most participants found heterogeneous study plans acceptable for learning any topic as long as these plans maintain high instructional quality.
Fareedah Alsaad, Abdussalam Alawini
FIE2
2023 Assessing Student Learning Across Various Database Query Languages
abstract
Previous research has shown that students encounter difficulties when learning database systems and their corresponding languages. Researchers have categorized these challenges into syntax and semantic errors and have identified common error types and overall learning obstacles among students. However, most existing studies have primarily focused on quantitatively assessing students‘ overall performance in an aggregated manner’ which may overlook valuable insights into individual-level knowledge transfer. In this study, we scrutinized over 250,000 submissions to query language programming assignments, their corresponding error messages, and the performance data of 702 students who took a database course in the Fall 2022 semester at the University of Illinois Urbana-Champaign to gain a comprehensive overview of each student's performance. We followed each student's progress in semantic and syntax errors across three query languages to determine their overall learning experience and whether knowledge transfer had occurred. Consequently, we discovered that many students may still encounter difficulties when transferring their knowledge from one language to another, despite having already learned and practiced the same abstract data operation concepts in one language. On the other hand, the majority of students were able to reduce syntax errors through practice in one language, but the rate of improvement varied among individuals. This study seeks to investigate two key aspects: the potential transfer of abstract data operation concepts among different database languages, and the possibility of a decrease in syntax errors through consistent practice within a single query language.
Zepei Li, Sophia Yang, Kathryn I. Cunningham, Abdussalam Alawini
FIE4
2023 Comparison of Student Learning Outcomes Among SQL Problem-Solving Patterns
abstract
Structured Query Language (SQL) plays a pivotal role in the effective management of relational databases and is a key skill across domains that engage with database systems, including research, development, and business management. However, mastering SQL can be challenging. To comprehend the approaches employed by students when solving SQL problems and address the challenges they faced during the learning process, our study analyzes submissions from the Database Systems course at the University of Illinois Urbana-Champaign during the Fall 2022 semester. We extend prior research involving line chart visualizations that facilitate instructors in identifying struggling students and understanding their submission behaviors. Yet, we acknowledge the limitations of this approach in providing timely feedback and actionable insights due to the sheer volume of visualizations. To address this, we developed an innovative technique using global sequence alignment scores and regular expression algorithms to compress student submission sequences. Our approach reveals submission patterns and pattern elements, leading to recommendations for instructors to enhance database education. By integrating student performance data, such as the number of submission attempts on a particular SQL problem and whether the student arrived at a correct final solution query, we aim to empirically support these recommendations, thereby enabling instructors to more accurately differentiate between struggling and excelling students.
Sophia Yang, Geoffrey L. Herman, Abdussalam Alawini
FIE3
2023 Uncovering Patterns of SQL Errors in Student Assignments: A Comparative Analysis of Different Assignment Types
abstract
Structured Query Language (SQL) is an essential skill to acquire for those who interact with databases, such as researchers, developers, and people involved in businesses. However, the challenges that these users face while learning SQL requires further research. In particular, the types of errors that students encounter on various assignment types or under exam conditions are an area that we are interested in to determine an optimal arrangement of coursework materials for improved learning. In this paper, we analyze 156,513 student SQL submissions to homework assignments, collaborative assignments, and exams of the Database Systems course available to 730 upper-level undergraduate and graduate students offered in the Fall 2022 semester at the University of Illinois Urbana-Champaign. We look at the ratio of syntax and semantic errors, and correct submissions for each of these assignment problem types as well as the most frequent syntax error codes. We visualize our data findings and draw recommendations for future coursework arrangements from the comparisons between the assignment types for a more effective acquisition of SQL as a skill. We found that although students most commonly encountered syntax error codes 1064 and 1054 regardless of the assignment type, they made more syntax errors (and fewer semantic errors) on exam problems compared with homework and collaborative assignment problems. We recommend instructors place a higher emphasis on non-timed SQL programming problems, targeted syntax drills during instruction, and syntax support during exams.
Sophia Yang, Zepei Li, Geoffrey L. Herman, Kathryn I. Cunningham, Abdussalam Alawini
FIE5
2023 Preparing Computer Science Education PhD Students: Our Process
abstract
Training the growing number of Computer Science Education (CSEd) PhD students is a pressing concern for our community. To meet the needs of our CSEd PhD students at University of Illinois Urbana-Champaign, we have developed a new course designed to strengthen students’ foundation in relevant fields. Through a collaborative process, we developed a reading list that covers the educational theory and perspectives that most inform our own work, as well as concepts that prepare our graduates to engage with the broader CSEd community.
Kathryn I. Cunningham, Colleen M. Lewis, Geoffrey L. Herman, Craig B. Zilles, Abdussalam Alawini
ICER (2)5
2022 The Effects of Teaching Modality on Collaborative Learning: A Controlled Study
abstract
This Research Full Paper presents our findings of studying the effects of teaching modality on collaborative learning by comparing data from two sections of a Database Systems course offered simultaneously, with one offered fully face-to-face in a classroom setting while the other is offered online through a flipped-classroom model. Both sections utilized a collaborative learning approach where students work on group activities for part of the class meeting. Since the two sections were almost identical except for the teaching modality, we are provided with a unique opportunity to study the effect of teaching modalities on collaborative learning. As part of this study, we analyze four crucial data sources: 1) student performance data from the grade book 2) student performance data from the online learning management platform 3) an end-of-semester survey given by the instructor and 4) an end-of-semester survey given by the university. We extract insights on the impact of teaching modalities on collaborative learning in order to identify factors that can enhance collaborative learning. We also study the effect of teaching modalities on students’ performance. We visualize our findings to differentiate between the two modalities, and draw on the strengths of each section to establish recommendations for the instructors for course improvement efforts.
Sophia Yang, Yongjoo Park, Abdussalam Alawini
FIE3
2021 Topic Transitions in MOOCs: An Analysis Study
Fareedah Alsaad, Thomas Reichel, Abdussalam Alawini
EDM4
2021 Echelon: An AI Tool for Clustering Student-Written SQL Queries
abstract
As part of teaching SQL, instructors often rely on auto-grading systems for marking students' assignments. However, such systems lack essential insights into the approaches students use to solve these assignments, allowing subtle flaws in student intuition to go unseen. Further, manual analysis of students' code submissions ranges from costly to impossible, depending on the assessments' frequency. In this paper, we present a system capable of extracting features that instructors deem significant from students' SQL queries and using them to generate clusters that capture the key approaches taken. To supplement this, we project the extracted information to an interactive dashboard and demonstrate its usefulness in allowing database systems professors and teaching staff to quickly identify trends in students' solutions.
Matthew Weston, Haorong Sun, Geoffrey L. Herman, Hisham Benotman, Abdussalam Alawini
FIE5
2021 Insights from Student Solutions to MongoDB Homework Problems
abstract
We analyze submissions for homework assignments of 527 students in an upper-level database course offered at the University of Illinois at Urbana-Champaign. The ability to query databases is becoming a crucial skill for technology professionals and academics. Although we observe a large demand for teaching database skills, there is little research on database education. Also, despite the industry's continued demand for NoSQL databases, we have virtually no research on the matter of how students learn NoSQL databases, such as MongoDB. In this paper, we offer an in-depth analysis of errors committed by students working on MongoDB homework assignments over the course of two semesters. We show that as students use more advanced MongoDB operators, they make more Reference errors. Additionally, when students face a new functionality of MongoDB operators, such as \texttt\$group operator, they usually take time to understand it but do not make the same errors again in later problems. Finally, our analysis suggests that students struggle with advanced concepts for a comparable amount of time. Our results suggest that instructors should allocate more time and effort for the discussed topics in our paper.
Ridha Alkhabaz, Seth Poulsen, Abdussalam Alawini
ITiCSE (1)4
2021 A Quantitative Analysis of Student Solutions to Graph Database Problems
abstract
As data grow both in size and in connectivity, the interest to use graph databases in the industry has been proliferating. However, there has been little research on graph database education. In response to the need to introduce college students to graph databases, this paper is the first to analyze students' errors in homework submissions of queries written in Cypher, the query language for Neo4j---the most prominent graph database. Based on 40,093 student submissions from homework assignments in an upper-level computer science database course at one university, this paper provides a quantitative analysis of students' learning when solving graph database problems. The data shows that students struggle the most to correctly use Cypher's WITH clause to define variable names before referencing in the WHERE clause and these errors persist over multiple homework problems requiring the same techniques, and we suggest a further improvement on the classification of syntactic errors.
Seth Poulsen, Ridha Alkhabaz, Abdussalam Alawini
ITiCSE (1)4
2021 Analyzing Patterns in Student SQL Solutions via Levenshtein Edit Distance
abstract
Structured Query Language (SQL), the standard language for relational database management systems, is an essential skill for software developers, data scientists, and professionals who need to interact with databases. SQL is highly structured and presents diverse ways for learners to acquire this skill. However, despite the significance of SQL to other related fields, little research has been done to understand how students learn SQL as they work on homework assignments. In this paper, we analyze students' SQL submissions to homework problems of the Database Systems course offered at the University of Illinois at Urbana-Champaign. For each student, we compute the Levenshtein Edit Distances between every submission and their final submission to understand how students reached their final solution and how they overcame any obstacles in their learning process. Our system visualizes the edit distances between students' submissions to a SQL problem, enabling instructors to identify interesting learning patterns and approaches. These findings will help instructors target their instruction in difficult SQL areas for the future and help students learn SQL more effectively.
Sophia Yang, Ziyuan Wei, Geoffrey L. Herman, Abdussalam Alawini
L@S4
2021 A Quantitative Analysis of Student Solutions to Graph Database Queries
abstract
As data grow both in size and in connectivity, the interest to use graph databases in industry has been growing rapidly. However, there has been little research on graph database education. In response to the need to introduce college students to graph databases, this paper is the first to analyze students' errors in their submissions writing different types of queries in Cypher, the query language for Neo4j--the most prominent graph database. Based on 40,093 student submission from homework assignments in an upper-level computer science database course at University of Illinois at Urbana-Champaign, this paper provides qualitative insights and quantitative analysis about students' learning when solving graph database problems. The data shows that writing more complex queries initially takes students more time and more attempts to write more complex queries correctly. Additionally, students struggle to correctly use Cypher's WITH clause to define variable names before referencing in the WHERE clause, and these errors persist over multiple homework problems requiring the same techniques.
Seth Poulsen, Ridha Alkhabaz, Abdussalam Alawini
SIGCSE4
2020 Unsupervised Approach for Modeling Content Structures of MOOCs
Fareedah Alsaad, Abdussalam Alawini
EDM2
2020 Insights from Student Solutions to SQL Homework Problems
abstract
We analyze the submissions of 286 students as they solved Structured Query Language (SQL) homework assignments for an upper-level databases course. Databases and the ability to query them are becoming increasingly essential for not only computer scientists but also business professionals, scientists, and anyone who needs to make data-driven decisions. Despite the increasing importance of SQL and databases, little research has documented student difficulties in learning SQL. We replicate and extend prior studies of students' difficulties with learning SQL. Students worked on and submitted their homework through an online learning management system with support for autograding of code. Students received immediate feedback on the correctness of their solutions and had approximately a week to finish writing eight to ten queries. We categorized student submissions by the type of error, or lack thereof, that students made, and whether the student was eventually able to construct a correct query. Like prior work, we find that the majority of student mistakes are syntax errors. In contrast with the conclusions of prior work, we find that some students are never able to resolve these syntax errors to create valid queries. Additionally, we find that students struggle the most when they need to write SQL queries related to GROUP BY and correlated subqueries. We suggest implications for instruction and future research.
Seth Poulsen, Liia Butler, Abdussalam Alawini, Geoffrey L. Herman
ITiCSE3
2019 Fine-Grained Provenance for Matching & ETL
abstract
Data provenance tools capture the steps used to produce analyses. However, scientists must choose among workflow provenance systems, which allow arbitrary code but only track provenance at the granularity of files; provenance APIs, which provide tuple-level provenance, but incur overhead in all computations; and database provenance tools, which track tuple-level provenance through relational operators and support optimization, but support a limited subset of data science tasks. None of these solutions are well suited for tracing errors introduced during common ETL, record alignment, and matching tasks - for data types such as strings, images, etc. Scientists need new capabilities to identify the sources of errors, find why different code versions produce different results, and identify which parameter values affect output. We propose PROVision, a provenance-driven troubleshooting tool that supports ETL and matching computations and traces extraction of content within data objects. PROVision extends database-style provenance techniques to capture equivalences, support optimizations, and enable selective evaluation. We formalize our extensions, implement them in the PROVision system, and validate their effectiveness and scalability for common ETL and matching tasks.
Abdussalam Alawini, Zachary G. Ives
ICDE2
2019 ProvCite: Provenance-based Data Citation
abstract
As research products expand to include structured datasets, the challenge arises of how to automatically generate citations to the results of arbitrary queries against such datasets. Previous work explored this problem in the context of conjunctive queries and views using a Rewriting-Based Model (RBM). However, an increasing number of scientific queries are aggregate, e.g. statistical summaries of the underlying data, for which the RBM cannot be easily extended. In this paper, we show how a Provenance-Based Model (PBM) can be leveraged to 1) generate citations to conjunctive as well as aggregate queries and views; 2) associate citations with individual result tuples to enable arbitrary subsets of the result set to be cited ( fine-grained citations ); and 3) be optimized to return citations in acceptable time. Our implementation of PBM in ProvCite shows that it not only handles a larger class of queries and views than RBM, but can outperform it when restricted to conjunctive views in some cases.
Yinjun Wu, Abdussalam Alawini, Daniel Deutch, Tova Milo, Susan B. Davidson
Proc. VLDB Endow.2
2018 Data Citation: Giving Credit Where Credit is Due
abstract
An increasing amount of information is being published in structured databases and retrieved using queries, raising the question of how query results should be cited. Since there are a large number of possible queries over a database, one strategy is to specify citations to a small set of frequent queries - citation views - and use these to construct citations to other "general" queries. We present three approaches to implementing citation views and describe alternative policies for the joint, alternate and aggregated use of citation views. Extensive experiments using both synthetic and realistic citation views and queries show the trade-offs between the approaches in terms of the time to generate citations, as well as the size of the resulting citation. They also show that the choice of policy has a huge effect both on performance and size, leading to useful guidelines for what policies to use and how to specify citation views.
Yinjun Wu, Abdussalam Alawini, Susan B. Davidson, Gianmaria Silvello
SIGMOD Conference2
2017 Automating Data Citation in CiteDB
abstract
An increasing amount of information is being collected in structured, evolving, curated databases, driving the question of how information extracted from such datasets via queries should be cited. While several databases say how data should be cited for web-page views of the database, they leave it to users to manually construct the citations. Furthermore, they do not say how data extracted by queries other than web-page views -- general queries -- should be cited. This demo shows how citations can be specified for a small set of views of the database, and used to automatically generate citations for general queries against the database.
Abdussalam Alawini, Susan B. Davidson, Yinjun Wu
Proc. VLDB Endow.1
2015 Towards automated prediction of relationships among scientific datasets
abstract
Before scientists can analyze, publish, or share their data, they often need to determine how their datasets are related. Determining relationships helps scientists identify the most complete version of a dataset, detect versions of datasets that complement each other, and determine multiple datasets that overlap. In previous work, we showed how observable relationships between two datasets help scientists recall their original derivation connection. While that work helped with identifying relationships between two datasets, it is infeasible for scientists to use it for finding relationships between all possible pairs in a large collection of datasets. In order to deal with larger numbers of datasets, we are extending our methodology with a relationship-prediction system, ReDiscover, a tool to identify pairs from a collection of datasets that are most likely related and the relationship between them. We report on the initial design of ReDiscover, which uses machine-learning methods such as Conditional Random Fields and Support Vector Machines to the relationship-discovery problem. Our preliminarily evaluation shows that ReDiscover predicted relationships with an average accuracy of 87%.
Abdussalam Alawini, David Maier 0001, Kristin Tufte, Bill Howe, Rashmi Nandikur
SSDBM1
2014 Helping scientists reconnect their datasets
abstract
It seems inevitable that the datasets associated with a research project proliferate over time: collaborators may extend datasets with new measurements and new attributes, new experimental runs result in new files with similar structures, and subsets of data are extracted for independent analysis. As these "residual" datasets begin to accrete over time, scientists can lose track of the derivation history that connects them, complicating data sharing, provenance tracking, and scientific reproducibility. In this paper, focusing on data in spreadsheets, we consider how observable relationships between two datasets can help scientists recall their original derivation connection. For instance, if dataset A is wholly contained in dataset B, B may be a more recent version of A and should be preferred when archiving or publishing.
Abdussalam Alawini, David Maier 0001, Kristin Tufte, Bill Howe
SSDBM1