VLDB 2026 Research / reviewers in the wild / expert
Hoda Khalil
dblp:205/8781
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2022
0000-0002-3459-616XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Systematic Analysis of Public Transit Data Availability in CanadaabstractRegional authorities will publish public transit route and timetable service offerings through the General Transit Feed Specification (GTFS) as a standard format. Systematically collected GTFS data can be used to study the structure, organization, and availability of public transit services nationally. In this work, we provide a systematic approach to collection of public transit data from various sources, preparation of a high-quality GTFS data inventory and analysis of the public transit offerings lending to numerous insights about transit availability at various geographic scales. Using Canada as a case study, we collected GTFS for the 213 candidate census subdivisions ([CSDs] representing cities, towns, municipalities, etc.) with populations greater than 20,000. These data were then cleaned and leveraged along with CSD-specific census statistics to comprehensively compare public transit offerings between provinces/territories and across CSD types. We determined that, despite a systematic collection process, the majority of CSDs lack official GTFS data, certain provinces are under- or over-represented in our analysis. We further proposed using the median of transit stop spatial density as a national baseline revealing that provinces such as Québec are severely lacking in public transit offerings. GTFS data analysis is insightful for understanding the current state and progress in urban public transportation, which is highly relevant to the United Nation’s Sustainability Development Goals. Our aggregated dataset and open-sourced codebase are publicly available at: github.com/chazingtheinfinite/canada-transit-study. Kevin Dick, Azizul Hasan, Jamil Dergham, A. J. Clarke, Hoda Khalil, Gabriel A. Wainer |
IEEE Big Data | 5 |
| 2022 | Canadian Jobs amid a Pandemic: Examining the Relationship between Professional Industry and Salary to Regional Key Performance IndicatorsabstractThe COVID-19 pandemic has contributed to un-precedented rates of unemployment and greater uncertainty in the job market. There is a growing need for data-driven tools and analyses to better inform the public on trends within the job market. In particular, obtaining a “snapshot” of available employment opportunities mid-pandemic promises insights to inform policy and support retraining programs. In this work, we combine data scraped from the Canadian Job Bank and Numbeo globally crowd-sourced repository to explore the relationship between job postings during a global pandemic and Key Performance Indicators (e.g. quality of life [QOL] index, cost of living) for major cities across Canada. This analysis aims to help Canadians make informed career decisions, collect a “snapshot” of the Canadian employment opportunities amid a pandemic, and inform job seekers in identifying the correct fit between the desired lifestyle of a city and their career. We collected a new high-quality dataset of job postings from jobbank.gc.ca obtained with the use of ethical web scraping and performed exploratory data analysis on this dataset to identify job opportunity trends. When optimizing for average salary of job openings with QOL, affordability, cost of living, and traffic indices, it was found that Edmonton, AB consistently scores higher than the mean, and is therefore an attractive place to move. Furthermore, we identified optimal provinces to relocate to with respect to individual skill levels. It was determined that Ajax, Marathon, and Chapleau, ON are each attractive cities for IT professionals, construction workers, and healthcare workers respectively when maximizing average salary. Finally, we publicly release our scraped dataset as a mid-pandemic snapshot of Canadian employment opportunities and present a public web application that provides an interactive visual interface that summarizes our findings for the general public and the broader research community. Rahul Anilkumar, Benjamin Melone, Michael Patsula, Christophe Tran, Christopher Wang, Kevin Dick, Hoda Khalil, Gabriel A. Wainer |
COMPSAC | 7 |
| 2022 | Elucidation of the Relationship Between a Song's Spotify Descriptive Metrics and its Popularity on Various PlatformsabstractThe music industry and personal music consumption have evolved dramatically with the advent of streaming plat-forms. In this evolving landscape, there is considerable interest in understanding what factors contribute to a song's popularity. Extrinsic (i.e. non-acoustic) features of a given song, such as the record label, and/or intrinsic (i.e. acoustic) features such as its energy may contribute to popularity on a given digital platform. In this work, we, for the first time, sought to systematically study how a song's Spotify acoustic descriptive features correlated with popularity metrics on various Internet platforms. Since each platform defines “popularity” according to platform-specific metrics, a large-scale correlation-based analysis was generated. The digital platforms considered in this article are Google Trends, WhoSampled, TikTok, Twitter, YouTube, and the Billboard Top-100. Platform-specific scrapers were created and all data was aggregated with the Spotify Echo Nest dataset of descriptive acoustic metrics. While the majority of correlations were unre-markable considering both Spearman and Pearson coefficients, a number of corroborating and contradictory findings resulted, with notable implications for acoustic features on various digital platforms. Notably, the YouTube view count was found to be positively correlated to the Spotify song popularity (p = 0.822), year (p = 0.600), and energy (p = 0.455) and moderately negatively correlated to accousticness (p = −0.542) and instrumentalness (p = −0.345). All reproducing code and aggregated data from this work are open-source for use by the broader research community. Laura Colley, Andrew Dybka, Adam Gauthier, Jacob Laboissonniere, Alexandre Mougeot, Nayeeb Mowla, Kevin Dick, Hoda Khalil, Gabriel A. Wainer |
COMPSAC | 8 |
| 2020 | Cell-DEVS for Social Phenomena ModelingabstractMotivated by the need for formal methods as well as supporting tools to model and simulate social systems, we propose cellular discrete-event system specification as a formalism for modeling social systems. We also propose the use of a toolkit that implements the formalism of cellular discrete-event system specifications to implement and visualize models of social systems. We present examples of social system models that are different in sizes, nature, and rules controlling the interactions within those systems. We show that cellular discrete-event system specification with its unique features can successfully deal with the shortcoming of other modeling techniques. In addition, we show that together with its supporting toolkit, cellular discrete-event system specification is suitable for modeling, simulating, implementing, and visualizing social systems. Hoda Khalil, Gabriel A. Wainer |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2017 | State-Based Tests Suites Automatic Generation Tool (STAGE-1)abstractState diagrams are widely used to model software artifacts, making state-based testing an interesting research topic. When conducting research on state-based testing for evaluating different testing criteria, often there is a need to devise numerous test suites in a systematic way according to selection criteria such as all-edges, all-transition-pairs, or the transition tree (W-method). Moreover, one also needs to satisfy each criterion in as many ways as possible to account for possible stochastic phenomena within each criterion. The main issue is then: how to automate the generation of as many, or even all, the different test suites for each criterion? This paper presents the first part of a framework, an automation tool chain that generates test trees from a state machine diagram, extracts test cases from the generated trees, and composes a test suite from each generated tree. This tool is the first to generate all possible distinctive trees using depth and breadth first graph traversal algorithms. The tool chain should be of interest to researchers in state-based testing as well as practitioners who are interested in alternative adequate test suites especially for comparing the effectiveness of the different test suites satisfying one criterion and the effectiveness of the other different criteria. Hoda Khalil, Yvan Labiche |
COMPSAC (1) | 1 |
| 2017 | Finding All Breadth First Full Spanning Trees in a Directed GraphabstractThis paper proposes an algorithm that is particularly concerned with generating all possible distinct spanning trees that are based on breadth-first-search directed graph traversal. The generated trees span all edges and vertices of the original directed graph. The algorithm starts by generating an initial tree, and then generates the rest of the trees using elementary transformations. It runs in O(E+T) time where E is the number of edges and T is the number of generated trees. In the worst-case scenario, this is equivalent to O (E+En/Nn) time complexity where N is the number of nodes in the original graph. The algorithm requires O(T) space. However, possible modifications to improve the algorithm space complexity are suggested. Furthermore, experiments are conducted to evaluate the algorithm performance and the results are listed. Hoda Khalil, Yvan Labiche |
COMPSAC (2) | 1 |
| 2017 | On FSM-Based Testing: An Empirical Study: Complete Round-Trip Versus Transition TreesabstractFinite state machines being intuitively understandable and suitable for modeling in many domains, they are adopted by many software designers. Therefore, testing systems that are modeled with state machines has received genuine attention. Among the studied testing strategies are complete round-trip paths and transition trees that cover round-trip paths in a piece wise manner. We present an empirical study that aims at comparing the effectiveness of the complete round-trip paths test suites to the transition trees test suites in one hand, and comparing the effectiveness of the different techniques used to generate transition trees (breadth first traversal, depth first traversal, and random traversal) on the other hand. We also compare the effectiveness of all the testing trees generated using each single traversal criterion. This is done through conducting an empirical evaluation using four case studies from different domains. Effectiveness is evaluated with mutants. Experimental results are presented and analyzed. Hoda Khalil, Yvan Labiche |
ISSRE | 1 |