VLDB 2026 Research / reviewers in the wild / expert
Carol Wong
dblp:67/9141
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Large Language Models for Analyzing and Improving Method Names in Scientific CodeabstractBackground: Research scientists increasingly rely on implementing software to support their research. While previous research has examined the impact of identifier names on program comprehension in traditional programming environments, limited work has explored this area in scientific software, especially regarding the quality of method names in the code. Aims: This study aims to evaluate the effectiveness of four popular Large Language Models (LLMs) in analyzing grammatical patterns and suggesting improvements for method names in scientific code. Method: We conducted an analysis of 496 method names extracted from Python-based Jupyter Notebooks, using LLMs to assess their ability to provide appraisals and recommendations. Results: Our findings show that the LLMs are somewhat effective in analyzing these method names and generally follow good naming practices, like starting method names with verbs. However, their inconsistent handling of domain-specific terminology and only moderate agreement with human annotations indicate that automated suggestions require human evaluation. Conclusions: This work provides foundational insights for improving the quality of scientific code through AI automation. Gunnar Larsen, Carol Wong, Anthony Peruma |
ESEM | 2 |
| 2025 | Identifier Name Similarities: An Exploratory StudyabstractBackground: Identifier names, which comprise a significant portion of the codebase, are the cornerstone of effective program comprehension. However, research has shown that poorly chosen names can significantly increase cognitive load and hinder collaboration. Even names that appear readable in isolation may lead to misunderstandings in contexts when they closely resemble other names in either structure or functionality. Aims: This exploratory study aims to investigate the occurrence of similar identifier names in projects and develop a taxonomy categorizing name similarities. Method: Five open-source Java projects were analyzed using automated extraction and manual review. Results: Our findings reveal a taxonomy comprising seven categories of identifier name similarities. Conclusions: We envision our initial taxonomy providing researchers with a platform to analyze and evaluate the impact of identifier name similarity on code comprehension, maintainability, and collaboration among developers, while also allowing for further refinement and expansion of the taxonomy. Carol Wong, Mai Abe, Silvia De Benedictis, Marissa Halim, Anthony Peruma |
ESEM | 1 |
| 2025 | Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research ScientistsabstractScientific software, defined as computer programs, scripts, or code used in scientific research, data analysis, modeling, or simulation, has become central to modern research. However, there is limited research on the readability and understandability of scientific code, both of which are vital for effective collaboration and reproducibility in scientific research. This study surveys 57 research scientists from various disciplines to explore their programming backgrounds, practices, and the challenges they face regarding code readability. Our findings reveal that most participants learn programming through self-study or on-the-job training, with$\mathbf{5 7. 9 \%}$lacking formal instruction in writing readable code. Scientists mainly use Python and$R$, relying on comments and documentation for readability. While most consider code readability essential for scientific reproducibility, they often face issues with inadequate documentation and poor naming conventions, with challenges including cryptic names and inconsistent conventions. Our findings also show low adoption of code quality tools and a trend towards utilizing large language models to improve code quality. These findings offer practical insights into enhancing coding practices and supporting sustainable development in scientific software. Alyssia Chen, Carol Wong, Bonita Sharif, Anthony Peruma |
ICPC | 2 |
| 2025 | Method Names in Jupyter Notebooks: An Exploratory StudyabstractMethod names play an important role in communicating the purpose and behavior of their functionality. Research has shown that high-quality names significantly improve code comprehension and the overall maintainability of software. However, these studies primarily focus on naming practices in traditional software development. There is limited research on naming patterns in Jupyter Notebooks, a popular environment for scientific computing and data analysis. In this exploratory study, we analyze the naming practices found in 691 methods across 384 Jupyter Notebooks, focusing on three key aspects: naming style conventions, grammatical composition, and the use of abbreviations and acronyms. Our findings reveal distinct characteristics of notebook method names, including a preference for conciseness and deviations from traditional naming patterns. We identified 68 unique grammatical patterns, with only 55.57 % of methods beginning with a verb. Further analysis revealed that half of the methods with return statements do not start with a verb. We also found that 30.39 % of method names contain abbreviations or acronyms, representing mathematical or statistical terms and image processing concepts, among others. We envision our findings contributing to developing specialized tools and techniques for evaluating and recommending high-quality names in scientific code and creating educational resources tailored to the notebook development community. Carol Wong, Gunnar Larsen, Rocky Huang, Bonita Sharif, Anthony Peruma |
ICPC | 1 |
| 2011 | Ensemble learning algorithms for classification of mtDNA into haplogroupsabstractClassification of mitochondrial DNA (mtDNA) into their respective haplogroups allows the addressing of various anthropologic and forensic issues. Unique to mtDNA is its abundance and non-recombining uni-parental mode of inheritance; consequently, mutations are the only changes observed in the genetic material. These individual mutations are classified into their cladistic haplogroups allowing the tracing of different genetic branch points in human (and other organisms) evolution. Due to the large number of samples, it becomes necessary to automate the classification process. Using 5-fold cross-validation, we investigated two classification techniques on the consented database of 21 141 samples published by the Genographic project. The support vector machines (SVM) algorithm achieved a macro-accuracy of 88.06% and micro-accuracy of 96.59%, while the random forest (RF) algorithm achieved a macro-accuracy of 87.35% and micro-accuracy of 96.19%. In addition to being faster and more memory-economic in making predictions, SVM and RF are better than or comparable to the nearest-neighbor method employed by the Genographic project in terms of prediction accuracy. Carol Wong, Yuran Li, Chih Lee, Chun-Hsi Huang |
Briefings Bioinform. | 1 |
| 1995 | A mobile robot that recognizes peopleabstractIn order for mobile robots to interact effectively with people they will have to recognize faces. We describe a robot system that finds people, approaches them and then recognizes them. The system uses a variety of techniques: color vision is used to find people; vision and sonar sensors are used to approach them; a template-based pattern recognition algorithm is used to isolate the face; and a neural network is used to recognize the face. All of these processes are controlled using an intelligent robot architecture that sequences and monitors the robot's actions. We present the results of many experimental runs using an actual mobile robot finding and recognizing up to six different people. Carol Wong, David Kortenkamp, Mark Speich |
ICTAI | 1 |