VLDB 2026 Research / reviewers in the wild / expert
Hazeline U. Asuncion
dblp:78/3305
· DBLP profile ↗
18ranked-venue papers
7as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Keyword Extraction From Specification Documents for Planning Security MechanismsabstractSoftware development companies heavily invest both time and money to provide post-production support to fix security vulnerabilities in their products. Current techniques identify vulnerabilities from source code using static and dynamic analyses. However, this does not help integrate security mechanisms early in the architectural design phase. We develop VDocScan, a technique for predicting vulnerabilities based on specification documents, even before the development stage. We evaluate VDocScan using an extensive dataset of CVE vulnerability reports mapped to over 3600 product documentations. An evaluation of 8 CWE vulnerability pillars shows that even interpretable whitebox classifiers predict vulnerabilities with up to 61.1% precision and 78% recall. Further, using strategies to improve the relevance of extracted keywords, addressing class imbalance, segregating products into categories such as Operating Systems, Web applications, and Hardware, and using blackbox ensemble models such as the random forest classifier improves the performance to 96% precision and 91.1% recall. The high precision and recall shows that VDocScan can anticipate vulnerabilities detected in a product's lifetime ahead of time during the Design phase to incorporate necessary security mechanisms. The performance is consistently high for vulnerabilities with the mode of introduction: architecture and design. Jeffy Jahfar Poozhithara, Hazeline U. Asuncion, Brent Lagesse |
ICSE | 2 |
| 2022 | Towards Lightweight Detection of Design Patterns in Source CodeabstractIdentifying which design patterns exist in source code helps maintenance engineers better understand source code and determine if new requirements can be satisfied.Automated techniques for finding design patterns generally require much time to label training datasets or to specify rules/queries for each pattern, and is difficult to extend support to secure design patterns (SDPs) and combination patterns.To address these challenges, we introduce PatternScout, a technique for automatic generation of SPARQL queries from UML Class diagrams and Sequence diagrams.These queries are used to detect patterns in the source code.Our results indicate that PatternScout can detect object-oriented design patterns (OODP) with accuracy that is comparable or better than existing techniques.It can also generate queries for SDPs that can be represented as UML Class diagrams. Jeffy Jahfar Poozhithara, Hazeline U. Asuncion, Brent Lagesse |
SEKE | 2 |
| 2017 | Data Provenance for Multi-Agent ModelsabstractMulti-agent simulations are useful for exploring collective patterns of individual behavior in social, biological, economic, network, and physical systems. However, there is no provenance support for multi-agent models (MAMs) in a distributed setting. To this end, we introduce ProvMASS, a novel approach to capture provenance of MAMs in a distributed memory by combining inter-process identification, lightweight coordination of in-memory provenance storage, and adaptive provenance capture. ProvMASS is built on top of the Multi-Agent Spatial Simulation (MASS) library, a framework that combines multi-agent systems with large-scale fine-grained agent-based models, or MAMs. Unlike other environments supporting MAMs, MASS parallelizes simulations with distributed memory, where agents and spatial data are shared application resources. We evaluate our approach with provenance queries to support three use cases and performance measures. Initial results indicate that our approach can support various provenance queries for MAMs at reasonable performance overhead. Delmar B. Davis, Jonathan Featherston, Munehiro Fukuda, Hazeline U. Asuncion |
eScience | 4 |
| 2017 | A Multi-agent Parallel Approach to Analyzing Large Climate Data SetsabstractDespite various cloud technologies that have parallelized and scaled up big data analysis, they target data mostly in texts which are easy to partition and thus easy to map over a cluster system. Therefore, their parallelization do not necessarily cover scientific structured data such as NetCDF or need additional, user-provided tools to convert the original data into specific formats. To facilitate user-intuitive parallelization of such scientific data analysis, this paper presents an agent-based approach that instantiates distributed arrays over a cluster system, maintains structured scientific data in these arrays, deploys many mobile agents over the arrays to perform computational actions on data, and collects necessary results. To demonstrate the practicability of our agent-based approach, we focused on climate change research and implemented a web-interfaced climate analysis, using the MASS (multi-agent spatial simulation) library. In this paper, we show practical advantages of, performance improvements by, and challenges for our agent-based approach in structured data analysis. Jason Woodring, Matthew Sell, Munehiro Fukuda, Hazeline U. Asuncion, Eric P. Salathe |
ICDCS | 4 |
| 2017 | Mapping Features to Source Code through Product Line Architecture: Traceability and ConformanceabstractExisting software product line approaches often develop and evolve product line features, architecture, and source code independently, which makes it difficult to manage the relationship and conformance between these artifacts. This paper presents a novel approach using the architecture as a pivot to address this problem. It consists of a modeling mechanism that integrates features specification into an architectural model, and an architecture-implementation mapping mechanism that combines code generation with annotation processing. The approach can trace a product line feature to the architecture and source code, and automatically update the architecture and source code to maintain their conformance when feature changes occur. We implemented an Eclipse-based toolset to support the approach, and conducted a case study with the Apache Solr open-source system. The result shows that our approach is both applicable and capable to support the development and evolution of real-world variations of a software system. Cuong Cu, Hazeline U. Asuncion |
ICSA | 3 |
| 2017 | BrainGrid+Workbench: High-performance/high-quality neural simulationabstractAvailability of affordable hardware that in effect enables desktop supercomputing has enabled more ambitious neural simulations driven by more complex software. However, this opportunity comes with costs, in terms of long learning curves to take advantage of the performance possibilities of idiosyncratic, architecturally heterogenous hardware and decreasing ability to be confident in the quality of simulation results. This paper describes a new neural simulation and software/data provenance framework that reduces the difficulty of taking full advantage of GPU computing and increases investigator confidence that simulations results are valid. Michael Stiber, Fumitaka Kawasaki, Delmar B. Davis, Hazeline U. Asuncion, Jewel Yun-Hsuan Lee, Destiny Boyer |
IJCNN | 4 |
| 2016 | Improving data provenance reconstruction via a multi-level funneling approachabstractThe ease with which data can be created, copied, modified, and deleted over the Internet has made it increasingly difficult to determine the source of web data. Data provenance, which provides information about the origin and lineage of a dataset, assists in determining its genuineness and trustworthiness. Several data provenance techniques record provenance when the data is created or modified. However, many existing datasets have no recorded provenance. Provenance Reconstruction techniques attempt to generate an approximate provenance in these datasets. Current reconstruction techniques require timing metadata to reconstruct provenance. In thats paper, we improve our multi-funneling technique, which combines existing techniques, including topic modeling, longest common subsequence, and genetic algorithm to achieve higher accuracy in reconstructing provenance without requiring timing metadata. In addition, we introduce novel funnels that are customized to the provided datasets, which further boosts precision and recall rates. We evaluated our approach with various experiments and compare the results of our approach with existing techniques. Finally, we present lessons learned, including the applicability of our approach to other datasets. Subha Vasudevan, William Pfeffer, Delmar B. Davis, Hazeline U. Asuncion |
eScience | 4 |
| 2015 | WIP: Provenance Support for Interdisciplinary Research on the North Creek WetlandsabstractResearchers working in the North Creek Wetlands are faced with the task of gathering and managing large amounts of data. This interdisciplinary group of researchers also require data provenance to ensure the integrity of their collected data. Currently, they record notes in the wetlands with pen and paper, transcribe these notes, and combine them with the other data they collected, such as spatial and image data. This process is error-prone and can be time consuming. Current provenance techniques also do not focus on supporting provenance from data collection to data processing and do not provide a flexible means of capturing provenance across different software tools and platforms. Our technique and our tool support, ProvEco System, leverages various technologies such as mobile devices, an off-the-shelf geographic information system software, an enterprise-level search facility, and open data standards to provide an interoperable provenance system. Our preliminary results indicate that our data collection application is easy to use and can assist with automatically reconstructing provenance for image files. We also offer generalizable lessons learned for developing a provenance system in other contexts. Jonathan Mason, Morteza Chini, Karen Potts, Nathan Duncan, Delmar B. Davis, Hazeline U. Asuncion |
e-Science | 7 |
| 2014 | Tracing Domain Data Concepts in Layered Applications
Mohammed Daubal, Nathan Duncan, Delmar B. Davis, Hazeline U. Asuncion |
SEKE | 4 |
| 2013 | Using Change Entries to Collect Software Project Information
Hazeline U. Asuncion, Macneil Shonle, Robert Porter, Karen Potts, Nathan Duncan, William Joseph Matthies Jr. |
SEKE | 1 |
| 2013 | Automated data provenance capture in spreadsheets, with case studies
Hazeline U. Asuncion |
Future Gener. Comput. Syst. | 1 |
| 2012 | Work in progress: Sustainable projects for software engineering courses: Collaborating with technology coursesabstractTeaching Software Engineering (SE) based on “real-world” projects engages students with practical application of software engineering concepts-students develop a deeper interest in the project deliverables while they acquire the skills of critically analyzing the problem and determining the best course of action. It can be challenging to find and maintain a reliable stream of suitable software projects that match learning outcomes, technical scope, and academic calendar of an SE class. On the other hand, a typical university campus has many nonComputer Science (CS) technology classes that require their students to study, understand, and evaluate existing software applications in specific areas. With purposeful coordination, these non-CS technology classes can serve as effective source of real projects for SE classes. This paper describes our experience of collaborating with the Education Program in their Technology in Education course. While we encountered some challenges, our experience has demonstrated that it is indeed mutually beneficial and rewarding for students in both courses. We offer recommendations on choosing non-CS technology classes and logistical guidelines to ensure the success of such collaborations. Hazeline U. Asuncion, Robin Lynn Angotti, Kelvin Sung |
FIE | 1 |
| 2012 | A Holistic Approach to Software Traceability
Hazeline U. Asuncion, Richard N. Taylor |
SEKE | 1 |
| 2011 | In Situ Data Provenance Capture in SpreadsheetsabstractThe capture of data provenance is a fundamentally important task in eScience. While provenance can be captured using techniques such as scientific workflows, typically these techniques do not trace internal data manipulations that occur within off-the-shelf analysis tools. Yet it is still essential to capture data provenance within such environments. This paper discusses an in situ provenance approach for spreadsheet data in MS Excel, a commonly used analysis environment among scientists. We describe the design and implementation of an Excel tool that captures provenance unobtrusively in the background, allows for user annotations, provides undo/redo functionality at various levels of task granularity, and presents the captured provenance in an accessible format to support a range of provenance queries for analysis. We also present several motivating use case scenarios and a user evaluation which suggests that our approach is both efficient and useful to scientists. Hazeline U. Asuncion |
eScience | 1 |
| 2011 | Presenting Software License Conflicts through Argumentation
Thomas A. Alspaugh, Hazeline U. Asuncion, Walt Scacchi |
SEKE | 2 |
| 2010 | Software traceability with topic modelingabstractSoftware traceability is a fundamentally important task in software engineering. The need for automated traceability increases as projects become more complex and as the number of artifacts increases. We propose an automated technique that combines traceability with a machine learning technique known as topic modeling. Our approach automatically records traceability links during the software development process and learns a probabilistic topic model over artifacts. The learned model allows for the semantic categorization of artifacts and the topical visualization of the software system. To test our approach, we have implemented several tools: an artifact search tool combining keyword-based search and topic modeling, a recording tool that performs prospective traceability, and a visualization tool that allows one to navigate the software architecture and view semantic topics associated with relevant artifacts and architectural components. We apply our approach to several data sets and discuss how topic modeling enhances software traceability, and vice versa. Categories and Subject Descriptors Hazeline U. Asuncion, Arthur U. Asuncion, Richard N. Taylor |
ICSE (1) | 1 |
| 2009 | Intellectual Property Rights Requirements for Heterogeneously-Licensed SystemsabstractHeterogeneously-licensed systems pose new challenges to analysts and system architects. Appropriate intellectual property rights must be available for the installed system, but without unnecessarily restricting other requirements, the system architecture, and the choice of components both initially and as it evolves. Such systems are increasingly common and important in e-business, game development, and other domains. Our semantic parameterization analysis of open-source licenses confirms that while most licenses present few roadblocks, reciprocal licenses such as the GNU General Public License produce knotty constraints that cannot be effectively managed without analysis of the system's license architecture. Our automated tool supports intellectual property requirements management and license architecture evolution. We validate our approach on an existing heterogeneously-licensed system. Thomas A. Alspaugh, Hazeline U. Asuncion, Walt Scacchi |
RE | 2 |
| 2007 | An end-to-end industrial software traceability toolabstractTraceability is an important aspect of software development that is often required by various professional standards and government agencies. Yet current industrial approaches do not typically address end-to-end traceability. Moreover, many industry projects become entangled in process overhead and fail to derive much benefit from current traceability solutions. This paper presents a successful end-to-end software traceability tool developed at Wonderware, a software development company and a business unit of Invensys Systems, Inc. This process-oriented approach achieves comprehensive traceability and supports the entire software development life cycle by focusing on both requirements traceability and process traceability. We offer new perspectives in analyzing the problem as well as general traceability guidelines. These guidelines have emerged from the experience of implementing and deploying the traceability tool within actual company constraints. We discuss encouraging results and point to the advantages gained in using our approach. Hazeline U. Asuncion, Frédéric François, Richard N. Taylor |
ESEC/SIGSOFT FSE | 1 |