VLDB 2026 Research / reviewers in the wild / expert
Philippe Charland
dblp:30/1570
· DBLP profile ↗
27ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0003-4051-9942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 5 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Security and privacy · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIN: Boosting binary code embedding by normalizing function inlinings
Mohammadhossein Amouei, Benjamin C. M. Fung, Philippe Charland |
J. Syst. Softw. | 3 |
| 2026 | PAC-X: Fuzzy Explainable AI for Multiclass Malware DetectionabstractResearchers often approach malware detection as a binary classification problem. However, evidence indicates that malware can belong to multiple families simultaneously, and malicious files frequently exhibit numerous benign features. Attackers exploit this by embedding malicious intent within benign features, making malware detection a problem better suited for fuzzy systems. Furthermore, providing explainability for such fuzzy classification remains a significant challenge, requiring specialized Explainable AI (XAI) frameworks. Existing XAI approaches offer insights into model decisions but are vulnerable to adversarial attacks that manipulate features to mislead models. To address these issues, we propose PAC-X, a novel XAI framework for malware detection. PAC-X integrates the Conditional Attention Neural Network (CAN-Net) to deliver comprehensive multi-fuzzy-class explainability and employs Contextual Fuzzy Clustering (CFC) to extract contextual insights from training data. This framework is resilient to adversarial manipulations, maintaining reliable and interpretable explanations even under adversarial conditions designed to mislead the model. Through extensive evaluations on diverse malware datasets, PAC-X demonstrates superior explainability and robustness compared to state-of-the-art XAI methods. It provides a critical advancement in cybersecurity by addressing the complexities of evasive malware detection and enabling a deeper interpretation of multi-class malware characteristics. Mohd Saqib, Benjamin C. M. Fung, Philippe Charland |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | ProvSpider: A Robust and Universal Toolkit for Binary Provenance Analysis Using Deep LearningabstractBinary provenance analysis recovers essential information, such as architecture, structure, and toolchain, from executables lacking reliable metadata. This is crucial for reverse engineering. However, provenance recovery from binaries is highly challenging, due to three key factors: (1) binaries span diverse CPU architectures; (2) Raw byte sequences are often extremely long without clear boundaries; and (3) Compilation alters control flow, register usage, and memory layout, obscuring the original code structure and complicating analysis. To address these challenges, we propose a novel and robust analysis toolset, namely ProvSpider, to identify segment boundaries, types of segments as well as target CPU architectures, bitness, and endianness based on code-only sections. ProvSpider is built based on a convolutional neural network (CNN) to learn local execution patterns. We embed byte sequences into eight-dimensional vectors to capture bytes’ global dependencies. The gating mechanism after convolutional layers filters out noise and keeps most representative features. At last, the sliding window divides lengthy byte sequences into fixedlength processable chunks. Our model achieves high accuracy in all five analysis tasks, significantly outperforming the state-of-the-art models. By providing a universal and robust approach, ProvSpider lays the foundation for advancing binary provenance analysis, facilitating future improvements in binary analysis and reverse engineering. Zhiwei Fu, Hanbo Yu, Steven H. H. Ding, Furkan Alaca, Philippe Charland |
NCA | 6 |
| 2025 | Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionabstractCybersecurity and software research have crossed paths with modern deep learning research for a few years. The power of large language models (LLMs) in particular has intrigued us to apply them to understanding binary code. In this paper, we investigate some of the many ways LLMs can be applied to binary code similarity detection, as it is a significantly more difficult task compared to source code similarity detection due to the sparsity of information and less meaningful syntax. It also has great practical implications, such as vulnerability and malware detection. We find that pretrained LLMs are mostly capable of detecting similar binary code, even with a zero-shot setting. Our main contributions and findings are to provide several supervised fine-tuning methods that, when combined, significantly surpass zero-shot LLMs and state-of-the-art binary code similarity detection methods. Specifically, we up-train the model through data augmentation, translation-style causal learning, LLM2Vec, and cumulative GTE loss. With a complete ablation study, we show that our training method can transform a generic language model into a powerful binary similarity expert, and is also robust and general enough for cross-optimization, cross-architecture, and cross-obfuscation detection. Litao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland |
NeurIPS | 5 |
| 2025 | MalGPT: A Generative Explainable Model for Malware Binaries
Mohd Saqib, Benjamin C. M. Fung, Steven H. H. Ding, Philippe Charland |
ECML/PKDD (4) | 4 |
| 2025 | NeuroYara: Learning to Rank for Yara Rules Generation Through Deep Language Modeling and Discriminative N-Gram EncodingabstractSignature-based malware detection methods are recognized for their simplicity, explainability, and efficiency. One of the most commonly used tools is Yara, which provides the syntax for crafting malware signatures. However, while developing high-quality Yara rules requires significant expertise in malware analysis, training such skilled analysts can be both resource-intensive and time-consuming. While a few works have been conducted to automate the generation of signatures, signatures generated by those works typically underperform the manually generated ones. In addition, these automated methods often depend on large static databases of hard-coded byte n-grams to minimize false positives. Instead of storing a large non-inclusive database to score byte n-grams, we propose a novel architecture utilizing two learning to rank neural networks to understand the underlying effectiveness and correlations among n-grams extracted for rule construction. This approach provides better flexibility and coverage of possible n-grams while reducing the required storage size from several GBs to only 10MBs. Combining these two models with a hierarchical density-based clustering method allows us to group multiple n-grams into logical conditions as Yara rules of higher quality. Experimental results show that our framework, NeuroYara, reduces the resources invested by analysts while generating rules with a low false-positive rate outperforming existing tools and manually-generated rules. Ziad Mansour, Weihan Ou, Steven H. H. Ding, Mohammad Zulkernine, Philippe Charland |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Obfuscated Clone Search in JavaScript based on Reinforcement Subsequence LearningabstractFinding similar code is important for software engineering, defense of intellectual property, and security, and one of the increasingly common ways adversaries use to defeat the detection of similar code is through obfuscations such as code transformation and scattering the code they wish to hide among long sequences. Moving code far enough apart poses a specific challenge for solutions with localized features (e.g., n-grams), or attention mechanisms as the code parts are distributed beyond the local context window. We introduce a neural network solution pattern called “Cybertron” that addresses this problem by utilizing reinforcement learning to train a code abstraction and summarization function; this converts arbitrarily long code into fixed-length real vectors in a way that is optimized for similarity search. The key to the design is the smart selection of important elements of the code and abstraction to preserve semantic function while minimizing syntactic feature information. We evaluated the approach on a three-challenge benchmark of obfuscated JavaScript, a scripting language that is commonly obfuscated and for which code-mixing is a rising challenge. The evaluation shows our approach identifies obfuscated code within even large scripts with an AUC of 78%, which outperforms current state-of-the-art sequence models by 7–35%. Leo Song, Steven H. H. Ding, Yuan Tian 0008, Li Tao Li, Weihan Ou, Philippe Charland, Andrew Walenstein |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | Mecha: A Neural-Symbolic Open-Set Homogeneous Decision Fusion Approach for Zero-Day Malware Similarity DetectionabstractWith increasing numbers of novel malware each year, tools are required for efficient and accurate variant matching under the same family, for the purpose of effective proactive threat detection, retro-hunting, and attack campaign tracking. All of the state-of-the-art Deep Learning (DL) approaches assume that the incoming samples originate from known families and incorrectly identify novel families. Additionally, most of the existing solutions that leverage the Siamese Neural Network architecture either rely on pair-wise comparisons or computationally expensive preprocessing steps that are not scalable to a real-world malware triage volume requirement. We propose a different route, Mecha, a Neural-Symbolic Machine Learning (ML) system for malware variant matching and zero-day family detection. Mecha is comprised of an embedding network trained in two different scenarios for byte string embedding and an open-set approximate nearest neighbour algorithm for variant matching and zero-day detection. Our embedding network uses triplet loss for embedding generation and reinforcement-based Expectation Maximization (EM) learning for full deployment optimization. We conduct multiple in-sample and out-of-sample experiments to demonstrate the model's generalizability toward novel variants and families. We also show that Mecha can detect samples outside the known set of malware samples with an accuracy greater than 0.990. Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Furkan Alaca, Philippe Charland, Andrew Walenstein |
IEEE Trans. Software Eng. | 5 |
| 2024 | Dynamic Neural Control Flow Execution: an Agent-Based Deep Equilibrium Approach for Binary Vulnerability DetectionabstractSoftware vulnerabilities are a challenge in cybersecurity. Manual security patches are often difficult and slow to be deployed, while new vulnerabilities are created. Binary code vulnerability detection is less studied and more complex compared to source code, and this has important practical implications. Deep learning has become an efficient and powerful tool in the security domain, where it provides end-to-end and accurate prediction. Modern deep learning approaches learn the program semantics through sequence and graph neural networks, using various intermediate representation of programs, such as abstract syntax trees (AST) or control flow graphs (CFG). Due to the complex nature of program execution, the output of an execution depends on the many program states and inputs. Also, a CFG generated from static analysis can be an overestimation of the true program flow. Moreover, the size of programs often does not allow a graph neural network with fixed layers to aggregate global information. To address these issues, we propose DeepEXE, an agent-based implicit neural network that mimics the execution path of a program. We use reinforcement learning to enhance the branching decision at every program state transition and create a dynamic environment to learn the dependency between a vulnerability and certain program states. An implicitly defined neural network enables nearly infinite state transitions until convergence, which captures the structural information at a higher level. The experiments are conducted on two semi-synthetic and two real-world datasets. We show that DeepEXE is an accurate and efficient method and outperforms the state-of-the-art vulnerability detection methods. Li Tao Li, Steven H. H. Ding, Andrew Walenstein, Philippe Charland, Benjamin C. M. Fung |
CIKM | 4 |
| 2024 | GAGE: Genetic Algorithm-Based Graph Explainer for Malware AnalysisabstractMalware analysts often prefer reverse engineering using Call Graphs, Control Flow Graphs (CFGs), and Data Flow Graphs (DFGs), which involves the utilization of black-box Deep Learning (DL) models. The proposed research introduces a structured pipeline for reverse engineering-based analysis, offering promising results compared to state-of-the-art methods and providing high-level interpretability for malicious code blocks in subgraphs. We propose the Canonical Executable Graph (CEG) as a new representation of Portable Executable (PE) files, uniquely incorporating syntactical and semantic information into its node embeddings. At the same time, edge features capture structural aspects of PE files. This is the first work to present a PE file representation encompassing syntactical, semantic, and structural characteristics, whereas previous efforts typically focused solely on syntactic or structural properties. Furthermore, recognizing the limitations of existing graph explanation methods within Explainable Artificial Intelligence (XAI) for malware analysis, primarily due to the specificity of malicious files, we introduce Genetic Algorithm-based Graph Explainer (GAGE). GAGE operates on the CEG, striving to identify a precise subgraph relevant to predicted malware families. Through experiments and comparisons, our proposed pipeline exhibits substantial improvements in model robustness scores and discriminative power compared to the previous benchmarks. Furthermore, we have successfully used GAGE in practical applications on real-world data, producing meaningful insights and interpretability. This research offers a robust solution to enhance cybersecurity by delivering a transparent and accurate understanding of malware behaviour. Moreover, the proposed algorithm is specialized in handling graph-based data, effectively dissecting complex content and isolating influential nodes. Mohd Saqib, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ICDE | 3 |
| 2024 | AsmDocGen: Generating Functional Natural Language Descriptions for Assembly Codeabstract"This study explores the field of software reverse engineering through the lens of code summarization, which involves generating informative and concise summaries of code functionality. A significant aspect of this research is the application of assembly code summarization in malware analysis, highlighting its critical role in understanding and mitigating potential security threats. Although there have been recent efforts to develop code summarization techniques for high-level programming languages, to the best of our knowledge, this study is the first attempt to generate comments for assembly code. For this purpose, we first built a carefully curated dataset of assembly function-comment pairs. We then focused on automatic assembly code summarization using transfer learning with pre-trained natural language processing (NLP) models, including BERT, DistilBERT, RoBERTa, and CodeBERT. The results of our experiments show a notable advantage of Code- BERT: despite its initial training on high-level programming languages alone, it excels in learning assembly language, outperforming other pre-trained NLP models."@eng Jesia Quader Yuki, Mohammadhossein Amouei, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ICSOFT | 4 |
| 2024 | VulEXplaineR: XAI for Vulnerability Detection on Assembly Code
Samaneh Mahdavifar, Mohd Saqib, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ECML/PKDD (9) | 4 |
| 2023 | GenTAL: Generative Denoising Skip-gram Transformer for Unsupervised Binary Code Similarity DetectionabstractBinary code similarity detection serves a critical role in cybersecurity. It alleviates the huge manual effort required in the reverse engineering process for malware analysis and vulnerability detection, where the original source code is often not available. Most of the existing solutions focus on a manual feature engineering process and customized code matching algorithms that are inefficient and inaccurate. Recent deep learning-based solutions embed the semantics of binary code into a latent space through supervised contrastive learning. However, one cannot cover all the possible forms in the training set to learn the variance of the same semantics. In this paper, we propose an unsupervised model aiming to learn the intrinsic representation of assembly code semantics. Specifically, we propose a Transformer-based auto-encoder like language model for the low-level assembly code grammar to capture the abstract semantic representation. By coupling a Transformer encoder and a skip-gram style loss design, it can learn a compact representation that is robust against different compilation options. We conduct experiments on four different block-level code similarity tasks. It shows that our method is more robust compared to the state-of-the-art solutions. Li Tao Li, Steven H. H. Ding, Philippe Charland |
IJCNN | 3 |
| 2023 | VulANalyzeR: Explainable Binary Vulnerability Detection with Multi-task Learning and Attentional Graph ConvolutionabstractSoftware vulnerabilities have been posing tremendous reliability threats to the general public as well as critical infrastructures, and there have been many studies aiming to detect and mitigate software defects at the binary level. Most of the standard practices leverage both static and dynamic analysis, which have several drawbacks like heavy manual workload and high complexity. Existing deep learning-based solutions not only suffer to capture the complex relationships among different variables from raw binary code but also lack the explainability required for humans to verify, evaluate, and patch the detected bugs. We propose VulANalyzeR, a deep learning-based model, for automated binary vulnerability detection, Common Weakness Enumeration-type classification, and root cause analysis to enhance safety and security. VulANalyzeR features sequential and topological learning through recurrent units and graph convolution to simulate how a program is executed. The attention mechanism is integrated throughout the model, which shows how different instructions and the corresponding states contribute to the final classification. It also classifies the specific vulnerability type through multi-task learning as this not only provides further explanation but also allows faster patching for zero-day vulnerabilities. We show that VulANalyzeR achieves better performance for vulnerability detection over the state-of-the-art baselines. Additionally, a Common Vulnerability Exposure dataset is used to evaluate real complex vulnerabilities. We conduct case studies to show that VulANalyzeR is able to accurately identify the instructions and basic blocks that cause the vulnerability even without given any prior knowledge related to the locations during the training phase. Litao Li, Steven H. H. Ding, Yuan Tian 0008, Benjamin C. M. Fung, Philippe Charland, Weihan Ou, Leo Song, Congwei Chen |
ACM Trans. Priv. Secur. | 5 |
| 2022 | Adversarial Variational Modality Reconstruction and Regularization for Zero-Day Malware Variants Similarity DetectionabstractMatching malware variants in the same malware family has always been a significant challenge for Cyber Threat Intelligence (CTI). For zero-day malware that does not belong to an existing family, a timely matching of its variants is essential for effective threat tracing and prompt response to the cyber incident. However, malware variants are of diverse forms that make them difficult to match. Additionally, the information extracted from a given malware sample is inaccurate, especially on zero-day malware. Existing malware solutions only focus on detecting known malware or find if two samples are similar without creating any reusable representation of the samples. In this paper, we propose the first practical and efficient solution for zero-day malware variant matching with reconstruction. By combining multi-modality learning and a Siamese-based structure, our model can navigate across different modalities and match zero-day variants. To address the missing or noisy modality issue, we propose a Conditional Variable Autoencoder with a Generative Adversarial Network for heightened resolution. We trained the model on 100,000 malware triplet pairs. Our experiments on real-world noisy samples show that the model out-performs the state-of-the-art and can accurately match not only zero-day malware, but also out-of-sample benign binaries of the same category. Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Philippe Charland, Andrew Walenstein, Litao Li |
ICDM | 4 |
| 2022 | DyAdvDefender: An instance-based online machine learning model for perturbation-trial-based black-box adversarial defense
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland |
Inf. Sci. | 3 |
| 2021 | A Novel and Dedicated Machine Learning Model for Malware Classification
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding |
ICSOFT | 3 |
| 2021 | I-MAD: Interpretable malware detector using Galaxy Transformer
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding |
Comput. Secur. | 3 |
| 2019 | Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationabstractReverse engineering is a manually intensive but necessary technique for understanding the inner workings of new malware, finding vulnerabilities in existing systems, and detecting patent infringements in released software. An assembly clone search engine facilitates the work of reverse engineers by identifying those duplicated or known parts. However, it is challenging to design a robust clone search engine, since there exist various compiler optimization options and code obfuscation techniques that make logically similar assembly functions appear to be very different. A practical clone search engine relies on a robust vector representation of assembly code. However, the existing clone search approaches, which rely on a manual feature engineering process to form a feature vector for an assembly function, fail to consider the relationships between features and identify those unique patterns that can statistically distinguish assembly functions. To address this problem, we propose to jointly learn the lexical semantic relationships and the vector representation of assembly functions based on assembly code. We have developed an assembly code representation learning model \emph{Asm2Vec}. It only needs assembly code as input and does not require any prior knowledge such as the correct mapping between assembly functions. It can find and incorporate rich semantic relationships among tokens appearing in assembly code. We conduct extensive experiments and benchmark the learning model with state-of-the-art static and dynamic clone search approaches. We show that the learned representation is more robust and significantly outperforms existing methods against changes introduced by obfuscation and optimizations. Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland |
IEEE Symposium on Security and Privacy | 3 |
| 2016 | Kam1n0: MapReduce-based Assembly Clone Search for Reverse EngineeringabstractAssembly code analysis is one of the critical processes for detecting and proving software plagiarism and software patent infringements when the source code is unavailable. It is also a common practice to discover exploits and vulnerabilities in existing software. However, it is a manually intensive and time-consuming process even for experienced reverse engineers. An effective and efficient assembly code clone search engine can greatly reduce the effort of this process, since it can identify the cloned parts that have been previously analyzed. The assembly code clone search problem belongs to the field of software engineering. However, it strongly depends on practical nearest neighbor search techniques in data mining and databases. By closely collaborating with reverse engineers and Defence Research and Development Canada (DRDC), we study the concerns and challenges that make existing assembly code clone approaches not practically applicable from the perspective of data mining. We propose a new variant of LSH scheme and incorporate it with graph matching to address these challenges. We implement an integrated assembly clone search engine called Kam1n0. It is the first clone search engine that can efficiently identify the given query assembly function's subgraph clones from a large assembly code repository. Kam1n0 is built upon the Apache Spark computation framework and Cassandra-like key-value distributed storage. A deployed demo system is publicly available. Extensive experimental results suggest that Kam1n0 is accurate, efficient, and scalable for handling large volume of assembly code. Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland |
KDD | 3 |
| 2012 | Semantic Web - The Missing Link in Global Source Code Analysis?abstractThere has been an ongoing trend towards open and shared source code that is published on the Internet in large software repositories to support collaborative development processes. While traditional source code analysis techniques perform well in single project contexts, new types of global source code analysis techniques are slowly introduced to address the analysis of global distributed and often incomplete source code. In this article, we discuss how the Semantic Web, an enabling technology for these emerging source code analysis domains, can support a standardized, formal, and semantic rich representation to model these corpora. We also illustrate how inference services can be used to provide support for emerging source code analysis approaches on this data, such as search, call graph construction, and clone detection. Iman Keivanloo, Juergen Rilling, Philippe Charland |
COMPSAC | 3 |
| 2011 | Reasoning about Global Clones: Scalable Semantic Clone DetectionabstractThe Semantic Web is slowly transforming the Web as we know it into a machine understandable pool of information that can be consumed and reasoned about by various clients. Source code is no exception to this trend and various communities have proposed standards to share code as linked data. With the availability of large amounts of open source code published in publicly accessible repositories, the introduction of massive horizontal scaling frameworks, and cloud computing infrastructures, a new era of software mining across information silos is reshaping the software engineering landscape. Given these technological advances, analyzing code at a global scale, across systems, projects and organizational boundaries, becomes feasible. In this paper, we introduce a clone detection algorithm and its implementation that can scale to such large global datasets, by modeling clones using description logic and applying a horizontal scaling Semantic Web reasoner. We demonstrate how our simple feature vector that only uses control statements, data types and method calls, can yield results similar to other popular clone detection tools. Our approach does not only allow us to reliably identify clones in a global context. By using a semantic reasoner, it also allows us to expand clone detection to a new class of semantic clones. We have compared our algorithm to some of the leading clone detection tools (DECKARD, CCFinder, JCD, and Simian) in order to validate our approach and show the differences in detected clones and performance. Philipp Schügerl, Juergen Rilling, Philippe Charland |
COMPSAC | 3 |
| 2011 | Quality Validation through Pattern Detection - A Semantic Web PerspectiveabstractGiven the ongoing trend towards the globalization of software systems, open networks, and distributed platforms, validating non-functional requirements and quality becomes essential. Our research addresses this challenge from two different perspectives: (1) the integration of knowledge and tool resources through Semantic Web technologies as part of our SE-PAD environment, in order to reduce or eliminate existing traditional information and analysis silos. (2) The ability to reason upon linked resources to infer both explicit and implicit patterns to support the validation of quality aspects. We have applied our SE-PAD environment for the detection of security and design patterns, as well as the violations of secure programming guidelines. David Walsh 0003, Philipp Schügerl, Juergen Rilling, Philippe Charland |
COMPSAC | 4 |
| 2011 | SeClone - A Hybrid Approach to Internet-Scale Real-Time Code Clone SearchabstractReal-time code clone search is an emerging family of clone detection research that aims at finding clone pairs matching an input code fragment in fractions of a second. For these techniques to meet actual real world requirements, they have to be scalable and provide a short response time. Our research presents a hybrid clone search approach using source code pattern indexing, information retrieval clustering, and Semantic Web reasoning to respectively achieve short response time, handle false positives, and support automated grouping/querying. Iman Keivanloo, Juergen Rilling, Philippe Charland |
ICPC | 3 |
| 2009 | A Contextual Guidance Approach to Software SecurityabstractWith the ongoing trend towards the globalization of software systems and their development, components in these systems might not only work together, but may end up evolving independently from each other. Modern IDEs have started to incorporate support for these highly distributed environments, by adding new collaborative features. As a result, assessing and controlling system quality (e.g. security concerns) during system evolution in these highly distributed systems become a major challenge. In this research, we introduce a unified ontological representation that integrates best security practices in a context-aware tool implementation. As part of our approach, we integrate information from traditional static source code analysis with semantic rich structural information in a unified ontological representation. We illustrate through several use cases how our approach can support the evolvability of software systems from a security quality perspective. Philipp Schügerl, David Walsh 0003, Juergen Rilling, Philippe Charland |
COMPSAC (2) | 4 |
| 2009 | Beyond generated software documentation - A web 2.0 perspectiveabstractOver the last decades, software engineering processes have constantly evolved to reflect cultural, social, technological, and organizational changes, which are often a direct result of the Internet. The introduction of the Web 2.0 resulted in further changes creating an interactive, community driven platform. However, these ongoing changes have yet to be reflected in the way we document software systems. Documentation generators, like Doxygen and its derivatives (Javadoc, Natural Docs, etc.) have become the de-facto industry standards for creating external technical software documentation from source code. However, the inter-woven representation of source code and documentation within a source code editor limits the ability of these approaches to provide rich media, internationalization, and interactive content. In this paper, we combine the functionality of a Web browser with a source code editor to provide source code documentation with rich media content. The paper presents our fully functional implementation of the editor within the Eclipse framework. Philipp Schügerl, Juergen Rilling, Philippe Charland |
ICSM | 3 |
| 2008 | A survey and evaluation of tool features for understanding reverse-engineered sequence diagramsabstractAbstract Sequence diagrams can be valuable aids to software understanding. However, they can be extremely large and hard to understand in spite of using modern tool support. Consequently, providing the right set of tool features is important if the tools are to help rather than hinder the user. This paper surveys research and commercial sequence diagram tools to determine the features they provide to support program understanding. Although there has been significant effort in developing these tools, many of them have not been evaluated using human subjects. To begin to address this gap, a preliminary study was performed with a specially designed sequence diagram tool that implements the features found during the survey. On the basis of an analysis of the study results, we discuss the features that were found to be useful and relate these to the tasks performed. It concludes by proposing how future tools can be improved to better support the exploration of large sequence diagrams. Copyright © 2008 Crown in the right of Canada. Published by John Wiley & Sons, Ltd. Chris Bennett, Del Myers, Margaret-Anne D. Storey, Daniel M. Germán, D. Ouellet, Martin Salois, Philippe Charland |
J. Softw. Maintenance Res. Pract. | 7 |