EDBT 2026 Demo / reviewers in the wild / expert
Georgia Frantzeskou
dblp:34/3260
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 50% Program analysis · 50% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › mining software repositories › source-code mining
code authorship attribution |
0.1 | 1 | 2006 | Effective identification of source code authors using byte-level information · ICSE 2006 |
Program analysis
source code analysis |
0.1 | 1 | 2006 | Effective identification of source code authors using byte-level information · ICSE 2006 |
Digital forensics and information hiding
authorship attribution |
0.0 | 1 | 2006 | Effective identification of source code authors using byte-level information · ICSE 2006 |
Methods — techniques the papers use, named apart from their topics
n-gram profiling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Author Identification in Imbalanced Sets of Source Code SamplesabstractSimilarly to natural language texts, source code documents can be distinguished by their style. Source code author identification can be viewed as a text classification task given that samples of known authorship by a set of candidate authors are available. Although very promising results have been reported for this task, the evaluation of existing approaches avoids focusing on the class imbalance problem and its effect on the performance. In this paper, we present a systematic experimental study of author identification in skewed training sets where the training samples are unequally distributed over the candidate authors. Two representative author identification methods are examined, one follows the profile-based paradigm (where a single representation is produced for all the available training samples per author) and the other follows the instance-based paradigm (where each training sample has its own individual representation). We examine the effect of the source code representation on the performance of these methods and show that the profile-based method is better able to handle cases of highly skewed training sets while the instance-based method is a better choice in balanced or slightly-skewed training sets. Evangelos Chatzicharalampous, Georgia Frantzeskou, Efstathios Stamatatos |
ICTAI | 2 |
| 2012 | A methodology to assess the impact of design patterns on software quality
Apostolos Ampatzoglou, Georgia Frantzeskou, Ioannis Stamelos |
Inf. Softw. Technol. | 2 |
| 2008 | Examining the significance of high-level programming features in source code author classification
Georgia Frantzeskou, Stephen G. MacDonell, Efstathios Stamatatos, Stefanos Gritzalis |
J. Syst. Softw. | 1 |
| 2006 | Effective identification of source code authors using byte-level informationabstractSource code author identification deals with the task of identifying the most likely author of a computer program, given a set of predefined author candidates. This is usually .based on the analysis of other program samples of undisputed authorship by the same programmer. There are several cases where the application of such a method could be of a major benefit, such as authorship disputes, proof of authorship in court, tracing the source of code left in the system after a cyber attack, etc. We present a new approach, called the SCAP (Source Code Author Profiles) approach, based on byte-level n-gram profiles in order to represent a source code author's style. Experiments on data sets of different programming-language (Java or C++) and varying difficulty (6 to 30 candidate authors) demonstrate the effectiveness of the proposed approach.A comparison with a previous source code authorship identification study based on more complicated information shows that the SCAP approach is language independent and that n-gram author profiles are better able to capture the idiosyncrasies of the source code authors. Moreover, the SCAP approach is able to deal surprisingly well with cases where only a limited amount of very short programs per programmer is available for training. It is also demonstrated that the effectiveness of the proposed model is not affected by the absence of comments in the source code, a condition usually met in cyber-crime cases. Georgia Frantzeskou, Efstathios Stamatatos, Stefanos Gritzalis, Sokratis K. Katsikas |
ICSE | 1 |