VLDB 2026 Research / reviewers in the wild / expert
Martin Steinebach
dblp:s/MartinSteinebach
· DBLP profile ↗
49ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-0240-0388ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 36 · 11 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCOPE: A Proof-of-Concept Framework for Measuring Alignment in Data-Scarce Cybergrooming Domainsabstract285 Jeong-Eun Choi, Shiying Fan, Martin Steinebach |
COMPSAC | 3 |
| 2026 | Systematic Analysis of the Unintentional CSAM-Generation-Potential of Text-to-Image ModelsabstractThe rapid advancements of generative text-to-image (T2I) models and their open accessibility have enabled users to generate high-quality, photorealistic images of humans. Ethical challenges, particularly the deliberate generation of child sexual abuse material (CSAM), have been widely recognized. By contrast, the unintentional creation of such content has received little scholarly attention. The legal risks associated with this phenomenon nevertheless pose a significant threat to the increasing number of users of generative models. To investigate this issue, we conduct a comprehensive systematic evaluation of the potential of state-of-the-art T2I models to generate CSAM against users’ intentions. We systematically generate datasets with prompts specifying adult subjects. Using age estimation models, we analyze the datasets regarding age compliance across different visual demographic properties and prompt variations. Our findings show that the six examined prominent T2I models generate images depicting underage individuals despite explicit adult-oriented prompts. Across various dataset settings, Stable Diffusion 3.5 Large and Qwen-Image generate the highest proportion of persons classified as underage in our experiments. We share insights and strategies to mitigate the risk of generating CSAM. Nicolas Göller, Martin Steinebach |
WACV | 2 |
| 2025 | Disinformation Analysis on Telegram: A Metadata-Centered, Privacy-Aware Datasetabstract2298 Jeong-Eun Choi, Karla Schäfer, York Yannikos, Martin Steinebach |
IEEE Big Data | 4 |
| 2025 | AI-Generated Text Detection Using RoBERTa: A Generalizability and Explainability AnalysisabstractWith the rise of AI-generated text, the need for efficient detectors that perform well on various kinds of text generated by different models and prompts is increasing. We trained and evaluated three detectors: fine-tuned RoBERTa, trained an RoBERTa based adapter and applied adapter fusion on an AI-generated text classification task. The detectors are tested for generalisation on three unseen datasets containing various generation models, generation types, and text styles. All three detectors performed well on the three test sets, outperforming the baselines introduced with the test sets. We found that completely AI-generated text is easier to detect than text that has been manipulated, paraphrased, or rewritten. In contrast, text of less than 50 words was harder to detect than longer text. While texts generated through translating, paraphrasing, and rewriting were better recognized by adapter fusion in most settings; fine-tuned RoBERTa yielded the best overall results. Using transformers-interpret as explainability method and POS-tagging, conjunctions (CC) were identified as characteristics of AI-generated text, whereby personal pronouns (PRP), verbs (present tense; VBP) and modals (MD) were identified as indicators for human-written text, independent of generation models and detectors viewed. Karla Schäfer, Martin Steinebach |
COMPSAC | 2 |
| 2025 | Generalization in Audio Deepfake Detection: Evaluating ASR Encoder-Based Feature ExtractionabstractAudio deepfakes are artificially generated audio recordings that are used to create inauthentic imitations of a person's utterances. In the context of the increasing utilisation of artificial intelligence, the recognition of these recordings is becoming increasingly important. Thereby, the generalization of the detectors is a major challenge in audio deepfake detection (ADD). We analysed six feature representations: MFCC, LFCC, Wav2Vec2.0 and three pre-trained encoders of automatic speech recognition (ASR) systems (Whisper, SpeechT5, Canary) for ADD. For evaluating their generalizability, we trained the models on the ASVspoof 2019 LA train set and tested on the in-the-wild (ITW) test set. Furthermore, we evaluated the training time of the models, analysed the effect of additional (newer) training data and performed a correlation analysis. The ASR encoders were outperformed by using simple MFCC features, achieving an EER of 17.96% on the ITW test set with the lowest training time (7 min. per epoch). The best results were calculated using Wav2Vec 2.0 (+MFCC) with an EER of 17.60%, but using additional training data, and a training time of 135 min. per epoch. Wav2Vec also showed the highest correlations in varying training settings, indicating the best predictability of the model's performance. Karla Schäfer, Martin Steinebach |
DSAA | 2 |
| 2025 | The Sound of Language: A Bilingual Analysis of Voice Conversion and Text-to-Speech SynthesisabstractWith the rise of audio deepfakes, there is an increasing need for comprehensive studies on their generation methods, especially regarding their quality. Areas such as languages beyond English and Chinese, as well as comparisons between voice conversion (VC) and text-to-speech synthesis (TTS), remain underexplored. In our study, we generated samples in English and German using 10 recent VC and TTS methods, including two publicly accessible online tools. We compared these samples using various evaluation methods to gain insights into their quality across different factors. Our analysis indicates that TTS performs slightly better than VC, with minor differences between English and German data. Interestingly, in VC, the gender of the source speaker has minimal influence on the generated samples. Instead, the cross-gender factor appears to affect VC. For both VC and TTS, the target speaker samples used for generation seem to influence the quality of the generated samples. Jeong-Eun Choi, Karla Schäfer, Martin Steinebach |
ICASSP | 3 |
| 2025 | Machine Learning-Based Detection of AI-Generated Text via Stylistic and Statistical Feature ModelingabstractThrough the advances of large-language models (LLMs) AI- generated text can be created with ease. But, these tools can also pose a threat, e.g. through the creation of disinformation. In this work, we analysed texts generated by three LLMs: GPT-3.5, LLaMA3, and Qwen from the CUDRT dataset. We extracted 220 stylistic and statistical features of human and AI-generated text using the LFTK library. First, we analysed the features using the pearson correlation. Second, we trained five machine learning models and tested the classifiers on detecting completely AI-generated, polished, rewritten texts, and summaries created by AI. We calculated an F1-score of 90%+ for the text generated entirely by AI, depending on the LLM used. We found that AI-generated texts, independent of LLM, can be identified through a high kuperman age, i.e. high word complexity, whereby human-written texts are written with higher lexical variation and richness. We provide an explanation for the classification results and a comparison with RoBERTa (fine-tuned). Karla Schäfer, Martin Steinebach |
TrustCom | 2 |
| 2024 | Unveiling the Darkness: Analysing Organised Crime on the Wall Street Market Darknet Marketplace using PGP Public KeysabstractDarknet marketplaces (DNMs) are digital platforms for e-commerce that are primarily used to trade illegal and illicit products. They incorporate technological advantages for privacy protection and contribute to the growth of cybercriminal activities. In the past, researchers have explored methods to investigate multiple identities of vendors covering different DNMs. Leaving aside phenomena such as malicious forgery of identities or Sybil attacks, usernames and their corresponding PGP public keys are used to build brands around users and are considered a trusted method of vendor authentication across DNMs. Shiying Fan, Paul Moritz Ranly, Lukas Graner, Inna Vogel, Martin Steinebach |
ARES | 5 |
| 2024 | Natural Language Steganography by ChatGPTabstractNatural language steganography as well as natural language watermarking have been challenging because of the complexity and lack of noise in natural language. But with the advent of LLMs like ChatGPT, controlled synthesis of written language has become available. In this work, we show how ChatGPT can be utilized to generate synthetic texts of a given topic that act as stego covers for hidden messages. Martin Steinebach |
ARES | 1 |
| 2024 | Comparative Analysis of Voice Conversion in German
Karla Schäfer, Jeong-Eun Choi, Martin Steinebach |
ICPR (22) | 3 |
| 2023 | An Analysis of PhotoDNAabstractPhotoDNA is a popular hash used to combat CSAM. So far, only limited information has been provided about this hash in terms of its performance. In this paper, we provide an overview of robustness and false positives, as well as some basic observations about its structure. We show that it is robust against typical image processing such as lossy compression. On the other hand, it is only of limited robustness against cropping. We also give some suggestions for improving the algorithm or its use. Martin Steinebach |
ARES | 1 |
| 2022 | Using Telegram as a carrier for image steganography: Analysing Telegrams API limitsabstractTelegram is a messaging platform with millions of users per month. For this reason, it is a possible vector for steganographic messages. We investigate the feasibility of using Telegram as a messenger service for images with steganographic content, specifically we use F5 as a proof of concept. We evaluate the optimal resolution and quality settings to achieve the highest possible payload size. In order to support longer message transfers over Telegram, we design a cover channel with a regular schedule of images to have a high bandwidth. We found that the optimal resolution for message transfers is 2560x2560 at JPEG quality settings of 82. And that this configuration allows us to send an average of 81 kilobytes of data per image. Niklas Bunzel, Tobias Chen, Martin Steinebach |
ARES | 3 |
| 2022 | Towards Image Hashing Robust Against Cropping and RotationabstractImage recognition is an important mechanism used in various scenarios. In the context of multimedia forensics, its most significant task is to automatically detect already known child and adolescent pornography in a large set of images. For this purpose, numerous methods based on robust hashing and feature extraction are already known, and recently also supported by machine learning. However, in general, these methods are either only partially robust to changes such as rotation and pruning, or they require a large amount of data and computation. We present a method based on a simple block hash that is efficient to compute and memory efficient. To be robust against cropping and rotation, we combine the method with image segmentation and a method to normalize the rotation of the objects. Our evaluation shows that the method produces results comparable to much more complex approaches, but requires fewer resources. Martin Steinebach, Tiberius Berwanger, Huajian Liu |
ARES | 1 |
| 2022 | Data Acquisition on a Large Darknet MarketplaceabstractDarknet marketplaces in the Tor network are popular places to anonymously buy and sell various kinds of illegal goods. Previous research on marketplaces ranged from analyses of type, availability and quality of goods to methods for identifying users. Although many darknet marketplaces exist, their lifespan is usually short, especially for very popular marketplaces that are in focus of law enforcement agencies. York Yannikos, Julian Heeger, Martin Steinebach |
ARES | 3 |
| 2021 | Discovery of Single-Vendor Marketplace Operators in the Tor-NetworkabstractIn the Tor-network are many single-vendor marketplace web sites with a wide range of offers. Some of these vendor websites could be hosted by the same operators. In this paper, a method is presented to find out similarities between these vendor websites to discover possible operational structures between them. In order to accomplish this, similarity values are determined between the darknet websites by combining various features from the different categories structure, content and metadata. A dataset is determined by a first execution of the method and manual validation. Based on this data set, important features are extracted using decision trees. The features of the category structure HTML-Tag, HTML-Class, HTML-DOM-Tree as well as the metadata features File Content and Links-To have proven to be particularly important and can very effectively highlight similarities between darknet web sites. Supported by the similarity detection method, it was found that only 49% of 258 single-vendor marketplaces were unique, i.e. no similar sites existed. In addition, it was possible to find several duplicates of vendor websites, which made up 20%. Fabian Brenner, Florian Platzer, Martin Steinebach |
ARES | 3 |
| 2021 | exHide: Hiding Data within the exFAT File SystemabstractRecently, steganographic techniques for hiding data in file system metadata gained focus. Tools for commonly used file systems were published but the exFAT file system did not get much attention – probably because its structure provides only few suitable locations to hide data. In this work we present two approaches to hide data in the exFAT file system. While the first approach is more flexible regarding embedding locations, it is rather fragile and provides a lower embedding rate. The second approach, called exHide, has stricter requirements for embedding, but is rather robust and provides a reasonable embedding rate. We describe the design of both approaches, evaluate them, and discuss their weaknesses and advantages. Julian Heeger, York Yannikos, Martin Steinebach |
ARES | 3 |
| 2021 | Comparison of Cyber Attacks on Services in the Clearnet and Darknet
York Yannikos, Quang Anh Dang, Martin Steinebach |
IFIP Int. Conf. Digital Forensics | 3 |
| 2020 | Privacy-enhanced robust image hashing with bloom filtersabstractRobust image hashes are used to detect known illegal images, even after image processing. This is, for example, interesting for a forensic investigation, or for a company to protect their employees and customers by filtering content. The disadvantage of robust hashes is that they leak structural information of the pictures, which can lead to privacy issues. Our scientific contribution is to extend a robust image hash with privacy protection. We thus introduce and discuss such a privacy-preserving concept. The approach uses a probabilistic data structure - known as Bloom filter - to store robust image hashes. Bloom filter store elements by mapping hashes of each element to an internal data structure. We choose a cryptographic hash function to one-way encrypt and store elements. The privacy of the inserted elements is thus protected. We evaluate our implementation, and compare it to its underlying robust image hashing algorithm. Thereby, we show the cost with respect to error rates for introducing a privacy protection into robust hashing. Finally, we discuss our approach's results and usability, and suggest possible future improvements. Uwe Breidenbach, Martin Steinebach, Huajian Liu |
ARES | 2 |
| 2020 | Non-blind steganalysisabstractThe increasing digitization offers new ways, possibilities and needs for a secure transmission of information. Steganography and its analysis constitute an essential part of IT-Security. In this work we show how methods of blind-steganalysis can be improved to work in non-blind scenarios. The main objective was to examine how to take advantage of the knowledge of reference images to maximize the accuracy-rate of the analysis. Niklas Bunzel, Martin Steinebach, Huajian Liu |
ARES | 2 |
| 2020 | Detecting double compression and splicing using benfords first digit lawabstractDetecting image forgeries in JPEG encoded images has been a research topic in the field of media forensics for a long time. Until today, it still holds a high importance as tools to create convincing manipulations of images have become more and more accessible to the public, which in return might be used to e.g. generate fake news. In this paper, a passive forensic detection framework to detect image manipulations is proposed based on compression artefacts and Benfords First Digit Law. It incorporates a supervised approach to reconstruct the compression history as well as provides an un-supervised detection approach to detect double compression for unknown quantization tables. The implemented algorithms were able to achieve high AUC values when classifying high quality images exceeding similar state-of-the-art methods. Raphael Antonius Frick, Huajian Liu, Martin Steinebach |
ARES | 3 |
| 2020 | Critical traffic analysis on the tor networkabstractTor is a widely-used anonymity network with more than two million daily users. A special feature of Tor is the hidden service architecture. Hidden services are a popular method for anonymous communication or sharing web contents anonymously. A specialty in Tor is that all data packets that are sent are structured completely identical for security reasons. They are encrypted using the TLS protocol and have a fixed size of exactly 512 bytes. In an earlier implementation, Tor was an example of networks without generated traffic noise to make traffic analysis more difficult. In this work we describe a method to deanonymize any hidden service on Tor based on traffic analysis, which is a threat to anonymity online. This method allows an attacker with modest resources to deanonymize any hidden services in less than 12.5 days. Florian Platzer, Marcel Schäfer, Martin Steinebach |
ARES | 3 |
| 2019 | Improved Manipulation Detection with Convolutional Neural Network for JPEG ImagesabstractJPEG images are ubiquitously used in most real-world online and mobile applications, where uncompressed images are not available from the very beginning when digital images are generated by digital camera or smartphone. In this paper, an improved manipulation detection scheme for JPEG images with convolutional neural network is proposed which works better in practical scenarios. No uncompressed or lossless compressed images are used for network training and testing. All images are stored in JPEG format before and after any kind of manipulation. The proposed scheme is also able to detect manipulation even if the manipulated images are compressed with different JPEG quality factors from the training images. Experimental results demonstrate that the proposed scheme significantly outperforms the existing method under practical conditions of real-world applications. Huajian Liu, Martin Steinebach, Kathrin Schölei |
ARES | 2 |
| 2019 | Fake News Detection by Image Montage RecognitionabstractFake news have been a problem for multiple years now and in addition to this "fake images" that accompany them are becoming increasingly a problem too. The aim of such fake images is to back up the fake message itself and make it appear authentic. For this purpose, more and more images such as photo-montages are used, which have been spliced from several images. This can be used to defame people by putting them in unfavorable situations or the other way around as propaganda by making them appear more important. In addition, montages may have been altered with noise and other manipulations to make an automatic recognition more difficult. In order to take action against such montages and still detect them automated, a concept based on feature detection is developed. Furthermore, an indexing of the features is carried out by means of a nearest neighbor algorithm in order to be able to quickly compare a high number of images. Afterwards, images suspected to be a montage are reviewed by a verifier. This concept is implemented and evaluated with two feature detectors. Even montages that have been manipulated with different methods are identified as such in an average of 100 milliseconds with a probability of mostly over 90%. Martin Steinebach, Karol Gotkowski, Huajian Liu |
ARES | 1 |
| 2019 | Privacy and Robust HashesabstractWithin a forensic examination of a computer for illegal image content, robust hashing can be used to detect images even after they have been altered. Here the perceptible properties of an image are used to create the hash values. Whether an image has the same content is determined by a distance function. Cryptographic hash functions, on the other hand, create a unique bit-sensitive value. With these, no similarity measurement is possible, since only with exact agreement a picture is found. A minimal change in the image results in a completely different cryptographic hash value. However, the robust hashes have an big disadvantage: hash values can reveal something about the structure of the picture. This results in a data protection leak. The advantage of a cryptographic hash function is in turn that its values do not allow any conclusions about the structure of an image. The aim of this work is to develop a procedure for which combines the advantages of both hashing functions. Martin Steinebach, Sebastian Lutz, Huajian Liu |
ARES | 1 |
| 2019 | Detection and Analysis of Tor Onion ServicesabstractTor onion services can be accessed and hosted anonymously on the Tor network. We analyze the protocols, software types, popularity and uptime of these services by collecting a large amount of .onion addresses. Websites are crawled and clustered based on their respective language. In order to also determine the amount of unique websites a de-duplication approach is implemented. To achieve this, we introduce a modular system for the real-time detection and analysis of onion services. Address resolution of onion services is realized via descriptors that are published to and requested from servers on the Tor network that volunteer for this task. We place a set of 20 volunteer servers on the Tor network in order to collect .onion addresses. The analysis of the collected data and its comparison to previous research provides new insights into the current state of Tor onion services and their development. The service scans show a vast variety of protocols with a significant increase in the popularity of anonymous mail servers and Bitcoin clients since 2013. The popularity analysis shows that the majority of Tor client requests is performed only for a small subset of addresses. The overall data reveals further that a large amount of permanent services provide no actual content for Tor users. A significant part consists instead of bots, services offered via multiple domains, or duplicated websites for phishing attacks. The total amount of onion services is thus significantly smaller than current statistics suggest. Martin Steinebach, Marcel Schäfer, Alexander Karakuz, Katharina Brandl, York Yannikos |
ARES | 1 |
| 2018 | Channel SteganalysisabstractThe rise of social networks during the last 10 years has created a situation in which up to 100 million new images and photographs are uploaded and shared by users every day. This environment poses an ideal background for those who wish to communicate covertly by the use of steganography. It also creates a new set of challenges for steganalysts, who have to shift their field of work away from a purely scientific laboratory environment and into a diverse real-world scenario, while at the same time having to deal with entirely new problems, such as the detection of steganographic channels or the impact that even a low false positive rate has when investigating the millions of images which are shared every day on social networks. Martin Steinebach, Andre Ester, Huajian Liu |
ARES | 1 |
| 2018 | New authentication concept using certificates for big data analytic toolsabstractCompanies analyse large amounts of data on clusters of machines, using big data analytic tools such as Apache Spark and Apache Flink to analyse the data. Big data analytic tools are mainly tested regarding speed and reliability. Efforts about Security and thus authentication are spent only at second glance. In such big data analytic tools, authentication is achieved with the help of the Kerberos protocol that is basically built as authentication on top of big data analytic tools. However, Kerberos is vulnerable to attacks, and it lacks providing high availability when users are all over the world. To improve the authentication, this work presents first an analysis of the authentication in Hadoop and the data analytic tools. Second, we propose a concept to deploy Transport Layer Security (TLS) not only for the security of data transportation but as well for authentication within the big data tools. This is done by establishing the connections using certificates with a short lifetime. Paul J. E. Velthuis, Marcel Schäfer, Martin Steinebach |
ARES | 3 |
| 2018 | Monitoring Product Sales in Darknet ShopsabstractAnonymity networks and hidden services like those accessible in Tor, also called the "darknet", in combination with cryptocurrencies like bitcoin provide a relatively safe environment for criminal online activities. While this is a challenge for law enforcement, it brings opportunities for researchers to monitor these activities as they are often not really hidden but rather obfuscated and/or anonymized. In this paper we discuss such a monitoring approach for product sales in the darknet. We collect bitcoin addresses and data about product offerings in a number of shops run as hidden services in Tor. We then analyze transactions in the bitcoin blockchain that can be mapped to specific product sales in these shops. York Yannikos, Annika Schäfer, Martin Steinebach |
ARES | 3 |
| 2017 | Forensic Image Inspection Assisted by Deep LearningabstractInvestigations on the charge of possessing child pornography usually require manual forensic image inspection in order to collect evidence. When storage devices are confiscated, law enforcement authorities are hence often faced with massive image datasets which have to be screened within a limited time frame. As the ability to concentrate and time are highly limited factors of a human investigator, we believe that intelligent algorithms can effectively assist the inspection process by rearranging images based on their content. Thus, more relevant images can be discovered within a shorter time frame, which is of special importance in time-critical investigations of triage character. Felix Mayer, Martin Steinebach |
ARES | 2 |
| 2017 | Towards Imperceptible Natural Language Watermarking for GermanabstractWatermarking natural language is still a challenge in the domain of digital watermarking. Here, only the textual information must be used as a cover. No format changes or modified illustrations are accepted. Still, natural language watermarking (NLW) has some important applications, especially in leakage tracking, where a small set of individually marked copies of a confidently text is distributed. Properties of watermarking schemes such as imperceptibility, blindness or adaptability to non-English languages are of importance here. In order to address these three simultaneously, we present a blind NLW scheme, consisting of four independent embedding methods, which operate on the phonetical, morphological, lexical and syntactical layer of German texts. An evaluation based on 1,645 assessments provided by 131 test persons reveals promising results. Oren Halvani, Martin Steinebach, Lukas Graner |
IH&MMSec | 2 |
| 2015 | A ROI-based self-embedding method with high recovery capabilityabstractIn this paper, a novel block-wise fragile image watermarking algorithm for tampering localization and recovery is proposed. The image is divided into Region of Interest (ROI) and Region of Non Interest (RONI). Considering the ROI-based self-embedding problem as a special erasure channel, fountain code is applied in our method to deal with the reference symbols loss. And to minimize quality degradation in ROI, the reference symbols for recovery are only embedded into RONI blocks. Theoretical analysis shows that the result is nearly optimal, and the experimental results demonstrate the proposed method can offer low payload and high tamper tolerance. And the quality of both the watermarked image and the reconstructed image is high. Hongliang Cai, Huajian Liu, Martin Steinebach |
ICASSP | 3 |
| 2015 | A novel image secret sharing scheme with meaningful sharesabstractIn this paper a novel (t, n) threshold image secret sharing scheme is proposed. Based on the idea that there is close connection between secret sharing and coding theory, coding method on GF(2m) is applied in our scheme instead of the classical Lagrange's interpolation method in order to deal with the fidelity loss problem in the recovery. All the generated share images are meaningful and the size of each share image is the same as the secret image. The analysis proves our scheme is perfect and ideal and also has high security. The experiment results demonstrate that all the shares have high quality and the secret image can be recovered exactly. Hongliang Cai, Huajian Liu, Qizhao Yuan, Martin Steinebach |
ICASSP | 4 |
| 2015 | Universal Threshold Calculation for Fingerprinting Decoders using Mixture ModelsabstractCollusion attacks on watermarked media copies are commonly countered by probabilistically generated fingerprinting codes and appropriate tracing algorithms. The latter calculates accusation scores representing the suspiciousness of the fingerprints. In a 'detect many' scenario a threshold decides which scores are associated to the colluders. This work proposes a universal method to calculate thresholds for different decoders solely with knowledge of the accusation scores from the actual attack. Applying mixture models on the scores, the threshold is set up satisfying the selected error probabilities. It is independent from the fingerprint generation and can be applied at any decoder. Also no knowledge about the number of attackers or their strategy is needed. Marcel Schäfer, Sebastian Mair 0001, Waldemar Berchtold, Martin Steinebach |
IH&MMSec | 4 |
| 2014 | An Efficient Intrinsic Authorship Verification Scheme Based on Ensemble LearningabstractAuthorship Verification is an important sub discipline of digital text forensics. Its goal is to decide, if two texts are written by the same author or not. We present an efficient Authorship Verification scheme based on an ensemble of K-Nearest Neighbor classifiers, where each classifier generates a decision regarding a feature category. Our scheme provides many benefits such as, for instance, the independence of linguistic resources like thesauruses or language models. Furthermore, it can handle different Indo-European languages as for instance English, German, Spanish, Greek, Dutch, Swedish or French. Another benefit is the low runtime, due to the fact that deep linguistic processing (tagging, chunking, parsing, etc.) is not taken into account. Moreover, our scheme can easily be modified for example by replacing the involved distance function, the acceptance criterion or the used features including their parameters. The proposed scheme is evaluated against the publicly available PAN-2013 Author Identification (AI) test corpus, where it was ranked as the second-best in the top ten list, as well as against five other test corpora, compiled by our own. We show in our experiments that it is possible to achieve promising results, even when using a fixed setting of parameters and features across seven different languages. Oren Halvani, Martin Steinebach |
ARES | 2 |
| 2014 | Efficient Cropping-Resistant Robust Image HashingabstractA digital forensics examiner often has to deal with large amounts of multimedia content during an investigation. One important part of such an investigation is to identify illegal material like pictures containing child pornography. Robust image hashing is an effective technique to help identifying known illegal images even after the original images were modified by applying various image processing operations. However, some specific operations lead to increased false negative rates when using robust image hashing. One of the most challenging operations today is image cropping. In this work we introduce an approach to counter cropping operations on images by combining image segmentation and efficient block mean image hashing. We show that we are able to achieve high detection rates for images where cropping operations where applied on the original known source. This further improves the robustness of our image hashing approach. Martin Steinebach, Huajian Liu, York Yannikos |
ARES | 1 |
| 2014 | A new method for image reconstruction using self-embeddingabstractIn this paper we propose a new method for content reconstruction using self-embedding technology. As the content reconstruction problem can be regarded as a special kind of erasure channel, we use fountain code which has good performance in erasure channel to generate reference symbols for reconstruction. Theoretical analysis of success bound about maximal tamper rate is given and is verified by the Monte Carlo simulation. The experiments of a specific scheme show that our method can reduce the payload and improve the quality of the watermarked image, while still achieving high quality reconstruction and high tamper tolerance. Hongliang Cai, Huajian Liu, Martin Steinebach |
ICASSP | 3 |
| 2014 | Data Corpora for Digital Forensics Education and Research
York Yannikos, Lukas Graner, Martin Steinebach, Christian Winter 0001 |
IFIP Int. Conf. Digital Forensics | 3 |
| 2013 | Towards a Process Model for Hash Functions in Digital Forensics
Frank Breitinger, Huajian Liu, Christian Winter 0001, Harald Baier, Alexey Rybalchenko, Martin Steinebach |
ICDF2C | 6 |
| 2013 | FaceHash: Face Detection and Robust Hashing
Martin Steinebach, Huajian Liu, York Yannikos |
ICDF2C | 1 |
| 2013 | Automating Video File Carving and Content Identification
York Yannikos, Nadeem Ashraf, Martin Steinebach, Christian Winter 0001 |
IFIP Int. Conf. Digital Forensics | 3 |
| 2013 | Hash-Based File Content Identification Using Distributed Systems
York Yannikos, Jonathan Schluessler, Martin Steinebach, Christian Winter 0001, Kalman Graffi |
IFIP Int. Conf. Digital Forensics | 3 |
| 2013 | Leakage detection and tracing for databasesabstractThis work presents a new approach for hash database individualization by blending pre-calculated dummy hashes for each user to their databases. This individualization is necessary to detect a leak after the distribution of for example Whitelists or Blacklists. The proposed code is based on collusion secure fingerprinting codes and provides resistance against intuitive attacks for databases as well as attacks for which users collaborate. Both, code generation and tracing requires minimal effort and the distributor is able to control the robustness desired. Advantages compared to salting are improved effort and leakage detection without access to the hash database. A combination of salting and the proposed approach is possible. Waldemar Berchtold, Marcel Schäfer, Martin Steinebach |
IH&MMSec | 3 |
| 2013 | Natural language watermarking for german textsabstractIn this paper we present four informed natural language watermark embedding methods, which operate on the lexical and syntactic layer of German texts. Our scheme provides several benefits in comparison to state-of-the-art approaches, as for instance that it is not relying on complex NLP operations like full sentence parsing, word sense disambiguation, named entity recognition or semantic role parsing. Even rich lexical resources (e.g. WordNet or the Collins thesaurus), which play an essential role in many previous approches, are unnecessary for our system. Instead, our methods require only a Part-Of-Speech Tagger, simple wordlists that act as black- and whitelists and a trained classifier, which automatically predicts the ability of potential lexical or syntactic patterns to carry portions of the watermark message. Besides this, a part of the proposed methods can be easily adapted into other Indo-European languages, since the grammar rules the methods rely on are not restricted only to the German language. Because the methods perform only lexical and minor syntactic transformations, the watermarked text is not affected by grammatical distortion and simultaneously the meaning of the text is preserved in 82.14% of the cases. Oren Halvani, Martin Steinebach, Patrick Wolf |
IH&MMSec | 2 |
| 2013 | 3D watermarking in the context of video gamesabstractThis paper proposes a novel digital watermarking algorithm for 3D mesh models that, while developed with focus on video games, operates on a generalized mesh model representation. It combines spectral mesh compression with modifications of the vertex norm distribution in the spatial domain. Furthermore, it uses an optimizer based on evolutionary strategies in order to find the best embedding for the individual 3D model and watermark message. The proposed method is a blind or oblivious watermarking scheme and the embedding is dependent on a secret key. Results from a currently top-selling video game show that the proposed scheme provides satisfactory robustness against various transformations while also retaining good visual quality. Daniel Trick, Waldemar Berchtold, Marcel Schäfer, Martin Steinebach |
MMSP | 4 |
| 2011 | Fast and adaptive tracing strategies for 3-secure fingerprint watermarking codesabstractFingerprinting codes are mechanisms to increase the security of transaction watermarking. Digital transaction watermarking is an accepted mechanism to discourage illegal distribution of multimedia. Here copies of the same content are distributed with individual markings. Simple but effective attacks on transaction watermarking are collusion attacks where multiple individualized copies of the work are compared in order to detect and attack the watermark positions and thus create a counterfeited watermark. Marcel Schäfer, Waldemar Berchtold, Martin Steinebach |
Digital Rights Management Workshop | 3 |
| 2011 | Robust Hashing for Efficient Forensic Analysis of Image Sets
Martin Steinebach |
ICDF2C | 1 |
| 2009 | On the Reliability of Cell Phone Camera Fingerprint Recognition
Martin Steinebach, Mohamed El Ouariachi, Huajian Liu, Stefan Katzenbeisser 0001 |
ICDF2C | 1 |
| 2006 | Digital Watermarking for Image Authentication with LocalizationabstractIn this paper we propose a novel watermarking scheme for image authentication. A high localization of tampering detection is achieved by applying a random permutation process where every embedded watermark bit verifies random image positions instead of a local image block. Thereby the resolution of tampering detection is significantly improved in comparison to existing solutions while keeping the payload low. Furthermore, the proposed scheme doesn't embed the watermark locally but distributes it into the suitable embedding wavelet coefficients, avoiding embedding in smooth regions. Therefore, the scheme is intrinsically secure to block-based local attacks and retains high fidelity of the watermarked image. Scalable sensitivity of tampering detection is also enabled in the authentication process. Experimental results demonstrate the performance and effectiveness of the scheme for image authentication. Huajian Liu, Martin Steinebach |
ICIP | 2 |
| 2001 | Joint watermarking of audio-visual dataabstractBoth audio and video watermarking enable copyright protection with owner or customer authentication and the detection of media manipulations. The available watermarking technology concentrates on single media like audio or video. But the typical multimedia stream consists of both video and audio data. Our goal is to provide a solution with robust and fragile aspects to guarantee authentication and integrity by using watermarks in combination with content information. To achieve this, video stream parsing capabilities have to be added to existing watermarking algorithms. We propose to extract the audio and video content, called feature, and embed the content features with a robust watermarking scheme into the audio and video data. These features of audio and video data have to be identified. Watermarking payload and feature data requirements have to be compared. Our goal is to describe our current state of work in audio and video processing. We introduce a new approach for a/v content security. We embed watermarks both in video and audio channels of a MPEG system stream to increase security against attacks on content integrity. Jana Dittmann, Martin Steinebach |
MMSP | 2 |