Christoph Meinel

dblp:m/CMeinel · DBLP profile ↗
← Back
41ranked-venue papers in the field
2as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 12Other / Interdisciplinary · 10 (1 first)Information Retrieval & Web Search · 9 (1 first)Data Mining & Knowledge Discovery · 7Knowledge Engineering, Semantic Web & Information Systems · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2024 Identifying Personal Identifiable Information (PII) in Unstructured Text: A Comparative Study on Transformers
Md Hasan Shahriar, Anne V. D. M. Kayem, David Reich, Christoph Meinel
DEXA (2)4
2023 Enabling PII Discovery in Textual Data via Outlier Detection
Anne V. D. M. Kayem, Christoph Meinel
DEXA (2)3
2022 A Comprehensive Review of Anomaly Detection in Web Logs
abstract
Anomaly detection is a significant problem that has been researched within diverse research areas and application domains, especially in the area of web-based internet services or cybersecurity. Many anomaly detection techniques have been developed for specific application domains, while others are more generic. The Log files of Web-server give insight into the state of web-server and applications running on it and enable the detection of abnormal incidents or behavior. This paper focuses on particularly web-server HTTP logs to the problems of Web-server Log Anomaly Detection (WLAD) due to their own nature and features and aims to provide a brief review of different Data-driven techniques to get to the bottom of recent studies and developments made in the context of WLAD. Moreover, in this paper, the literature related to webserver logs analysis, as well as other closely related to the WLAD topic, are taken into consideration for review. We have classified existing techniques into different categories based on the underlying approach adopted. When applying a particular technique, these assumptions can be used as guidelines to assess the method’s effectiveness in this area. We also provide a basic security anomaly detection approach for each category and compare the existing methods as variants of the basic technique. Further, we identify the cons and pros of the current practices for each category. We also discuss the computational complexity of the methods, which is an essential issue in the domain of Big Data.
Mehryar Majd, Pejman Najafi, Seyed Ali Alhosseini, Feng Cheng 0002, Christoph Meinel
BDCAT5
2022 CoK: A Survey of Privacy Challenges in Relation to Data Meshes
Nikolai Podlesny, Anne V. D. M. Kayem, Christoph Meinel
DEXA (1)3
2022 Virtual machines pre-copy live migration cost modeling and prediction: a survey
abstract
Abstract Live migration is an essential feature in virtual infrastructure and cloud computing datacenters. Using live migration, virtual machines can be online migrated from a physical machine to another with negligible service interruption. Load balance, power saving, dynamic resource allocation, and high availability algorithms in virtual data-centers and cloud computing environments are dependent on live migration. Live migration process has six phases that result in live migration cost. Several papers analyze and model live migration costs for different hypervisors, different kinds of workloads and different models of analysis. In addition, there are also many other papers that provide prediction techniques for live migration costs. It is a challenge for the reader to organize, classify, and compare live migration overhead research papers due to the broad focus of the papers in this domain. In this survey paper, we classify, analyze, and compare different papers that cover pre-copy live migration cost analysis and prediction from different angels to show the contributions and the drawbacks of each study. Papers classification helps the readers to get different studies details about a specific live migration cost parameter. The classification of the paper considers the papers’ research focus, methodology, the hypervisors, and the cost parameters. Papers analysis helps the readers to know which model can be used for which hypervisor and to know the techniques used for live migration cost analysis and prediction. Papers comparison shows the contributions, drawbacks, and the modeling differences by each paper in a table format that simplifies the comparison. Virtualized Data-center and cloud computing clusters admins can also make use of this paper to know which live migration cost prediction model can fit for their environments.
Mohamed Esam Elsaid, Hazem M. Abbas, Christoph Meinel
Distributed Parallel Databases3
2021 Enabling Co-owned Image Privacy on Social Media via Agent Negotiation
abstract
Social media has become a popular communication platform on which shared content such as images form a large part of the communicated data. Yet, shared images can reveal sensitive information in the sense that the data after its publication remains accessible. Existing studies provide mechanisms to modify co-owned images for user privacy but require that every user involved be online in order to reach an agreement. In cases where users are offline at the time when the image is posted, no privacy agreement can be reached. Having a method of reaching a privacy agreement even when some of the users in the co-owned image are offline is useful in enforcing individual privacy settings vis-a-vis the co-owned image. In this paper, we present a multi-agent negotiation model that enforces individual privacy settings with respect to co-owned images even when the users are offline. Our multi-agent model includes three components, namely a coordinator agent, predictor agent, and filtering algorithm. The coordinator agent collects users’ opinions vis-a-vis a co-owned image to form an image that expresses the opinions of the involved users. The predictor agent supports the expression of offline user opinions, while the filtering algorithm removes privacy-violating information with respect to recent user opinions. Results from our proof-of-concept implementation indicate that improved efficiency in terms of privacy decisions can be achieved by employing agents to support offline user decisions regarding shared content.
Farzad Nourmohammadzadeh Motlagh, Anne V. D. M. Kayem, Nikolai Podlesny, Christoph Meinel
iiWAS4
2019 Towards Identifying De-anonymisation Risks in Distributed Health Data Silos
Nikolai Podlesny, Anne V. D. M. Kayem, Christoph Meinel
DEXA (1)3
2019 Incrementally updating unary inclusion dependencies in dynamic data
Nuhad Shaabani, Christoph Meinel
Distributed Parallel Databases2
2018 Improving the Efficiency of Inclusion Dependency Detection
abstract
The detection of all inclusion dependencies (INDs) in an unknown dataset is at the core of any data profiling effort. Apart from the discovery of foreign key relationships, INDs can help perform data integration, integrity checking, schema (re-)design, and query optimization. With the advent of Big Data, the demand increases for efficient INDs discovery algorithms that can scale with the input data size. To this end, we propose S-indd++ as a scalable system for detecting unary INDs in large datasets. S-indd++ applies a new stepwise partitioning technique that helps discard a large number of attributes in early phases of the detection by processing the first partitions of smaller sizes. S-indd++ also extends the concept of the attribute clustering to decide which attributes to be discarded based on the clustering result of each partition. Moreover, in contrast to the state-of-the-art, S-indd++ does not require the partition to fit into the main memory- which is a highly appreciable property in the face of the ever growing datasets. We conducted an exhaustive evaluation of S-indd ++ by applying it to large datasets with thousands attributes and more than 266 million tuples. The results show the high superiority of S-indd++ over the state-of-the-art. S-indd++ reduced up to 50~% of the runtime in comparison with Binder, and up to 98~% in comparison with S-indd.
Nuhad Shaabani, Christoph Meinel
CIKM2
2018 Personality Exploration System for Online Social Networks: Facebook Brands As a Use Case
abstract
User-generated content on social media platforms is a rich source of latent information about individual variables. Crawling and analyzing this content provides a new approach for enterprises to personalize services and put forward product recommendations. In the past few years, brands made a gradual appearance on social media platforms for advertisement, customers support and public relation purposes and by now it became a necessity throughout all branches. This online identity can be represented as a brand personality that reflects how a brand is perceived by its customers. We exploited recent research in text analysis and personality detection to build an automatic brand personality prediction model on top of the (Five-Factor Model) and (Linguistic Inquiry and Word Count) features extracted from publicly available benchmarks. The proposed model reported significant accuracy in predicting specific personality traits form brands. For evaluating our prediction results on actual brands, we crawled the Facebook API for 100k posts from the most valuable brands' pages in the USA and we visualize exemplars of comparison results and present suggestions for future directions.
Raad Bin Tareaf, Philipp Berger 0001, Patrick Hennig, Christoph Meinel
WI4
2017 Clustering Heuristics for Efficient t-closeness Anonymisation
Anne V. D. M. Kayem, Christoph Meinel
DEXA (2)2
2017 Incremental Discovery of Inclusion Dependencies
abstract
Inclusion dependencies form one of the most fundamental classes of integrity constraints. Their importance in classical data management is reinforced by modern applications such as data profiling, data cleaning, entity resolution and schema matching. Their discovery in an unknown dataset is at the core of any data analysis effort. Therefore, several research approaches have focused on their efficient discovery in a given, static dataset. However, none of these approaches are appropriate for applications on dynamic datasets, such as transactional datasets, scientific applications, and social network. In these cases, discovery techniques should be able to efficiently update the inclusion dependencies after an update in the dataset, without reprocessing the entire dataset.
Nuhad Shaabani, Christoph Meinel
SSDBM2
2016 Automated k-Anonymization and l-Diversity for Shared Data Privacy
Anne V. D. M. Kayem, C. T. Vester, Christoph Meinel
DEXA (1)3
2016 Detecting Maximum Inclusion Dependencies without Candidate Generation
Nuhad Shaabani, Christoph Meinel
DEXA (2)2
2016 A Journey of Bounty Hunters: Analyzing the Influence of Reward Systems on StackOverflow Question Response Times
abstract
Question and Answering (Q&A) platforms are an important source for information and a first place to go when searching for help. Q&A sites, like StackOverflow (SO), use reward systems to incentivize users to answer fast and accurately. In this paper we study and predict the response time for those questions on StackOverflow, that benefit from an additional incentive through so called bounties. Shaped by different motivations and rules these questions perform unlike regular questions. As our key finding we note that topic related factors provide a much stronger evidence than previously found factors for these questions. Finally, we compare models based on these features predicting the response time in the context of bounty questions.
Philipp Berger 0001, Patrick Hennig, Tom Bocklisch, Tom Herold, Christoph Meinel
WI5
2015 Scalable Inclusion Dependency Discovery
Nuhad Shaabani, Christoph Meinel
DASFAA (1)2
2015 Does Multilevel Semantic Representation Improve Text Categorization?
Cheng Wang 0002, Haojin Yang 0001, Christoph Meinel
DEXA (1)3
2015 Hot spot detection - An interactive cluster heat map for sentiment analysis
abstract
The blogosphere allows analysts to track opinions and sentiments of individuals, groups or the general public with large sample sizes regarding many topics. Essential for the sentiment analysis are visualizations. The visual understanding of large corpora's sentiment is far more effective than relying on textual representations of the analyzed content. Users are very interested in changes in the public opinion. Thus, the identification of patterns is of high interest. In this paper, we propose a cluster heat map visualization for sentiment visualization that displays the sentiment development of various related terms over time intervals. As we want to encourage the discovery of patterns over multiple related topics, we apply an ordering algorithm based on dimensionality reduction to the cluster heat map and improve upon the ordering algorithm to enable fast pattern recognition.
Patrick Hennig, Philipp Berger 0001, Maximilian Brehm, Bastien Grasnick, Jonathan Herdt, Christoph Meinel
DSAA6
2015 An Improved System For Real-Time Scene Text Recognition
abstract
In this paper we showcase a system for real-time text detection and recognition. We apply deep features created by Convolutional Neural Networks (CNNs) for both text detection and word recognition task. For text detection we follow the common localization-verification scheme which already shown its excellent ability in numerous previous work. In text localization stage, textual regions are roughly detected by using a MSERs (Maximally Stable Extremal Regions) detector with high recall rate. False alarms are then eliminated by using a CNNs classifier, and remaining text regions are further grouped into words. In the word recognition stage, we developed an skeleton-based text binarization method for segmenting text from its background. A CNNs based recognizer is then applied for recognizing character. The initial experiments show the powerful ability of deep features for text classification comparing with commonly used visual features. Our current implementation demonstrates real-time performance for recognizing scene text by using a standard PC with webcam.
Haojin Yang 0001, Cheng Wang 0002, Xiaoyin Che, Sheng Luo 0002, Christoph Meinel
ICMR5
2014 Geographic focus detection using multiple location taggers
abstract
Being able to identify locations associated to a Web resource is essential for providing location-based Web applications. However, geographical information in Web documents is rarely supplied in a machine-readable way and therefore not easily discoverable. As a consequence, it is necessary to extract geographical keywords from Web documents and to associate locations with them. This method is called location tagging. In this paper we present a location tagging approach for unstructured documents which utilizes multiple external location providers. Detected locations are ranked according to their relevance for the document, in order to identify a document's geographical focus, which is its most representative location. We present an exemplary implementation of our proposed approach using two location providers and evaluate our method's applicability.
Philipp Berger 0001, Patrick Hennig, Dustin Glaeser, Hauke Klement, Christoph Meinel
ASONAM5
2014 Accelerate the detection of trends by using sentiment analysis within the blogosphere
abstract
Information about upcoming trends is considered to be a valuable source of knowledge for both, companies and individuals. A large number of market analysts working at monitoring a particular business field, with many employing manual methods to do so. Since the amount of available data on the internet is far too high for humans to monitor, which carries a major risk of substantial amount of information being missed, the necessity arose to detect emerging trends automatically. Weblogs are an important medium to publish information and discuss certain topics. The web platform BlogIntelligence analyzes and visualizes the content and interconnection of blogs in the blogosphere. One area of focus is the detection of trends over a period of time, which is especially helpful for product vendors. But even more interesting, views expressed in weblog posts influence the reader's opinion. Integrating the strength and direction of expressed sentiments enhances the trend detection significantly. In this work, we introduce an approach to enrich the trend detection with sentiment analysis.
Patrick Hennig, Philipp Berger 0001, Claudia Lehmann, Andrina Mascher, Christoph Meinel
ASONAM5
2014 Exploring emotions over time within the blogosphere
abstract
A lot of research efforts are going on in the area of mining emotions within the world wide web. The BlogIntelligence application is analyzing tons of blog posts and extracts emotions out of this big amount of data. Therefore we thought about how to visualize these emotions in a very meaningful way. While we applied a smart map as a proven technique, we overcame conceptual and technical challenges to provide a feasible utility.
Patrick Hennig, Philipp Berger 0001, Christoph Meinel, Lukas Pirl, Lukas Schulze
DSAA3
2013 Web Mining Accelerated with In-Memory and Column Store Technology
Patrick Hennig, Philipp Berger 0001, Christoph Meinel
ADMA (1)3
2013 Identifying Domain Experts in the Blogosphere - Ranking Blogs Based on Topic Consistency
abstract
Current ranking algorithms, such as Page Rank, Technorati authority, and BI-Impact, favor blogs that report on a diversity of topics since those attract a large audience and thus more visitors, links, and comments. On the other side, niche blogs with a very specific topic only attract a small audience and thus have only a small reach. This results in a low ranking from today's blog retrieval systems. We argue that the consistency of a blog, i.e. how focused an author reports on a single topic, is a sign for expert knowledge. To find these blogs is particular important for other domain experts to identify blogs that they would like to follow and stay in active contact. To ease the retrieval of expert blogs, i.e. to separate them from the mass of blogs that report on random topics, we introduce a metric for blogs based on topic consistency. We divide the consistency ranking in four different aspects: (1) intra-post, (2) inter-post, (3) intra-blog, and (4) inter-blog consistency. By evaluating the metric with a test data set of 12,000 crawled blogs, we demonstrate the plausibility of our approach.
Philipp Berger 0001, Patrick Hennig, Christoph Meinel
Web Intelligence3
2012 Collaboratecom Special Issue Analyzing Distributed Whiteboard Interactions
abstract
We present the digital whiteboard system Tele-Board, which automatically captures all interactions made on the all-digital whiteboard and thus offers possibilities for a fast interpretation of usage characteristics. Analyzing team work at whiteboards is a time-consuming and error-prone process if manual interpretation techniques are applied. In a case study, we demonstrate how to conduct and analyze whiteboard experiments with the help of our system. The study investigates the role of video compared to an audio-only connection for distributed work settings. With the simplified analysis of communication data, we can prove that the video teams were more active than the audio teams and the distribution of whiteboard interaction between team members was more balanced. This way, an automatic analysis can not only support manual observations and codings, but also give insights that cannot be achieved with other systems. Beyond the overall view on one sessions focusing on key figures, it is also possible to find out more about the internal structure of a session.
Lutz Gericke, Raja Gumienny, Christoph Meinel
Int. J. Cooperative Inf. Syst.3
2010 Visualizing Blog Archives to Explore Content- and Context-Related Interdependencies
abstract
There has been virtually little in the way of user interfaces designed for the exploration and information gathering from large weblog datasets to allow for an integrated and aggregated knowledge collection and information analysis tool. Users have to rely on their own capability to find, select or filter entries and navigate through a blog archive. For weblogs with a large collection of entries this task easily becomes tedious, since current blog interfaces lack fundamental support for facilitating the exploration of their archives. A solution to this problem could be POSTCONNECT, a mature blog-archive visualization tool presented in this paper.
Justus Bross, Patrick Schilf, Christoph Meinel
Web Intelligence3
2009 X-Tracking the Changes of Web Navigation Patterns
Long Wang 0002, Christoph Meinel
PAKDD2
2009 Telling experts from spammers: expertise ranking in folksonomies
abstract
With a suitable algorithm for ranking the expertise of a user in a collaborative tagging system, we will be able to identify experts and discover useful and relevant resources through them. We propose that the level of expertise of a user with respect to a particular topic is mainly determined by two factors. Firstly, an expert should possess a high quality collection of resources, while the quality of a Web resource depends on the expertise of the users who have assigned tags to it. Secondly, an expert should be one who tends to identify interesting or useful resources before other users do. We propose a graph-based algorithm, SPEAR (SPamming-resistant Expertise Analysis and Ranking), which implements these ideas for ranking users in a folksonomy. We evaluate our method with experiments on data sets collected from Delicious.com comprising over 71,000 Web documents, 0.5 million users and 2 million shared bookmarks. We also show that the algorithm is more resistant to spammers than other methods such as the original HITS algorithm and simple statistical measures.
Michael G. Noll, Ching-man Au Yeung, Nicholas Gibbins, Christoph Meinel, Nigel Shadbolt
SIGIR4
2008 Who Reads and Writes the Social Web? A Security Architecture for Web 2.0 Applications
abstract
The World Wide Web has changed during the last decade. The so-called Web 2.0 enables inexperienced users to become worldwide publishers. Most often these users also don't have any idea about how to protect their own user-generated content or how to trust in content provided by aggregated and syndicated services. Public key infrastructures, digital signatures, and reputation services are well established but hard to understand and to handle for the layperson. We propose an efficient and user-friendly security architecture based on the popular tagging paradigm that connects user-defined tags with security policies, rules, and social network information to ensure access control, data integrity, and confidence also in derived and syndicated data.
Matthias Quasthoff, Harald Sack, Christoph Meinel
ICIW3
2008 The Metadata Triumvirate: Social Annotations, Anchor Texts and Search Queries
abstract
In this paper, we study and compare three different but related types of metadata about Web documents: social annotations provided by readers of Web documents, hyperlink anchor text provided by authors of Web documents, and search queries of users trying to find Web documents. We introduce a large research data set called CABS120k, which we have created for this study from a variety of information sources such as AOL500k, the Open Directory Project, del.icio.us/Yahoo!, Google and the WWW in general. We use this data set to investigate several characteristics of said metadata including length, novelty, diversity, and similarity and discuss theoretical and practical implications.
Michael G. Noll, Christoph Meinel
Web Intelligence2
2007 Mining the Students' Learning Interest in Browsing Web-Streaming Lectures
abstract
Web-streaming lectures overcome the space and time barriers between learning and teaching, but bring higher requirements on the learning feedback of students when they browse lectures. In this paper, we discover the students learning interest from their usage data in Web-based learning environment by using multi data mining methods. The learning interests are expressed in six questions, which were asked by the teachers. We use simple statistics, associate rules mining, multi linear regression and similarity comparing to answer different questions. The usage data of online learners are heterogeneous, including HTTP server logs and REAL Helix Universal logs, and these heterogeneous usage data are transformed into students browsing profiles. We implement our work on our Web-based learning environment: tele-TASK. The mined results help teachers to know their students clearly and adjust their teaching schedules efficiently
Long Wang 0002, Christoph Meinel
CIDM2
2007 Authors vs. readers: a comparative study of document metadata and content in the www
abstract
Collaborative tagging describes the process by which many users add metadata in the form of unstructured keywords to shared content. The recent practical success of web services with such a tagging component like Flickr or del.icio.us has provided a plethora of user-supplied metadata about web content for everyone to leverage.
Michael G. Noll, Christoph Meinel
ACM Symposium on Document Engineering2
2007 Semantic Composition of Lecture Subparts for a Personalized e-Learning
Naouel Karam, Serge Linckels, Christoph Meinel
ESWC3
2007 Why HTTPS Is Not Enough - A Signature-Based Architecture for Trusted Content on the Social Web
abstract
Easy to use, interactive web applications accumulating data from heterogeneous sources represent a recent trend on the World Wide Web, referred to as the Social Web. There however, security standards are often disregarded in favor of interface design or brand new features. This prevents the new services from gaining ground in the enterprise, in medical or e-government environments. We propose the deployment of XML Digital Signatures on web content and demonstrate how an architecture enabling for various security properties would look like. The solution proposed will benefit from the research on security engineering in Service-Oriented Architectures and thus allows for an in-depth analysis on the results.
Matthias Quasthoff, Harald Sack, Christoph Meinel
Web Intelligence3
2007 Detecting the Changes ofWeb Students' Learning Interest
abstract
In this paper, we discover the changes of students' learning interest from their usage data in web-based learning environment. Due to the effects on each other of the changes in Web students and Web lectures, we seek a method that integrates the changes in both sides to measure the changes of learning interest. We implement our work on our Web-based learning environment: tele-TASK. The mined results help teachers to know their students clearly and adjust their teaching schedules efficiently.
Long Wang 0002, Christoph Meinel
Web Intelligence2
2006 Building Content Clusters Based on Modelling Page Pairs
Christoph Meinel, Long Wang 0002
APWeb1
2004 Behaviour Recovery and Complicated Pattern Definition in Web Usage Mining
Long Wang 0002, Christoph Meinel
ICWE2
2004 Automatic Interpretation of Natural Language for a Multimedia E-learning Tool
Serge Linckels, Christoph Meinel
ICWE2
2000 Logging and Signing Document-Transfers on the WWW-A Trusted Third Party Gateway
abstract
We discuss a service that aims to make quoting of online documents, "Web contents" easy and provable. For that reason we report the conception of a gateway that works as a trusted third party (TTP) service which is based on a public key infrastructure (PKI). The developed service consists of the signing of any data-transmission that was done via the TTP-gateway. After the data-transfer a set of data can be requested from the used gateway that is signed with the TTP-gateways private key. This signed set of data contains for each request that was processed by the gateway at least three components. Those are the request from the client, the reply from the server and finally the signature of the (TTP) server. Storing this signed data the recipient at the client side can provide it to other parties suitable for a latter verification of the data transfer. The TTP server generates automatically verifiable statements of the kind "this request resulted in that response". Now anyone that trusts the chosen TTP-gateways statements will be able to verify the data-transfer by the use of the trusted third parties certified public key. Furthermore we describe a prototype implementation of such a service using HTTP. Finally a possible employment of the TTP-gateway is discussed.
Andreas Heuer 0002, Frank Losemann, Christoph Meinel
WISE3
1994 On the Complexity of Analysis and Manipulation of Boolean Functions in Terms of Decision Graphs
Jordan Gergov, Christoph Meinel
Inf. Process. Lett.2
1990 Logic VS. Complexity Theoretic Properties of the Graph Accessibility Problem for Directed Graphs of Bounded Degree
Christoph Meinel
Inf. Process. Lett.1