Katarzyna Musial

dblp:11/5835 · also Katarzyna Musial-Gabrys · DBLP profile ↗
← Back
53ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 16 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Twinning Complex Networked Systems: Data-Driven Calibration of the mABCD Synthetic Graph Generator
Piotr Bródka, Michal Czuba, Bogumil Kaminski, Lukasz Krainski, Katarzyna Musial, Pawel Pralat, Mateusz Stolarski
WAW5
2025 Stochastic Block Models for Complex Network Analysis: A Survey
abstract
Complex networks enable to represent and characterize the interactions between entities in various complex systems which widely exist in the real world and usually generate vast amounts of data about all the elements, their behaviors and interactions over time. The studies concentrating on new network analysis approaches and methodologies are vital because of the diversity and ubiquity of complex networks. The stochastic block model (SBM), based on Bayesian theory, is a statistical network model. SBMs are essential tools for analyzing complex networks since SBMs have the advantages of interpretability, expressiveness, flexibility and generalization. Thus, designing diverse SBMs and their learning algorithms for various networks has become an intensively researched topic in network analysis and data mining. In this article, we review, in a comprehensive and in-depth manner, SBMs for different types of networks (i.e., model extensions), existing methods (including parameter estimation and model selection) for learning optimal SBMs for given networks and SBMs combined with deep learning. Finally, we provide an outlook on the future research directions of SBMs.
Xueyan Liu 0001, Wenzhuo Song, Katarzyna Musial, Yang Li 0030, Xuehua Zhao, Bo Yang 0002
ACM Trans. Knowl. Discov. Data3
2024 On taking advantage of opportunistic meta-knowledge to reduce configuration spaces for automated machine learning
David Jacob Kedziora, Tien-Dung Nguyen 0002, Katarzyna Musial, Bogdan Gabrys
Expert Syst. Appl.3
2023 Machine learning for administrative health records: A systematic review of techniques and applications
Adrian Wilkins-Caruana, Madhushi Niluka Bandara, Katarzyna Musial, Daniel R. Catchpoole, Paul J. Kennedy
Artif. Intell. Medicine3
2023 Inferring actual treatment pathways from patient records
Adrian Wilkins-Caruana, Madhushi Niluka Bandara, Katarzyna Musial, Daniel R. Catchpoole, Paul J. Kennedy
J. Biomed. Informatics3
2022 Skills Taught vs Skills Sought: Using Skills Analytics to Identify the Gaps between Curriculum and Job Markets
Alireza Ahadi, Kirsty Kitto, Marian-Andrei Rizoiu, Katarzyna Musial
EDM4
2022 NATS-Bench: Benchmarking NAS Algorithms for Architecture Topology and Size
abstract
Neural architecture search (NAS) has attracted a lot of attention and has been illustrated to bring tangible benefits in a large number of applications in the past few years. Architecture topology and architecture size have been regarded as two of the most important aspects for the performance of deep learning models and the community has spawned lots of searching algorithms for both of those aspects of the neural architectures. However, the performance gain from these searching algorithms is achieved under different search spaces and training setups. This makes the overall performance of the algorithms incomparable and the improvement from a sub-module of the searching model unclear. In this paper, we propose NATS-Bench, a unified benchmark on searching for both topology and size, for (almost) any up-to-date NAS algorithm. NATS-Bench includes the search space of 15,625 neural cell candidates for architecture topology and 32,768 for architecture size on three datasets. We analyze the validity of our benchmark in terms of various criteria and performance comparison of all candidates in the search space. We also show the versatility of NATS-Bench by benchmarking 13 recent state-of-the-art NAS algorithms on it. All logs and diagnostic information trained using the same setup for each candidate are provided. This facilitates a much larger community of researchers to focus on developing better NAS algorithms in a more comparable and computationally effective environment. All codes are publicly available at: https://xuanyidong.com/assets/projects/NATS-Bench.
Xuanyi Dong, Lu Liu 0019, Katarzyna Musial, Bogdan Gabrys
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 An insight into network structure measures and number of driver nodes
abstract
Control of complex networks is one of the most challenging open problems within network science. One view says that we can only claim to fully understand a network if we have the ability to influence or control it and predict the results of the employed control mechanisms. The area of control and controllability has progressed notably in the past ten years with several frameworks proposed namely, structural, exact, and physical. With continuing advancement in the area, the need to develop effective and efficient control methods that provide robust control is increasingly critical. The ultimate responsibility for controlling the network lies with the set of driver nodes that, according to the classical definition of the control theory of complex systems, can steer the network from any given state to a desired final state. To be able to develop better control mechanisms, we need to understand the relationship between different network structures and the number of driver nodes needed to control a given structure. This will allow understanding of which networks might be easier to control and the resources needed to control them. In this paper, we present a systematic study that builds an understanding of how network profiles (random (R), small-world (SW), scale-free (SF)) influence the number of driver nodes needed for control. Additionally, we also consider real social networks and identify their driver nodes set to further expand the discussion. We mean to find a correlation between network structure measures and number of driver nodes. Our results show that there is in fact a strong relationship between these.
Abida Sadaf, Luke Mathieson, Katarzyna Musial
ASONAM3
2021 Exploring Opportunistic Meta-knowledge to Reduce Search Spaces for Automated Machine Learning
abstract
Machine learning (ML) pipeline composition and optimisation have been studied to seek multi-stage ML models' i.e. preprocessor-inclusive, that are both valid and well-performing. These processes typically require the design and traversal of complex configuration spaces consisting of not just individual ML components and their hyperparameters, but also higher-level pipeline structures that link these components together. Optimisation efficiency and resulting ML-model accuracy both suffer if this pipeline search space is unwieldy and excessively large; it becomes an appealing notion to avoid costly evaluations of poorly performing ML components ahead of time. Accordingly, this paper investigates whether, based on previous experience, a pool of available classifiers/regressors can be preemptively culled ahead of initiating a pipeline composition/opti-misation process for a new ML problem, i.e. dataset. The previous experience comes in the form of classifier/regressor accuracy rankings derived, with loose assumptions, from a substantial but non-exhaustive number of pipeline evaluations; this meta-knowledge is considered ‘opportunistic’. Numerous experiments with the AutoWeka4MCPS package, including ones leveraging similarities between datasets via the relative landmarking method, show that, despite its seeming unreliability, opportunistic meta-knowledge can improve ML outcomes. However, results also indicate that the culling of classifiers/regressors should not be too severe either. In effect, it is better to search through a ‘top tier’ of recommended predictors than to pin hopes onto one previously supreme performer.
Tien-Dung Nguyen 0002, David Jacob Kedziora, Katarzyna Musial, Bogdan Gabrys
IJCNN3
2021 AutoWeka4MCPS-AVATAR: Accelerating automated machine learning pipeline composition and optimisation
Tien-Dung Nguyen 0002, Katarzyna Musial, Bogdan Gabrys
Expert Syst. Appl.2
2021 A Scalable Redefined Stochastic Blockmodel
abstract
Stochastic blockmodel (SBM) is a widely used statistical network representation model, with good interpretability, expressiveness, generalization, and flexibility, which has become prevalent and important in the field of network science over the last years. However, learning an optimal SBM for a given network is an NP-hard problem. This results in significant limitations when it comes to applications of SBMs in large-scale networks, because of the significant computational overhead of existing SBM models, as well as their learning methods. Reducing the cost of SBM learning and making it scalable for handling large-scale networks, while maintaining the good theoretical properties of SBM, remains an unresolved problem. In this work, we address this challenging task from a novel perspective of model redefinition. We propose a novel redefined SBM with Poisson distribution and its block-wise learning algorithm that can efficiently analyse large-scale networks. Extensive validation conducted on both artificial and real-world data shows that our proposed method significantly outperforms the state-of-the-art methods in terms of a reasonable trade-off between accuracy and scalability. 1
Xueyan Liu 0001, Bo Yang 0002, Hechang Chen, Katarzyna Musial, Hongxu Chen 0002, Yang Li 0030, Wanli Zuo
ACM Trans. Knowl. Discov. Data4
2021 A block-based generative model for attributed network embedding
Xueyan Liu 0001, Bo Yang 0002, Wenzhuo Song, Katarzyna Musial, Wanli Zuo, Hongxu Chen 0002, Hongzhi Yin
World Wide Web4
2020 Topic Enhanced Sentiment Spreading Model in Social Networks Considering User Interest
abstract
Emotion is a complex emotional state, which can affect our physiology and psychology and lead to behavior changes. The spreading process of emotions in the text-based social networks is referred to as sentiment spreading. In this paper, we study an interesting problem of sentiment spreading in social networks. In particular, by employing a text-based social network (Twitter) , we try to unveil the correlation between users' sentimental statuses and topic distributions embedded in the tweets, then to automatically learn the influence strength between linked users. Furthermore, we introduce user interest to refine the influence strength. We develop a unified probabilistic framework to formalize the problem into a topic-enhanced sentiment spreading model. The model can predict users' sentimental statuses based on their historical emotional status, topic distributions in tweets and social structures. Experiments on the Twitter dataset show that the proposed model significantly outperforms several alternative methods in predicting users' sentimental status. We also discover an intriguing phenomenon that positive and negative sentiment is more relevant to user interest than neutral ones. Our method offers a new opportunity to understand the underlying mechanism of sentimental spreading in online social networks.
Xiaobao Wang, Di Jin 0001, Katarzyna Musial, Jianwu Dang 0001
AAAI3
2020 Curriculum profile: modelling the gaps between curriculum and the job market
Aleksandr Gromov, Andrei Maslennikov, Nik Dawson, Katarzyna Musial, Kirsty Kitto
EDM4
2020 AVATAR - Machine Learning Pipeline Evaluation Using Surrogate Model
abstract
The evaluation of machine learning (ML) pipelines is essential during automatic ML pipeline composition and optimisation. The previous methods such as Bayesian-based and genetic-based optimisation, which are implemented in Auto-Weka, Auto-sklearn and TPOT, evaluate pipelines by executing them. Therefore, the pipeline composition and optimisation of these methods requires a tremendous amount of time that prevents them from exploring complex pipelines to find better predictive models. To further explore this research challenge, we have conducted experiments showing that many of the generated pipelines are invalid, and it is unnecessary to execute them to find out whether they are good pipelines. To address this issue, we propose a novel method to evaluate the validity of ML pipelines using a surrogate model (AVATAR). The AVATAR enables to accelerate automatic ML pipeline composition and optimisation by quickly ignoring invalid pipelines. Our experiments show that the AVATAR is more efficient in evaluating complex pipelines in comparison with the traditional evaluation approaches requiring their execution.
Tien-Dung Nguyen 0002, Tomasz Maszczyk, Katarzyna Musial, Marc-André Zöller, Bogdan Gabrys
IDA3
2020 Biomedical Named-Entity Recognition by Hierarchically Fusing BioBERT Representations and Deep Contextual-Level Word-Embedding
abstract
Text mining in the biomedical domain is increasingly important as the volume of biomedical documents increases. Thanks to advances in natural language processing (NLP), extracting valuable information from the biomedical literature is gaining popularity among researchers, and deep learning has enabled the development of effective biomedical text mining models. However, directly applying advancements in NLP to biomedical sources often yields unsatisfactory results, due to a word distribution drift from the general language domain corpora to specific biomedical corpora, and this drift introduces linguistic ambiguities. To overcome these challenges, this paper presents a novel method for biomedical named entity-recognition (BioNER) through hierarchically fusing representations from BioBERT, which is trained on biomedical corpora and Deep contextual-level word embeddings to handle the linguistic challenges within biomedical literature. Proposed text representation is then fed to attention-based Bi-directional Long Short Term Memory (BiLSTM) with Conditional random field (CRF) for the BioNER task. The experimental analysis shows that our proposed end-to-end methodology outperforms existing state-of-the-art methods for the BioNER task.
Usman Naseem, Katarzyna Musial, Peter W. Eklund, Mukesh Prasad
IJCNN2
2020 Towards Improved Deep Contextual Embedding for the identification of Irony and Sarcasm
abstract
Humans use tonal stress and gestural cues to reveal negative feelings that are expressed ironically using positive or intensified positive words when communicating vocally. However, in textual data, like posts on social media, cues on sentiment valence are absent, thus making it challenging to identify the true meaning of utterances, even for the human reader. For a given post, an intelligent natural language processing system should be able to identify whether a post is ironic/sarcastic or not. Recent work confirms the difficulty of detecting sarcastic/ironic posts. To overcome challenges involved in the identification of sentiment valence, this paper presents the identification of irony and sarcasm in social media posts through transformer-based deep, intelligent contextual embedding - T-DICE - which improves noise within contexts. It solves the language ambiguities such as polysemy, semantics, syntax, and words sentiments by integrating embeddings. T-DICE is then forwarded to attention-based Bidirectional Long Short Term Memory (BiLSTM) to find out the sentiment of a post. We report the classification performance of the proposed network on benchmark datasets for #irony & #sarcasm. Results demonstrate that our approach outperforms existing state-of-the-art methods.
Usman Naseem, Muhammad Imran Razzak, Peter W. Eklund, Katarzyna Musial
IJCNN4
2020 Multi-level Graph Convolutional Networks for Cross-platform Anchor Link Prediction
abstract
Cross-platform account matching plays a significant role in social network analytics, and is beneficial for a wide range of applications. However, existing methods either heavily rely on high-quality user generated content (including user profiles) or suffer from data insufficiency problem if only focusing on network topology, which brings researchers into an insoluble dilemma of model selection. In this paper, to address this problem, we propose a novel framework that considers multi-level graph convolutions on both local network structure and hypergraph structure in a unified manner. The proposed method overcomes data insufficiency problem of existing work and does not necessarily rely on user demographic information. Moreover, to adapt the proposed method to be capable of handling large-scale social networks, we propose a two-phase space reconciliation mechanism to align the embedding spaces in both network partitioning based parallel training and account matching across different social networks. Extensive experiments have been conducted on two large-scale real-life social networks. The experimental results demonstrate that the proposed method outperforms the state-of-the-art models with a big margin.
Hongxu Chen 0002, Hongzhi Yin, Xiangguo Sun, Tong Chen 0005, Bogdan Gabrys, Katarzyna Musial
KDD6
2020 Towards skills-based curriculum analytics: can we automate the recognition of prior learning?
abstract
In an era that will increasingly depend upon lifelong learning, the LA community will need to facilitate the movement and sharing of data and information across institutional and geographic boundaries. This will help us to recognise prior learning (RPL) and to personalise the learner experience. Here, we explore the utility of skills-based curriculum analytics and how it might facilitate the process of awarding RPL between two institutions. We explore the potential utility of combining natural language processing and skills taxonomies to map between subject descriptions for these two different institutions, presenting two algorithms we have developed to facilitate RPL and evaluating their performance. We draw attention to some of the issues that arise, listing areas that we consider ripe for future work in a surprisingly underexplored area.
Kirsty Kitto, Nikhil Sarathy, Aleksandr Gromov, Ming Liu 0007, Katarzyna Musial, Simon Buckingham Shum
LAK5
2020 Transformer based Deep Intelligent Contextual Embedding for Twitter sentiment analysis
Usman Naseem, Muhammad Imran Razzak, Katarzyna Musial, Muhammad Imran 0001
Future Gener. Comput. Syst.3
2020 ModMRF: A modularity-based Markov Random Field method for community detection
abstract
Complex networks are widely used in the research of social and biological fields. Analyzing real community structure in networks is the key to the study of complex networks. Modularity optimization is one of the most popular techniques in community detection. However, due to its greedy characteristic, it leads to a large number of incorrect partitions and more communities than in reality. Existing methods use the modularity as a Hamiltonian at the finite temperature to solve the above problem. Nevertheless, modularity is not formalized as a statistical model in the method, which makes many statistical inference methods limited and cannot be used. Moreover, the method uses the sum-product version of belief propagation (BP) and its performance is not as good as the max-sum version, since it calculates per-variable marginal probabilities rather than the joint probability . To address these issues, we propose a novel Markov Random Field (MRF) method by formalizing modularity as an energy function based on the rich structures of MRF to represent properties and constraints of this problem, and use the max-sum BP to infer model parameters. In order to analyze our method and compare it with existing methods, we conducted experiments on both real-world and synthetic networks with ground-truth of communities, showing that the new method outperforms the state-of-the-art methods.
Di Jin 0001, Yue Song 0001, Dongxiao He, Zhiyong Feng 0002, Shizhan Chen, Katarzyna Musial
Neurocomputing8
2020 Semi-supervised stochastic blockmodel for structure analysis of signed networks
abstract
Finding hidden structural patterns is a critical problem for all types of networks, including signed networks. Among all of the methods for structural analysis of complex network, stochastic blockmodel (SBM) is an important research tool because it is flexible and can generate networks with many different types of structures. However, most existing SBM learning methods for signed networks are unsupervised, leading to poor performance in terms of finding hidden structural patterns, especially when handling noisy and sparse networks. Learning SBM in a semi-supervised way is a promising avenue for overcoming the above difficulty. In this type of model, a small number of labelled nodes and a large number of unlabelled nodes, coupled with their network structures, are simultaneously used to train SBM. We propose a novel semi-supervised signed stochastic blockmodel and its learning algorithm based on variational Bayesian inference, with the goal of discovering both assortative (the nodes connect more densely in same clusters than that in different clusters) and disassortative (the nodes link more sparsely in same clusters than that in different clusters) structures from signed networks. The proposed model is validated through a number of experiments wherein it compared with the state-of-the-art methods using both synthetic and real-world data. The carefully designed tests, allowing to account for different scenarios, show our method outperforms other approaches existing in this space. It is especially relevant in the case of noisy and sparse networks as they constitute the majority of the real-world networks.
Xueyan Liu 0001, Wenzhuo Song, Katarzyna Musial, Xuehua Zhao, Wanli Zuo, Bo Yang 0002
Knowl. Based Syst.3
2019 Community Detection in Social Networks Considering Topic Correlations
abstract
Network contents including node contents and edge contents can be utilized for community detection in social networks. Thus, the topic of each community can be extracted as its semantic information. A plethora of models integrating topic model and network topologies have been proposed. However, a key problem has not been resolved that is the semantic division of a community. Since the definition of community is based on topology, a community might involve several topics. To ach
Yingkui Wang, Di Jin 0001, Katarzyna Musial, Jianwu Dang 0001
AAAI3
2019 Emotional Contagion-Based Social Sentiment Mining in Social Networks by Introducing Network Communities
abstract
The rapid development of social media services has facilitated the communication of opinions through online news, blogs, microblogs, instant-messages, and so on. This article concentrates on the mining of readers' social sentiments evoked by social media materials. Existing methods are only applicable to a minority of social media like news portals with emotional voting information, while ignore the emotional contagion between writers and readers. However, incorporating such factors is challenging since the learned hidden variables would be very fuzzy (because of the short and noisy text in social networks). In this paper, we try to solve this problem by introducing a high-order network structure, i.e. communities. We first propose a new generative model called Community-Enhanced Social Sentiment Mining (CESSM), which 1) considers the emotional contagion between writers and readers to capture precise social sentiment, and 2) incorporates network communities to capture coherent topics. We then derive an inference algorithm based on Gibbs sampling. Empirical results show that, CESSM achieves significantly superior performance against the state-of-the-art techniques for text sentiment classification and interestingness in social sentiment mining.
Xiaobao Wang, Di Jin 0001, Mengquan Liu, Dongxiao He, Katarzyna Musial, Jianwu Dang 0001
CIKM5
2019 DICE: Deep Intelligent Contextual Embedding for Twitter Sentiment Analysis
abstract
The sentiment analysis of the social media-based short text (e.g., Twitter messages) is very valuable for many good reasons, explored increasingly in different communities such as text analysis, social media analysis, and recommendation. However, it is challenging as tweet-like social media text is often short, informal and noisy, and involves language ambiguity such as polysemy. The existing sentiment analysis approaches are mainly for document and clean textual data. Accordingly, we propose a Deep Intelligent Contextual Embedding (DICE), which enhances the tweet quality by handling noises within contexts, and then integrates four embeddings to involve polysemy in context, semantics, syntax, and sentiment knowledge of words in a tweet. DICE is then fed to a Bi-directional Long Short Term Memory (BiLSTM) network with attention to determine the sentiment of a tweet. The experimental results show that our model outperforms several baselines of both classic classifiers and combinations of various word embedding models in the sentiment analysis of airline-related tweets.
Usman Naseem, Katarzyna Musial
ICDAR2
2018 The Impact of Social Versus Individual Learning for Agents' Risk Perception During Epidemics
abstract
Epidemics have always been a source of concern to people, both at the individual and government level. To fight outbreaks effectively, we need advanced tools that enable us to understand the factors that influence the spread of life-threatening diseases.
Shaheen A. Abdulkareem, Ellen-Wien Augustijn-Beckers, Katarzyna Musial, Yaseen T. Mustafa, Tatiana Filatova
eScience3
2018 NetSim - The framework for complex network generator
abstract
Networks are everywhere and their many types, including social networks, the Internet, food webs etc., have been studied for the last few decades. However, in real-world networks, it’s hard to find examples that can be easily comparable, i.e. have the same density or even number of nodes and edges. We propose a flexible and extensible NetSim framework to understand how properties in different types of networks change with varying number of edges and vertices. Our approach enables to simulate three classical network models (random, small-world and scale-free) with easily adjustable model parameters and network size. To be able to compare different networks, for a single experimental setup we kept the number of edges and vertices fixed across the models. To understand how they change depending on the number of nodes and edges we ran over 30,000 simulations and analysed different network characteristics that cannot be derived analytically. Two of the main findings from the analysis are that the average shortest path does not change with the density of the scale-free network but changes for small-world and random networks; the apparent difference in mean betweenness centrality of the scale-free network compared with random and small-world networks.
Akanda Wahid-Ul-Ashraf, Marcin Budka, Katarzyna Musial
KES3
2018 Robust Detection of Communities with Multi-semantics in Large Attributed Networks
Di Jin 0001, Ziyang Liu 0004, Dongxiao He, Bogdan Gabrys, Katarzyna Musial
KSEM (1)5
2018 Adaptive community detection incorporating topology and content in social networks✰
abstract
In social network analysis , community detection is a basic step to understand the structure and function of networks. Some conventional community detection methods may have limited performance because they merely focus on the networks’ topological structure . Besides topology, content information is another significant aspect of social networks. Although some state-of-the-art methods started to combine these two aspects of information for the sake of the improvement of community partitioning, they often assume that topology and content carry similar information. In fact, for some examples of social networks, the hidden characteristics of content may unexpectedly mismatch with topology. To better cope with such situations, we introduce a novel community detection method under the framework of non-negative matrix factorization (NMF). Our proposed method integrates topology as well as content of networks and has an adaptive parameter (with two variations) to effectively control the contribution of content with respect to the identified mismatch degree. Based on the disjoint community partition result, we also introduce an additional overlapping community discovery algorithm, so that our new method can meet the application requirements of both disjoint and overlapping community detection. The case study using real social networks shows that our new method can simultaneously obtain the community structures and their corresponding semantic description , which is helpful to understand the semantics of communities. Related performance evaluations on both artificial and real networks further indicate that our method outperforms some state-of-the-art methods while exhibiting more robust behavior when the mismatch between topology and content is observed.
Meng Qin 0002, Di Jin 0001, Kai Lei, Bogdan Gabrys, Katarzyna Musial
Knowl. Based Syst.5
2017 A Community Bridge Boosting Social Network Link Prediction Model
abstract
Link prediction in social networks is a very challenging research problem. The majority of existing approaches are based on the assumption that a given network evolves following a single phenomenon, e.g. "rich get richer" or "friend of my friend is my friend". However, dynamics of network dynamic changes over time and different parts of the network evolve in different manner. Because of that, we hypothesise that the prediction accuracy can be improved by providing different treatment to different nodes and links. Building on that assumption, we propose a Community Bridge Boosting Prediction Model (CBBPM) that treats certain bridge nodes differently depending on their structural position. For such bridge nodes their similarity score obtained using traditional link-based prediction methods is boosted. By doing so the importance of these nodes is increased and at the same time ensuring that the CBBPM can be used with any existing link prediction method. Our experimental results show that such bridge node similarity boosting mechanism can improve the accuracy of traditional link prediction methods.
Fei Gao 0009, Katarzyna Musial, Bogdan Gabrys
ASONAM2
2017 Adaptive Community Detection Incorporating Topology and Content in Social Networks
abstract
In social network analysis, community detection is a basic step to understand the structure, function and semantics of networks. Some conventional community detection methods may have limited performance because they merely focus on topological structure of networks. In addition to topology, content information is another significant aspect of social networks. Some state-of-the-art methods started to combine these two aspects of information, but they often assume that topology and content share the same characteristics. However, for some examples of social networks, content may mismatch with topological structure. In order to better cope with such situations, we introduce a novel community detection method under the framework of non-negative matrix factorization (NMF). Our proposed method integrates topology and content of networks, and introduces a novel adaptive parameter for controlling the contribution of content with respect to the identified mismatch degree between the topological and content information. The case study using real social networks show that our new method can simultaneously obtain community partition and the corresponding semantic descriptions. Experiments on both artificial networks and real social networks further indicate that our method outperforms some state-of-the-art methods while exhibiting more robust behaviour when the mismatch topological and content information is observed.
Meng Qin 0002, Di Jin 0001, Dongxiao He, Bogdan Gabrys, Katarzyna Musial
ASONAM5
2017 Using Centrality Measures to Predict Helpfulness-Based Reputation in Trust Networks
abstract
In collaborative Web-based platforms, user reputation scores are generally computed according to two orthogonal perspectives: (a) helpfulness-based reputation (HBR) scores and (b) centrality-based reputation (CBR) scores. In HBR approaches, the most reputable users are those who post the most helpful reviews according to the opinion of the members of their community. In CBR approaches, a “who-trusts-whom” network—known as a trust network —is available and the most reputable users occupy the most central position in the trust network, according to some definition of centrality. The identification of users featuring large HBR scores is one of the most important research issue in the field of Social Networks, and it is a critical success factor of many Web-based platforms like e-marketplaces, product review Web sites, and question-and-answering systems. Unfortunately, user reviews/ratings are often sparse, and this makes the calculation of HBR scores inaccurate. In contrast, CBR scores are relatively easy to calculate provided that the topology of the trust network is known. In this article, we investigate if CBR scores are effective to predict HBR ones, and, to perform our study, we used real-life datasets extracted from CIAO and Epinions (two product review Web sites) and Wikipedia and applied five popular centrality measures—Degree Centrality, Closeness Centrality, Betweenness Centrality, PageRank and Eigenvector Centrality—to calculate CBR scores. Our analysis provides a positive answer to our research question: CBR scores allow for predicting HBR ones and Eigenvector Centrality was found to be the most important predictor. Our findings prove that we can leverage trust relationships to spot those users producing the most helpful reviews for the whole community.
Pasquale De Meo, Katarzyna Musial, Domenico Rosaci, Giuseppe M. L. Sarnè, Lora Aroyo
ACM Trans. Internet Techn.2
2016 Hybrid structure-based link prediction model
abstract
In network science several topology-based link prediction methods have been developed so far. The classic social network link prediction approach takes as an input a snapshot of a whole network. However, with human activities behind it, this social network keeps changing. In this paper, we consider link prediction problem as a time-series problem and propose a hybrid link prediction model that combines eight structure-based prediction methods and self-adapts the weights assigned to each included method. To test the model, we perform experiments on two real world networks with both sliding and growing window scenarios. The results show that our model outperforms other structure-based methods when both precision and recall of the prediction results are considered.
Fei Gao King's, Katarzyna Musial
ASONAM2
2016 Pricing Options with Portfolio-Holding Trading Agents in Direct Double Auction
abstract
Options constitute integral part of modern financial trades, and are priced according to the risk associated with buying or selling certain asset in future. Financial literature mostly concentrates on risk-neutral methods of pricing options such as Black-Scholes model. However, it is an emerging field in option pricing theory to use trading agents with utility functions to determine the option's potential payoff for the agent. In this paper, we use one of such methodologies developed by Othman and Sandholm to design portfolio-holding agents that are endowed with popular option portfolios such as bullish spread, butterfly spread, straddle, etc to price options. Agents use their portfolios to evaluate how buying or selling certain option would change their current payoff structure, and form their orders based on this information. We also simulate these agents in a multi-unit direct double auction. The emerging prices are compared to risk-neutral prices under different market conditions. Through an appropriate endowment of option portfolios to agents, we can also mimic market conditions where the population of agents are bearish, bullish, neutral or non-neutral in their beliefs.
Sarvar Abdullaev, Peter McBurney, Katarzyna Musial
ECAI3
2014 Direct Exchange Mechanisms for Option Pricing
Sarvar Abdullaev, Peter McBurney, Katarzyna Musial
EUMAS3
2013 Active learning and inference method for within network classification
abstract
In relational learning tasks such as within network classification the main problem arises from the inference of nodes' labels based on the the ground true labels of remaining nodes. The problem becomes even harder if the nodes from initial network do not have any labels assigned and they have to be acquired. However, labels of which nodes should be obtained in order to provide fair classification results? Active learning and inference is a practical framework to study this problem. The method for active learning and inference in within network classification based on node selection is proposed in the paper. Based on the structure of the network it is calculated the utility score for each node, the ranking is formulated and for selected nodes the labels are acquired. The paper examines several distinct proposals for utility scores and selection methods reporting their impact on collective classification results performed on various real-world networks.
Tomasz Kajdanowicz, Radoslaw Michalski, Katarzyna Musial, Przemyslaw Kazienko
ASONAM3
2013 What kind of network are you?: using local and global characteristics in network categorisation tasks
abstract
The amount of research done in the area of real--world networked systems is rapidly growing. Everybody knows what six degrees of separation or small--world phenomenon are. Scientists very easily give labels to the networks they analyse. If it has power law node degree distribution then it has to be scale--free network or if there is high clustering coefficient then it must be small--world network. These simplifications, although convenient, are not always very useful from the perspective of understanding phenomena existing within the network. In this paper we decided to go back to the basics and investigate whether analysis of one single measure is enough to describe a network. We analyse both local and global characteristics in order to discover the "true" nature of a network. Not only using local and/or global measures can lead to different classification of a network but we also show how significantly different interpretation can result from analysing the same data by building network models as directed/undirected and/or weighted/binary graphs.
Katarzyna Musial, Bogdan Gabrys, Marcin Buczko
ASONAM1
2013 Creation and growth of online social network - How do social networks evolve?
abstract
Social networks are an example of complex systems consisting of nodes that can interact with each other and based on these activities the social relations are defined. The dynamics and evolution of social networks are very interesting but at the same time very challenging areas of research. In this paper the formation and growth of one of such structures extracted from data about human activities within online social networking system is investigated. Dynamics of both local and global characteristics are studied. Analysis of the dynamics of the network growth showed that it changes over time—from random process to power-law growth. The phase transition between those two is clearly visible. In general, node degree distribution can be described as the scale-free but it does not emerge straight from the beginning. Social networks are known to feature high clustering coefficient and friend-of-a-friend phenomenon. This research has revealed that in online social network, although the clustering coefficient grows over time, it is lower than expected. Also the friend-of-a-friend phenomenon is missing. On the other hand, the length of the shortest paths is small starting from the beginning of the network existence so the small-world phenomenon is present. The unique element of the presented study is that the data, from which the online social network was extracted, represents interactions between users from the beginning of the social networking site existence. The system, from which the data was obtained, enables users to interact using different communication channels and it gives additional opportunity to investigate multi-relational character of human relations.
Katarzyna Musial, Marcin Budka, Krzysztof Juszczyszyn
World Wide Web1
2013 Social networks on the Internet
abstract
The rapid development and expansion of the Internet and the social–based services comprised by the common Web 2.0 idea provokes the creation of the new area of research interests, i.e. social networks on the Internet called also virtual or online communities. Social networks can be either maintained and presented by social networking sites like MySpace , LinkedIn or indirectly extracted from the data about user interaction, activities or achievements such as emails, chats, blogs, homepages connected by hyperlinks, commented photos in multimedia sharing system, etc. A social network is the set of human beings or rather their digital representations that refer to the registered users who are linked by relationships extracted from the data about their activities, common communication or direct links gathered in the internet–based systems. Both digital representations named in the paper internet identities as well as their relationships can be characterized in many different ways. Such diversity yields for building a comprehensive and coherent view onto the concept of internet–based social networks. This survey provides in–depth analysis and classification of social networks existing on the Internet together with studies on selected examples of different virtual communities.
Katarzyna Musial, Przemyslaw Kazienko
World Wide Web1
2012 A Probabilistic Approach to Structural Change Prediction in Evolving Social Networks
abstract
We propose a predictive model of structural changes in elementary sub graphs of social network based on Mixture of Markov Chains. The model is trained and verified on a dataset from a large corporate social network analyzed in short, one day-long time windows, and reveals distinctive patterns of evolution of connections on the level of local network topology. We argue that the network investigated in such short timescales is highly dynamic and therefore immune to classic methods of link prediction and structural analysis, and show that in the case of complex networks, the dynamic sub graph mining may lead to better prediction accuracy. The experiments were carried out on the logs from the Wroclaw University of Technology mail server.
Krzysztof Juszczyszyn, Adam Gonczarek, Jakub M. Tomczak, Katarzyna Musial, Marcin Budka
ASONAM4
2011 The Dynamic Structural Patterns of Social Networks Based on Triad Transitions
abstract
In modern social networks built from the data collected in various computer systems we observe constant changes corresponding to external events or the evolution of underlying organizations. In this work we present a new approach to the description and quantifying evolutionary patterns of social networks illustrated with the data from the Enron email dataset. We propose the discovery of local network connection patterns (in this case: triads of nodes), measuring their transitions during network evolution and present the preliminary results of this approach. We define the Triad Transition Matrix (TTM) containing the probabilities of transitions between triads, then we show how it can help to discover the dynamic patterns of network evolution. Also, we analyse the roles performed by different triads in the network evolution by the creation of triad transition graph built from the TTM, which allows us to characterize the tendencies of structural changes in the investigated network. The future applications of our approach are also proposed and discussed.
Krzysztof Juszczyszyn, Marcin Budka, Katarzyna Musial
ASONAM3
2011 Multidimensional Social Network: Model and Analysis
Przemyslaw Kazienko, Katarzyna Musial, Elzbieta Kukla, Tomasz Kajdanowicz, Piotr Bródka
ICCCI (1)2
2011 Multidimensional Social Network in the Social Recommender System
abstract
All online sharing systems gather data that reflects users' collective behavior and their shared activities. This data can be used to extract different kinds of relationships which can be grouped into layers and which are basic components of the multidimensional social network (MSN) proposed in the paper. The layers are created on the basis of two types of relations between humans, i.e., direct and object-based ones which, respectively, correspond to either social or semantic links between individuals. For better understanding of the complexity of the social network structure, layers and their profiles were identified and studied on two, spanned in time, snapshots of the `Flickr' population. Additionally, for each layer, a separate strength measure was proposed. The experiments on the `Flickr' photo sharing system revealed that the relationships between users result either from semantic links between objects they operate on or from social connections of these users. Moreover, the density of the social network increases in time. The second part of this paper is devoted to building a social recommender system that supports the creation of new relations between users in a multimedia sharing system. Its main goal is to generate personalized suggestions that are continuously adapted to users' needs depending on the personal weights assigned to each layer in the MSN. The conducted experiments confirmed the usefulness of the proposed model.
Przemyslaw Kazienko, Katarzyna Musial, Tomasz Kajdanowicz
IEEE Trans. Syst. Man Cybern. Part A2
2009 Motif-Based Analysis of Social Position Influence on Interconnection Patterns in Complex Social Network
abstract
Motifs are small subgraphs showing statistically significant occurrence in given network. Motif analysis helps to insight into the local topology and functions of complex networks. The social position measure is interpreted as the importance of the node (user) within the network. We propose to fuse motif analysis with the social position assessment by colouring the nodes according to the measured position. As the distribution of discovered coloured motifs is utilized to mine the interconnection patterns between nodes, the results allow us to evaluate the influence of social position on the local topology of network connections. The experiment was carried out on the large social network derived from email communication.
Katarzyna Musial, Krzysztof Juszczyszyn
ACIIDS1
2009 Molecular dynamics modelling of the temporal changes in complex networks
abstract
The dynamic of complex social networks is nowadays one of the research areas of growing importance. The knowledge about the temporal changes of the network topology and characteristics is crucial in networked communication systems in which accurate predictions are important. In this paper a physics-inspired method to track the changes within complex social network is proposed. This method is based on the dynamic molecular modelling technique used in physics for simulation of large sets of interacting particles. The data for the conducted research was derived from e-mail communication within big company (Wroclaw University of Technology). From this information the social network of employees was extracted. The created social network was utilized to evaluate the methodology of social network dynamics modelling proposed by authors.
Krzysztof Juszczyszyn, Anna Musial, Katarzyna Musial, Piotr Bródka
IEEE Congress on Evolutionary Computation3
2009 Properties of Bridge Nodes in Social Networks
Katarzyna Musial, Krzysztof Juszczyszyn
ICCCI1
2009 Efficiency of Node Position Calculation in Social Networks
Piotr Bródka, Katarzyna Musial, Przemyslaw Kazienko
KES (2)2
2009 Structural Changes in an Email-Based Social Network
Krzysztof Juszczyszyn, Katarzyna Musial
KES-AMSTA2
2008 Patterns of Interactions in Complex Social Networks Based on Coloured Motifs Analysis
Katarzyna Musial, Krzysztof Juszczyszyn, Bogdan Gabrys, Przemyslaw Kazienko
ICONIP (2)1
2008 Local Topology of Social Network Based on Motif Analysis
Krzysztof Juszczyszyn, Przemyslaw Kazienko, Katarzyna Musial
KES (2)3
2008 Recommendation of Multimedia Objects Based on Similarity of Ontologies
Przemyslaw Kazienko, Katarzyna Musial, Krzysztof Juszczyszyn
KES (1)2
2008 Mining Personal Social Features in the Community of Email Users
Przemyslaw Kazienko, Katarzyna Musial
SOFSEM2
2006 Social Capital in Online Social Networks
Przemyslaw Kazienko, Katarzyna Musial
KES (2)2