Ross Maciejewski

dblp:81/5349 · DBLP profile ↗
← Back
16ranked-venue papers in the field
0as first author
5since 2021 · last 2024
0000-0001-8803-6355ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6Data Mining & Knowledge Discovery · 5Information Retrieval & Web Search · 4Database Systems & Data Management · 1
YearPublicationVenuePosition
2024 IDNet: A Novel Identity Document Dataset via Few-Shot and Quality-Driven Synthetic Data Generation
abstract
Effective fraud detection and analysis of government-issued identity documents, such as passports, driver’s licenses, and identity cards, are essential in thwarting identity theft and bolstering security on online platforms. The accuracy of training fraud detection and analysis tools depends on the availability of extensive and diverse identity document datasets. However, current publicly available benchmark datasets for identity document analysis, including MIDV-500, MIDV-2020, and FMIDV, fall short in several aspects: they offer a limited number of samples of ten European country document types, cover insufficient varieties of fraud patterns, and seldom include alterations in critical personal identifying fields such as portrait images, limiting their utility in training models capable of detecting realistic frauds while preserving privacy. In response to these shortcomings, our research introduces a new benchmark dataset, IDNet, designed to advance privacy-preserving fraud detection efforts, synthesized by integrating the generative models and a Bayesian optimization approach. The IDNet dataset comprises 837, 060 images of synthetically generated identity documents, totaling approximately 490 gigabytes, categorized into 20 types from 10 U.S. states and 10 European countries, which is the largest identity document dataset publicly available today. We evaluated the fidelity and utility of IDNet to demonstrate the effectiveness of our unique synthetic data generation method. We also presented two use cases of the dataset, illustrating how it can aid in training privacy-preserving fraud detection methods, and facilitating the generation of camera and video capturing of identity documents.
Lulu Xie, Yancheng Wang 0001, Soham Nag, Rajeev Goel, Niranjan Erappa Narayana Swamy, Yingzhen Yang, Chaowei Xiao, Jonathan Prisby, Ross Maciejewski, Jia Zou 0001
IEEE Big Data10
2023 Fairness-Aware Clique-Preserving Spectral Clustering of Temporal Graphs
abstract
With the widespread development of algorithmic fairness, there has been a surge of research interest that aims to generalize the fairness notions from the attributed data to the relational data (graphs). The vast majority of existing work considers the fairness measure in terms of the low-order connectivity patterns (e.g., edges), while overlooking the higher-order patterns (e.g., k-cliques) and the dynamic nature of real-world graphs. For example, preserving triangles from graph cuts during clustering is the key to detecting compact communities; however, if the clustering algorithm only pays attention to triangle-based compactness, then the returned communities lose the fairness guarantee for each group in the graph. Furthermore, in practice, when the graph (e.g., social networks) topology constantly changes over time, one natural question is how can we ensure the compactness and demographic parity at each timestamp efficiently. To address these problems, we start from the static setting and propose a spectral method that preserves clique connections and incorporates demographic fairness constraints in returned clusters at the same time. To make this static method fit for the dynamic setting, we propose two core techniques, Laplacian Update via Edge Filtering and Searching and Eigen-Pairs Update with Singularity Avoided. Finally, all proposed components are combined into an end-to-end clustering framework named F-SEGA, and we conduct extensive experiments to demonstrate the effectiveness, efficiency, and robustness of F-SEGA.
Dongqi Fu, Dawei Zhou 0003, Ross Maciejewski, Arie Croitoru, Marcus Boyd, Jingrui He
WWW3
2022 InfoFair: Information-Theoretic Intersectional Fairness
abstract
Algorithmic fairness is becoming increasingly important in data mining and machine learning. Among others, a foundational notation is group fairness. The vast majority of the existing works on group fairness, with a few exceptions, primarily focus on debiasing with respect to a single sensitive attribute, despite the fact that the co-existence of multiple sensitive attributes (e.g., gender, race, marital status, etc.) in the real-world is commonplace. As such, methods that can ensure a fair learning outcome with respect to all sensitive attributes of concern simultaneously need to be developed. In this paper, we study the problem of information-theoretic intersectional fairness (InfoFair), where statistical parity, a representative group fairness measure, is guaranteed among demographic groups formed by multiple sensitive attributes of interest. We formulate it as a mutual information minimization problem and propose a generic end-to-end algorithmic framework to solve it. The key idea is to leverage a variational representation of mutual information, which considers the variational distribution between learning outcomes and sensitive attributes, as well as the density ratio between the variational and the original distributions. Our proposed framework is generalizable to many different settings, including other statistical notions of fairness, and could handle any type of learning task equipped with a gradientbased optimizer. Empirical evaluations in the fair classification task on three real-world datasets demonstrate that our proposed framework can effectively debias the classification results with minimal impact to the classification accuracy.
Jian Kang 0008, Tiankai Xie, Xintao Wu, Ross Maciejewski, Hanghang Tong
IEEE Big Data4
2022 DISCO: Comprehensive and Explainable Disinformation Detection
abstract
Disinformation refers to false information deliberately spread to influence the general public, and the negative impact of disinformation on society can be observed in numerous issues, such as political agendas and manipulating financial markets. In this paper, we identify prevalent challenges and advances related to automated disinformation detection from multiple aspects and propose a comprehensive and explainable disinformation detection framework called DISCO. It leverages the heterogeneity of disinformation and addresses the opaqueness of prediction. Then we provide a demonstration of DISCO on a real-world fake news detection task with satisfactory detection accuracy and explanation. The demo video and source code of DISCO is now publicly available https://github.com/DongqiFu/DISCO. We expect that our demo could pave the way for addressing the limitations of identification, comprehension, and explainability as a whole.
Dongqi Fu, Yikun Ban, Hanghang Tong, Ross Maciejewski, Jingrui He
CIKM4
2022 Meta-Learned Metrics over Multi-Evolution Temporal Graphs
abstract
Graph metric learning methods aim to learn the distance metric over graphs such that similar (e.g., same class) graphs are closer and dissimilar (e.g., different class) graphs are farther apart. This is of critical importance in many graph classification applications such as drug discovery and epidemics categorization. Most, if not all, graph metric learning techniques consider the input graph as static, and largely ignore the intrinsic dynamics of temporal graphs. However, in practice, a graph typically has heterogeneous dynamics (e.g., microscopic and macroscopic evolution patterns). As such, labeling a temporal graph is usually expensive and also requires background knowledge. To learn a good metric over temporal graphs, we propose a temporal graph metric learning framework, Temp-GFSM. With only a few labeled temporal graphs, Temp-GFSM outputs a good metric that can accurately classify different temporal graphs and be adapted to discover new subspaces for unseen classes. Each proposed component in Temp-GFSM answers the following questions: What patterns are evolving in a temporal graph? How to weigh these patterns to represent the characteristics of different temporal classes? And how to learn the metric with the guidance from only a few labels? Finally, the experimental results on real-world temporal graph classification tasks from various domains show the effectiveness of our Temp-GFSM.
Dongqi Fu, Liri Fang, Ross Maciejewski, Vetle I. Torvik, Jingrui He
KDD3
2020 Enhancing Collective Estimates by Aggregating Cardinal and Ordinal Inputs
abstract
There are many factors that affect the quality of data received from crowdsourcing, including cognitive biases, varying levels of expertise, and varying subjective scales. This work investigates how the elicitation and integration of multiple modalities of input can enhance the quality of collective estimations. We create a crowdsourced experiment where participants are asked to estimate the number of dots within images in two ways: ordinal (ranking) and cardinal (numerical) estimates. We run our study with 300 participants and test how the efficiency of crowdsourced computation is affected when asking participants to provide ordinal and/or cardinal inputs and how the accuracy of the aggregated outcome is affected when using a variety of aggregation methods. First, we find that more accurate ordinal and cardinal estimations can be achieved by prompting participants to provide both cardinal and ordinal information. Second, we present how accurate collective numerical estimates can be achieved with significantly fewer people when aggregating individual preferences using optimization-based consensus aggregation models. Interestingly, we also find that aggregating cardinal information may yield more accurate ordinal estimates.
Ryan Kemmer, Yeawon Yoo, Adolfo R. Escobedo, Ross Maciejewski
HCOMP4
2020 InFoRM: Individual Fairness on Graph Mining
abstract
Algorithmic bias and fairness in the context of graph mining have largely remained nascent. The sparse literature on fair graph mining has almost exclusively focused on group-based fairness notation. However, the notion of individual fairness, which promises the fairness notion at a much finer granularity, has not been well studied. This paper presents the first principled study of Individual Fairness on gRaph Mining (InFoRM). First, we present a generic definition of individual fairness for graph mining which naturally leads to a quantitative measure of the potential bias in graph mining results. Second, we propose three mutually complementary algorithmic frameworks to mitigate the proposed individual bias measure, namely debiasing the input graph, debiasing the mining model and debiasing the mining results. Each algorithmic framework is formulated from the optimization perspective, using effective and efficient solvers, which are applicable to multiple graph mining tasks. Third, accommodating individual fairness is likely to change the original graph mining results without the fairness consideration. We conduct a thorough analysis to develop an upper bound to characterize the cost (i.e., the difference between the graph mining results with and without the fairness consideration). We perform extensive experimental evaluations on real-world datasets to demonstrate the efficacy and generality of the proposed methods.
Jian Kang 0008, Jingrui He, Ross Maciejewski, Hanghang Tong
KDD3
2020 Crowd Teaching with Imperfect Labels
abstract
The need for annotated labels to train machine learning models led to a surge in crowdsourcing - collecting labels from non-experts. Instead of annotating from scratch, given an imperfect labeled set, how can we leverage the label information obtained from amateur crowd workers to improve the data quality? Furthermore, is there a way to teach the amateur crowd workers using this imperfect labeled set in order to improve their labeling performance? In this paper, we aim to answer both questions via a novel interactive teaching framework, which uses visual explanations to simultaneously teach and gauge the confidence level of the crowd workers.
Yao Zhou 0003, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan 0001, Jingrui He
WWW3
2019 ORIGIN: Non-Rigid Network Alignment
abstract
Network alignment is a fundamental task in many high-impact applications. Most of the existing approaches either explicitly or implicitly consider the alignment matrix as a linear transformation to map one network to another, and might overlook the complicated alignment relationship across networks. On the other hand, node representation learning based alignment methods are hampered by the incomparability among the node representations of different networks. In this paper, we propose a unified semi-supervised deep model (ORIGIN) that simultaneously finds the non-rigid network alignment and learns node representations in multiple networks in a mutually beneficial way. The key idea is to learn node representations by the effective graph convolutional networks, which subsequently enable us to formulate network alignment as a point set alignment problem. The proposed method offers two distinctive advantages. First (node representations), unlike the existing graph convolutional networks that aggregate the node information within a single network, we can effectively aggregate the auxiliary information from multiple sources, achieving far-reaching node representations. Second (network alignment), guided by the high-quality node representations, our proposed non-rigid point set alignment approach overcomes the bottleneck of the linear transformation assumption. We conduct extensive experiments that demonstrate the proposed non-rigid alignment method is (1) effective, outperforming both the state-of-the-art linear transformation-based methods and node representation based methods, and (2) efficient, with a comparable computational time between the proposed multi-network representation learning component and its single-network counterpart.
Hanghang Tong, Jiejun Xu, Yifan Hu 0001, Ross Maciejewski
IEEE BigData5
2019 Multilevel Network Alignment
abstract
Network alignment, which aims to find the node correspondence across multiple networks, is a fundamental task in many areas, ranging from social network analysis to adversarial activity detection. The state-of-the-art in the data mining community often view the node correspondence as a probabilistic cross-network node similarity, and thus inevitably introduce an O(n2) lower bound on the computational complexity. Moreover, they might ignore the rich patterns (e.g., clusters) accompanying the real networks. In this paper, we propose a multilevel network alignment algorithm (Moana) which consists of three key steps. It first efficiently coarsens the input networks into their structured representations, and then aligns the coarsest representations of the input networks, followed by the interpolations to obtain the alignment at multiple levels including the node level at the finest granularity. The proposed coarsen-align-interpolate method bears two key advantages. First, it overcomes the O(n2) lower bound, achieving a linear complexity. Second, it helps reveal the alignment between rich patterns of the input networks at multiple levels (e.g., node, clusters, super-clusters, etc.). Extensive experimental evaluations demonstrate the efficacy of the proposed algorithm on both the node-level alignment and the alignment among rich patterns (e.g., clusters) at different granularities.
Hanghang Tong, Ross Maciejewski, Tina Eliassi-Rad
WWW3
2018 Source Free Domain Adaptation Using an Off-the-Shelf Classifier
abstract
With the advancements in many data mining and machine learning tasks, together with the availability of large-scale annotated data sets, there have been an increasing number of off-the-shelf tools for addressing these tasks, like Stanford NLP Toolkit and Caffe Model Zoo. However, many of these tasks are time-evolving in nature due to, e.g., the emergence of new features and the change of class conditional distribution of features. As a result, the off-the-shelf tools are not able to adapt to such changes and will suffer from sub-optimal performance in the target application. In this paper, we propose a generic framework named AOT for adapting the outputs from an off-the-shelf tool to accommodate the changes in the learning task. It considers two major types of changes, i.e., label deficiency and distribution shift, and aims to maximally boost the performance of the off-the-shelf tool in the target domain, with the help of a limited number of target domain labeled examples. Furthermore, we propose an iterative algorithm to solve the resulting optimization problem, and we demonstrate the superior performance of the proposed AOT framework on text and image data sets.
Arun Reddy Nelakurthi, Ross Maciejewski, Jingrui He
IEEE BigData2
2018 Motif-Preserving Dynamic Local Graph Cut
abstract
Modeling and characterizing high-order connectivity patterns are essential for understanding many complex systems, ranging from social networks to collaboration networks, from finance to neuroscience. However, existing works on high-order graph clustering assume that the input networks are static. Consequently, they fail to explore the rich high-order connectivity patterns embedded in the network evolutions, which may play fundamental roles in real applications. For example, in financial fraud detection, detecting loops formed by sequenced transactions helps identify money laundering activities; in emerging trend detection, star-shaped structures showing in a short burst may indicate novel research topics in citation networks. In this paper, we bridge this gap by proposing a local graph clustering framework that captures structure-rich subgraphs, taking into consideration the information of high-order structures in temporal networks. In particular, our motif-preserving dynamic local graph cut framework (MOTLOC) is able to model various user-defined temporal network structures and find clusters with minimum conductance in a polylogarithmic time complexity. Extensive empirical evaluations on synthetic and real networks demonstrate the effectiveness and efficiency of our MOTLOC framework.
Dawei Zhou 0003, Jingrui He, Hasan Davulcu, Ross Maciejewski
IEEE BigData4
2017 User-guided Cross-domain Sentiment Classification
abstract
Sentiment analysis has been studied for decades, and it is widely used in many real applications such as media monitoring. In sentiment analysis, when addressing the problem of limited labeled data from the target domain, transfer learning, or domain adaptation, has been successfully applied, which borrows information from a relevant source domain with abundant labeled data to improve the prediction performance in the target domain. The key to transfer learning is how to model the relatedness among different domains. For sentiment analysis, a common practice is to assume similar sentiment polarity for the common keywords shared by different domains. However, existing methods largely overlooked the human factor, i.e., the users who expressed such sentiment. In this paper, we address this problem by explicitly modeling the human factor related to sentiment classification. In particular, we assume that the content generated by the same user across different domains is biased in the same way in terms of the sentiment polarity. In other words, optimistic/pessimistic users demonstrate consistent sentiment patterns, no matter what the context is. To this end, we propose a new graph-based approach named U-Cross, which models the relatedness of different domains via both the shared users and keywords. It is non-parametric and semi-supervised in nature. Furthermore, we also study the problem of shared user selection to prevent ‘negative transfer’. In the experiments, we demonstrate the effectiveness of U-Cross by comparing it with existing state-of-the-art techniques on three real data sets.
Arun Reddy Nelakurthi, Hanghang Tong, Ross Maciejewski, Nadya Bliss, Jingrui He
SDM3
2015 Understanding hotspots: a topological visual analytics approach
abstract
Analysis of spatio-temporal event data is of central importance in many domains of science and policy making. Current visualization methods rely on animation, small multiples, and space-time cubes to enable spatio-temporal data exploration. These methods require the user to remember state spaces or deal with layout occlusions when exploring their data. To overcome such issues, we propose a novel visualization technique for such data that applies the topological notion of Reeb graphs to identify hotspots as areas of relatively high event density within kernel density estimates. We illustrate that the topological identification of hotspots proposed in this paper is able to elucidate lifetime, properties, and relationships of hotspots by visualizing their temporal evolution based on the spatio-temporal Reeb graph. To validate our approach, we demonstrate our method on an epidemiological and a crime dataset. The resulting visualizations assist users in quickly identifying and comprehending important dates, events, hotspot properties, and relationships between hotspots.
Jonas Lukasczyk, Ross Maciejewski, Christoph Garth, Hans Hagen
SIGSPATIAL/GIS2
2013 A novel visual analytics approach for clustering large-scale social data
abstract
Social data refers to data individuals create that is knowingly and voluntarily shared by them and is an exciting avenue into gaining insight into interpersonal behaviors and interaction. However, such data is large, heterogeneous and often incomplete, properties that make the analysis of such data extremely challenging. One common method of exploring such data is through cluster analysis, which can enable analysts to find groups of related users, behaviors and interactions. This paper presents a novel visual analysis approach for detecting clusters within large-scale social networks by utilizing a divide-analyze-recombine scheme that sequentially performs data partitioning, subset clustering and result recombination within an integrated visual interface. A case study on a microblog messaging data (with 4.8 millions users) is used to demonstrate the feasibility of this approach and comparisons are also provided to illustrate the performance benefits of this approach with respect to existing solutions.
Zhangye Wang, Juanxia Zhou, Jiyuan Liao, Wei Chen 0001, Ross Maciejewski
IEEE BigData6
2013 Understanding Twitter data with TweetXplorer
abstract
In the era of big data it is increasingly difficult for an analyst to extract meaningful knowledge from a sea of information. We present TweetXplorer, a system for analysts with little information about an event to gain knowledge through the use of effective visualization techniques. Using tweets collected during Hurricane Sandy as an example, we will lead the reader through a workflow that exhibits the functionality of the system.
Fred Morstatter, Shamanth Kumar, Huan Liu 0001, Ross Maciejewski
KDD4