EDBT 2026 Demo / reviewers in the wild / expert
Malik Magdon-Ismail
dblp:53/1994
· DBLP profile ↗
21ranked-venue papers in the field
3as first author
3since 2021 · last 2024
0000-0001-7327-7770ORCID · reported
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 13 (2 first)Information Retrieval & Web Search · 4Big Data, Cloud & Distributed Data Systems · 2Other / Interdisciplinary · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Natural Language Processing for Extracting Rich Disease Data Aligned To Satellite Meteorological DataabstractGlobal climate change is redefining our understanding of how diseases spread. In Sri Lanka, vector-borne diseases such as dengue fever historically surged during the monsoon seasons when temperatures were high enough for mosquito eggs to hatch. Unfortunately, due to rising temperatures and more erratic rainfall patterns, mosquito eggs can now hatch year-round making outbreaks increasingly unpredictable, leading to an alarming rise in hospitalizations and deaths. More data is needed to adapt our response to these diseases in an increasingly warmer world. In the contemporary landscape, a wealth of disease information is available, yet accessibility remains limited due to unstructured data formats such as PDFs. Therefore, converting unstructured disease reports into structured formats is necessary for effectively leveraging data. This paper introduces a comprehensive framework for collecting unstructured disease reports and transforming them into analyzable formats. By creating separate models tailored to each data format, we can ensure accuracy compared to general models. These straightforward models enhance accessibility and empower other researchers to use our tools. The returned structured data can then be harnessed for analysis, statistical purposes, and informing evidence-based public health interventions, thus facilitating more informed decision-making in healthcare. We deploy this framework to produce geospatial data for Sri Lanka and Brazil for many different conditions and align these data with satellite environmental data, providing for the first time a structured, aligned powerful dataset for disease modeling. Mahi Pasarkar, Junseob Kim, Eoin O'Gara, Alan Zhang, Malik Magdon-Ismail, Thilanka Munasinghe, Jiaqi Weng, David Qiu, Ethan Cruz, Jennifer C. Wei, Ashan Pathirana |
IEEE Big Data | 5 |
| 2024 | Graph Representation Learning for Dengue ForecastingabstractThe global expansion of the dengue belt, driven by climate change and increased urbanization, has led to a significant rise in dengue cases worldwide (1). Early warning systems (EWS) coupled with prompt public health response mechanisms are crucial in mitigating dengue-related morbidity and mortality globally. In Sri Lanka, dengue transmission occurs year-round with two peaks correlating to the southwest monsoon from May to September and the northeast monsoon from October to January (2). The presence of multiple dengue virus serotypes (DENV1–4) complicates epidemiological patterns, as sequential infections with different serotypes can increase the risk of severe disease manifestations detected by surveillance systems (3). Understanding and integrating these virological dynamics, vector dynamics, and real-time surveillance data are essential for developing effective EWS and targeted public health interventions. We propose the use of Graph Neural Networks (GNNs) as an EWS. Using Earth observational data from NASA’s global satellites and dengue incidence data from Sri Lanka’s Ministry of Health, we developed traditional and graph-based EWS to forecast dengue cases across Sri Lanka’s 25 districts between 2013 and 2022. We demonstrate empirically that GNNs incorporating spatiotemporal relations significantly outperform traditional EWS models such as Autoregressive Integrated Moving Average (ARIMA), Random Forest, and Long Short-Term Memory (LSTM). Our source code is available on GitHub. Jiaqi Weng, David Qiu, Ethan Cruz, Malik Magdon-Ismail, Thilanka Munasinghe, Jennifer C. Wei, Ashan Pathirana, Mahi Pasarkar |
IEEE Big Data | 4 |
| 2023 | Learning Network Dynamics from Noisy Steady StatesabstractWe present efficient algorithms to learn the parameters governing the dynamics of networked agents, given equilibrium steady state data. A key feature of our methods is the ability to learn without seeing the dynamics, using only the steady states. A key to the efficiency of our approach is the use of mean-field approximations to tune the parameters within a nonlinear least squares (NLS) framework. Our results on real networks demonstrate the accuracy of our approach in two ways. Using the learned parameters, we can: (i) Recover more accurate estimates of the true steady states when the observed steady states are noisy. (ii) Predict evolution to new equilibrium steady states after perturbations to the network topology. Yanna Ding, Jianxi Gao, Malik Magdon-Ismail |
ASONAM | 3 |
| 2020 | NoisyCUR: An Algorithm for Two-Cost Budgeted Matrix Completion
Alex Gittens, Malik Magdon-Ismail |
ECML/PKDD (1) | 3 |
| 2019 | The intrinsic scale of networks is smallabstractWe define the intrinsic scale at which a network begins to reveal its identity as the scale at which subgraphs in the network (created by a random walk) are distinguishable from similar sized subgraphs in a perturbed copy of the network. We conduct an extensive study of intrinsic scale for several networks, ranging from structured (e.g. road networks) to ad-hoc and unstructured (e.g. crowd sourced information networks), to biological. We find: (a) The intrinsic scale is surprisingly small (7-20 vertices), even though the networks are many orders of magnitude larger. (b) The intrinsic scale quantifies "structure" in a network - networks which are explicitly constructed for specific tasks have smaller intrinsic scale. (c) The structure at different scales can be fragile (easy to disrupt) or robust. Malik Magdon-Ismail, Kshiteesh Hegde |
ASONAM | 1 |
| 2017 | NP-hardness and inapproximability of sparse PCA
Malik Magdon-Ismail |
Inf. Process. Lett. | 1 |
| 2016 | Network classification using adjacency matrix embeddings and deep learningabstractWe study a natural problem: Given a small piece of a large parent network, is it possible to identify the parent network? We approach this problem from two perspectives. First, using several “sophisticated” or “classical” network features that have been developed over decades of social network study. These features measure aggregate properties of the network and have been found to take on distinctive values for different types of network, at the large scale. By using these classical features within a standard machine learning framework, we show that one can identify large parent networks from small (even 8-node) subgraphs. Second, we present a novel adjacency matrix embedding technique which converts the small piece of the network into an image and, within a deep learning framework, we are able to obtain prediction accuracies upward of 80%, which is comparable to or slightly better than the performance from classical features. Our approach provides a new tool for topology-based prediction which may be of interest in other network settings. Our approach is plug and play, and can be used by non-domain experts. It is an appealing alternative to the often arduous task of creating domain specific features using domain expertise. Philip Watters, Malik Magdon-Ismail |
ASONAM | 3 |
| 2016 | Manipulation among the Arbiters of Collective Intelligence: How Wikipedia Administrators Mold Public OpinionabstractOur reliance on networked, collectively built information is a vulnerability when the quality or reliability of this information is poor. Wikipedia, one such collectively built information source, is often our first stop for information on all kinds of topics; its quality has stood up to many tests, and it prides itself on having a “neutral point of view.” Enforcement of neutrality is in the hands of comparatively few, powerful administrators. In this article, we document that a surprisingly large number of editors change their behavior and begin focusing more on a particular controversial topic once they are promoted to administrator status. The conscious and unconscious biases of these few, but powerful, administrators may be shaping the information on many of the most sensitive topics on Wikipedia; some may even be explicitly infiltrating the ranks of administrators in order to promote their own points of view. In addition, we ask whether administrators who change their behavior in this suspicious manner can be identified in advance. Neither prior history nor vote counts during an administrator’s election are useful in doing so, but we find that an alternative measure, which gives more weight to influential voters, can successfully reject these suspicious candidates. This second result has important implications for how we harness collective intelligence: even if wisdom exists in a collective opinion (like a vote), that signal can be lost unless we carefully distinguish the true expert voter from the noisy or manipulative voter. Sanmay Das, Allen Lavoie, Malik Magdon-Ismail |
ACM Trans. Web | 3 |
| 2015 | Actions Are Louder than Words in Social MediaabstractWe study the relationship between the level of chatter on a social medium (like Twitter) and the level of the observed actions related to the chatter. For example, in a disaster, how does relief-donation chatter on Twitter correlate with the dollar amount received? One hypothesis is that a fraction of those who act will also tweet about it, which implies linear scaling, action ∝ chatter. On the other hand, if there is a contagion effect (those who tweet about donation incite others to donate) and these incited donors tend to be "quiet" and not broadcast their actions, then we expect superlinear scaling, Rostyslav Korolov, Justin Peabody, Allen Lavoie, Sanmay Das, Malik Magdon-Ismail, William A. Wallace |
ASONAM | 5 |
| 2014 | A note on sparse least-squares regression
Christos Boutsidis, Malik Magdon-Ismail |
Inf. Process. Lett. | 2 |
| 2014 | Random Projections for Linear Support Vector MachinesabstractLet X be a data matrix of rank ρ, whose rows represent n points in d -dimensional space. The linear support vector machine constructs a hyperplane separator that maximizes the 1-norm soft margin. We develop a new oblivious dimension reduction technique that is precomputed and can be applied to any input matrix X . We prove that, with high probability, the margin and minimum enclosing ball in the feature space are preserved to within ϵ-relative error, ensuring comparable generalization as in the original space in the case of classification. For regression, we show that the margin is preserved to ϵ-relative error with high probability. We present extensive experiments with real and synthetic data to support our theory. Saurabh Paul, Christos Boutsidis, Malik Magdon-Ismail, Petros Drineas |
ACM Trans. Knowl. Discov. Data | 3 |
| 2013 | Deconstructing centrality: thinking locally and ranking globally in networksabstractWe examine whether the prominence of individuals in different social networks is determined by their position in their local network or by how the community to which they belong relates to other communities. To this end, we introduce two new measures of centrality, both based on communities in the network: local and community centrality. Community centrality is a novel concept that we introduce to describe how central one's community is within the whole network. We introduce an algorithm to estimate the distance between communities and use it to find the centrality of communities. Using data from several social networks, we show that community centrality is able to capture the importance of communities in the whole network. We then conduct a detailed study of different social networks and determine how various global measures of prominence relate to structural centrality measures. Our measures deconstruct global centrality along local and community dimensions. In some cases, prominence is determined almost exclusively by local information, while in others a mix of local and community centrality matters. Our methodology is a step toward understanding of the processes that contribute to an actor's prominence in a network. Sibel Adali, Malik Magdon-Ismail |
ASONAM | 3 |
| 2013 | Manipulation among the arbiters of collective intelligence: how wikipedia administrators mold public opinionabstractOur reliance on networked, collectively built information is a vulnerability when the quality or reliability of this information is poor. Wikipedia, one such collectively built information source, is often our first stop for information on all kinds of topics; its quality has stood up to many tests, and it prides itself on having a "Neutral Point of View". Enforcement of neutrality is in the hands of comparatively few, powerful administrators. We find a surprisingly large number of editors who change their behavior and begin focusing more on a particular controversial topic once they are promoted to administrator status. The conscious and unconscious biases of these few, but powerful, administrators may be shaping the information on many of the most sensitive topics on Wikipedia; some may even be explicitly infiltrating the ranks of administrators in order to promote their own points of view. Neither prior history nor vote counts during an administrator's election can identify those editors most likely to change their behavior in this suspicious manner. We find that an alternative measure, which gives more weight to influential voters, can successfully reject these suspicious candidates. This has important implications for how we harness collective intelligence: even if wisdom exists in a collective opinion (like a vote), that signal can be lost unless we carefully distinguish the true expert voter from the noisy or manipulative voter. Sanmay Das, Allen Lavoie, Malik Magdon-Ismail |
CIKM | 3 |
| 2013 | iHypR: Prominence ranking in networks of collaborations with hyperedgesabstractWe present a new algorithm called iHypR for computing prominence of actors in social networks of collaborations. Our algorithm builds on the assumption that prominent actors collaborate on prominent objects, and prominent objects are naturally grouped into prominent clusters or groups (hyperedges in a graph). iHypR makes use of the relationships between actors, objects, and hyperedges to compute a global prominence score for the actors in the network. We do not assume the hyperedges are given in advance. Hyperedges computed by our method can perform as well or even better than “true” hyperedges. Our algorithm is customized for networks of collaborations, but it is generally applicable without further tuning. We show, through extensive experimentation with three real-life data sets and multiple external measures of prominence, that our algorithm outperforms existing well-known algorithms. Our work is the first to offer such an extensive evaluation. We show that unlike most existing algorithms, the performance is robust across multiple measures of performance. Further, we give a detailed study of the sensitivity of our algorithm to different data sets and the design choices within the algorithm that a user may wish to change. Our article illustrates the various trade-offs that must be considered in computing prominence in collaborative social networks. Sibel Adali, Malik Magdon-Ismail |
ACM Trans. Knowl. Discov. Data | 2 |
| 2012 | Communities and Balance in Signed Networks: A Spectral ApproachabstractDiscussion based websites like Epinions.com and Slashdot.com allow users to identify both friends and foes. Such networks are called Signed Social Networks and mining communities of like-minded users from these networks has potential value. We extend existing community detection algorithms that work only on unsigned networks to be applicable to signed networks. In particular, we develop a spectral approach augmented with iterative optimization. We use our algorithms to study both communities and structural balance. Our results indicate that modularity based communities are distinct from structurally balanced communities. Pranay Anchuri, Malik Magdon-Ismail |
ASONAM | 2 |
| 2012 | Identifying Long Lived Social Communities Using Structural PropertiesabstractWe present a two step procedure to identify long lasting communities, or evolutions, in social networks. First, we use axiomatic foundations to `rigorously' establish shorter, strongly-connected evolutions. In the second step, we use heuristics to combine these shorter evolutions to form longer evolutions. We apply the procedure on data generated from two networks - the DBLP co-authorship database and Live Journal blog data. We visually validate our algorithms by examining the topic evolution of the associated documents. Our results demonstrate that our algorithms, based solely on structural properties of the data (who interacts with whom), are able to track thematic trends in the literature. We then use a machine learning framework to identify the structural features of the early stages of a community's evolution are most useful for predicting the lifetime of the community. We find that (in order) size, intensity and stability are the most important features. Mark K. Goldberg, Malik Magdon-Ismail |
ASONAM | 2 |
| 2012 | Actions speak as loud as words: predicting relationships from social behavior dataabstractIn recent years, new studies concentrating on analyzing user personality and finding credible content in social media have become quite popular. Most such work augments features from textual content with features representing the user's social ties and the tie strength. Social ties are crucial in understanding the network the people are a part of. However, textual content is extremely useful in understanding topics discussed and the personality of the individual. We bring a new dimension to this type of analysis with methods to compute the type of ties individuals have and the strength of the ties in each dimension. We present a new genre of behavioral features that are able to capture the "function" of a specific relationship without the help of textual features. Our novel features are based on the statistical properties of communication patterns between individuals such as reciprocity, assortativity, attention and latency. We introduce a new methodology for determining how such features can be compared to textual features, and show, using Twitter data, that our features can be used to capture contextual information present in textual features very accurately. Conversely, we also demonstrate how textual features can be used to determine social attributes related to an individual. Sibel Adali, Fred Sisenda, Malik Magdon-Ismail |
WWW | 3 |
| 2012 | A Model for Information Growth in Collective Wisdom ProcessesabstractCollaborative media such as wikis have become enormously successful venues for information creation. Articles accrue information through the asynchronous editing of users who arrive both seeking information and possibly able to contribute information. Most articles stabilize to high-quality, trusted sources of information representing the collective wisdom of all the users who edited the article. We propose a model for information growth which relies on two main observations: (i) as an article’s quality improves, it attracts visitors at a faster rate (a rich-get-richer phenomenon); and, simultaneously, (ii) the chances that a new visitor will improve the article drops (there is only so much that can be said about a particular topic). Our model is able to reproduce many features of the edit dynamics observed on Wikipedia; in particular, it captures the observed rise in the edit rate, followed by 1/ t decay. Despite differences in the media, we also document similar features in the comment rates for a segment of the LiveJournal blogosphere. Sanmay Das, Malik Magdon-Ismail |
ACM Trans. Knowl. Discov. Data | 2 |
| 2011 | Prominence Ranking in Graphs with Community Structure
Sibel Adali, Malik Magdon-Ismail, Jonathan T. Purnell |
ICWSM | 3 |
| 2010 | A Permutation Approach to ValidationabstractWe give a permutation approach to validation (estimation of out-sample error). One typical use of validation is model selection. We establish the legitimacy of the proposed permutation complexity by proving a uniform bound on the out-sample error, similar to a VC-style bound. We extensively demonstrate this approach experimentally on synthetic data, standard data sets from the UCI-repository, and a novel diffusion data set. The out-of-sample error estimates are comparable to cross validation (CV); yet, the method is more efficient and robust, being less susceptible to overfitting during model selection. Malik Magdon-Ismail, Konstantin Mertsalov |
SDM | 1 |
| 2009 | Models of Communication Dynamics for Simulation of Information DiffusionabstractWe study information diffusion in real-life and synthetic dynamic networks, using well known threshold and cascade models of diffusion. Our test-bed is the communication network of the LiveJournal Blogosphere. We observe that the dynamic and static versions of the Blogograph, yield very different behaviors of the diffusion. It was earlier discovered that the communication dynamics of the Blogograph is quite high - over 60% of the links each week were not present in the previous week, though the size of the node set is relatively stable. Our models of the Blogograph evolution reproduce general stable statistics of the real-life Blogograph. We discover that the diffusion footprint on our models closely approximate the diffusion footprint of the real-life dynamic network. Konstantin Mertsalov, Malik Magdon-Ismail, Mark K. Goldberg |
ASONAM | 2 |