EDBT 2026 Demo / reviewers in the wild / expert
James P. Bagrow
dblp:28/7564
· DBLP profile ↗
11ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0002-4614-0792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evolving Form and Function: Dual-Objective Optimization in Neural Symbolic Regression Networksabstract[RETRACTED]Data increasingly abounds, but distilling their underlying relationships down to something interpretable remains challenging. One approach is genetic programming, which `symbolically regresses' a data set down into an equation. However, symbolic regression (SR) faces the issue of requiring training from scratch for each new dataset. To generalize across all datasets, deep learning techniques have been applied to SR. These networks, however, are only able to be trained using a symbolic objective: NN-generated and target equations are symbolically compared. But this does not consider the predictive power of these equations, which could be measured by a behavioral objective that compares the generated equation's predictions to actual data. Here we introduce a method that combines gradient descent and evolutionary computation to yield neural networks that minimize the symbolic and behavioral errors of the equations they generate from data. As a result, these evolved networks are shown to generate more symbolically and behaviorally accurate equations than those generated by networks trained by state-of-the-art gradient based neural symbolic regression methods. We hope this method suggests that evolutionary algorithms, combined with gradient descent, can improve SR results by yielding equations with more accurate form and function. Amanda Bertschinger, James P. Bagrow, Josh C. Bongard |
GECCO | 2 |
| 2023 | Revisiting Stylized Facts for Modern Stock MarketsabstractIn 2001, Rama Cont introduced a now-widely used set of ‘stylized facts’ to synthesize empirical studies of financial time series, resulting in 11 qualitative properties presumed to be universal to all financial markets. Here, we replicate Cont’s analyses for a convenience sample of stocks drawn from the U.S. stock market following a fundamental shift in market regulation. Our study relies on the same authoritative data as that used by the U.S. regulator. We find conclusive evidence in the modern market for eight of Cont’s original facts, while we find weak support for one additional fact and no support for the remaining two. Our study represents the first test of the original set of 11 stylized facts against a consistent set of stocks, therefore providing insight into how Cont’s stylized facts should be viewed in the context of modern stock markets. Ethan Ratliff-Crain, Colin Michael Van Oort, James P. Bagrow, Matthew T. K. Koehler, Brian F. Tivnan |
IEEE Big Data | 3 |
| 2023 | The Metric is the Message: Benchmarking Challenges for Neural Symbolic Regression
Amanda Bertschinger, Q. Tyrell Davis, James P. Bagrow, Josh C. Bongard |
ECML/PKDD (4) | 3 |
| 2023 | Node Placement to Maximize Reliability of a Communication Network with Application to Satellite SwarmsabstractThe structure of a mobile ad hoc network changes dynamically based on node positioning. We consider a setting in which nodes can communicate if they are within a prescribed distance of one another, giving rise to a communication network. An example is a swarm of small satellites that cooperate to perform tasks; such swarms are likely to become commonplace in space missions. In this paper, we consider the problem of adding a new node or repositioning a current node in the network while optimizing a given network parameter such as network reliability. Although there are infinitely many locations to place the new node in space, there are only finitely many possible changes to the communication network. We provide an algorithm that enumerates all possible network changes in time O($n$2log n) or O($n$3log n), for networks in 2- or 3-dimensional Euclidean space, respectively. We apply the proposed algorithm to a satellite swarm formation planning problem, where the goal is to maximize network reliability. Calum Buchanan, James P. Bagrow, M. Puck Rombach, Hamid R. Ossareh |
SMC | 2 |
| 2022 | The OCEAN mailing list data set: Network analysis spanning mailing lists and code repositoriesabstractCommunication surrounding the development of an open source project largely occurs outside the software repository itself. Historically, large communities often used a collection of mailing lists to discuss the different aspects of their projects. Multimodal tool use, with software development and communication happening on different channels, complicates the study of open source projects as a sociotechnical system. Here, we combine and standardize mailing lists of the Python community, resulting in 954,287 messages from 1995 to the present. We share all scraping and cleaning code to facilitate reproduction of this work, as well as smaller datasets for the Golang (122,721 messages), Angular (20,041 messages) and Node.js (12,514 messages) communities. To showcase the usefulness of these data, we focus on the CPython repository and merge the technical layer (which GitHub account works on what file and with whom) with the social layer (messages from unique email addresses) by identifying 33% of GitHub contributors in the mailing list data. We then explore correlations between the valence of social messaging and the structure of the collaboration network. We discuss how these data provide a laboratory to test theories from standard organizational science in large open source projects. Melanie Warrick, Samuel F. Rosenblatt, Jean-Gabriel Young, Amanda Casari, Laurent Hébert-Dufresne, James P. Bagrow |
MSR | 6 |
| 2022 | A Review and Framework for Modeling Complex Engineered System Development ProcessesabstractDeveloping complex engineered systems (CES) poses significant challenges for engineers, managers, designers, and businesspeople alike due to the inherent complexity of the systems and contexts involved. Furthermore, experts have expressed great interest in building a foundation of “theories” that describe the patterns underlying how development process qualities shape system outcomes. This article contributes to that foundation in two ways. First, it identifies the core elements of CES development processes (CESDPs) through a literature review. Then, it proposes the ComplEX System Integrated Utilities Model (CESIUM), a novel framework for exploring how numerous system and development process characteristics may affect the performance of CES. CESIUM creates abstract representations of a system architecture, the corresponding engineering organization, and the new product development process through which the organization designs the system. It does so by representing the system as a network of interdependent artifacts designed by agents. Simulated agents iteratively design their artifacts through optimization and share information with other agents, thereby advancing the CES toward a solution. Hence, this article makes it possible for researchers to compare how development process characteristics shape system outcomes. John Meluso, Jesse Austin-Breneman, James P. Bagrow, Laurent Hébert-Dufresne |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Which contributions count? Analysis of attribution in open sourceabstractOpen source software projects usually acknowledge contributions with text files, websites, and other idiosyncratic methods. These data sources are hard to mine, which is why contributorship is most frequently measured through changes to repositories, such as commits, pushes, or patches. Recently, some open source projects have taken to recording contributor actions with standardized systems; this opens up a unique opportunity to understand how community-generated notions of contributorship map onto codebases as the measure of contribution. Here, we characterize contributor acknowledgment models in open source by analyzing thousands of projects that use a model called All Contributors to acknowledge diverse contributions like outreach, finance, infrastructure, and community management. We analyze the life cycle of projects through this model's lens and contrast its representation of contributorship with the picture given by other methods of acknowledgment, including GitHub's top committers indicator and contributions derived from actions taken on the platform. We find that community-generated systems of contribution acknowledgment make work like idea generation or bug finding more visible, which generates a more extensive picture of collaboration. Further, we find that models requiring explicit attribution lead to more clearly defined boundaries around what is and is not a contribution. Jean-Gabriel Young, Amanda Casari, Katie McLaughlin, Milo Z. Trujillo, Laurent Hébert-Dufresne, James P. Bagrow |
MSR | 6 |
| 2018 | Efficient Crowd Exploration of Large Networks: The Case of Causal AttributionabstractAccurately and efficiently crowdsourcing complex, open-ended tasks can be difficult, as crowd participants tend to favor short, repetitive "microtasks". We study the crowdsourcing of large networks where the crowd provides the network topology via microtasks. Crowds can explore many types of social and information networks, but we focus on the network of causal attributions, an important network that signifies cause-and-effect relationships. We conduct experiments on Amazon Mechanical Turk (AMT) testing how workers can propose and validate individual causal relationships and introduce a method for independent crowd workers to explore large networks. The core of the method, Iterative Pathway Refinement, is a theoretically-principled mechanism for efficient exploration via microtasks. We evaluate the method using synthetic networks and apply it on AMT to extract a large-scale causal attribution network. Worker interactions reveal important characteristics of causal perception and the generated network data can help improve our understanding of causality and causal inference. Daniel Berenberg, James P. Bagrow |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2017 | Which friends are more popular than you?: Contact strength and the friendship paradox in social networksabstractThe friendship paradox states that in a social network, egos tend to have lower degree than their alters, or, "your friends have more friends than you do". Most research has focused on the friendship paradox and its implications for information transmission, but treating the network as static and unweighted. Yet, people can dedicate only a finite fraction of their attention budget to each social interaction: a high-degree individual may have less time to dedicate to individual social links, forcing them to modulate the quantities of contact made to their different social ties. Here we study the friendship paradox in the context of differing contact volumes between egos and alters, finding a connection between contact volume and the strength of the friendship paradox. The most frequently contacted alters exhibit a less pronounced friendship paradox compared with the ego, whereas less-frequently contacted alters are more likely to be high degree and give rise to the paradox. We argue therefore for a more nuanced version of the friendship paradox: "your closest friends have slightly more friends than you do", and in certain networks even: "your best friend has no more friends than you do". We demonstrate that this relationship is robust, holding in both a social media and a mobile phone dataset. These results have implications for information transfer and influence in social networks, which we explore using a simple dynamical model. James P. Bagrow, Christopher M. Danforth, Lewis Mitchell |
ASONAM | 1 |
| 2016 | What we write about when we write about causality: Features of causal statements across large-scale social discourseabstractIdentifying and communicating relationships between causes and effects is important for understanding our world, but is affected by language structure, cognitive and emotional biases, and the properties of the communication medium. Despite the increasing importance of social media, much remains unknown about causal statements made online. To study real-world causal attribution, we extract a large-scale corpus of causal statements made on the Twitter social network platform as well as a comparable random control corpus. We compare causal and control statements using statistical language and sentiment analysis tools. We find that causal statements have a number of significant lexical and grammatical differences compared with controls and tend to be more negative in sentiment than controls. Causal statements made online tend to focus on news and current events, medicine and health, or interpersonal relationships, as shown by topic models. By quantifying the features and potential biases of causality communication, this study improves our understanding of the accuracy of information and opinions found online. Thomas C. McAndrew, Josh C. Bongard, Christopher M. Danforth, Peter Sheridan Dodds, Paul Hines, James P. Bagrow |
ASONAM | 6 |
| 2011 | More Voices Than Ever? Quantifying Media Bias in Networks
Yu-Ru Lin, James P. Bagrow, David Lazer |
ICWSM | 2 |