Alberto M. Segre

dblp:200/7923 · also Alberto Maria Segre · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
6since 2021 · last 2025
0000-0002-8886-6559ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 11 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2025 TempoBiGen: A Curated Generative Model for Healthcare Mobility Logs with Visit Duration
Hieu Vu, Alberto M. Segre, Bijaya Adhikari
ECML/PKDD (9)2
2023 Continually-Adaptive Representation Learning Framework for Time-Sensitive Healthcare Applications
abstract
Continual learning has emerged as a powerful approach to address the challenges of non-stationary environments, allowing machine learning models to adapt to new data while retaining the previously acquired knowledge. In time-sensitive healthcare applications, where entities such as physicians, hospital rooms, and medications exhibit continuous changes over time, continual learning holds great promise, yet its application remains relatively unexplored. This paper aims to bridge this gap by proposing a novel framework, i.e., Continually-Adaptive Representation Learning, designed to adapt representations in response to changing data distributions in evolving healthcare applications. Specifically, the proposed approach develops a continual learning strategy wherein the context information (e.g., interactions) of healthcare entities is exploited to continually identify and retrain the representations of those entities whose context evolved over time. Moreover, different from existing approaches, the proposed approach leverages the valuable patient information present in clinical notes to generate accurate and robust healthcare embeddings. Notably, the proposed continually-adaptive representations are have practical benefits in low-resource clinical settings where it is difficult to training machine learning models from scratch to accommodate the newly available data streams. Experimental evaluations on real-world healthcare datasets demonstrate the effectiveness of our approach in time-sensitive healthcare applications such as Clostridioides difficile (C.diff) Infection (CDI) incidence prediction task and medical intensive care unit transfer prediction task.
Akash Choudhuri, Hankyu Jang, Alberto M. Segre, Philip Polgreen, Kishlay Jha, Bijaya Adhikari
CIKM3
2022 Near-Optimal Spectral Disease Mitigation in Healthcare Facilities
abstract
Healthcare associated infections (HAIs) impose a substantial burden, both on patients and on the healthcare system. Designing effective strategies by using interventions such as vaccination, isolation, cleaning, mobility modification, etc., to reduce HAI spread is an important computational challenge. Spectral approaches are quite useful for modeling and solving problems of reducing disease spread over contact networks, but they have not been used for disease-spread models and contact networks that are specific for HAIs. Our main contribution in this paper is to close this gap. We make 3 specific contributions. (i) We present the first epidemic threshold results on temporal bipartite networks, i.e., a time-varying sequence of bipartite people-location network, for the Susceptible-Infected-Susceptible (SIS) model. (ii) We leverage our epidemic threshold result to pose the HAI mitigation problem as minimizing the spectral radius of the system matrix, while removing few nodes or edges. We present a scalable combinatorial algorithm that provides approximation guarantees. (iii) Through extensive experiments on actual healthcare contact networks derived from operations data from the University of Iowa Hospitals and Clinics, Carilion Clinic, and several other healthcare facilities, we show that our algorithm consistently outperforms a number of baselines (random, degree, top-k, eigen centrality) both in terms of reducing the spectral radius of the system matrix and in terms of reducing infections.
Masahiro Kiji, D. M. Hasibul Hasan, Alberto M. Segre, Sriram V. Pemmaraju, Bijaya Adhikari
ICDM3
2022 AdaAX: Explaining Recurrent Neural Networks by Learning Automata with Adaptive States
abstract
Recurrent neural networks (RNN) are widely used for handling sequence data. However, their black-box nature makes it difficult for users to interpret the decision-making process. We propose a new method to construct deterministic finite automata to explain RNN. In an automaton, states are abstracted from hidden states produced by the RNN, and the transitions represent input symbols. Thus, users can follow the paths of transitions, called patterns, to understand how a prediction is produced. Existing methods for extracting automata partition the hidden state space at the beginning of the extraction, which often leads to solutions that are either inaccurate or too large in size to comprehend. Unlike previous methods, our approach allows the automata states to be formed adaptively during the extraction. Instead of defining patterns on pre-determined clusters, our proposed model, AdaAX, identifies small sets of hidden states determined by patterns with finer granularity in data. Then these small sets are gradually merged to form states, allowing users to trade fidelity for lower complexity. Experiments show that our automata can achieve higher fidelity while being significantly smaller in size than baseline methods on synthetic and complex real datasets.
Dat Hong, Alberto M. Segre
KDD2
2021 The Impact of Pair Programming on College Students' Interest, Perceptions, and Achievement in Computer Science
abstract
Active and collaborative learning has shown considerable promise for improving student outcomes and reducing group disparities. As one common form of collaborative learning, pair programming is an adapted work practice implemented widely in higher education computing programs. In the classroom setting, it typically involves two computer science students working together on the same programming assignment. The present study examined a cluster-randomized trial of 1,198 undergraduates in 96 lab sections. Overall, pair programming had no significant effect on students’ course performance; subject matter interest; plans for future coursework; or their confidence, comfort, and anxiety with computer science. These findings were consistent across various student characteristics, except that students with favorable pretest scores exhibited negative effects from pair programming.
Nicholas A. Bowman, Lindsay Jarratt, K. C. Culver, Alberto M. Segre
ACM Trans. Comput. Educ.4
2021 COVID-19 modeling and non-pharmaceutical interventions in an outpatient dialysis unit
abstract
This paper describes a data-driven simulation study that explores the relative impact of several low-cost and practical non-pharmaceutical interventions on the spread of COVID-19 in an outpatient hospital dialysis unit. The interventions considered include: (i) voluntary self-isolation of healthcare personnel (HCPs) with symptoms; (ii) a program of active syndromic surveillance and compulsory isolation of HCPs; (iii) the use of masks or respirators by patients and HCPs; (iv) improved social distancing among HCPs; (v) increased physical separation of dialysis stations; and (vi) patient isolation combined with preemptive isolation of exposed HCPs. Our simulations show that under conditions that existed prior to the COVID-19 outbreak, extremely high rates of COVID-19 infection can result in a dialysis unit. In simulations under worst-case modeling assumptions, a combination of relatively inexpensive interventions such as requiring surgical masks for everyone, encouraging social distancing between healthcare professionals (HCPs), slightly increasing the physical distance between dialysis stations, and-once the first symptomatic patient is detected-isolating that patient, replacing the HCP having had the most exposure to that patient, and relatively short-term use of N95 respirators by other HCPs can lead to a substantial reduction in both the attack rate and the likelihood of any spread beyond patient zero. For example, in a scenario with R0 = 3.0, 60% presymptomatic viral shedding, and a dialysis patient being the infection source, the attack rate falls from 87.8% at baseline to 34.6% with this intervention bundle. Furthermore, the likelihood of having no additional infections increases from 6.2% at baseline to 32.4% with this intervention bundle.
Hankyu Jang, Philip Polgreen, Alberto M. Segre, Sriram V. Pemmaraju
PLoS Comput. Biol.3
2019 Evaluating architectural changes to alter pathogen dynamics in a dialysis unit: for the CDC MInD-healthcare group
abstract
This paper presents a high-fidelity agent-based simulation of the spread of methicillin-resistant Staphylococcus aureus (MRSA), a serious hospital acquired infection, within the dialysis unit at the University of Iowa Hospitals and Clinics (UIHC). The simulation is based on ten days of fine-grained healthcare worker (HCW) movement and interaction data collected from a sensor mote instrumentation of the dialysis unit by our research group in the fall of 2013. The simulation layers a detailed model of MRSA pathogen transfer, die-off, shedding, and infection on top of agent interactions obtained from data. The specific question this paper focuses on is whether there are simple, inexpensive architectural or process changes one can make in the dialysis unit to reduce the spread of MRSA? We evaluate two architectural changes of the nurses' station: (i) splitting the central nurses' station into two smaller distinct nurses' stations, and (ii) doubling the surface area of the nursing station. The first architectural change is modeled as a graph partitioning problem on a HCW contact network obtained from our HCW movement data. Somewhat counter-intuitively, our results suggest that the first architectural modification and the resulting reduction in HCW-HCW contacts has little to no effect on the spread of MRSA and may in fact lead to an increase in MRSA infection counts in some cases. In contrast, the second modification leads to a substantial reduction - between 12% and 22% for simulations with different parameters - in the number of patients infected by MRSA. These results suggest that the dynamics of an environmentally mediated infection such as MRSA may be quite different from that of infections whose spread is not substantially affected by the environment (e.g., respiratory infections or influenza).
Hankyu Jang, Samuel Justice, Philip Polgreen, Alberto M. Segre, Daniel K. Sewell, Sriram V. Pemmaraju
ASONAM4
2019 How Prior Programming Experience Affects Students' Pair Programming Experiences and Outcomes
abstract
Pair programming is a collaborative learning approach in computer science in which students (or employees) work closely with a partner on the same programming task. A long-standing question within pair programming is whether certain combinations of students lead to greater learning, effort, and/or performance. Earlier studies have explored the role of prior programming experience, including the discrepancy between partners' experience, as a potentially important factor in shaping these outcomes. However, the previous findings are highly inconsistent, which may result from divergent (and often suboptimal) ways of defining previous experience or skill, problems with self-selection into pairs, and small sample sizes that often yield nonsignificant results. The present study sought to improve on all of these limitations through an examination of 587 undergraduates who each participated in three different randomly assigned pairings. Not surprisingly, students' own programming experience was positively related to understanding concepts from lab, confidence in the finished product, and overall interest in computer science. However, students who worked with a more experienced partner actually had poorer outcomes, including lower effort exerted on the assignment, perceptions that their partner gave more effort than they did, less time in the driving role (i.e., typing out the assignment), lower understanding of concepts from lab, and less interest in computer science overall. The partner's experience is unrelated to other outcomes, including confidence in the finished assignment, feeling productive during lab session, and average grades received during the pairing. These results provide important considerations for the assignment of students to pairs.
Nicholas A. Bowman, Lindsay Jarratt, K. C. Culver, Alberto M. Segre
ITiCSE4
2019 A Large-Scale Experimental Study of Gender and Pair Composition in Pair Programming
abstract
The proportion of women in computer science majors is currently lower than in any other STEM major. Various studies have sought to explain---and ultimately find ways to reduce---gender disparities in computer science participation and persistence. Pair programming has been proposed as a practice that may not only promote outcomes overall within college and workplace environments, but also help diminish isolation and boost the confidence of women in computer science. Some promising results have been obtained for women in pair programming, but the findings are not consistent across studies, and the limitations of previous research make it difficult to draw strong conclusions. The present study examined 969 undergraduates in several introductory computer science courses who engaged in three different pairings throughout the semester. All pairings were randomly assigned, so the findings reflect the causal influence of the gender pair characteristics on a variety of student outcomes. Overall, having a female partner led to several positive outcomes relative to having a male partner; these included greater lab section attendance as well as greater confidence in the finished product and confidence in the solution for the pair programming assignment. The advantages of having a female partner were occasionally greater for female students than for male students. Overall, the significant findings were most pronounced in the course intended for computer science majors. These results offer evidence for the educational benefits of pair programming for promoting women's participation in computer science, as well as the need for careful consideration of pair composition.
Lindsay Jarratt, Nicholas A. Bowman, K. C. Culver, Alberto M. Segre
ITiCSE4
2010 Virtual Agents Based Simulation for Training Healthcare Workers in Hand Hygiene Procedures
Jeffrey W. Bertrand, Sabarish V. Babu, Philip Polgreen, Alberto M. Segre
IVA4
2008 MLIP: using multiple processors to compute the posterior probability of linkage
abstract
BACKGROUND: Localization of complex traits by genetic linkage analysis may involve exploration of a vast multidimensional parameter space. The posterior probability of linkage (PPL), a class of statistics for complex trait genetic mapping in humans, is designed to model the trait model complexity represented by the multidimensional parameter space in a mathematically rigorous fashion. However, the method requires the evaluation of integrals with no functional form, making it difficult to compute, and thus further test, develop and apply. This paper describes MLIP, a multiprocessor two-point genetic linkage analysis system that supports statistical calculations, such as the PPL, based on the full parameter space implicit in the linkage likelihood. RESULTS: The fundamental question we address here is whether the use of additional processors effectively reduces total computation time for a PPL calculation. We use a variety of data - both simulated and real - to explore the question "how close can we get?" to linear speedup. Empirical results of our study show that MLIP does significantly speed up two-point log-likelihood ratio calculations over a grid space of model parameters. CONCLUSION: Observed performance of the program is dependent on characteristics of the data including granularity of the parameter grid space being explored and pedigree size and structure. While work continues to further optimize performance, the current version of the program can already be used to efficiently compute the PPL. Thanks to MLIP, full multidimensional genome scans are now routinely being completed at our centers with runtimes on the order of days, not months or years.
Manika Govil, Alberto M. Segre, Veronica J. Vieland
BMC Bioinform.2
2007 Predicting the Impact of an Electronic Health Record on Practice Patterns Using Computational Modeling and Simulation
Thomas R. Clancy, Connie White-Delaney, Alberto M. Segre, Kathleen M. Carley, Andrew Kuziak, Hwanjo Yu
AMIA3
2007 Fast Computation of Human Genetic Linkage
abstract
Genetic linkage analysis is a recombinant technology used for mapping disease genes on the genome, based on genotypic and phenotypic data collected from families that have affected members. The LOD score is a commonly used statistic in genetic linkage analysis. LOD scores are computed assuming specific values for genetic parameters. However, for complex disorders the specified parameter values are often unknown. One way to address this issue is to maximize the LOD score over all genetic parameters to get a maximum LOD score, or MOD score. Another way is to integrate the LOD score across the genetic parameters to form a posterior probability of linkage, or PPL. Both methods require calculation of large numbers of LOD scores under different sets of parameter values. These calculations may be very time-consuming and can form a significant bottleneck in disease gene mapping. The motivation for this work is to speed up the computation of large numbers of LOD scores in linkage analysis. Instead of the usual LOD calculation where the likelihood of a pedigree under each set of parameter values is computed based on traversing the pedigree, the likelihood of the pedigree is computed here as an algebraic expression that can be optimized and reused. This optimized likelihood expression can be evaluated an arbitrary number of times for LOD scores under different values of the genetic parameters, resulting in much faster speeds. Our initial results show that this approach can speed up the traditional genetic linkage computation by 10~1200 times.
Hongling Wang, Alberto M. Segre, Yungui Huang, Jeffrey R. O'Connell, Veronica J. Vieland
BIBE2
2006 A High Throughput Approach to Combinatorial Search on Grids
abstract
Current distributed combinatorial search algorithms assume the use of managed or reserved resources. However, grid resources are shared and exhibit highly dynamic availability. Accommodating these resources in runtime collaboration for distributed search applications is a challenge. We work on nagging, a naturally scalable and fault-tolerant distributed search paradigm, and propose a high throughput collaboration approach, NoG (nagging on grid), that is continuously adaptive to dynamic resource availability. Dynamic scheduling and collaboration tree grafting algorithms are devised to handle dynamic join and leave of grid resources
Yan Liu 0009, Alberto M. Segre, Shaowen Wang 0001
HPDC2
2006 Privacy-Preserving Data Set Union
Alberto M. Segre, Andrew Wildenberg, Veronica J. Vieland
Privacy in Statistical Databases1
2002 Bayesian Classification of Respiratory Disease and Asthma with Administrative Data Sets
Kirk T. Phillips, Alberto M. Segre
AMIA2
2002 Nagging: A scalable fault-tolerant paradigm for distributed search
Alberto M. Segre, Sean L. Forman, Giovanni Resta, Andrew Wildenberg
Artif. Intell.1
2002 A Distributed Learning Algorithm for Bayesian Inference Networks
abstract
We present a new distributed algorithm for computing the minimum description length (MDL) in learning Bayesian inference networks from data. Our learning algorithm exploits both properties of the MDL-based score metric and a distributed, asynchronous, adaptive search technique called nagging. Nagging is intrinsically fault-tolerant, has dynamic load balancing features, and scales well. We demonstrate the viability, effectiveness, and scalability of our approach empirically with several experiments using networked machines. More specifically, we show that our distributed algorithm can provide optimal solutions for larger problems as well as good solutions for Bayesian networks of up to 150 variables.
Wai Lam, Alberto M. Segre
IEEE Trans. Knowl. Data Eng.2
2001 OAMulator: a teaching resource to introduce computer architecture concepts
abstract
The OAMulator is a Web-based resource to support the teaching of instruction set architecture, assembly languages, memory, addressing, high-level programming, and compilation. The tool is based on a simple, virtual CPU architecture, called the One Address Machine. A compiler allows us to take programs written in a special programming language, called OAMPL, and transform them into OAM assembly. An OAM assembler/emulator interprets and executes OAM assembly code produced by the compiler or written directly by students. The OAMulator is targeted at non-CS students who take introductory courses in information technology or information systems. Such students are normally exposed to concepts of computer hardware and software, but it is difficult for them to make the connection between the two. The OAMulator takes the mystery out of CPU architecture by letting students gain confidence with the concepts of compilers and binary execution. The Web-based deployment allows students to work on problems in convenient locations, at their own pace, and with rewarding interaction.
Filippo Menczer, Alberto M. Segre
ACM J. Educ. Resour. Comput.2
1997 Distributed Data Mining of Probabilistic Knowledge
abstract
We present a distributed approach to data mining of a knowledge representation scheme known as Bayesian belief networks which are capable of dealing with uncertain knowledge. We make use of a machine learning paradigm and a distributed asynchronous search technique to achieve the task of distributed knowledge discovery from data. Our approach boasts a number of features, including dynamic load balancing and fault tolerance. Empirical experiments have been conducted to illustrate its feasibility, solving large scale Bayesian network discovery problems with multiple workstations.
Wai Lam, Alberto M. Segre
ICDCS2
1997 Nagging: A Distributed, Adversarial Search-Pruning Technique Applied to First-Order Inference
David B. Sturgill, Alberto M. Segre
J. Autom. Reason.2
1996 Nonparametric Statistical Methods for Experimental Evaluations of Speedup Learning
Geoffrey J. Gordon, Alberto M. Segre
ICML2
1996 Exploratory Analysis of Speedup Learning Data Using Epectation Maximization
Alberto M. Segre, Geoffrey J. Gordon, Charles Elkan
Artif. Intell.1
1994 Using Hundreds of Workstations to Solve First-Order Logic Problems
Alberto M. Segre, David B. Sturgill
AAAI1
1994 A Novel Asynchronous Parallelism Scheme for First-Order Logic
David B. Sturgill, Alberto M. Segre
CADE2
1994 Getting the Most from Flawed Theories
Moshe Koppel, Alberto M. Segre, Ronen Feldman
ICML2
1994 A High-Performance Explanation-Based Learning Algorithm
Alberto M. Segre, Charles Elkan
Artif. Intell.1
1994 Bias-Driven Revision of Logical Domain Theories
abstract
The theory revision problem is the problem of how best to go about revising a deficient domain theory using information contained in examples that expose inaccuracies. In this paper we present our approach to the theory revision problem for propositional domain theories. The approach described here, called PTR, uses probabilities associated with domain theory elements to numerically track the ``flow'' of proof through the theory. This allows us to measure the precise role of a clause or literal in allowing or preventing a (desired or undesired) derivation for a given example. This information is used to efficiently locate and repair flawed elements of the theory. PTR is proved to converge to a theory which correctly classifies all examples, and shown experimentally to be fast and accurate even for deep theories.
Moshe Koppel, Ronen Feldman, Alberto M. Segre
J. Artif. Intell. Res.3
1993 Sholom M. Weiss and Casimir A. Kulikowski, Computer Systems That Learn
Alberto M. Segre, Geoffrey J. Gordon
Artif. Intell.1
1993 Bounded-Overhead Caching for Definite-Clause Theorem Proving
Alberto M. Segre, Daniel Scharstein
J. Autom. Reason.1
1992 On Combining Multiple Speedup Techniques
Alberto M. Segre
ML1
1991 Incremental Refinement of Approximate Domain Theories
Ronen Feldman, Alberto M. Segre, Moshe Koppel
ML2
1991 A Critical Look at Experimental Evaluations of EBL
Alberto M. Segre, Charles Elkan, Alexander Russell
Mach. Learn.1
1987 On the Operationality/Generality Trade-off in Explanation-based Learning
Alberto M. Segre
IJCAI1
1985 Explanation-based manipulator learning: Acquisition of planning ability through observation
abstract
This paper describes a robot manipulator system currently under development which learns from observation. The system improves its problem-solving capabilities through the acquisition of task-related concepts. The system observes manipulator command sequences that solve problems currently beyond its own panning abilities. General problem-solving schemata are automatically constructed via a knowledge-based analysis of how the observed command sequence achieved the goal. This learning technique is based on explanatory schema acquisition. It is a knowledge-based approach, requiring sufficient background knowledge to understand the observed sequence. The acquired schemata serve two purposes: they allow the system to solve problems that were previously unsolvable, and they aid in the understanding of later observations.
Alberto M. Segre, Gerald DeJong
ICRA1
1983 An Expert System For The Production Of Phoneme Strings From Unmarked English Text Using Machine-Induced Rules
Alberto M. Segre, Bruce Arne Sherwood, Wayne B. Dickerson
EACL1