VLDB 2026 Research / reviewers in the wild / expert
Luis M. Rocha
dblp:r/LuisMateusRocha · also Luis Mateus Rocha
· DBLP profile ↗
32ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-9402-887XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | myAURA : a personalized health library for epilepsy management via knowledge graph sparsification and visualizationabstractOBJECTIVES: Report the development of the patient-centered myAURA application and suite of methods designed to aid epilepsy patients, caregivers, and clinicians in making decisions about self-management and care. MATERIALS AND METHODS: myAURA rests on an unprecedented collection of epilepsy-relevant heterogeneous data resources, such as biomedical databases, social media, and electronic health records (EHRs). We use a patient-centered biomedical dictionary to link the collected data in a multilayer knowledge graph (KG) computed with a generalizable, open-source methodology. RESULTS: Our approach is based on a novel network sparsification method that uses the metric backbone of weighted graphs to discover important edges for inference, recommendation, and visualization. We demonstrate by studying drug-drug interaction from EHRs, extracting epilepsy-focused digital cohorts from social media, and generating a multilayer KG visualization. We also present our patient-centered design and pilot-testing of myAURA, including its user interface. DISCUSSION: The ability to search and explore myAURA's heterogeneous data sources in a single, sparsified, multilayer KG is highly useful for a range of epilepsy studies and stakeholder support. CONCLUSION: Our stakeholder-driven, scalable approach to integrating traditional and nontraditional data sources enables both clinical discovery and data-powered patient self-management in epilepsy and can be generalized to other chronic conditions. Rion Brattig Correia, Jordan C. Rozum, Leonard E. Cross, Jack Felag, Michael Gallant, Bruce W. Herr, Aehong Min, Jon Sanchez-Valle, Deborah Stungis Rocha, Alfonso Valencia, Katy Börner, Wendy Miller, Luis M. Rocha |
J. Am. Medical Informatics Assoc. | 15 |
| 2025 | CANA v1.0.0: efficient quantification of canalization in automata networksabstractSUMMARY: The biomolecular networks underpinning cell function exhibit canalization, or the buffering of fluctuations required to function in a noisy environment. We present a new major release of CANA, v1.0.0, an open-source Python package for understanding canalization in automata network models, discrete dynamical systems in which activation of biomolecular entities (e.g. transcription of genes) is modeled as the activity of coupled automata. One understudied putative mechanism for canalization is the functional equivalence of biomolecular regulators (e.g. among the transcription factors for a gene). We study this mechanism using the theory of symmetry in discrete functions. We present a new exact method, schematodes, for finding maximal symmetry groups among the inputs to discrete functions, and integrate it into CANA. The schematodes method substantially outperforms the inexact method of previous CANA versions both in speed and accuracy. We apply CANA v1.0.0 to study symmetry in 74 experimentally supported automata network models from the Cell Collective (CC) repository. The symmetry distribution is significantly different in the CC than in random automata with the same in-degree (connectivity) and bias (average output) (Kolmogorov-Smirnov test, P ≪ .001). Its spread is much wider than in a null model (IQR 0.31 versus IQR 0.20 with equal medians), demonstrating that the CC is enriched in functions with extreme symmetry or asymmetry. AVAILABILITY AND IMPLEMENTATION: CANA source is on https://github.com/CASCI-lab/CANA and is installable via pip install cana. Source for schematodes is on https://github.com/CASCI-lab/schematodes. Analysis scripts are on https://github.com/CASCI-lab/symmetryInCellCollective. Austin M. Marcus, Jordan C. Rozum, Herbert Sizek, Luis M. Rocha |
Bioinform. | 4 |
| 2025 | Focused digital cohort selection from social media using the metric backbone of biomedical knowledge graphsabstractSocial media data allows researchers to construct large digital cohorts - groups of users who post health-related content - to study the interplay between human behavior and medical treatment. Identifying the users most relevant to a specific health problem is, however, a challenge in that social media sites vary in the generality of their discourse. While X (formerly Twitter), Instagram, and Facebook cater to wide ranging topics, Reddit subgroups and dedicated patient advocacy forums trade in much more specific, biomedically-relevant discourse. To filter relevant users on any social media, we have developed a general method and tested it on epilepsy discourse. We analyzed the text from posts by users who mention epilepsy drugs at least once in the general-purpose social media sites X and Instagram, the epilepsy-focused Reddit subgroup (r/Epilepsy), and the Epilepsy Foundation of America (EFA) forums. We used a curated medical terminology dictionary to generate a knowledge graph (KG) from each social media site, whereby nodes represent terms, and edge weights denote the strength of association between pairs of terms in the collected text. Our method is based on computing the metric backbone of each KG, which yields the (sparsified) subgraph of edges that participate in shortest paths. By comparing the subset of users who contribute to the backbone to the subset who do not, we show that epilepsy-focused social media users contribute to the KG backbone in much higher proportion than do general-purpose social media users. Furthermore, using human annotation of Instagram posts, we demonstrate that users who do not contribute to the backbone are much more likely to use dictionary terms in a manner inconsistent with their biomedical meaning and are rightly excluded from the cohort of interest. Our metric backbone approach, thus, has several benefits: it yields focused user cohorts who engage in discourse relevant to a targeted biomedical problem; unlike engagement-based approaches, it can retain low-engagement users who nonetheless contribute meaningful biomedical insights and filter out very vocal users who contribute no relevant content, it is parameter-free, algebraically principled, does not require classifiers or human-curation, and is simple to compute with the open-source code we provide. Jack Felag, Jordan C. Rozum, Rion Brattig Correia, Luis M. Rocha |
J. Biomed. Informatics | 6 |
| 2023 | Understanding Contexts and Challenges of Information Management for Epilepsy CareabstractEpilepsy is a common chronic neurological disease. People with epilepsy (PWE) and their caregivers face several challenges related to their epilepsy management, including quality of care, care coordination, side effects, and stigma management. The sociotechnical issues of the information management contexts and challenges for epilepsy care may be mitigated through effective information management. We conducted 4 focus groups with 5 PWE and 7 caregivers to explore how they manage epilepsy-related information and the challenges they encountered. Primary issues include challenges of finding the right information, complexities of tracking and monitoring data, and limited information sharing. We provide a framework that encompasses three attributes - individual epilepsy symptoms and health conditions, information complexity, and circumstantial constraints. We suggest future design implications to mitigate these challenges and improve epilepsy information management and care coordination. Aehong Min, Wendy Miller, Luis M. Rocha, Katy Börner, Rion Brattig Correia, Patrick C. Shih |
CHI | 3 |
| 2023 | Contact networks have small metric backbones that maintain community structure and are primary transmission subgraphsabstractThe structure of social networks strongly affects how different phenomena spread in human society, from the transmission of information to the propagation of contagious diseases. It is well-known that heterogeneous connectivity strongly favors spread, but a precise characterization of the redundancy present in social networks and its effect on the robustness of transmission is still lacking. This gap is addressed by the metric backbone, a weight- and connectivity-preserving subgraph that is sufficient to compute all shortest paths of weighted graphs. This subgraph is obtained via algebraically-principled axioms and does not require statistical sampling based on null-models. We show that the metric backbones of nine contact networks obtained from proximity sensors in a variety of social contexts are generally very small, 49% of the original graph for one and ranging from about 6% to 20% for the others. This reflects a surprising amount of redundancy and reveals that shortest paths on these networks are very robust to random attacks and failures. We also show that the metric backbone preserves the full distribution of shortest paths of the original contact networks-which must include the shortest inter- and intra-community distances that define any community structure-and is a primary subgraph for epidemic transmission based on pure diffusion processes. This suggests that the organization of social contact networks is based on large amounts of shortest-path redundancy which shapes epidemic spread in human populations. Thus, the metric backbone is an important subgraph with regard to epidemic spread, the robustness of social networks, and any communication dynamics that depend on complex network shortest paths. Rion Brattig Correia, Alain Barrat, Luis M. Rocha |
PLoS Comput. Biol. | 3 |
| 2022 | On the feasibility of dynamical analysis of network models of biochemical regulationabstractTo the Editor, A recent article by Weidner et al. (2021) presents a method to extract graph properties that are predictive of the dynamical behavior of multivariate, discrete models of biochemical regulation. In other words, a method that uses only features from the structure of network interactions to predict which nodes are most involved in automata network dynamics. However, the authors claim that dynamical analysis of large automata network models is ‘not even feasible’. To make sure that others are not discouraged from working on this problem, it is important to clarify that effective dynamical analysis of automata network models, to the contrary, is feasible. Unlike what is suggested in the article, graph-based analysis of static features is not the only analytical avenue for large systems biology models of regulation and signaling dynamics because there are dynamical methods that are, indeed, scalable. By scalable we mean that the computational complexity of methods employed to analyze multivariate dynamical systems in regard to their dynamical behavior (e.g. controllability, convergence to attractors, robustness to perturbations, etc.) is manageable. That is, a given method is scalable if results can be computed in finite (and reasonable) time, with finite memory, for a system of a reasonable size—ideally for network models in Systems Biology, up to thousands of nodes. There has been much interest recently in predicting multivariate dynamics from static network structure alone, especially in regard to the controllability of systems biology models of gene regulation, signaling and cellular differentiation (Fiedler et al., 2013; Liu et al., 2011; Nacher and Akutsu, 2013; Zanudo et al., 2017). These are quite welcome methods because we often lack information about the underlying causal interaction dynamics. However, two very popular (and scalable) methods in network science lead to very erroneous predictions of the subsets of (driver) variables that control dynamics (Gates and Rocha, 2016). The most accurate of these methods are based on feedback vertex set theory (Fiedler et al., 2013; Zanudo et al., 2017), which does not scale well and can only make predictions about the entire ensemble of dynamical systems that fits the same static interaction graph (Gates et al., 2021). Another putative reason for pursuing structure-only methods is that even when the underlying interaction dynamics of each variable is known, it is not feasible to compute the dynamical (or attractor) landscape of the entire multivariate system when the interaction network is sufficiently large. This prevents us from exhaustively enumerating all possible interventions that can control the dynamics from one attractor basin to another. The most important scalability constraint in the analysis of automata networks is the number of node variables, n, as the dynamical landscape of such systems is comprised of sn possible configurations of states s (Gates and Rocha, 2016). For instance, a well-known Boolean network model of intracellular signaling networks in generic fibroblasts is comprised of 130 nodes (Helikar et al., 2008), thus it can be in one of 2130 possible state configurations—a dynamical landscape that is too large to be exhaustively searched. Importantly, it is also true that enumeration of all possible interventions is infeasible in a simple graph of sufficient size. The computational complexity of finding all possible subsets of the set of nodes, or generating the powerset, is at least o(2n) (Moore, 1971). Therefore, exhaustive search of all possible interventions in any network is ultimately infeasible for large graphs whether one uses structure- or dynamics-based methods. Computing the full dynamics of a multivariate system certainly adds to the complexity of exhaustively searching all possible interventions. But many software tools exist—and are collected in repositories such as the CoLoMoTo Consortium (Naldi et al., 2015)—that allow for the identification of attractors in automata networks without full enumeration of their dynamical landscapes, such as PyBoolNet (Klarner et al., 2017) and Boolink (Karanam et al., 2021). In particular, recent developments in attractor identification algorithms and code (PyStableMotifs) now allow the dynamical analysis of thousands of networks, some with over 15 000 nodes (Rozum et al., 2021). From another angle, a novel update scheme for asynchronous automata networks has been shown to reduce the complexity of attractor identification, enabling the modeling and dynamical analysis of genome-scale networks (Paulevé et al., 2020). Moreover, various computational approaches have been developed and applied to extract the key drivers of collective dynamics of biochemical network models without going through every possible subset of nodes, much less the entire dynamical landscape, in a brute force manner (Biane and Delaplace, 2019; Hari et al., 2021; Rozum et al., 2021; Su and Pang, 2020; Zañudo and Albert, 2015). Indeed, scalable methods exist that remove the redundancy of the dynamics of each variable (micro-level) to allow for a characterization of the entire causal macro-level dynamics, in both complete (Marques-Pita and Rocha, 2013) and probabilistic (Gates et al., 2021) manners. These scalable methods, and the associated software tool CANA (Correia et al., 2018), provide causal graph representations of automata networks that synthesize both structure and dynamics. They are exhaustive in the sense that they preserve all effective interactions of the micro-level dynamics—only redundant interactions are disregarded. Thus, it is very feasible to analyze systems biology models without disregarding their dynamics, allowing the precise study of any putative intervention that controls the dynamics just as easily as structure-only methods, but with additional accuracy afforded by information about the dynamics. The author is indebted to the anonymous reviewers for their most valuable comments and references which have considerably strengthened this letter. The work was partially funded by the National Institutes of Health, National Library of Medicine Program, grant no. 01LM011945-01, and by NSF-NRT grant no. 1735095 ‘Interdisciplinary Training in Complex Networks and Systems.’ The author thanks Deborah Rocha for thorough line editing. Conflict of Interest: none declared. Luis M. Rocha |
Bioinform. | 1 |
| 2021 | Just In Time: Challenges and Opportunities of First Aid Care Information Sharing for Supporting Epileptic Seizure ResponseabstractThere are over three million people living with epilepsy in the U.S. People with epilepsy experience multiple daily challenges such as seizures, social isolation, social stigma, experience of physical and emotional symptoms, medication side effects, cognitive and memory deficits, care coordination difficulties, and risks of sudden unexpected death. In this work, we report findings collected from 3 focus groups of 11 people with epilepsy and caregivers and 10 follow-up questionnaires. We found that these participants feel that most people do not know how to deal with seizures. To improve others' abilities to respond safely and appropriately to someone having seizures, people with epilepsy and caregivers would like to share and educate the public about their epilepsy conditions, reduce common misconceptions about seizures and prevent associated stigma, and get first aid help from the public when needed. Considering social stigma, we propose design implications of future technologies for effective delivery of appropriate first aid care information to bystanders around individuals with epilepsy when they experience a seizure. Aehong Min, Wendy Miller, Luis M. Rocha, Katy Börner, Rion Brattig Correia, Patrick C. Shih |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2014 | Designing a Minimalist Socially Aware Robotic Agent for the HomeabstractWe present a minimalist social robot that relies on long timeseries of low resolution data such as mechanical vibration, temperature, lighting, sounds and collisions.Our goal is to develop an experimental system for growing socially situated robotic agents whose behavioral repertoire is subsumed by the social order of the space.To get there we are designing robots that use their simple sensors and motion feedback routines to recognize different classes of human activity and then associate to each class a range of appropriate behaviors.We use the Katie Family of robots, built on the iRobot Create platform, an Arduino Uno, and a Raspberry Pi.We describe its sensor abilities and exploratory tests that allow us to develop hypotheses about what objects (sensor data) correspond to something known and observable by a human subject.We use machine learning methods to classify three social scenarios from over a hundred experiments, demonstrating that it is possible to detect social situations with high accuracy, using the low-resolution sensors from our minimalist robot. Matthew R. Francisco, Ian B. Wood, Selma Sabanovic, Luis M. Rocha |
ALIFE | 4 |
| 2014 | Structure and Dynamics Affect the Controlability of Complex Systems: A Preliminary StudyabstractComplex systems are typically understood as large nonlinear multivariate systems. Their organization and behavior are commonly modeled by representations such as graphs and automata networks. Graphs, where nodes representing variables lack intrinsic dynamics, capture the structure or organization of complex systems. The simplest way to study multi-variate dynamics, is to allow network nodes to have states and update them with automata; for instance, Boolean networks (BN) are canonical models of complex systems and exhibit a wide range of dynamical behaviors [1]. Alexander J. Gates, Luis M. Rocha |
ALIFE | 2 |
| 2013 | An integrated pharmacokinetics ontology and corpus for text miningabstractBACKGROUND: Drug pharmacokinetics parameters, drug interaction parameters, and pharmacogenetics data have been unevenly collected in different databases and published extensively in the literature. Without appropriate pharmacokinetics ontology and a well annotated pharmacokinetics corpus, it will be difficult to develop text mining tools for pharmacokinetics data collection from the literature and pharmacokinetics data integration from multiple databases. DESCRIPTION: A comprehensive pharmacokinetics ontology was constructed. It can annotate all aspects of in vitro pharmacokinetics experiments and in vivo pharmacokinetics studies. It covers all drug metabolism and transportation enzymes. Using our pharmacokinetics ontology, a PK-corpus was constructed to present four classes of pharmacokinetics abstracts: in vivo pharmacokinetics studies, in vivo pharmacogenetic studies, in vivo drug interaction studies, and in vitro drug interaction studies. A novel hierarchical three level annotation scheme was proposed and implemented to tag key terms, drug interaction sentences, and drug interaction pairs. The utility of the pharmacokinetics ontology was demonstrated by annotating three pharmacokinetics studies; and the utility of the PK-corpus was demonstrated by a drug interaction extraction text mining analysis. CONCLUSIONS: The pharmacokinetics ontology annotates both in vitro pharmacokinetics experiments and in vivo pharmacokinetics studies. The PK-corpus is a highly valuable resource for the text mining of pharmacokinetics parameters and drug interactions. Heng-Yi Wu, Shreyas D. Karnik, Abhinita Subhadarshini, Santosh Philips, Chienwei Chiang, Malaz Boustani, Luis M. Rocha, Sara K. Quinney, David A. Flockhart, Lang Li 0001 |
BMC Bioinform. | 10 |
| 2012 | Correction: A linear classifier based on entity recognition tools and a statistical approach to method extraction in the protein-protein interaction literatureabstractAbstract Correction to A. Lourenço, M. Conover, A. Wong, A. Nematzadeh, F. Pan, H. Shatkay, and L.M. Rocha."A Linear Classifier Based on Entity Recognition Tools and a Statistical Approach to Method Extraction in the Protein-Protein Interaction Literature". BMC Bioinformatics 2011, 12(Suppl 8):S12. doi: http://10.1186/1471-2105-12-S8-S12 . Anália Lourenço, Michael D. Conover, Azadeh Nematzadeh, Fengxia Pan, Hagit Shatkay, Luis M. Rocha |
BMC Bioinform. | 7 |
| 2011 | Schema redescription in cellular automata: Revisiting emergence in complex systemsabstractWe present a method to eliminate redundancy in the transition tables of Boolean automata: schema redescription with two symbols. One symbol is used to capture redundancy of individual input variables, and another to capture permutability in sets of input variables: fully characterizing the canalization present in Boolean functions. Two-symbol schemata explain aspects of the behaviour of automata networks that the characterization of their emergent patterns does not capture. We use our method to compare two well-known cellular automata rules for the density classification task: GKL and GP. We show that despite having very different emergent behaviour, these rules are very similar. Indeed, GKL is a special case of GP. Therefore, we demonstrate that it is more feasible to compare cellular automata via schema redescriptions of their rules, than by looking at their emergent behaviour, leading us to question the tendency in complexity research to pay much more attention to emergent patterns than to local (micro-level) interactions. Manuel Marques-Pita, Luis M. Rocha |
ALIFE | 2 |
| 2011 | The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full textabstractBACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows. Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu |
BMC Bioinform. | 21 |
| 2011 | A linear classifier based on entity recognition tools and a statistical approach to method extraction in the protein-protein interaction literatureabstractBACKGROUND: We participated, as Team 81, in the Article Classification and the Interaction Method subtasks (ACT and IMT, respectively) of the Protein-Protein Interaction task of the BioCreative III Challenge. For the ACT, we pursued an extensive testing of available Named Entity Recognition and dictionary tools, and used the most promising ones to extend our Variable Trigonometric Threshold linear classifier. Our main goal was to exploit the power of available named entity recognition and dictionary tools to aid in the classification of documents relevant to Protein-Protein Interaction (PPI). For the IMT, we focused on obtaining evidence in support of the interaction methods used, rather than on tagging the document with the method identifiers. We experimented with a primarily statistical approach, as opposed to employing a deeper natural language processing strategy. In a nutshell, we exploited classifiers, simple pattern matching for potential PPI methods within sentences, and ranking of candidate matches using statistical considerations. Finally, we also studied the benefits of integrating the method extraction approach that we have used for the IMT into the ACT pipeline. RESULTS: For the ACT, our linear article classifier leads to a ranking and classification performance significantly higher than all the reported submissions to the challenge in terms of Area Under the Interpolated Precision and Recall Curve, Mathew's Correlation Coefficient, and F-Score. We observe that the most useful Named Entity Recognition and Dictionary tools for classification of articles relevant to protein-protein interaction are: ABNER, NLPROT, OSCAR 3 and the PSI-MI ontology. For the IMT, our results are comparable to those of other systems, which took very different approaches. While the performance is not very high, we focus on providing evidence for potential interaction detection methods. A significant majority of the evidence sentences, as evaluated by independent annotators, are relevant to PPI detection methods. CONCLUSIONS: For the ACT, we show that the use of named entity recognition tools leads to a substantial improvement in the ranking and classification of articles relevant to protein-protein interaction. Thus, we show that our substantially expanded linear classifier is a very competitive classifier in this domain. Moreover, this classifier produces interpretable surfaces that can be understood as "rules" for human understanding of the classification. We also provide evidence supporting certain named entity recognition tools as beneficial for protein-interaction article classification, or demonstrating that some of the tools are not beneficial for the task. In terms of the IMT task, in contrast to other participants, our approach focused on identifying sentences that are likely to bear evidence for the application of a PPI detection method, rather than on classifying a document as relevant to a method. As BioCreative III did not perform an evaluation of the evidence provided by the system, we have conducted a separate assessment, where multiple independent annotators manually evaluated the evidence produced by one of our runs. Preliminary results from this experiment are reported here and suggest that the majority of the evaluators agree that our tool is indeed effective in detecting relevant evidence for PPI detection methods. Regarding the integration of both tasks, we note that the time required for running each pipeline is realistic within a curation effort, and that we can, without compromising the quality of the output, reduce the time necessary to extract entities from text for the ACT pipeline by pre-selecting candidate relevant text using the IMT pipeline. Anália Lourenço, Michael D. Conover, Azadeh Nematzadeh, Fengxia Pan, Hagit Shatkay, Luis M. Rocha |
BMC Bioinform. | 7 |
| 2010 | Collective Classification of Biomedical Articles using T-Cell Cross-regulation
Alaa Abi-Haidar, Luis M. Rocha |
ALIFE | 2 |
| 2010 | BioDR: Semantic indexing networks for biomedical document retrievalabstractIn Biomedical research, retrieving documents that match an interesting query is a task performed quite frequently. Typically, the set of obtained results is extensive containing many non-interesting documents and consists in a flat list, i.e., not organized or indexed in any way. This work proposes BioDR, a novel approach that allows the semantic indexing of the results of a query, by identifying relevant terms in the documents. These terms emerge from a process of Named Entity Recognition that annotates occurrences of biological terms (e.g. genes or proteins) in abstracts or full-texts. The system is based on a learning process that builds an Enhanced Instance Retrieval Network (EIRN) from a set of manually classified documents, regarding their relevance to a given problem. The resulting EIRN implements the semantic indexing of documents and terms, allowing for enhanced navigation and visualization tools, as well as the assessment of relevance for new documents. Anália Lourenço, Rafael Carreira, Daniel Glez-Peña, José Ramón Méndez 0001, Sónia Carneiro, Luis M. Rocha, Fernando Díaz 0001, Eugénio C. Ferreira, Isabel Rocha, Florentino Fernández Riverola, Miguel Rocha 0001 |
Expert Syst. Appl. | 6 |
| 2010 | Classification of Protein-Protein Interaction Full-Text Documents Using Text and Citation Network FeaturesabstractWe participated (as Team 9) in the Article Classification Task of the Biocreative II.5 Challenge: binary classification of full-text documents relevant for protein-protein interaction. We used two distinct classifiers for the online and offline challenges: 1) the lightweight Variable Trigonometric Threshold (VTT) linear classifier we successfully introduced in BioCreative 2 for binary classification of abstracts and 2) a novel Naive Bayes classifier using features from the citation network of the relevant literature. We supplemented the supplied training data with full-text documents from the MIPS database. The lightweight VTT classifier was very competitive in this new full-text scenario: it was a top-performing submission in this task, taking into account the rank product of the Area Under the interpolated precision and recall Curve, Accuracy, Balanced F-Score, and Matthew's Correlation Coefficient performance measures. The novel citation network classifier for the biomedical text mining domain, while not a top performing classifier in the challenge, performed above the central tendency of all submissions, and therefore indicates a promising new avenue to investigate further in bibliome informatics. Artemy Kolchinsky, Alaa Abi-Haidar, Jasleen Kaur 0002, Ahmed Abdeen Hamed, Luis M. Rocha |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2009 | Literature mining on pharmacokinetics numerical data: A feasibility study
Sara K. Quinney, Stephen D. Hall, Luis M. Rocha, Lang Li 0001 |
J. Biomed. Informatics | 6 |
| 2008 | Adaptive Spam Detection Inspired by the Immune System
Alaa Abi-Haidar, Luis M. Rocha |
ALIFE | 2 |
| 2008 | Conceptual Structure in Cellular Automata - The Density Classification Task
Manuel Marques-Pita, Luis M. Rocha |
ALIFE | 2 |
| 2008 | The Role of Conceptual Structure in Designing Cellular Automata to Perform Collective Computation
Manuel Marques-Pita, Melanie Mitchell, Luis M. Rocha |
UC | 3 |
| 2007 | Agent-Based Model of Genotype EditingabstractEvolutionary algorithms rarely deal with ontogenetic, non-inherited alteration of genetic information because they are based on a direct genotype-phenotype mapping. In contrast, several processes have been discovered in nature which alter genetic information encoded in DNA before it is translated into amino-acid chains. Ontogenetically altered genetic information is not inherited but extensively used in regulation and development of phenotypes, giving organisms the ability to, in a sense, re-program their genotypes according to environmental cues. An example of post-transcriptional alteration of gene-encoding sequences is the process of RNA Editing. Here we introduce a novel Agent-based model of genotype editing and a computational study of its evolutionary performance in static and dynamic environments. This model builds on our previous Genetic Algorithm with Editing, but presents a fundamentally novel architecture in which coding and non-coding genetic components are allowed to co-evolve. Our goals are: (1) to study the role of RNA Editing regulation in the evolutionary process, (2) to understand how genotype editing leads to a different, and novel evolutionary search algorithm, and (3) the conditions under which genotype editing improves the optimization performance of traditional evolutionary algorithms. We show that genotype editing allows evolving agents to perform better in several classes of fitness functions, both in static and dynamic environments. We also present evidence that the indirect genotype/phenotype mapping resulting from genotype editing leads to a better exploration/exploitation compromise of the search process. Therefore, we show that our biologically-inspired model of genotype editing can be used to both facilitate understanding of the evolutionary role of RNA regulation based on genotype editing in biology, and advance the current state of research in Evolutionary Computation. Chien-Feng Huang, Jasleen Kaur 0002, Ana Gabriela Maguitman, Luis M. Rocha |
Evol. Comput. | 4 |
| 2005 | Tracking extrema in dynamic environments using a coevolutionary agent-based model of genotype editionabstractTypical applications of evolutionary optimization in static environments involve the approximation of the extrema of functions. For dynamic environments, the interest is not to locate the extrema but to follow it as closely as possible. This paper compares the extrema-tracking performance of a traditional Genetic Algorithm and a coevolutionary agent-based model of Genotype Editing (ABMGE). This model is constructed using several genetic editing characteristics that are gleaned from the RNA editing system as observed in several organisms. The incorporation of editing mechanisms provides a means for artificial agents with genetic descriptions to gain greater phenotypic plasticity. By allowing the family of editors and the genotypes of agents to co-evolve using the re-generation of editors as a control switch for environmental changes, the artificial agents in ABMGE can discover proper editors to facilitate the tracking of the extrema in dynamic environments. We will show that this agent-based model, together with a coevolutionary mechanism, is more adaptive and robust than the GA. We expect the framework proposed in this paper to advance the current state of research of Evolutionary Computation in dynamic environments. Chien-Feng Huang, Luis M. Rocha |
GECCO | 2 |
| 2005 | MyLibrary@LANL: Proximity and Semi-metric Networks for a Collaborative and Recommender Web ServiceabstractWe describe a network approach to building recommendation systems for a Web service. We employ two different types of weighted graphs in our analysis and development: proximity graphs, a type of fuzzy graphs based on a co-occurrence probability, and semi-metric distance graphs, which do not observe the triangle inequality of Euclidean distances. Both types of graphs are used to develop intelligent recommendation and collaboration systems for the MyLibrary@LANL Web service, a user-centered front-end to the Los Alamos National Laboratory's digital library collections and Web resources. Luis M. Rocha, Tiago Simas, Andreas Rechtsteiner, Mariella Di Giacomo, Richard Luce |
Web Intelligence | 1 |
| 2005 | Introduction to the Special Issue: Embodied and Situated CognitionabstractJanuary 01 2005 Introduction to the Special Issue: Embodied and Situated Cognition In Special Collection: CogNet Fernando Almeida e Costa, Fernando Almeida e Costa Centre for Computational, Neuroscience and Robotics, Evolutionary and Adaptive, Systems Group, Cogs/Informatics—School of Science and Technology, University of Sussex, Brighton, BN1 9QH, U.K. [email protected] Search for other works by this author on: This Site Google Scholar Luis Mateus Rocha Luis Mateus Rocha School of Informatics and Cognitive Science Program, Indiana University, 1900 East Tenth Street, Bloomington, IN 47406 [email protected] Search for other works by this author on: This Site Google Scholar Author and Article Information Fernando Almeida e Costa Centre for Computational, Neuroscience and Robotics, Evolutionary and Adaptive, Systems Group, Cogs/Informatics—School of Science and Technology, University of Sussex, Brighton, BN1 9QH, U.K. [email protected] Luis Mateus Rocha School of Informatics and Cognitive Science Program, Indiana University, 1900 East Tenth Street, Bloomington, IN 47406 [email protected] Online Issn: 1530-9185 Print Issn: 1064-5462 © 2005 Massachusetts Institute of Technology2005 Artificial Life (2005) 11 (1-2): 5–11. https://doi.org/10.1162/1064546053279035 Cite Icon Cite Permissions Share Icon Share Facebook Twitter LinkedIn MailTo Views Icon Views Article contents Figures & tables Video Audio Supplementary Data Peer Review Search Site Citation Fernando Almeida e Costa, Luis Mateus Rocha; Introduction to the Special Issue: Embodied and Situated Cognition. Artif Life 2005; 11 (1-2): 5–11. doi: https://doi.org/10.1162/1064546053279035 Download citation file: Ris (Zotero) Reference Manager EasyBib Bookends Mendeley Papers EndNote RefWorks BibTex toolbar search Search Dropdown Menu toolbar search search input Search input auto suggest filter your search All ContentAll JournalsArtificial Life Search Advanced Search This content is only available as a PDF. © 2005 Massachusetts Institute of Technology2005 Article PDF first page preview Close Modal You do not currently have access to this content. Fernando Almeida e Costa, Luis M. Rocha |
Artif. Life | 2 |
| 2005 | Material Representations: From the Genetic Code to the Evolution of Cellular AutomataabstractWe present a new definition of the concept of representation for cognitive science that is based on a study of the origin of structures that are used to store memory in evolving systems. This study consists of novel computer experiments in the evolution of cellular automata to perform nontrivial tasks as well as evidence from biology concerning genetic memory. Our key observation is that representations require inert structures to encode information used to construct appropriate dynamic configurations for the evolving system. We propose criteria to decide if a given structure is a representation by unpacking the idea of inert structures that can be used as memory for arbitrary dynamic configurations. Using a genetic algorithm, we evolved cellular automata rules that can perform nontrivial tasks related to the density task (or majority classification problem) commonly used in the literature. We present the particle catalogs of the new rules following the computational mechanics framework. We discuss if the evolved cellular automata particles may be seen as representations according to our criteria. We show that while they capture some of the essential characteristics of representations, they lack an essential one. Our goal is to show that artificial life can be used to shed new light on the computation-versus-dynamics debate in cognitive science, and indeed function as a constructive bridge between the two camps. Our definitions of representation and cellular automata experiments are proposed as a complementary approach, with both dynamics and informational modes of explanation. Luis M. Rocha, Wim Hordijk |
Artif. Life | 1 |
| 2005 | Protein annotation as term categorization in the gene ontology using word proximity networksabstractBACKGROUND: We participated in the BioCreAtIvE Task 2, which addressed the annotation of proteins into the Gene Ontology (GO) based on the text of a given document and the selection of evidence text from the document justifying that annotation. We approached the task utilizing several combinations of two distinct methods: an unsupervised algorithm for expanding words associated with GO nodes, and an annotation methodology which treats annotation as categorization of terms from a protein's document neighborhood into the GO. RESULTS: The evaluation results indicate that the method for expanding words associated with GO nodes is quite powerful; we were able to successfully select appropriate evidence text for a given annotation in 38% of Task 2.1 queries by building on this method. The term categorization methodology achieved a precision of 16% for annotation within the correct extended family in Task 2.2, though we show through subsequent analysis that this can be improved with a different parameter setting. Our architecture proved not to be very successful on the evidence text component of the task, in the configuration used to generate the submitted results. CONCLUSION: The initial results show promise for both of the methods we explored, and we are planning to integrate the methods more closely to achieve better results overall. Karin Verspoor, Judith D. Cohn, Cliff A. Joslyn, Susan M. Mniszewski, Andreas Rechtsteiner, Luis M. Rocha, Tiago Simas |
BMC Bioinform. | 6 |
| 2004 | A Systematic Study of Genetic Algorithms with Genotype Editing
Chien-Feng Huang, Luis M. Rocha |
GECCO (1) | 2 |
| 2004 | Study and analysis of workspace awareness in CDebate: a groupware application for collaborative debatesabstractAbstract In this paper, we study the workspace awareness in a groupware application allowing the development of an information task through collaborative debates. The application, called CDebate, is based on the APRI (Action–Perception–Reflection–Intention) model, which establishes a cognitive and motor states organization that occurs when humans are interacting with one another in a constructivist and collaborative learning situation. In CDebate, the interactions among students occur through a graphical language that reflects the mental operations appropriate for a debate. As an evaluation method, a conceptual framework, which provides a set of elements that give information about the up-to-the-moment knowledge about participants' location and actions, is used. The results of this study allow us to confirm that group awareness information, supported through a graphical language and a window showing the participants' presence (informal awareness), were sufficient for success in the collaborative learning situation. This experience could be useful for interface designers of groupware applications, in particular for collaborative debate interfaces. Manuel Romero Salcedo, César A. Osuna-Gómez, Leonid Sheremetov, Luis A. Villa-Vargas, Carlos Morales, Luis M. Rocha, Manuel Chi |
Interact. Comput. | 6 |
| 2003 | Exploration of RNA editing and design of robust genetic algorithmsabstractThis paper presents our computational methodology using genetic algorithms (GA) for exploring the nature of RNA editing. These models are constructed using several genetic editing characteristics that are gleaned from the RNA editing system as observed in several organisms. We have expanded the traditional genetic algorithm with artificial editing mechanisms as proposed by (Rocha, 1997). The incorporation of editing mechanisms provides a means for artificial agents with genetic descriptions to gain greater phenotypic plasticity, which is environmentally regulated. Our first implementations of these ideas have shed some light into the evolutionary implications of RNA editing. Based on these understandings, we demonstrate how to select proper RNA editors for designing more robust GAs, and the results show promising applications to real-world problems. We expect that the framework proposed both facilitate determining the evolutionary role of RNA editing in biology, and advance the current state of research in genetic algorithms. Chien-Feng Huang, Luis M. Rocha |
IEEE Congress on Evolutionary Computation | 2 |
| 2002 | Combination of evidence in recommendation systems characterized by distance functionsabstractRecommendation systems for different document networks (DN), such as the World Wide Web, digital libraries, or scientific databases, often make use of distance functions extracted from relationships among documents and between documents and semantic tags. The distance functions computed from these relations establish associative networks among items of the DN, and allow recommendation systems to identify relevant associations for individual users. The process of recommendation can be improved by integrating associative data from different sources. Thus, we are presented with a problem of combining evidence (about associations between items) from different sources characterized by distance functions. In this paper we summarize our work on: (1) inferring associations from semi-metric distance functions; and (2) combining evidence from different (distance) associative DN. Luis M. Rocha |
FUZZ-IEEE | 1 |
| 1998 | Towards a Formal Taxonomy of Hybrid Uncertainty Representations
Cliff A. Joslyn, Luis M. Rocha |
Inf. Sci. | 2 |