Elon Portugaly

dblp:69/1214 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 96% Computational science and engineering · 4%
Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 36% Knowledge representation and reasoning · 32% Reinforcement learning · 32%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
immunoinformatics
0.912025
Scalable Universal T-Cell Receptor Embeddings from Adaptive Immune Repertoires · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.212013
Counterfactual reasoning and learning systems: the example of computational advertising · J. Mach. Learn. Res. 2013
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.212013
Counterfactual reasoning and learning systems: the example of computational advertising · J. Mach. Learn. Res. 2013
Machine learning › Reinforcement learning
off-policy evaluation
0.212013
Counterfactual reasoning and learning systems: the example of computational advertising · J. Mach. Learn. Res. 2013
Bioinformatics and computational biology
hierarchical clustering
0.112008
Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space · ISMB 2008
Bioinformatics and computational biology › protein sequence analysis › protein family analysis
protein family clustering
0.112008
Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space · ISMB 2008
Bioinformatics and computational biology
protein sequence analysis
0.112008
Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space · ISMB 2008
Algorithms and data structures › clustering › hierarchical clustering
agglomerative clustering
0.112008
Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space · ISMB 2008
Algorithms and data structures
clustering
0.112008
Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space · ISMB 2008
Computational science and engineering
motor control
0.112005
Noise and the two-thirds power Law · NIPS 2005
Bioinformatics and computational biology › genomics
structural genomics
0.012002
Selecting targets for structural determination by navigating in a graph of protein families · Bioinform. 2002
Bioinformatics and computational biology
protein structure prediction
0.012000
Probabilities for having a new fold on the basis of a map of all protein sequences · RECOMB 2000
Machine learning › Probabilistic and Bayesian machine learning
noise modeling
0.012005
Noise and the two-thirds power Law · NIPS 2005
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein family
0.012002
Selecting targets for structural determination by navigating in a graph of protein families · Bioinform. 2002

Methods — techniques the papers use, named apart from their topics

random projection · 1.7glove · 1.7counterfactual reasoning · 0.2memory-constrained clustering · 0.2UPGMA · 0.2signal analysis · 0.1gaussian noise modeling · 0.1graph navigation · 0.0PSI-BLAST · 0.0probabilistic modeling · 0.0
YearPublicationVenuePosition
2025 Scalable Universal T-Cell Receptor Embeddings from Adaptive Immune Repertoires
abstract
T cells are a key component of the adaptive immune system, targeting infections, cancers, and allergens with specificity encoded by their T cell receptors (TCRs), and retaining a memory of their targets. High-throughput TCR repertoire sequencing captures a cross-section of TCRs that encode the immune history of any subject, though the data are heterogeneous, high dimensional, sparse, and mostly unlabeled. Sets of TCRs responding to the same antigen, *i.e.*, a protein fragment, co-occur in subjects sharing immune genetics and exposure history. Here, we leverage TCR co-occurrence across a large set of TCR repertoires and employ the GloVe (Pennington et al., 2014) algorithm to derive low-dimensional, dense vector representations (embeddings) of TCRs. We then aggregate these TCR embeddings to generate subject-level embeddings based on observed *subject-specific* TCR subsets. Further, we leverage random projection theory to improve GloVe's computational efficiency in terms of memory usage and training time. Extensive experimental results show that TCR embeddings targeting the same pathogen have high cosine similarity, and subject-level embeddings encode both immune genetics and pathogenic exposure history.
Paidamoyo Chapfuwa, Ilker Demirel, Lorenzo Pisani, Javier Zazo, Elon Portugaly, H. Jabran Zahid, Julia Greissl
ICLR5
2013 Counterfactual reasoning and learning systems: the example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero Candela, Denis Xavier Charles, David Maxwell Chickering, Elon Portugaly, Dipankar Ray, Patrice Y. Simard, Ed Snelson
J. Mach. Learn. Res.6
2010 Hidden Markov model speed heuristic and iterative HMM search procedure
abstract
BACKGROUND: Profile hidden Markov models (profile-HMMs) are sensitive tools for remote protein homology detection, but the main scoring algorithms, Viterbi or Forward, require considerable time to search large sequence databases. RESULTS: We have designed a series of database filtering steps, HMMERHEAD, that are applied prior to the scoring algorithms, as implemented in the HMMER package, in an effort to reduce search time. Using this heuristic, we obtain a 20-fold decrease in Forward and a 6-fold decrease in Viterbi search time with a minimal loss in sensitivity relative to the unfiltered approaches. We then implemented an iterative profile-HMM search method, JackHMMER, which employs the HMMERHEAD heuristic. Due to our search heuristic, we eliminated the subdatabase creation that is common in current iterative profile-HMM approaches. On our benchmark, JackHMMER detects 14% more remote protein homologs than SAM's iterative method T2K. CONCLUSIONS: Our search heuristic, HMMERHEAD, significantly reduces the time needed to score a profile-HMM against large sequence databases. This search heuristic allowed us to implement an iterative profile-HMM search method, JackHMMER, which detects significantly more remote protein homologs than SAM's T2K and NCBI's PSI-BLAST.
L. Steven Johnson, Sean R. Eddy, Elon Portugaly
BMC Bioinform.3
2008 Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein space
abstract
MOTIVATION: UPGMA (average linking) is probably the most popular algorithm for hierarchical data clustering, especially in computational biology. However, UPGMA requires the entire dissimilarity matrix in memory. Due to this prohibitive requirement, UPGMA is not scalable to very large datasets. APPLICATION: We present a novel class of memory-constrained UPGMA (MC-UPGMA) algorithms. Given any practical memory size constraint, this framework guarantees the correct clustering solution without explicitly requiring all dissimilarities in memory. The algorithms are general and are applicable to any dataset. We present a data-dependent characterization of hardness and clustering efficiency. The presented concepts are applicable to any agglomerative clustering formulation. RESULTS: We apply our algorithm to the entire collection of protein sequences, to automatically build a comprehensive evolutionary-driven hierarchy of proteins from sequence alone. The newly created tree captures protein families better than state-of-the-art large-scale methods such as CluSTr, ProtoNet4 or single-linkage clustering. We demonstrate that leveraging the entire mass embodied in all sequence similarities allows to significantly improve on current protein family clusterings which are unable to directly tackle the sheer mass of this data. Furthermore, we argue that non-metric constraints are an inherent complexity of the sequence space and should not be overlooked. The robustness of UPGMA allows significant improvement, especially for multidomain proteins, and for large or divergent families. AVAILABILITY: A comprehensive tree built from all UniProt sequence similarities, together with navigation and classification tools will be made available as part of the ProtoNet service. A C++ implementation of the algorithm is available on request.
Yaniv Loewenstein, Elon Portugaly, Menachem Fromer, Michal Linial
ISMB2
2006 EVEREST: automatic identification and classification of protein domains in all protein sequences
abstract
BACKGROUND: Proteins are comprised of one or several building blocks, known as domains. Such domains can be classified into families according to their evolutionary origin. Whereas sequencing technologies have advanced immensely in recent years, there are no matching computational methodologies for large-scale determination of protein domains and their boundaries. We provide and rigorously evaluate a novel set of domain families that is automatically generated from sequence data. Our domain family identification process, called EVEREST (EVolutionary Ensembles of REcurrent SegmenTs), begins by constructing a library of protein segments that emerge in an all vs. all pairwise sequence comparison. It then proceeds to cluster these segments into putative domain families. The selection of the best putative families is done using machine learning techniques. A statistical model is then created for each of the chosen families. This procedure is then iterated: the aforementioned statistical models are used to scan all protein sequences, to recreate a library of segments and to cluster them again. RESULTS: Processing the Swiss-Prot section of the UniProt Knoledgebase, release 7.2, EVEREST defines 20,230 domains, covering 85% of the amino acids of the Swiss-Prot database. EVEREST annotates 11,852 proteins (6% of the database) that are not annotated by Pfam A. In addition, in 43,086 proteins (20% of the database), EVEREST annotates a part of the protein that is not annotated by Pfam A. Performance tests show that EVEREST recovers 56% of Pfam A families and 63% of SCOP families with high accuracy, and suggests previously unknown domain families with at least 51% fidelity. EVEREST domains are often a combination of domains as defined by Pfam or SCOP and are frequently sub-domains of such domains. CONCLUSION: The EVEREST process and its output domain families provide an exhaustive and validated view of the protein domain world that is automatically generated from sequence data. The EVEREST library of domain families, accessible for browsing and download at 1, provides a complementary view to that provided by other existing libraries. Furthermore, since it is automatic, the EVEREST process is scalable and we will run it in the future on larger databases as well. The EVEREST source files are available for download from the EVEREST web site.
Elon Portugaly, Amir Harel, Nathan Linial, Michal Linial
BMC Bioinform.1
2005 Noise and the two-thirds power Law
abstract
The two-thirds power law, an empirical law stating an inverse non-linear relationship between the tangential hand speed and the curvature of its trajectory during curved motion, is widely acknowledged to be an invariant of upper-limb movement. It has also been shown to exist in eyemotion, locomotion and was even demonstrated in motion perception and prediction. This ubiquity has fostered various attempts to uncover the origins of this empirical relationship. In these it was generally attributed either to smoothness in hand- or joint-space or to the result of mechanisms that damp noise inherent in the motor system to produce the smooth trajectories evident in healthy human motion. We show here that white Gaussian noise also obeys this power-law. Analysis of signal and noise combinations shows that trajectories that were synthetically created not to comply with the power-law are transformed to power-law compliant ones after combination with low levels of noise. Furthermore, there exist colored noise types that drive non-power-law trajectories to power-law compliance and are not affected by smoothing. These results suggest caution when running experiments aimed at verifying the power-law or assuming its underlying existence without proper analysis of the noise. Our results could also suggest that the power-law might be derived not from smoothness or smoothness-inducing mechanisms operating on the noise inherent in our motor system but rather from the correlated noise which is inherent in this motor system.
Uri Maoz, Elon Portugaly, Tamar Flash, Yair Weiss
NIPS2
2002 Selecting targets for structural determination by navigating in a graph of protein families
abstract
MOTIVATION: A major goal in structural genomics is to enrich the catalogue of proteins whose 3D structures are known. In an attempt to address this problem we mapped over 10 000 proteins with solved structures onto a graph of all Swissprot protein sequences (release 36, approximately 73 000 proteins) provided by ProtoMap, with the goal of sorting proteins according to their likelihood of belonging to new superfamilies. We hypothesized that proteins within neighbouring clusters tend to share common structural superfamilies or folds. If true, the likelihood of finding new superfamilies increases in clusters that are distal from other solved structures within the graph. RESULTS: We defined an order relation between unsolved proteins according to their 'distance' from solved structures in the graph, and sorted approximately 48 000 proteins. Our list can be partitioned into three groups: approximately 35 000 proteins sharing a cluster with at least one known structure; approximately 6500 proteins in clusters with no solved structure but with neighbouring clusters containing known structures; and a third group contains the rest of the proteins, approximately 6100 (in 1274 clusters). We tested the quality of the order relation using thousands of recently solved structures that were not included when the order was defined. The tests show that our order is significantly better (P-value approximately 10(5)) than a random order. More interestingly, the order within the union of the second and third groups, and the order within the third group alone, perform better than random (P-values: 0.0008 and 0.15, respectively) and are better than alternative orders created using PSI-BLAST. Herein, we present a method for selecting targets to be used in structural genomics projects. AVAILABILITY: List of proteins to be used for targets selection combined with a set of biological filters for narrowing down potential targets is in http://www.protarget.cs.huji.ac.il.
Elon Portugaly, Ilona Kifer, Michal Linial
Bioinform.1
2000 Probabilities for having a new fold on the basis of a map of all protein sequences
abstract
It is a major problem in the study of protein structure to predict which proteins have new, currently unknown structural folds. In an attempt to address this problem we studied the location of all proteins with solved structures within the map of all known protein sequences provided by ProtoMap. The mutual distances in this map among solved structures are used to derive a probabilistic model from which we infer an estimate for the probability of an unsolved protein to have a new fold. The probabilities were based on data from SCOP release 1.37. The results were evaluated against the more recent SCOP pre-release 1.41. Our predicted probabilities for unsolved proteins to have a new fold are very well correlated with the proportion of new folds among recently released structures. Thus, information about the structure of proteins can be inferred from a global relational view of protein sequences. Finally, the same procedure was applied to estimate probabilities on the basis of SCOP 1.41. A list of the highest scoring proteins is provided: These are about 80 non-membranous proteins that belong to clusters with more than 5 proteins and achieve the highest probability to have a new fold. A rational selection for 3D determination of those targets is expected to accelerate the pace of new fold discovery.
Elon Portugaly, Michal Linial
RECOMB1