VLDB 2026 Research / reviewers in the wild / expert
Zohar Yakhini
dblp:16/7018
· DBLP profile ↗
53ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0002-0420-5412ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 1 first-author · 11 since 2021Theory of computation · 6 · 3 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Sequential Random Sampling Approach to PIR in DNA-based Data Storage
Chen Wang 0134, Eitan Yaakobi, Zohar Yakhini |
ISIT | 3 |
| 2025 | An Interpretation of Spearman Correlation via k-Subset PermutationsabstractDeveloped by Charles Spearman in the beginning of the 20 th century, Spearman correlation is a popular method of quantifying rank correlation in many data driven studies and projects. We suggest a new interpretation of Spearman correlation using$k$-subset permutations. We characterize the distribution of the Spearman correlation of a permutation obtained when starting with some initial permutation and then uniformly shuffling a random subset of$k$indices in that permutation, leaving all other entries unchanged. In this paper, we present some key aspects of the resulting distribution, including its expected value and variance. Specifically, we characterize the dependence on$k, n$, as well as, in some cases, on the initial permutation itself. Our results can potentially be useful in developing a stable statistics framework for Spearman correlation. Oriel Limor, Zohar Yakhini |
HPCC | 2 |
| 2025 | Constrained Coding for Composite DNA: Channel Capacity and Efficient ConstructionsabstractComposite DNA is a recent novel method to increase the information capacity of DNA-based data storage above the theoretical limit of 2 bits/symbol. In this method, every composite symbol does not store a single DNA nucleotide but a mixture of the four nucleotides in a predetermined ratio. By using different mixtures and ratios, the alphabet can be extended to have much more than four symbols in the naire approach. While this method enables higher data content per synthesis cycle, potentially reducing the DNA synthesis cost, it also imposes significant challenges for accurate DNA sequencing since the baselevel errors can easily change the mixture of bases and their ratio, resulting in changes to the composite symbols. With this motivation, we propose efficient constrained coding techniques to enforce the biological constraints, including the runlength-limited constraint and the GC-content constraint, into every DNA synthesized oligo, regardless of the mixture of bases in each composite letter and their corresponding ratio. Our contributions include computing the capacity of the constrained channel, constructing efficient encoders/decoders, and providing the best options for the composite letters to obtain capacityapproaching codes. For certain codes' parameters, our methods incur only one redundant symbol. Tuan Thanh Nguyen 0001, Chen Wang 0134, Kui Cai 0001, Yiwei Zhang 0018, Zohar Yakhini |
ISIT | 5 |
| 2025 | The Labeled Coupon Collector ProblemabstractWe generalize the well-known Coupon Collector Problem (CCP) in combinatorics. Our problem is to find the minimum and expected number of draws, with replacement, required to recover n distinctly labeled coupons, with each draw consisting of a random subset of k different coupons and a random ordering of their associated labels. We specify two variations of the problem, Type-I in which the set of labels is known at the start, and Type-II in which the set of labels is unknown at the start. We show that our problem can be viewed as an extension of the separating system problem introduced by Rényi and Katona, provide a full characterization of the minimum, and provide a numerical approach to finding the expectation using a Markov chain model, with special attention given to the case where two coupons are drawn at a time. Andrew Tan, Oriel Limor, Daniella Bar-Lev, Ryan Gabrys, Zohar Yakhini, Paul H. Siegel |
ITW | 5 |
| 2024 | Efficient Random Sampling from Very Large Databases
Idan Cohen, Aviv Yehezkel, Zohar Yakhini |
DEXA (1) | 3 |
| 2024 | Representing Information on DNA Using Patterns Induced by Enzymatic LabelingabstractEnzymatic DNA labeling is a powerful tool with applications in biochemistry, molecular biology, biotechnology, medical science, and genomic research. This paper contributes to the evolving field of DNA-based data storage by presenting a formal framework for modeling DNA labeling in strings, specifically tailored for data storage purposes. Our approach involves a known DNA molecule as a template for labeling, employing patterns induced by a set of designed labels to represent information. One hypothetical implementation can use CRISPR-Cas9 and gRNA reagents for labeling. Various aspects of the general labeling channel, including fixed-length labels, are explored, and upper bounds on the maximal size of the corresponding codes are given. The study includes the development of an efficient encoder-decoder pair that is proven optimal in terms of maximum code size under specific conditions. Daniella Bar-Lev, Tuvi Etzion, Eitan Yaakobi, Zohar Yakhini |
ISIT | 4 |
| 2024 | Error-Correcting Codes for Combinatorial Composite DNAabstractData storage in DNA is developing as a possible solution for archival digital data. Recently, to further increase the potential capacity of DNA-based data storage systems, the combinatorial composite DNA synthesis method was suggested. This approach extends the DNA alphabet by harnessing short DNA fragment reagents, known as shortmers. The shortmers are building blocks of the alphabet symbols, each consisting of a fixed number of shortmers. Thus, when information is read, it is possible that one of the shortmers that forms part of the composition of a symbol is missing and therefore the symbol cannot be determined. In this paper, we model this type of error as a type of asymmetric error and propose code constructions that can correct such errors in this setup. We also provide a lower bound on the redundancy of such error-correcting codes and give an explicit encoder and decoder for our construction. Our suggested error model is also supported by an analysis of data from actual experiments that produced DNA according to the combinatorial scheme. Lastly, we also provide a statistical evaluation of the probability of observing such error events, as a function of read depth. Omer Sabary, Inbal Preuss, Ryan Gabrys, Zohar Yakhini, Leon Anavy, Eitan Yaakobi |
ISIT | 4 |
| 2024 | Studying the Cycle Complexity of DNA SynthesisabstractStoring data in DNA is being explored as an efficient solution for archiving and in-object storage. Synthesis time and cost remain challenging, significantly limiting some applications at this stage. In this paper we investigate efficient synthesis, as it relates to cyclic synchronized synthesis technologies, such as photolithography. We define performance metrics related to the number of cycles needed for the synthesis of any fixed number of bits. We first expand on some results from the literature related to the channel capacity, addressing densities beyond those covered by prior work. This leads us to develop effective encoding achieving rate and capacity that are higher than previously reported. Finally, we analyze cost based on a parametric definition and determine some bounds and asymptotics. We investigate alphabet sizes that can be larger than 4, both for theoretical completeness and since practical approaches to such schemes were recently suggested and tested in the literature. Amit Zrihan, Eitan Yaakobi, Zohar Yakhini |
ITW | 3 |
| 2024 | HIPI: Spatially resolved multiplexed protein expression inferred from H&E WSIsabstractSolid tumors are characterized by complex interactions between the tumor, the immune system and the microenvironment. These interactions and intra-tumor variations have both diagnostic and prognostic significance and implications. However, quantifying the underlying processes in patient samples requires expensive and complicated molecular experiments. In contrast, H&E staining is typically performed as part of the routine standard process, and is very cheap. Here we present HIPI (H&E Image Interpretation and Protein Expression Inference) for predicting cell marker expression from tumor H&E images. We process paired H&E and CyCIF images taken from serial sections of colorectal cancers to train our model. We show that our model accurately predicts the spatial distribution of several important cell markers, on both held-out tumor regions as well as new tumor samples taken from different patients. Moreover, using only the tissue image morphology, HIPI is able to colocalize the interactions between different cell types, further demonstrating its potential clinical significance. Ron Zeira, Leon Anavy, Zohar Yakhini, Ehud Rivlin, Daniel Freedman |
PLoS Comput. Biol. | 3 |
| 2024 | Privacy Preserving Feature Selection for Sparse Linear RegressionabstractPrivacy-Preserving Machine Learning (PPML) provides protocols for learning and statistical analysis of data that may be distributed amongst multiple data owners (e.g., hospitals that own proprietary healthcare data), while preserving data privacy. The PPML literature includes protocols for various learning methods, including ridge regression. Ridge regression controls the L2 norm of the model, but does not aim to strictly reduce the number of non-zero coefficients, namely the L0 norm of the model. Reducing the number of non-zero coefficients (a form of feature selection) is important for avoiding overfitting, and for reducing the cost of using learnt models in practice. In this work, we develop a first privacy-preserving protocol for sparse linear regression under L0 constraints. The protocol addresses data contributed by several data owners (e.g., hospitals). Our protocol outsources the bulk of the computation to two non-colluding servers, using homomorphic encryption as a central tool. We provide a rigorous security proof for our protocol, where security is against semi-honest adversaries controlling any number of data owners and at most one server. We implemented our protocol, and evaluated performance with nearly a million samples and up to 40 features. Adi Akavia, Ben Galili, Hayim Shaul, Mor Weiss, Zohar Yakhini |
Proc. Priv. Enhancing Technol. | 5 |
| 2024 | Sequence Design and Reconstruction Under the Repeat Channel in Enzymatic DNA SynthesisabstractUsing synthetic DNA for data storage and for physical information encoding in labeling, tracing, and authentication applications is becoming more feasible as synthesis and reading technologies are improving. DNA in data storage applications has several advantages such as very high physical density and robustness. Some of the new synthesis technologies lead to repetition noise, consisting of sticky insertions and deletions in the resulting messages. In this paper, we address reconstruction algorithms for multiple trace communication channels with repetition (sticky insertion and deletion) noise. We prove correctness and analyze failure rates, both analytically and on simulated data. We identify a failure mechanism related to alternating stretches in the design sequence that leads to a potential bias in the data derived from reads (traces) and used for reconstruction. To minimize this effect we introduce alternating length limited codes (ALL codes) and analyze some of their properties. Roy Shafir, Omer Sabary, Leon Anavy, Eitan Yaakobi, Zohar Yakhini |
IEEE Trans. Commun. | 5 |
| 2023 | Efficient Privacy-Preserving Viral Strain Classification via k-mer Signatures and FHEabstractWith the development of sequencing technologies, viral strain classification - which is critical for many applications, including disease monitoring and control - has become widely deployed. Typically, a lab (client) holds a viral sequence, and requests classification services from a centralized repository of labeled viral sequences (server). However, such “classification as a service” raises privacy concerns. In this paper we propose a privacy-preserving viral strain classification protocol that allows the client to obtain classification services from the server, while maintaining complete privacy of the client's viral strains. The privacy guarantee is against active servers, and the correctness guarantee is against passive ones. We implemented our protocol and performed extensive benchmarks, showing that it obtains almost perfect accuracy (99.8%-100%) and microAUC (0.999), and high efficiency (amortized per-sequence client and server runtimes of 4.95ms and 0.53ms, respectively, and 0.21MB communication). In addition, we present an extension of our protocol that guarantees server privacy against passive clients, and provide an empirical evaluation showing that this extension provides the same high accuracy and microAUC, with amortized per sequences overhead of only a few milliseconds in client and server runtime, and 0.3MB in communication complexity. Along the way, we develop an enhanced packing technique in which two reals are packed in a single complex number, with support for homomorphic inner products of vectors of ciphertexts. We note that while similar packing techniques were used before, they only supported additions and multiplication by constants. Adi Akavia, Ben Galili, Hayim Shaul, Mor Weiss, Zohar Yakhini |
CSF | 5 |
| 2021 | Autoencoder Image Interpolation by Shaping the Latent SpaceabstractOne of the fascinating properties of deep learning is the ability of the network to reveal the underlying factors characterizing elements in datasets of different types. Autoencoders represent an effective approach for computing these factors. Autoencoders have been studied in the context of enabling interpolation between data points by decoding convex combinations of latent vectors. However, this interpolation often leads to artifacts or produces unrealistic results during reconstruction. We argue that these incongruities are due to the structure of the latent space and to the fact that such naively interpolated latent vectors deviate from the data manifold. In this paper, we propose a regularization technique that shapes the latent representation to follow a manifold that is consistent with the training images and that forces the manifold to be smooth and locally convex. This regularization not only enables faithful interpolation between data points, as we show herein but can also be used as a general regularization technique to avoid overfitting or to produce new samples for data augmentation. Alon Oring, Zohar Yakhini, Yacov Hel-Or |
ICML | 2 |
| 2021 | Sequence Reconstruction Under Stutter Noise in Enzymatic DNA SynthesisabstractSynthetic DNA is an attractive alternative for data storage media due to its high information density, low energy usage, and exceptional robustness. Enzymatic DNA synthesis was recently introduced to allow cost effective synthesis of longer DNA molecules for data storage. This method is characterized by stutter errors which are sticky insertions so that every base in the designed sequence may be synthesized more than once. In this work, we study the problem of reconstructing the original sequence from a set of noisy reads originating from the stuttering enzymatic synthesis. We present different reconstruction algorithms and analyze their expected success probability and error rate for three different scenarios that depend on the information which is known about the stutter errors. We evaluate algorithmic performance analytically as well as by using simulations. We are especially interested in characterizing the performance as a function of the read depth. Our findings can be used to evaluate the trade-offs between synthesis quality indicators and the sequencing depth required for reconstruction with high probability. In principle, the probability of reconstruction failure exponentially decays with the sequencing depth, as demonstrated in the study. We also analyze the use of error-correcting codes to improve the error performance. Roy Shafir, Omer Sabary, Leon Anavy, Eitan Yaakobi, Zohar Yakhini |
ITW | 5 |
| 2021 | On the stability of log-rank test under labeling errorsabstractMOTIVATION: Log-rank test is a widely used test that serves to assess the statistical significance of observed differences in survival, when comparing two or more groups. The log-rank test is based on several assumptions that support the validity of the calculations. It is naturally assumed, implicitly, that no errors occur in the labeling of the samples. That is, the mapping between samples and groups is perfectly correct. In this work, we investigate how test results may be affected when considering some errors in the original labeling. RESULTS: We introduce and define the uncertainty that arises from labeling errors in log-rank test. In order to deal with this uncertainty, we develop a novel algorithm for efficiently calculating a stability interval around the original log-rank P-value and prove its correctness. We demonstrate our algorithm on several datasets. AVAILABILITY AND IMPLEMENTATION: We provide a Python implementation, called LoRSI, for calculating the stability interval using our algorithm https://github.com/YakhiniGroup/LoRSI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ben Galili, Anat Samohi, Zohar Yakhini |
Bioinform. | 3 |
| 2021 | On the stability of log-rank test under labeling errors
Ben Galili, Anat Samohi, Zohar Yakhini |
Bioinform. | 3 |
| 2021 | Assessing heterogeneity in spatial data using the HTA index with applications to spatial transcriptomics and imagingabstractMOTIVATION: Tumour heterogeneity is being increasingly recognized as an important characteristic of cancer and as a determinant of prognosis and treatment outcome. Emerging spatial transcriptomics data hold the potential to further our understanding of tumour heterogeneity and its implications. However, existing statistical tools are not sufficiently powerful to capture heterogeneity in the complex setting of spatial molecular biology. RESULTS: We provide a statistical solution, the HeTerogeneity Average index (HTA), specifically designed to handle the multivariate nature of spatial transcriptomics. We prove that HTA has an approximately normal distribution, therefore lending itself to efficient statistical assessment and inference. We first demonstrate that HTA accurately reflects the level of heterogeneity in simulated data. We then use HTA to analyze heterogeneity in two cancer spatial transcriptomics datasets: spatial RNA sequencing by 10x Genomics and spatial transcriptomics inferred from H&E. Finally, we demonstrate that HTA also applies to 3D spatial data using brain MRI. In spatial RNA sequencing, we use a known combination of molecular traits to assert that HTA aligns with the expected outcome for this combination. We also show that HTA captures immune-cell infiltration at multiple resolutions. In digital pathology, we show how HTA can be used in survival analysis and demonstrate that high levels of heterogeneity may be linked to poor survival. In brain MRI, we show that HTA differentiates between normal ageing, Alzheimer's disease and two tumours. HTA also extends beyond molecular biology and medical imaging, and can be applied to many domains, including GIS. AVAILABILITY AND IMPLEMENTATION: Python package and source code are available at: https://github.com/alonalj/hta. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alona Levy-Jurgenson, Xavier Tekpli, Zohar Yakhini |
Bioinform. | 3 |
| 2021 | Erratum to: Assessing heterogeneity in spatial data using the HTA index with applications to spatial transcriptomics and imagingabstractBioinformatics (2021), https://doi.org/10.1093/bioinformatics/btab569 Upon the original publication of this manuscript, errors were noted. This erratum has been published to address the following changes introduced due to production error and that have subsequently been corrected: All URL links for the figures have been corrected to link to the appropriated figures. Under sub-section “2.3.1 Equal-weight regions”, the below equation has been updated as follows: The Publisher would like to apologize for these errors. Alona Levy-Jurgenson, Xavier Tekpli, Zohar Yakhini |
Bioinform. | 3 |
| 2021 | SOLQC: Synthetic Oligo Library Quality Control toolabstractMOTIVATION: Recent years have seen a growing number and an expanding scope of studies using synthetic oligo libraries for a range of applications in synthetic biology. As experiments are growing by numbers and complexity, analysis tools can facilitate quality control and support better assessment and inference. RESULTS: We present a novel analysis tool, called SOLQC, which enables fast and comprehensive analysis of synthetic oligo libraries, based on NGS analysis performed by the user. SOLQC provides statistical information such as the distribution of variant representation, different error rates and their dependence on sequence or library properties. SOLQC produces graphical reports from the analysis, in a flexible format. We demonstrate SOLQC by analyzing literature libraries. We also discuss the potential benefits and relevance of the different components of the analysis. AVAILABILITY AND IMPLEMENTATION: SOLQC is a free software for non-commercial use, available at https://app.gitbook.com/@yoav-orlev/s/solqc/. For commercial use please contact the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Omer Sabary, Yoav Orlev, Roy Shafir, Leon Anavy, Eitan Yaakobi, Zohar Yakhini |
Bioinform. | 6 |
| 2021 | miRNA normalization enables joint analysis of several datasets to increase sensitivity and to reveal novel miRNAs differentially expressed in breast cancerabstractDifferent miRNA profiling protocols and technologies introduce differences in the resulting quantitative expression profiles. These include differences in the presence (and measurability) of certain miRNAs. We present and examine a method based on quantile normalization, Adjusted Quantile Normalization (AQuN), to combine miRNA expression data from multiple studies in breast cancer into a single joint dataset for integrative analysis. By pooling multiple datasets, we obtain increased statistical power, surfacing patterns that do not emerge as statistically significant when separately analyzing these datasets. To merge several datasets, as we do here, one needs to overcome both technical and batch differences between these datasets. We compare several approaches for merging and jointly analyzing miRNA datasets. We investigate the statistical confidence for known results and highlight potential new findings that resulted from the joint analysis using AQuN. In particular, we detect several miRNAs to be differentially expressed in estrogen receptor (ER) positive versus ER negative samples. In addition, we identify new potential biomarkers and therapeutic targets for both clinical groups. As a specific example, using the AQuN-derived dataset we detect hsa-miR-193b-5p to have a statistically significant over-expression in the ER positive group, a phenomenon that was not previously reported. Furthermore, as demonstrated by functional assays in breast cancer cell lines, overexpression of hsa-miR-193b-5p in breast cancer cell lines resulted in decreased cell viability in addition to inducing apoptosis. Together, these observations suggest a novel functional role for this miRNA in breast cancer. Packages implementing AQuN are provided for Python and Matlab: https://github.com/YakhiniGroup/PyAQN. Shay Ben-Elazar, Miriam Ragle Aure, Kristin Jonsdottir, Suvi-Katri Leivonen, Vessela N. Kristensen, Emiel A. M. Janssen, Kristine Kleivi Sahlberg, Ole Christian Lingjærde, Zohar Yakhini |
PLoS Comput. Biol. | 9 |
| 2020 | IoT or NoT: Identifying IoT Devices in a Short Time ScaleabstractIn recent years the number of IoT devices in home networks has increased dramatically. Whenever a new device connects to the network, it must be quickly managed and secured using the relevant security mechanism or QoS policy. Thus a key challenge is to distinguish between IoT and NoT devices in a matter of minutes. Unfortunately, there is no clear indication of whether a device in a network is an IoT. In this paper, we propose different classifiers that identify a device as IoT or non-IoT, in a short time scale, and with high accuracy.Our classifiers were constructed using machine learning techniques on a seen (training) dataset and were tested on an unseen (test) dataset. They successfully classified devices that were not in the seen dataset with accuracy above 95%. The first classifier is a logistic regression classifier based on traffic features. The second classifier is based on features we retrieve from DHCP packets. Finally, we present a unified classifier that leverages the advantages of the other two classifiers. Anat Bremler-Barr, Haim Levy, Zohar Yakhini |
NOMS | 3 |
| 2017 | Mutual enrichment in aggregated ranked lists with applications to gene expression regulationabstractBioinformatics (2016) 32 (17), i464–i472. doi: 10.1093/bioinformatics/btw435 Comparison of MULSEA-m and mmHG Singletons. Performance results for 100 random instances with 50 lists total of 1000 elements each, 20 planted lists, VAR = 0.5. The pivot is generated by the cumulative model. The noise levels depicted are 0 to 5001000=0.5 In reviewing the published version of the above article the authors noticed that the legend for Figure 6 is in reverse order. Below please find the corrected figure and its caption. The article has now been corrected online. Dalia Cohn-Alperovich, Alona Rabner, Ilona Kifer, Yael Mandel-Gutfreund, Zohar Yakhini |
Bioinform. | 5 |
| 2017 | Optimizing Analytical Depth and Cost Efficiency of IEF-LC/MS ProteomicsabstractIEF LC-MS/MS is an analytical method that incorporates a two-step sample separation prior to MS identification of proteins. When analyzing complex samples this preparatory separation allows for higher analytical depth and improved quantification accuracy of proteins. However, cost and analysis time are greatly increased as each analyzed IEF fraction is separately profiled using LC-MS/MS. We propose an approach that selects a subset of IEF fractions for LC-MS/MS analysis that is highly informative in the context of a group of proteins of interest. Specifically, our method allows a significant reduction in cost and instrument time as compared to the standard protocol of running all fractions, with little compromise to coverage. We develop algorithmics to optimize the selection of the IEF fractions on which to run LC-MS/MS. We translate the fraction optimization task to Minimum Set Cover, a well-studied NP-hard problem. We develop heuristic solutions and compare them in terms of effectiveness and running times. We provide examples to demonstrate advantages and limitations of each algorithmic approach. Finally, we test our methodology by applying it to experimental data obtained from IEF LC-MS/MS analysis of yeast and human samples. We demonstrate the benefit of this approach for analyzing complex samples with a focus on different protein sets of interest. Ilona Kifer, Rui Mamede Branca, Amir Ben-Dor, Linhui Zhai, Janne Lehtiö, Zohar Yakhini |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2016 | Extending partial haplotypes to full genome haplotypes using chromosome conformation capture dataabstractMOTIVATION: Complex interactions among alleles often drive differences in inherited properties including disease predisposition. Isolating the effects of these interactions requires phasing information that is difficult to measure or infer. Furthermore, prevalent sequencing technologies used in the essential first step of determining a haplotype limit the range of that step to the span of reads, namely hundreds of bases. With the advent of pseudo-long read technologies, observable partial haplotypes can span several orders of magnitude more. Yet, measuring whole-genome-single-individual haplotypes remains a challenge. A different view of whole genome measurement addresses the 3D structure of the genome-with great development of Hi-C techniques in recent years. A shortcoming of current Hi-C, however, is the difficulty in inferring information that is specific to each of a pair of homologous chromosomes. RESULTS: In this work, we develop a robust algorithmic framework that takes two measurement derived datasets: raw Hi-C and partial short-range haplotypes, and constructs the full-genome haplotype as well as phased diploid Hi-C maps. By analyzing both data sets together we thus bridge important gaps in both technologies-from short to long haplotypes and from un-phased to phased Hi-C. We demonstrate that our method can recover ground truth haplotypes with high accuracy, using measured biological data as well as simulated data. We analyze the impact of noise, Hi-C sequencing depth and measured haplotype lengths on performance. Finally, we use the inferred 3D structure of a human genome to point at transcription factor targets nuclear co-localization. AVAILABILITY AND IMPLEMENTATION: The implementation available at https://github.com/YakhiniGroup/SpectraPh CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shay Ben-Elazar, Benny Chor, Zohar Yakhini |
Bioinform. | 3 |
| 2016 | Mutual enrichment in aggregated ranked lists with applications to gene expression regulationabstractMOTIVATION: It is often the case in biological measurement data that results are given as a ranked list of quantities-for example, differential expression (DE) of genes as inferred from microarrays or RNA-seq. Recent years brought considerable progress in statistical tools for enrichment analysis in ranked lists. Several tools are now available that allow users to break the fixed set paradigm in assessing statistical enrichment of sets of genes. Continuing with the example, these tools identify factors that may be associated with measured differential expression. A drawback of existing tools is their focus on identifying single factors associated with the observed or measured ranks, failing to address relationships between these factors. For example, a scenario in which genes targeted by multiple miRNAs play a central role in the DE signal but the effect of each single miRNA is too subtle to be detected, as shown in our results. RESULTS: We propose statistical and algorithmic approaches for selecting a sub-collection of factors that can be aggregated into one ranked list that is heuristically most associated with an input ranked list (pivot). We examine performance on simulated data and apply our approach to cancer datasets. We find small sub-collections of miRNA that are statistically associated with gene DE in several types of cancer, suggesting miRNA cooperativity in driving disease related processes. Many of our findings are consistent with known roles of miRNAs in cancer, while others suggest previously unknown roles for certain miRNAs. AVAILABILITY AND IMPLEMENTATION: Code and instructions for our algorithmic framework, MULSEA, are in: https://github.com/YakhiniGroup/MULSEAContact:[email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dalia Cohn-Alperovich, Alona Rabner, Ilona Kifer, Yael Mandel-Gutfreund, Zohar Yakhini |
Bioinform. | 5 |
| 2015 | ENViz: a Cytoscape App for integrated statistical analysis and visualization of sample-matched data with multiple data typesabstractAbstract Summary: ENViz (Enrichment Analysis and Visualization) is a Cytoscape app that performs joint enrichment analysis of two types of sample matched datasets in the context of systematic annotations. Such datasets may be gene expression or any other high-throughput data collected in the same set of samples. The enrichment analysis is done in the context of pathway information, gene ontology or any custom annotation of the data. The results of the analysis consist of significant associations between profiled elements of one of the datasets to the annotation terms (e.g. miR-19 was associated to the cell-cycle process in breast cancer samples). The results of the enrichment analysis are visualized as an interactive Cytoscape network. Availability and implementation: ENViz is publically available in the Cytoscape App Store (http://apps.cytoscape.org/apps/enviz). For additional information please visit the tool website: http://www.agilent.com/labs/research/compbio/enviz/ Contact: [email protected] Israel Steinfeld, Roy Navon, Michael L. Creech, Zohar Yakhini, Anya Tsalenko |
Bioinform. | 4 |
| 2014 | Optimizing analytical depth and cost efficiency of IEF-LC/MS proteomicsabstractIEF LC-MS/MS (Iso-Electric-Focusing Liquid-Chromatography Tandem-Mass-Spectrometry) is an analytical method that incorporates a two-step sample separation prior to MS identification of proteins. When analyzing complex samples this preparatory separation allows for higher analytical depth and improved quantification accuracy of proteins and PTMs. IEF fractionation builds upon isoelectric point (pI) differences between peptides or proteins. In standard IEF LC-MS/MS, each fraction is separately profiled using LC-MS/MS. The cost of the complete assay, therefore, is strongly dependent on the number of pI fractions being analyzed. Commonly, studies are focused on a specific group of proteins. We propose an approach that selects a subset of fractions for LC-MS/MS analysis that is highly informative in the context of the proteins of interest. Specifically, our method allows a significant reduction in cost and instrument time as compared to the standard protocol of running all fractions, with little compromise to coverage. We develop algorithmics to optimize the selection of the IEF fractions on which to run LC-MS/MS. We translate the fraction optimization task to Minimum Set Cover (MSC), a well-studied NP-hard problem. We develop heuristic solutions and compare them in terms of effectiveness and of practical running times. We provide examples to demonstrate advantages and limitations of each algorithmic approach. Finally, we test our methodology by applying it to experimental data obtained from IEF LC-MS/MS analysis of yeast and human samples. We demonstrate the benefit of this approach for analyzing complex samples with a focus on different protein sets of interest. Ilona Kifer, Amir Ben-Dor, Zohar Yakhini, Rui Mamede Branca, Janne Lehtiö |
BIBM | 3 |
| 2013 | Mutual Enrichment in Ranked Lists and the Statistical Assessment of Position Weight Matrix Motifs
Limor Leibovich, Zohar Yakhini |
WABI | 2 |
| 2012 | Dotted interval graphsabstractWe introduce a generalization of interval graphs, which we call Dotted Interval Graphs (DIG). A dotted interval graph is an intersection graph of arithmetic progressions (dotted intervals). Coloring of dotted interval graphs naturally arises in the context of high throughput genotyping. We study the properties of dotted interval graphs, with a focus on coloring. We show that any graph is a DIG, but that DIG d graphs, that is, DIGs in which the arithmetic progressions have a jump of at most d , form a strict hierarchy. We show that coloring DIG d graphs is NP-complete even for d = 2. For any fixed d , we provide a 5/6 d + o ( d ) approximation for the coloring of DIG d graphs. Finally, we show that finding the maximal clique in DIG d graphs is fixed parameter tractable in d . Yonatan Aumann, Moshe Lewenstein, Oren Melamud, Ron Y. Pinter, Zohar Yakhini |
ACM Trans. Algorithms | 5 |
| 2011 | Cancer Computational BiologyabstractEditorialIntroduction of high-throughput measurement technolo-gies combined with the increase of the scientific knowl-edge base, with respect to our understanding of cellularand biological processes, resulted in establishing compu-ter and information science as an important and funda-mental component of modern biology. High-throughputmeasurement technologies, such as microarray-basedprofiling, mass spectrometry screens, and high-through-put sequencing, give rise to several computational chal-lenges. On one hand, they require a rigorous approachto assay design. Scientists and technology developerswork on optimizing assay components so as to maxi-mize the information obtained through the measure-ment. On the other hand, the use of high-throughputmeasurement gives rise to large quantities of data thatneeds to be pre-processed and analyzed to obtain mean-ingful knowledge. This processing and analysis is per-formed on various levels - from pre-processing the rawdata, such as images from microarrays or raw sequencereads - to analyzing the data and to the discovery of bio-markers or other biologically meaningful characteristics.Measurement technology addresses several aspects ofcellular processes such as DNA, RNA, proteomics,metabolomics, epigenetics and pathways. This increasein the scientific knowledge base also leads to a centralrole played by data analysis and modeling, stronglygrounded in computational methods. Systems biology orintegrative biology approaches and network analysis areof specific importance in this context.The above is even further emphasized in the contextof cancer research. Samples are complex and heteroge-neous, and cancer related mechanisms involve manylayers of the process that leads from the genome to cel-lular function. One example of a specific need of canceris the study of large scale aberrations in the genome.CNVs (copy number variations) were recentlyrecognized as abundant in normal cell populations andas related to many other disease types but they are stilla hallmark of cancer [1,2]. Genomes in cancer cellsoften have a structure that allows them to bypassgrowth control cellular processes. Regions coding fortumor suppressor genes are often deleted and regionsharboring oncogenes may be amplified. This is the case,for example, for p16 and myc, respectively [3-5]. Rear-rangements, such as inversions and translocations, giverise to tumor-driving fusion products as in the case ofBCR-Abl and the Philadelphia Chromosome as well asin more recent findings implicating fusion structures insolid tumors. Cancer research therefore makes use ofdata analysis methods and tools that address interpreta-tion of copy number data and the understanding of theeffect of genome changes on transcriptome level as wellas proteome level profiles of tumors. Other specificcomputational needs of cancer research are related toepigenetic changes, somatic evolution, definition of genesets in the context of specific cancer types, and to drugsand data that measures the effects of drugs.Computational biologists focusing on cancer developmethods for the genome scale characterization of tumors,on various levels of the molecular process. Data analysismethods often rely on the analysis of high-throughputmeasurement data and they provide understanding of therelationship between various molecular characteristics ofcells. For example - how do genome structural aberra-tions and changes in copy number, a result of increasedgenome instability in cancer, affect the expression ofgenes and other functional elements such as miRNA, andhow do the latter changes affect the function of relatedproteins. Understanding of the association of genomiccharacteristics and clinical properties of primary tumorsamples, xenografts or cell lines contributes to persona-lized cancer medicine through the development of pre-dictive biomarkers of drug efficacy. Many researchprojects therefore aim to discover biomarkers, at eithergenome, transcriptome or proteome level that are prog-nostic of cancer progression or predictive of response tospecific therapeutic agents [6,7]. Cancer computationalbiology also focuses on analyzing molecules and Zohar Yakhini, Igor Jurisica |
BMC Bioinform. | 1 |
| 2009 | GOrilla: a tool for discovery and visualization of enriched GO terms in ranked gene listsabstractBACKGROUND: Since the inception of the GO annotation project, a variety of tools have been developed that support exploring and searching the GO database. In particular, a variety of tools that perform GO enrichment analysis are currently available. Most of these tools require as input a target set of genes and a background set and seek enrichment in the target set compared to the background set. A few tools also exist that support analyzing ranked lists. The latter typically rely on simulations or on union-bound correction for assigning statistical significance to the results. RESULTS: GOrilla is a web-based application that identifies enriched GO terms in ranked lists of genes, without requiring the user to provide explicit target and background sets. This is particularly useful in many typical cases where genomic data may be naturally represented as a ranked list of genes (e.g. by level of expression or of differential expression). GOrilla employs a flexible threshold statistical approach to discover GO terms that are significantly enriched at the top of a ranked gene list. Building on a complete theoretical characterization of the underlying distribution, called mHG, GOrilla computes an exact p-value for the observed enrichment, taking threshold multiple testing into account without the need for simulations. This enables rigorous statistical analysis of thousand of genes and thousands of GO terms in order of seconds. The output of the enrichment analysis is visualized as a hierarchical structure, providing a clear view of the relations between enriched GO terms. CONCLUSION: GOrilla is an efficient GO analysis tool with unique features that make a useful addition to the existing repertoire of GO enrichment tools. GOrilla's unique features and advantages over other threshold free enrichment tools include rigorous statistics, fast running time and an effective graphical representation. GOrilla is publicly available at: http://cbl-gorilla.cs.technion.ac.il Eran Eden, Roy Navon, Israel Steinfeld, Doron Lipson, Zohar Yakhini |
BMC Bioinform. | 5 |
| 2007 | Framework for Identifying Common Aberrations in DNA Copy Number Data
Amir Ben-Dor, Doron Lipson, Anya Tsalenko, Mark Reimers, Lars O. Baumbusch, Michael T. Barrett, John N. Weinstein, Anne-Lise Børresen-Dale, Zohar Yakhini |
RECOMB | 9 |
| 2007 | Optimization of probe coverage for high-resolution oligonucleotide aCGHabstractMOTIVATION: The resolution at which genomic alterations can be mapped by means of oligonucleotide aCGH (array-based comparative genomic hybridization) is limited by two factors: the availability of high-quality probes for the target genomic sequence and the array real-estate. Optimization of the probe selection process is required for arrays that are designed to probe specific genomic regions in very high resolution without compromising probe quality constraints. RESULTS: In this paper we describe a well-defined optimization problem associated with the problem of probe selection for high-resolution aCGH arrays. We propose the whenever possible in-cover as a formulation that faithfully captures the requirement of probe selection problem, and provide a fast randomized algorithm that solves the optimization problem in O(n logn) time, as well as a deterministic algorithm with the same asymptotic performance. We apply the method in a typical high-definition array design scenario and demonstrate its superiority with respect to alternative approaches. AVAILABILITY: Address requests to the authors. Doron Lipson, Zohar Yakhini, Yonatan Aumann |
Bioinform. | 2 |
| 2007 | Similarities and differences of gene expression in yeast stress conditionsabstractUNLABELLED: MOTIVATION AND METHODS: All living organisms and the survival of all cells critically depend on their ability to sense and quickly adapt to changes in the environment and to other stress conditions. We study stress response mechanisms in Saccharomyces cerevisiae by identifying genes that, according to very stringent criteria, have persistent co-expression under a variety of stress conditions. This is enabled through a fast clique search method applied to the intersection of several co-expression graphs calculated over the data of Gasch et al. This method exploits the topological characteristics of these graphs. RESULTS: We observe cliques in the intersection graphs that are much larger than expected under a null model of changing gene identities for different stress conditions but maintaining the co-expression topology within each one. Persistent cliques are analyzed to identify enriched function as well as enriched regulation by a small number of TFs. These TFs, therefore, characterize a universal and persistent reaction to stress response. We further demonstrate that the vertices (genes) of many cliques in the intersection graphs are co-localized in the yeast genome, to a degree far beyond the random expectation. Co-localization can hypothetically contribute to a quick co-ordinated response. We propose the use of persistent cliques in further study of properties of co-regulation. Oleg Rokhlenko, Ydo Wexler, Zohar Yakhini |
Bioinform. | 3 |
| 2007 | A supervised approach for identifying discriminating genotype patterns and its application to breast cancer dataabstractMOTIVATION: Large-scale association studies, investigating the genetic determinants of a phenotype of interest, are producing increasing amounts of genomic variation data on human cohorts. A fundamental challenge in these studies is the detection of genotypic patterns that discriminate individuals exhibiting the phenotype under study from individuals that do not possess it. The difficulty stems from the large number of single nucleotide polymorphism (SNP) combinations that have to be tested. The discrimination problem becomes even more involved when additional high-throughput data, such as gene expression data, are available for the same cohort. RESULTS: We have developed a graph theoretic approach for identifying discriminating patterns (DPs) for a given phenotype in a genotyped population. The method is based on representing the SNP data as a bipartite graph of individuals and their SNP states, and identifying fully connected subgraphs of this graph that relate individuals enriched for a given phenotypic group. The method can handle additional data types such as expression profiles of the genotyped population. It is reminiscent of biclustering approaches with the crucial difference that its search process is guided by the phenotype under consideration in a supervised manner. We tested our approach in simulations and on real data. In simulations, our method was able to retrieve planted patterns with high success rate. We then applied our approach to a dataset of 72 breast cancer patients with available gene expression profiles, genotyped over 695 SNPs. We detected several DPs that were highly significant with respect to various clinical phenotypes, and investigated the groups of patients and the groups of genes they defined. We found the patient groups to be highly enriched for other phenotypes and to display expression coherency among their profiles. The gene groups displayed functional coherency and involved genes with known role in cancer, providing additional support to their involvement. AVAILABILITY: The program is available upon request. Nir Yosef, Zohar Yakhini, Anya Tsalenko, Vessela N. Kristensen, Anne-Lise Børresen-Dale, Eytan Ruppin, Roded Sharan |
Bioinform. | 2 |
| 2007 | Semi-supervised class discovery using quantitative phenotypes - CVD as a case studyabstractGenomic studies typically focus on comparing disease to healthy population. In our work, various parameters, including peripheral blood mononuclear (PBM) cells expression profiling, were stratified solely from healthy subjects. To analyze the data we developed a semi-supervised class discovery method, constraining the search space to patterns that respect an order induced by the rich quantitative annotations. We show that our method is robust enough to detect known clinical parameters with accordance to expected values. We also use our method to elucidate cardiovascular disease (CVD) putative risk factors. One of the basic tasks in gene expression data analysis is finding differentially expressed genes between 2 classes (such as tumor vs. normal). Among the various methods for measuring differential expression (e.g. Student t-test), we focus on TnoM [ 1 ] which is a non-parametric statistical score that affords an exact p-value. When many partitions of the sample set are possible, one would like to assess the statistical significance of any partition considered, and to compare between partitions. In overabundance [ 2 ] analysis the exact p-value of the TNoM score is used to estimate the expected number of differentially expressed genes. By comparing to the actually observed number we can calculate the overabundance of differentially expressed genes. This quantity can be used as a figure of merit: higher overabundance indicating a more profound change in the cell state. Typical class discovery in gene expression data searches over all possible partitions of the set of samples and uses heuristic methods to do so [ 2 , 3 ]. Given any quantitative phenotype, we can constrain the search space to patterns that respect the order it induces on the set of samples. This approach reduces the search space from O(3 ) to O(n ) making the search feasible (Figure 1 ). IMT levels available for 42 subjects are presented. All threshold pairs of IMT levels were tested, each representing a partition of the samples to high IMT levels, low IMT levels and samples not used. Marked in red is the threshold pair with the highest overabundance of genes, giving rise to the partition of 27 samples with IMT values of 0.6–0.86 vs. 9 samples with IMT values of 0.92–1.05. We applied our method to PBMC gene expression profiling data, collected from 49 healthy subjects. Clinical, laboratory measurement and CVD prognostic indicators were also collected, adding more then 160 phenotypic parameters for each subject. One of the interesting phenotypic parameters is Carotid Intima-Media Thickness (IMT) [ 4 ], a CVD prognostic indicator. Using semi-supervised class discovery with the IMT values we received IMT threshold levels that are in agreement with the known prognosis values (Figure 1 ). The differentially expressed genes in this partition were enriched with GO terms related to vesicle-mediated transport (p < 10 ) and glycolysis (p < 10 ), giving mechanistic insights to the difference between the two cell states. Israel Steinfeld, Roy Navon, Diego Ardigò, Ivana Zavaroni, Zohar Yakhini |
BMC Bioinform. | 5 |
| 2007 | Discovering Motifs in Ranked Lists of DNA SequencesabstractComputational methods for discovery of sequence elements that are enriched in a target set compared with a background set are fundamental in molecular biology research. One example is the discovery of transcription factor binding motifs that are inferred from ChIP-chip (chromatin immuno-precipitation on a microarray) measurements. Several major challenges in sequence motif discovery still require consideration: (i) the need for a principled approach to partitioning the data into target and background sets; (ii) the lack of rigorous models and of an exact p-value for measuring motif enrichment; (iii) the need for an appropriate framework for accounting for motif multiplicity; (iv) the tendency, in many of the existing methods, to report presumably significant motifs even when applied to randomly generated data. In this paper we present a statistical framework for discovering enriched sequence elements in ranked lists that resolves these four issues. We demonstrate the implementation of this framework in a software application, termed DRIM (discovery of rank imbalanced motifs), which identifies sequence motifs in lists of ranked DNA sequences. We applied DRIM to ChIP-chip and CpG methylation data and obtained the following results. (i) Identification of 50 novel putative transcription factor (TF) binding sites in yeast ChIP-chip data. The biological function of some of them was further investigated to gain new insights on transcription regulation networks in yeast. For example, our discoveries enable the elucidation of the network of the TF ARO80. Another finding concerns a systematic TF binding enhancement to sequences containing CA repeats. (ii) Discovery of novel motifs in human cancer CpG methylation data. Remarkably, most of these motifs are similar to DNA sequence elements bound by the Polycomb complex that promotes histone methylation. Our findings thus support a model in which histone methylation and CpG methylation are mechanistically linked. Overall, we demonstrate that the statistical framework embodied in the DRIM software tool is highly effective for identifying regulatory sequence elements in a variety of applications ranging from expression and ChIP-chip to CpG methylation data. DRIM is publicly available at http://bioinfo.cs.technion.ac.il/drim. Eran Eden, Doron Lipson, Sivan Yogev, Zohar Yakhini |
PLoS Comput. Biol. | 4 |
| 2005 | Efficient Calculation of Interval Scores for DNA Copy Number Data Analysis
Doron Lipson, Yonatan Aumann, Amir Ben-Dor, Nathan Linial, Zohar Yakhini |
RECOMB | 5 |
| 2005 | A High-Throughput Approach for Associating microRNAs with Their Activity Conditions
Chaya Ben-Zaken Zilberstein, Michal Ziv-Ukelson, Ron Y. Pinter, Zohar Yakhini |
RECOMB | 4 |
| 2005 | Dotted interval graphs and high throughput genotyping
Yonatan Aumann, Moshe Lewenstein, Oren Melamud, Ron Y. Pinter, Zohar Yakhini |
SODA | 5 |
| 2005 | Designing optimally multiplexed SNP genotyping assays
Yonatan Aumann, Efrat Manisterski, Zohar Yakhini |
J. Comput. Syst. Sci. | 3 |
| 2004 | Finding approximate tandem repeats in genomic sequencesabstractAn efficient algorithm is presented for detecting approximate tandem repeats in genomic sequences. The algorithm is based on a flexible statistical model which allows a wide range of definitions of approximate tandem repeats. The ideas and methods underlying the algorithm are described and examined and its effectiveness on genomic data is demonstrated. Ydo Wexler, Zohar Yakhini, Yechezkel Kashi, Dan Geiger |
RECOMB | 2 |
| 2004 | Joint Analysis of DNA Copy Numbers and Gene Expression Levels
Doron Lipson, Amir Ben-Dor, Elinor Dehan, Zohar Yakhini |
WABI | 4 |
| 2003 | Towards optimally multiplexed applications of universal DNA tag systemsabstractWe study a design and optimization problem that occurs, for example, when single nucleotide polymorphisms (SNPs) are to be genotyped using a universal DNA tag array. The problem of optimizing the universal array to avoid disruptive cross-hybridization between universal components of the system was addressed in a previous work. However, cross-hybridization can also occur assay-specifically, due to unwanted complementarity involving assay-specific components. Here we examine the problem of identifying the most economic experimental configuration of the assay-specific components that avoids cross-hybridization. Our formalization translates this problem into the problem of covering the vertices of one side of a bipartite graph by a minimum number of balanced subgraphs of maximum degree 1. We show that the general problem is NP-complete. However, in the real biological setting the vertices that need to be covered have degrees bounded by d. We exploit this restriction and develop an O(d)-approximation algorithm for the problem. We also give an O(d)-approximation for a variant of the problem in which the covering subgraphs are required to be vertex-disjoint. In addition, we propose a stochastic model for the input data and use it to prove a lower bound on the cover size. We complement our theoretical analysis by implementing two heuristic approaches and testing their performance on simulated and real SNP data. Amir Ben-Dor, Tzvika Hartman, Benno Schwikowski, Roded Sharan, Zohar Yakhini |
RECOMB | 5 |
| 2003 | A Distance-Based Branch and Bound Feature Selection Algorithm
Ari Frank, Dan Geiger, Zohar Yakhini |
UAI | 3 |
| 2003 | Designing Optimally Multiplexed SNP Genotyping Assays
Yonatan Aumann, Efrat Manisterski, Zohar Yakhini |
WABI | 3 |
| 2002 | Discovering local structure in gene expression data: the order-preserving submatrix problemabstractThis paper concerns the discovery of patterns in gene expression matrices, in which each element gives the expression level of a given gene in a given experiment. Most existing methods for pattern discovery in such matrices are based on clustering genes by comparing their expression levels in all experiments, or clustering experiments by comparing their expression levels for all genes. Our work goes beyond such global approaches by looking for local patterns that manifest themselves when we focus simultaneously on a subset G of the genes and a subset T of the experiments. Specifically, we look for order-preserving submatrices (OPSMs), in which the expression levels of all genes induce the same linear ordering of the experiments (we show that the OPSM search problem is NP-hard in the worst case). Such a pattern might arise, for example, if the experiments in T represent distinct stages in the progress of a disease or in a cellular process, and the expression levels of all genes in G vary across the stages in the same way.We define a probabilistic model in which an OPSM is hidden within an otherwise random matrix. Guided by this model we develop an efficient algorithm for finding the hidden OPSM in the random matrix. In data generated according to the model the algorithm recovers the hidden OPSM with very high success rate. Application of the methods to breast cancer data seems to reveal significant local patterns.Our algorithm can be used to discover more than one OPSM within the same data set, even when these OPSMs overlap. It can also be adapted to handle relaxations and extensions of the OPSM condition. For example, we may allow the different rows of G x T to induce similar but not identical orderings of the columns, or we may allow the set T to include more than one representative of each stage of a biological process. Amir Ben-Dor, Benny Chor, Richard M. Karp, Zohar Yakhini |
RECOMB | 4 |
| 2002 | Designing Specific Oligonucleotide Probes for the Entire S. cerevisiae Transcriptome
Doron Lipson, Peter Webb, Zohar Yakhini |
WABI | 3 |
| 2001 | Class discovery in gene expression dataabstractRecent studies (Alizadeh et al, [1]; Bittner et al,[5]; Golub et al, [11]) demonstrate the discovery of putative disease subtypes from gene expression data. The underlying computational problem is to partition the set of sample tissues into statistically meaningful classes. In this paper we present a novel approach to class discovery and develop automatic analysis methods. Our approach is based on statistically scoring candidate partitions according to the overabundance of genes that separate the different classes. Indeed, in biological datasets, an overabundance of genes separating known classes is typically observed. we measure overabundance against a stochastic null model. This allows for highlighting subtle, yet meaningful, partitions that are supported on a small subset of the genes. Amir Ben-Dor, Nir Friedman, Zohar Yakhini |
RECOMB | 3 |
| 2000 | Tissue classification with gene expression profilesabstractConstantly improving gene expression profiling technologies are expected to provide understanding and insight into cancer related cellular processes. Gene expression data is also expected to significantly and in the development of efficient cancer diagnosis and classification platforms. In this work we examine two sets of gene expression data measured across sets of tumor and normal clinical samples One set consists of 2,000 genes, measured in 62 epithelial colon samples [1]. The second consists of ≈ 100,000 clones, measured in 32 ovarian samples (unpublished, extension of data set described in [26]). Amir Ben-Dor, Laurakay Bruhn, Nir Friedman, Iftach Nachman, Michèl Schummer, Zohar Yakhini |
RECOMB | 6 |
| 2000 | Universal DNA tag systems: a combinatorial design schemeabstractCustom-designed DNA arrays offer the possibility of simultaneously monitoring thousands of hybridization reactions These arrays show great potential for many medical and scientific applications such as polymorphism analysis and genotyping. Relatively high costs are associated with the need to specifically design and synthesize problem specific arrays. Recently, an alternative approach was suggested that utilizes fixed, universal arrays. This approach presents an interesting design problem—the arrays should contain as many probes as possible, while minimizing experimental errors caused by cross-hybridization. We use a simple thermodynamic model to cast this design problem in a formal mathematical framework. Employing new combinatorial ideas, we derive an efficient construction for the design problem, and prove that our construction is near-optimal. Amir Ben-Dor, Richard M. Karp, Benno Schwikowski, Zohar Yakhini |
RECOMB | 4 |
| 1999 | Clustering gene expression patternsabstractWith the advance of hybridization array technology researchers can measure expression levels of sets of genes across different conditions and over time.Analysis of data produced by such experiments offers potential insight into gene function and regulatory mechanisms.We describe the problem of clustering multi-condition gene expression patterns.We define an appropriate stochastic model of the input, and use this model for performance evaluations.We present a O(n(log(n))c)time algorithm that recovers cluster structures with high probability, in this model, where n is the number of genes.In addition to the theoretical treatment, we suggest a practical heuristic approach based on the same ideas.We demonstrate the algorithm's performance first on simulated data, and then on actual gene expression data. Amir Ben-Dor, Zohar Yakhini |
RECOMB | 2 |
| 1996 | On the Sample Complexity of Learning Bayesian Networks
Nir Friedman, Zohar Yakhini |
UAI | 2 |