EDBT 2026 Demo / reviewers in the wild / expert
Haim Schweitzer
dblp:s/HaimSchweitzer · also Haim Shvaytser
· DBLP profile ↗
9ranked-venue papers in the field
3as first author
4since 2021 · last 2024
—ORCID · none
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5 (1 first)Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The art of centering without centering for robust principal component analysis
Guihong Wan, Baokun He, Haim Schweitzer |
Data Min. Knowl. Discov. | 3 |
| 2021 | Edge Sparsification for Graphs via Meta-LearningabstractWe present a novel edge sparsification approach for semi-supervised learning on undirected and attributed graphs. The main challenge is to retain few edges while minimizing the loss of node classification accuracy. The task can be mathematically formulated as a bi-level optimization problem. We propose to use meta-gradients, which have traditionally been used in meta-learning, to solve the optimization problem, specifically, treating the graph adjacency matrix as hyperparameters to optimize. Experimental results show the effectiveness of the proposed approach. Remarkably, with the resulting sparse and light graph, in many cases the classification accuracy is significantly improved. Guihong Wan, Haim Schweitzer |
ICDE | 2 |
| 2021 | A Lookahead Algorithm for Robust Subspace RecoveryabstractA common task in the analysis of data is to compute an approximate embedding of the data in a low-dimensional subspace. The standard algorithm for computing this subspace is the well-known Principal Component Analysis (PCA). PCA can be extended to the case where some data points are viewed as “outliers” that can be ignored, allowing the remaining data points (inliers”) to be more tightly embedded. We develop a new algorithm that detects outliers so that they can be removed prior to applying PCA. The main idea is to rank each point by looking ahead and evaluating the change in the global PCA error if that point is considered as an outlier. Our technical contribution is showing that this lookahead procedure can be implemented efficiently, producing an accurate algorithm with running time not much above the running time of standard PCA algorithms. Guihong Wan, Haim Schweitzer |
ICDM | 2 |
| 2021 | A Fast Algorithm for Simultaneous Sparse Approximation
Guihong Wan, Haim Schweitzer |
PAKDD (3) | 2 |
| 2020 | Fast Distance Metrics in Low-dimensional Space for Neighbor Search ProblemsabstractWe consider popular dimension reduction techniques that project data on a low dimensional subspace. They include Principal Component Analysis, Column Subset Selection, and Johnson-Lindenstrauss projections. These techniques have been classically used to efficiently compute various approximations. We propose the following three-step procedure for enhancing the accuracy of such approximations: 1. Unknown quantities in the approximation are replaced with random variables. 2. The Maximum Entropy method is applied to infer the most likely probability distribution. 3. Expected values of the random variables are used to compute the enhanced estimates. Our use of the Maximum Entropy method requires knowledge of vector norms that can be easily computed during the dimension reduction. We demonstrate significant enhancements in average accuracy for Euclidean distance and Mahalanobis distance, and improvements in evaluating k-nearest neighbors and k-furthest neighbors by using the enhanced Euclidean distance formula. Guihong Wan, Crystal Maung, Haim Schweitzer |
ICDM | 4 |
| 2015 | Improved Greedy Algorithms for Sparse Approximation of a Matrix in Terms of Another MatrixabstractWe consider simultaneously approximating all the columns of a data matrix in terms of few selected columns of another matrix that is sometimes called “the dictionary”. The challenge is to determine a small subset of the dictionary columns that can be used to obtain an accurate prediction of the entire data matrix. Previously proposed greedy algorithms for this task compare each data column with all dictionary columns, resulting in algorithms that may be too slow when both the data matrix and the dictionary matrix are large. A recent approach for accelerating the run time requires large amounts of memory to keep temporary values during the run of the algorithm. We propose two new algorithms that can be used even when both the data matrix and the dictionary matrix are large. The first algorithm is exact, with output identical to some previously proposed greedy algorithms. It takes significantly less memory when compared to the current state-of-the-art, and runs much faster when the dictionary matrix is sparse. The second algorithm uses a low rank approximation to the data matrix to further improve the run time. The algorithms use new recursive formulas for computing the greedy selection criterion. The formulas enable decoupling most of the computations related to the data matrix from the computations related to the dictionary matrix. Crystal Maung, Haim Schweitzer |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1998 | Computational limitations of model-based recognitionabstractReliable object recognition is an essential part of most visual systems. Model-based approaches to object recognition use a database (a library) of modeled objects; for a given set of sensed data, the problem of model-based recognition is to identify and locate the objects from the library that are present in the data. We show that the complexity of model-based recognition depends very heavily on the number of object models in the library even if each object is modeled by a small number of discrete features. Specifically, deciding whether a discrete set of sensed data can be interpreted as transformed object models from a given library is NP-complete if the transformation is any combination of translation, rotation, scaling, and perspective projection. This suggests that efficient algorithms for model-based recognition must use additional structure to avoid the inherent computational difficulties. © 1998 John Wiley & Sons, Inc. Haim Schweitzer, Sanjeev R. Kulkarni |
Int. J. Intell. Syst. | 1 |
| 1997 | A Distributed Algorithm for Content Based Indexing of Images by Projections on Ritz Primary Images
Haim Schweitzer |
Data Min. Knowl. Discov. | 1 |
| 1985 | Fuzzy and probability vectors as elements of a vector space
Haim Schweitzer, Shmuel Peleg |
Inf. Sci. | 1 |