Sándor Szedmák

dblp:89/472 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-1469-2215ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Systems, architecture and hardware · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Databases, data management, data science and information retrieval · 2Theory of computation · 2
YearPublicationVenuePosition
2025 Efficient Microservice Monitoring Via Kernel Transformation and FFT Forecasting
abstract
Change point detection (CPD) is crucial for identifying abrupt shifts in time series data across various domains. In microservice architectures, where components operate at high velocity and scale, monitoring efficiency becomes particularly critical for maintaining system reliability while minimizing over-head. Efficient change point detection allows for timely root cause analysis, dynamic resource allocation, and targeted instance restarts, interventions that can prevent cascading failures across interconnected services. This paper presents a novel approach called MMR-FFT CPD that transforms high-dimensional mi-croservice telemetry data into a one-dimensional kernel-based representation through Maximum Margin Regression (MMR). The resulting 1D time series is then processed using Fast Fourier Transform (FFT) for efficient temporal pattern forecasting, which enables prediction of expected future values. By computing the correlation between these FFT-based predictions and the actual observed values, our method can identify deviations that indicate potential change points in microservice behavior. By tracking only these compact representations rather than raw metrics, our method dramatically reduces computational overhead and network traffic in distributed monitoring scenarios. We evaluated MMR-FFT CPD against multiple state-of-the-art algorithms. Our experiments demonstrate that with a 1D data transfer to the monitoring entity, we can maintain competitive detection accuracy, which is advantageous, particularly for high-dimensional data streams typical in modern distributed systems. The results provide valuable insights for implementing efficient, low-overhead monitoring in microservice environments with varying computational and network constraints, enabling more resnonsive and resilient microservice architectures.
Marianna Ojanen, Maryam Sabzevari, Sándor Szedmák
CLOUD3
2025 Scaling up drug combination surface prediction
abstract
Drug combinations are required to treat advanced cancers and other complex diseases. Compared with monotherapy, combination treatments can enhance efficacy and reduce toxicity by lowering the doses of single drugs-and there especially synergistic combinations are of interest. Since drug combination screening experiments are costly and time-consuming, reliable machine learning models are needed for prioritizing potential combinations for further studies. Most of the current machine learning models are based on scalar-valued approaches, which predict individual response values or synergy scores for drug combinations. We take a functional output prediction approach, in which full, continuous dose-response combination surfaces are predicted for each drug combination on the cell lines. We investigate the predictive power of the recently proposed comboKR method, which is based on a powerful input-output kernel regression technique and functional modeling of the response surface. In this work, we develop a scaled-up formulation of the comboKR, which also implements improved modeling choices: we (1) incorporate new modeling choices for the output drug combination response surfaces to the comboKR framework, and (2) propose a projected gradient descent method to solve the challenging pre-image problem that is traditionally solved with simple candidate set approaches. We provide thorough experimental analysis of comboKR 2.0 with three real-word datasets within various challenging experimental settings, including cases where drugs or cell lines have not been encountered in the training data. Our comparison with synergy score prediction methods further highlights the relevance of dose-response prediction approaches, instead of relying on simple scoring methods.
Riikka Huusari, Tianduanyi Wang, Sándor Szedmák, Diogo Dias, Tero Aittokallio, Juho Rousu
Briefings Bioinform.3
2024 Protein function prediction through multi-view multi-label latent tensor reconstruction
abstract
BACKGROUND: In last two decades, the use of high-throughput sequencing technologies has accelerated the pace of discovery of proteins. However, due to the time and resource limitations of rigorous experimental functional characterization, the functions of a vast majority of them remain unknown. As a result, computational methods offering accurate, fast and large-scale assignment of functions to new and previously unannotated proteins are sought after. Leveraging the underlying associations between the multiplicity of features that describe proteins could reveal functional insights into the diverse roles of proteins and improve performance on the automatic function prediction task. RESULTS: We present GO-LTR, a multi-view multi-label prediction model that relies on a high-order tensor approximation of model weights combined with non-linear activation functions. The model is capable of learning high-order relationships between multiple input views representing the proteins and predicting high-dimensional multi-label output consisting of protein functional categories. We demonstrate the competitiveness of our method on various performance measures. Experiments show that GO-LTR learns polynomial combinations between different protein features, resulting in improved performance. Additional investigations establish GO-LTR's practical potential in assigning functions to proteins under diverse challenging scenarios: very low sequence similarity to previously observed sequences, rarely observed and highly specific terms in the gene ontology. IMPLEMENTATION: The code and data used for training GO-LTR is available at https://github.com/aalto-ics-kepaco/GO-LTR-prediction .
Robert Ebo Armah-Sekum, Sándor Szedmák, Juho Rousu
BMC Bioinform.2
2024 Scalable variable selection for two-view learning tasks with projection operators
abstract
Abstract In this paper we propose a novel variable selection method for two-view settings, or for vector-valued supervised learning problems. Our framework is able to handle extremely large scale selection tasks, where number of data samples could be even millions. In a nutshell, our method performs variable selection by iteratively selecting variables that are highly correlated with the output variables, but which are not correlated with the previously chosen variables. To measure the correlation, our method uses the concept of projection operators and their algebra. With the projection operators the relationship, correlation, between sets of input and output variables can also be expressed by kernel functions, thus nonlinear correlation models can be exploited as well. We experimentally validate our approach, showing on both synthetic and real data its scalability and the relevance of the selected features.
Sándor Szedmák, Riikka Huusari, Tat Hong Duong Le, Juho Rousu
Mach. Learn.1
2022 Strain design optimization using reinforcement learning
abstract
Engineered microbial cells present a sustainable alternative to fossil-based synthesis of chemicals and fuels. Cellular synthesis routes are readily assembled and introduced into microbial strains using state-of-the-art synthetic biology tools. However, the optimization of the strains required to reach industrially feasible production levels is far less efficient. It typically relies on trial-and-error leading into high uncertainty in total duration and cost. New techniques that can cope with the complexity and limited mechanistic knowledge of the cellular regulation are called for guiding the strain optimization. In this paper, we put forward a multi-agent reinforcement learning (MARL) approach that learns from experiments to tune the metabolic enzyme levels so that the production is improved. Our method is model-free and does not assume prior knowledge of the microbe's metabolic network or its regulation. The multi-agent approach is well-suited to make use of parallel experiments such as multi-well plates commonly used for screening microbial strains. We demonstrate the method's capabilities using the genome-scale kinetic model of Escherichia coli, k-ecoli457, as a surrogate for an in vivo cell behaviour in cultivation experiments. We investigate the method's performance relevant for practical applicability in strain engineering i.e. the speed of convergence towards the optimum response, noise tolerance, and the statistical stability of the solutions found. We further evaluate the proposed MARL approach in improving L-tryptophan production by yeast Saccharomyces cerevisiae, using publicly available experimental data on the performance of a combinatorial strain library. Overall, our results show that multi-agent reinforcement learning is a promising approach for guiding the strain optimization beyond mechanistic knowledge, with the goal of faster and more reliably obtaining industrially attractive production levels.
Maryam Sabzevari, Sándor Szedmák, Merja Penttilä, Paula Jouhten, Juho Rousu
PLoS Comput. Biol.2
2021 Modeling drug combination effects via latent tensor reconstruction
abstract
MOTIVATION: Combination therapies have emerged as a powerful treatment modality to overcome drug resistance and improve treatment efficacy. However, the number of possible drug combinations increases very rapidly with the number of individual drugs in consideration, which makes the comprehensive experimental screening infeasible in practice. Machine-learning models offer time- and cost-efficient means to aid this process by prioritizing the most effective drug combinations for further pre-clinical and clinical validation. However, the complexity of the underlying interaction patterns across multiple drug doses and in different cellular contexts poses challenges to the predictive modeling of drug combination effects. RESULTS: We introduce comboLTR, highly time-efficient method for learning complex, non-linear target functions for describing the responses of therapeutic agent combinations in various doses and cancer cell-contexts. The method is based on a polynomial regression via powerful latent tensor reconstruction. It uses a combination of recommender system-style features indexing the data tensor of response values in different contexts, and chemical and multi-omics features as inputs. We demonstrate that comboLTR outperforms state-of-the-art methods in terms of predictive performance and running time, and produces highly accurate results even in the challenging and practical inference scenario where full dose-response matrices are predicted for completely new drug combinations with no available combination and monotherapy response measurements in any training cell line. AVAILABILITY AND IMPLEMENTATION: comboLTR code is available at https://github.com/aalto-ics-kepaco/ComboLTR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianduanyi Wang, Sándor Szedmák, Haishan Wang 0001, Tero Aittokallio, Tapio Pahikkala, Anna Cichonska, Juho Rousu
Bioinform.2
2020 Using Machine Learning for Decreasing State Uncertainty in Planning
abstract
We present a novel approach for decreasing state uncertainty in planning prior to solving the planning problem. This is done by making predictions about the state based on currently known information, using machine learning techniques. For domains where uncertainty is high, we define an active learning process for identifying which information, once sensed, will best improve the accuracy of predictions. We demonstrate that an agent is able to solve problems with uncertainties in the state with less planning effort compared to standard planning techniques. Moreover, agents can solve problems for which they could not find valid plans without using predictions. Experimental results also demonstrate that using our active learning process for identifying information to be sensed leads to gathering information that improves the prediction process.
Senka Krivic, Michael Cashmore, Daniele Magazzeni, Sándor Szedmák, Justus H. Piater
J. Artif. Intell. Res.4
2018 Liquid-chromatography retention order prediction for metabolite identification
abstract
Motivation: Liquid Chromatography (LC) followed by tandem Mass Spectrometry (MS/MS) is one of the predominant methods for metabolite identification. In recent years, machine learning has started to transform the analysis of tandem mass spectra and the identification of small molecules. In contrast, LC data is rarely used to improve metabolite identification, despite numerous published methods for retention time prediction using machine learning. Results: We present a machine learning method for predicting the retention order of molecules; that is, the order in which molecules elute from the LC column. Our method has important advantages over previous approaches: We show that retention order is much better conserved between instruments than retention time. To this end, our method can be trained using retention time measurements from different LC systems and configurations without tedious pre-processing, significantly increasing the amount of available training data. Our experiments demonstrate that retention order prediction is an effective way to learn retention behaviour of molecules from heterogeneous retention time data. Finally, we demonstrate how retention order prediction and MS/MS-based scores can be combined for more accurate metabolite identifications when analyzing a complete LC-MS/MS run. Availability and implementation: Implementation of the method is available at https://version.aalto.fi/gitlab/bache1/retention_order_prediction.git.
Eric Bach 0002, Sándor Szedmák, Céline Brouard, Sebastian Böcker, Juho Rousu
Bioinform.2
2018 Learning with multiple pairwise kernels for drug bioactivity prediction
abstract
Motivation: Many inference problems in bioinformatics, including drug bioactivity prediction, can be formulated as pairwise learning problems, in which one is interested in making predictions for pairs of objects, e.g. drugs and their targets. Kernel-based approaches have emerged as powerful tools for solving problems of that kind, and especially multiple kernel learning (MKL) offers promising benefits as it enables integrating various types of complex biomedical information sources in the form of kernels, along with learning their importance for the prediction task. However, the immense size of pairwise kernel spaces remains a major bottleneck, making the existing MKL algorithms computationally infeasible even for small number of input pairs. Results: We introduce pairwiseMKL, the first method for time- and memory-efficient learning with multiple pairwise kernels. pairwiseMKL first determines the mixture weights of the input pairwise kernels, and then learns the pairwise prediction function. Both steps are performed efficiently without explicit computation of the massive pairwise matrices, therefore making the method applicable to solving large pairwise learning problems. We demonstrate the performance of pairwiseMKL in two related tasks of quantitative drug bioactivity prediction using up to 167 995 bioactivity measurements and 3120 pairwise kernels: (i) prediction of anticancer efficacy of drug compounds across a large panel of cancer cell lines; and (ii) prediction of target profiles of anticancer compounds across their kinome-wide target spaces. We show that pairwiseMKL provides accurate predictions using sparse solutions in terms of selected kernels, and therefore it automatically identifies also data sources relevant for the prediction problem. Availability and implementation: Code is available at https://github.com/aalto-ics-kepaco. Supplementary information: Supplementary data are available at Bioinformatics online.
Anna Cichonska, Tapio Pahikkala, Sándor Szedmák, Heli Julkunen, Antti Airola, Markus Heinonen, Tero Aittokallio, Juho Rousu
Bioinform.3
2017 Decreasing Uncertainty in Planning with State Prediction
abstract
In real world environments the state is almost never completely known. Exploration is often expensive. The application of planning in these environments is consequently more difficult and less robust. In this paper we present an approach for predicting new information about a partially-known state. The state is translated into a partially-known multigraph, which can then be extended using machine-learning techniques. We demonstrate the effectiveness of our approach, showing that it enhances the scalability of our planners, and leads to less time spent on sensing actions.
Senka Krivic, Michael Cashmore, Daniele Magazzeni, Bram Ridder, Sándor Szedmák, Justus H. Piater
IJCAI5
2017 Utilising Kronecker Decomposition and Tensor-based Multi-view Learning to predict where people are looking in images
Kitsuchart Pasupa, Sándor Szedmák
Neurocomputing2
2016 Soft Kernel Target Alignment for Two-Stage Multiple Kernel Learning
Huibin Shen, Sándor Szedmák, Céline Brouard, Juho Rousu
DS2
2016 Robotic playing for hierarchical complex skill learning
abstract
In complex manipulation scenarios (e.g. tasks requiring complex interaction of two hands or in-hand manipulation), generalization is a hard problem. Current methods still either require a substantial amount of (supervised) training data and / or strong assumptions on both the environment and the task. In this paradigm, controllers solving these tasks tend to be complex. We propose a paradigm of maintaining simpler controllers solving the task in a small number of specific situations. In order to generalize to novel situations, the robot transforms the environment from novel situations into a situation where the solution of the task is already known. Our solution to this problem is to play with objects and use previously trained skills (basis skills). These skills can either be used for estimating or for changing the current state of the environment and are organized in skill hierarchies. The approach is evaluated in complex pick-and-place scenarios that involve complex manipulation. We further show that these skills can be learned by autonomous playing.
Simon Hangl, Emre Ugur, Sándor Szedmák, Justus H. Piater
IROS3
2016 Learning undirected graphical models using persistent sequential Monte Carlo
Hanchen Xiong, Sándor Szedmák, Justus H. Piater
Mach. Learn.2
2015 Multi-label Object Categorization Using Histograms of Global Relations
abstract
In this paper, we present an object categorization system capable of assigning multiple and related categories for novel objects using multi-label learning. In this system, objects are described using global geometric relations of 3D features. We propose using the Joint SVM method for learning and we investigate the extraction of hierarchical clusters as a higher-level description of objects to assist the learning. We make comparisons with other multi-label learning approaches as well as single-label approaches (including a state-of-the-art methods using different object descriptors). The experiments are carried out on a dataset of 100 objects belonging to 13 visual and action-related categories. The results indicate that multi-label methods are able to identify the relation between the dependent categories and hence perform categorization accordingly. It is also found that extracting hierarchical clusters does not lead to gain in the system's performance. The results also show that using histograms of global relations to describe objects leads to fast learning in terms of the number of samples required for training.
Wail Mustafa, Hanchen Xiong, Dirk Kraft, Sándor Szedmák, Justus H. Piater, Norbert Krüger
3DV4
2015 Can Computer Vision Problems Benefit from Structured Hierarchical Classification?
Thomas Hoyoux, Antonio Jose Rodríguez-Sánchez, Justus H. Piater, Sándor Szedmák
CAIP (2)4
2015 Learning missing edges via kernels in partially-known graphs
Senka Krivic, Sándor Szedmák, Hanchen Xiong, Justus H. Piater
ESANN2
2015 Learning to Predict Where People Look with Tensor-Based Multi-view Learning
Kitsuchart Pasupa, Sándor Szedmák
ICONIP (1)2
2015 Using structural bootstrapping for object substitution in robotic executions of human-like manipulation tasks
abstract
In this work we address the problem of finding replacements of missing objects that are needed for the execution of human-like manipulation tasks. This is a usual problem that is easily solved by humans provided their natural knowledge to find object substitutions: using a knife as a screwdriver or a book as a cutting board. On the other hand, in robotic applications, objects required in the task should be included in advance in the problem definition. If any of these objects is missing from the scenario, the conventional approach is to manually redefine the problem according to the available objects in the scene. In this work we propose an automatic way of finding object substitutions for the execution of manipulation tasks. The approach uses a logic-based planner to generate a plan from a prototypical problem definition and searches for replacements in the scene when some of the objects involved in the plan are missing. This is done by means of a repository of objects and attributes with roles, which is used to identify the affordances of the unknown objects in the scene. Planning actions are grounded using a novel approach that encodes the semantic structure of manipulation actions. The system was evaluated in a KUKA arm platform for the task of preparing a salad with successful results.
Alejandro Agostini, Mohamad Javad Aein, Sándor Szedmák, Eren Erdal Aksoy, Justus H. Piater, Florentin Wörgötter
IROS3
2015 SCurV: A 3D descriptor for object classification
abstract
3D Object recognition is one of the big problems in Computer Vision which has a direct impact in Robotics. There have been great advances in the last decade thanks to point cloud descriptors. These descriptors do very well at recognizing object instances in a wide variety of situations. Of great interest is also to know how descriptors perform in object classification tasks. With that idea in mind, we introduce a descriptor designed for the representation of object classes. Our descriptor, named SCurV, exploits 3D shape information and is inspired by recent findings from neurophysiology. We compute and incorporate surface curvatures and distributions of local surface point projections that represent flatness, concavity and convexity in a 3D object-centered and view-dependent descriptor. These different sources of information are combined in a novel and simple, yet effective, way of combining different features to improve classification results which can be extended to the combination of any type of descriptor. Our experimental setup compares SCurV with other recent descriptors on a large classification task. Using a large and heterogeneous database of 3D objects, we perform our experiments both on a classical, flat classification task and within a novel framework for hierarchical classification. On both tasks, the SCurV descriptor outperformed all other 3D descriptors tested.
Antonio Jose Rodríguez-Sánchez, Sándor Szedmák, Justus H. Piater
IROS2
2015 Scalable, accurate image annotation with joint SVMs and output kernels
Hanchen Xiong, Sándor Szedmák, Justus H. Piater
Neurocomputing2
2014 Towards Maximum Likelihood: Learning Undirected Graphical Models using Persistent Sequential Monte Carlo
Hanchen Xiong, Sándor Szedmák, Justus H. Piater
ACML2
2014 Joint SVM for Accurate and Fast Image Tagging
Hanchen Xiong, Sándor Szedmák, Justus H. Piater
ESANN2
2014 Towards Sparsity and Selectivity: Bayesian Learning of Restricted Boltzmann Machine for Early Visual Features
Hanchen Xiong, Sándor Szedmák, Antonio Jose Rodríguez-Sánchez, Justus H. Piater
ICANN2
2014 Knowledge propagation and relation learning for predicting action effects
abstract
Learning to predict the effects of actions applied to pairs of objects is a difficult task that requires learning complex relations with sparse, incomplete and noisy information. Our Knowledge Propagation approach propagates affordance predictions by exploiting similarities among object properties, action parameters and resulting effects. The knowledge is propagated in a graph where a missing edge, corresponding to an unknown interaction between two objects (nodes), is predicted via the superposition of all paths connecting those objects in the graph. The high complexity of affordance representation is addressed through the use of Maximum Margin Multi-Valued Regression (MMMVR), which scales well to complex problems of multiple layers. With increased diversity and size of object databases and the addition of other parametric combinatory actions, we expect to achieve complex systems that leverage learned structure for subsequent learning, achieving structural bootstrapping over lifelong development and learning. In this paper, we extend MMMVR for learning of paired-object affordances, i.e., for predicting the effects of actions applied to pairs of objects. In our experiments, we evaluated this method on a dataset composed of 83 objects and 83×83 interactions. We compared the prediction performance with standard classifiers that predict the effect category given the object pair's low-level features or single-object affordances. The experiments show that our proposed method achieves significantly higher prediction performance especially when supported with Active Learning.
Sándor Szedmák, Emre Ugur, Justus H. Piater
IROS1
2013 A Study of Point Cloud Registration with Probability Product Kernel Functions
abstract
3D point cloud registration is an essential problem in 3D object and scene understanding. In many realistic circumstances, however, because of noise during data acquisition and large motion between two point clouds, most existing approaches can hardly work satisfactorily without good initial alignment or manually marked correspondences. Inspired by the popular kernel methods in machine learning community, this paper puts forward a general point cloud registration framework by constructing kernel functions over 3D point clouds. More specifically, Gaussian mixtures Based on the point clouds are established and probability product kernel functions are exploited for the registration. To enhance the generality of the framework, SE(3) on-manifold optimization scheme is employed to compute the optimal motion. Experimental results show that our registration framework works robustly when many outliers are presented and motion between point clouds is relatively large, and compares favorably to related methods.
Hanchen Xiong, Sándor Szedmák, Justus H. Piater
3DV2
2012 Kernel-Mapping Recommender system algorithms
Mustansar Ali Ghazanfar, Adam Prügel-Bennett, Sándor Szedmák
Inf. Sci.3
2011 Incremental Kernel Mapping Algorithms for Scalable Recommender Systems
abstract
Recommender systems apply machine learning techniques for filtering unseen information and can predict whether a user would like a given item. Kernel Mapping Recommender (KMR) system algorithms have been proposed, which offer state-of-the-art performance. One potential drawback of the KMR algorithms is that the training is done in one step and hence they cannot accommodate the incremental update with the arrival of new data making them unsuitable for the dynamic environments. From this line of research, we propose a new heuristic, which can build the model incrementally without retraining the whole model from scratch when new data (item or user) are added to the recommender system dataset. Furthermore, we proposed a novel perceptron-type algorithm, which is a fast incremental algorithm for building the model that maintains a good level of accuracy and scales well with the data. We show empirically over two datasets that the proposed algorithms give quite accurate results while providing significant computation savings.
Mustansar Ali Ghazanfar, Sándor Szedmák, Adam Prügel-Bennett
ICTAI2
2011 Exploitation of Machine Learning Techniques in Modelling Phrase Movements for Machine Translation
Yizhao Ni, Craig Saunders, Sándor Szedmák, Mahesan Niranjan
J. Mach. Learn. Res.3
2010 The application of structured learning in natural language processing
Yizhao Ni, Craig Saunders, Sándor Szedmák, Mahesan Niranjan
Mach. Transl.3
2007 A metamorphosis of Canonical Correlation Analysis into multivariate maximum margin learning
Sándor Szedmák, Tijl De Bie, David R. Hardoon
ESANN1
2007 Synthesis of maximum margin and multiview learning using unlabeled data
Sándor Szedmák, John Shawe-Taylor
Neurocomputing1
2006 A Correlation Approach for Automatic Image Annotation
David R. Hardoon, Craig Saunders, Sándor Szedmák, John Shawe-Taylor
ADMA3
2006 Synthesis of maximum margin and multiview learning using unlabeled data
Sándor Szedmák, John Shawe-Taylor
ESANN1
2006 Kernel-Based Learning of Hierarchical Multilabel Classification Models
abstract
We present a kernel-based algorithm for hierarchical text classification where the documents are allowed to belong to more than one category at a time. The classification model is a variant of the Maximum Margin Markov Network framework, where the classification hierarchy is represented as a Markov tree equipped with an exponential family defined on the edges. We present an efficient optimization algorithm based on incremental conditional gradient ascent in single-example subspaces spanned by the marginal dual variables. The optimization is facilitated with a dynamic programming based algorithm that computes best update directions in the feasible set. Experiments show that the algorithm can feasibly optimize training sets of thousands of examples and classification hierarchies consisting of hundreds of nodes. Training of the full hierarchical model is as efficient as training independent SVM-light classifiers for each node. The algorithm's predictive accuracy was found to be competitive with other recently introduced hierarchical multi-category or multilabel classification learning algorithms.
Juho Rousu, Craig Saunders, Sándor Szedmák, John Shawe-Taylor
J. Mach. Learn. Res.3
2005 Learning hierarchical multi-category text classification models
abstract
We present a kernel-based algorithm for hierarchical text classification where the documents are allowed to belong to more than one category at a time. The classification model is a variant of the Maximum Margin Markov Network framework, where the classification hierarchy is represented as a Markov tree equipped with an exponential family defined on the edges. We present an efficient optimization algorithm based on incremental conditional gradient ascent in single-example subspaces spanned by the marginal dual variables. Experiments show that the algorithm can feasibly optimize training sets of thousands of examples and classification hierarchies consisting of hundreds of nodes. The algorithm's predictive accuracy is competitive with other recently introduced hierarchical multi-category or multilabel classification learning algorithms.
Juho Rousu, Craig Saunders, Sándor Szedmák, John Shawe-Taylor
ICML3
2005 Two view learning: SVM-2K, Theory and Practice
abstract
Kernel methods make it relatively easy to define complex highdimensional feature spaces. This raises the question of how we can identify the relevant subspaces for a particular learning task. When two views of the same phenomenon are available kernel Canonical Correlation Analysis (KCCA) has been shown to be an effective preprocessing step that can improve the performance of classification algorithms such as the Support Vector Machine (SVM). This paper takes this observation to its logical conclusion and proposes a method that combines this two stage learning (KCCA followed by SVM) into a single optimisation termed SVM-2K. We present both experimental and theoretical analysis of the approach showing encouraging results and insights.
Jason D. R. Farquhar, David R. Hardoon, Hongying Meng, John Shawe-Taylor, Sándor Szedmák
NIPS5
2004 Pareto-optimal patterns in logical analysis of data
Peter L. Hammer, Alexander Kogan, Bruno Simeone, Sándor Szedmák
Discret. Appl. Math.4
2004 Saturated systems of homogeneous boxes and the logical analysis of numerical data
Peter L. Hammer, Yanpei Liu, Bruno Simeone, Sándor Szedmák
Discret. Appl. Math.4
2004 Canonical Correlation Analysis: An Overview with Application to Learning Methods
abstract
We present a general method using kernel canonical correlation analysis to learn a semantic representation to web images and their associated text. The semantic space provides a common representation and enables a comparison between the text and images. In the experiments, we look at two approaches of retrieving images based on only their content from a text query. We compare orthogonalization approaches against a standard cross-representation retrieval technique known as the generalized vector space model.
David R. Hardoon, Sándor Szedmák, John Shawe-Taylor
Neural Comput.2