EDBT 2026 Demo / reviewers in the wild / expert
Panagiotis Patrinos
dblp:55/896 · also Panos Patrinos
· DBLP profile ↗
24ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-4824-7697ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantization-aware matrix factorization for low bit rate image compressionabstractLossy image compression is essential for efficient transmission and storage. Traditional compression methods mainly rely on discrete cosine transform (DCT) or singular value decomposition (SVD), both of which represent image data in continuous domains and, therefore, necessitate carefully designed quantizers. Notably, these methods consider quantization as a separate step, which prevents quantization errors from being incorporated into the compression process and degrades the reconstruction quality, particularly in SVD-based methods. To address this issue, we introduce a quantization-aware matrix factorization (QMF) to develop a novel lossy image compression method. QMF provides a low-rank representation of the image data as a product of two smaller matrices, with elements constrained to bounded integer values, thereby effectively integrating quantization with low-rank approximation. We propose an efficient, provably convergent iterative algorithm for QMF using a block coordinate descent scheme, with subproblems having closed-form solutions. Our experiments demonstrate that our method consistently outperforms JPEG at low bit rates below 0.25 bits per pixel. We also demonstrated that our method has an improved capability to preserve visual semantics compared to JPEG at low bit rates by evaluating an ImageNet pre-trained classifier on compressed images. The project is available at https://github.com/pashtari/qmf . Pooya Ashtari, Pourya Behmandpoor, Fateme Nateghi Haredasht, Jonathan H. Chen, Panagiotis Patrinos, Sabine Van Huffel |
Inf. Sci. | 5 |
| 2026 | A Deep Learning-Based Resource Allocator for Communication Networks With Dynamic User Utility DemandsabstractDeep learning (DL) based resource allocation (RA) has recently gained significant attention due to its performance efficiency. However, most related studies assume an ideal case where the number of users and their utility demands, e.g., data rate constraints, are fixed, and the designed DL-based RA scheme exploits a policy trained only for these fixed parameters. Consequently, computationally complex policy retraining is required whenever these parameters change. In this paper, we introduce a DL-based resource allocator (ALCOR) that allows users to adjust their utility demands freely, such as based on their application layer requirements. ALCOR employs deep neural networks (DNNs) as the policy in a time-sharing problem. The underlying optimization algorithm iteratively optimizes the on-off status of users to satisfy their utility demands in expectation. The policy performs unconstrained RA (URA)–—RA without considering user utility demands–—among active users to maximize the sum utility (SU) at each time instant. Depending on the chosen URA scheme, ALCOR can perform RA in either a centralized or distributed scenario. The derived convergence analyses provide theoretical guarantees for ALCOR’s convergence, and numerical experiments corroborate its effectiveness compared to meta-learning and reinforcement learning approaches. Pourya Behmandpoor, Mark Eisen, Panagiotis Patrinos, Marc Moonen |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Tight Analysis of Difference-of-Convex Algorithm (DCA) Improves Convergence Rates for Proximal Gradient DescentabstractWe investigate a difference-of-convex (DC) formulation where the second term is allowed to be weakly convex. We examine the precise behavior of a single iteration of the difference-of-convex algorithm (DCA), providing a tight characterization of the objective function decrease, distinguishing between six distinct parameter regimes. Our proofs, inspired by the performance estimation framework, are notably simplified compared to related prior research. We subsequently derive sublinear convergence rates for the DCA towards critical points, assuming at least one of the functions is smooth. Additionally, we explore the underexamined equivalence between proximal gradient descent (PGD) and DCA iterations, demonstrating how DCA, a parameter-free algorithm, without the need for a stepsize, serves as a tool for studying the exact convergence rates of PGD. Finally, we propose a method to optimize the DC decomposition to achieve optimal convergence rates, potentially transforming the subtracted function to become weakly convex. Teodor Rotaru, Panagiotis Patrinos, François Glineur |
AISTATS | 2 |
| 2025 | Nonlinearly Preconditioned Gradient Methods under Generalized SmoothnessabstractWe analyze nonlinearly preconditioned gradient methods for solving smooth minimization problems. We introduce a generalized smoothness property, based on the notion of abstract convexity, that is broader than Lipschitz smoothness and provide sufficient first- and second-order conditions. Notably, our framework encapsulates algorithms associated with the gradient clipping method and brings out novel insights for the class of $(L_0,L_1)$-smooth functions that has received widespread interest recently, thus allowing us to extend beyond already established methods. We investigate the convergence of the proposed method in both the convex and nonconvex setting. Konstantinos A. Oikonomidis, Jan Quan, Emanuel Laude, Panagiotis Patrinos |
ICML | 4 |
| 2025 | ExAMPC: the Data-Driven Explainable and Approximate NMPC with Physical InsightsabstractAmidst the surge in the use of Artificial Intelligence (AI) for control purposes, classical and model-based control methods maintain their popularity due to their transparency and deterministic nature. However, advanced controllers like Nonlinear Model Predictive Control (NMPC), despite proven capabilities, face adoption challenges due to their computational complexity and unpredictable closed-loop performance in complex validation systems. This paper introduces ExAMPC, a methodology bridging classical control and explainable AI by augmenting the NMPC with data-driven insights to improve the trustworthiness and reveal the optimization solution and closed-loop performance’s sensitivities to physical variables and system parameters. By employing a low-order spline embedding, we reduce the open-loop trajectory dimensionality by over 95%, and integrate it with SHAP and Symbolic Regression from eXplainable AI (XAI) for an approximate NMPC, enabling intuitive physical insights into the NMPC’s optimization routine. The prediction accuracy of the approximate NMPC is enhanced through physics-inspired continuous-time constraints penalties, reducing the predicted continuous trajectory violations by 93%. ExAMPC also enables accurate forecasting of the NMPC’s computational requirements with explainable insights on worst-case scenarios. Experimental validation on automated valet parking and autonomous racing with lap-time optimization, demonstrates the methodology’s practical effectiveness for potential real-world applications. Jean Pierre Allamaa, Panagiotis Patrinos, Tong Duy Son |
IROS | 2 |
| 2025 | Escaping saddle points without Lipschitz smoothness: the power of nonlinear preconditioningabstractWe study generalized smoothness in nonconvex optimization, focusing on $(L_0, L_1)$-smoothness and anisotropic smoothness. The former was empirically derived from practical neural network training examples, while the latter arises naturally in the analysis of nonlinearly preconditioned gradient methods. We introduce a new sufficient condition that encompasses both notions, reveals their close connection, and holds in key applications such as phase retrieval and matrix factorization. Leveraging tools from dynamical systems theory, we then show that nonlinear preconditioning—including gradient clipping—preserves the saddle point avoidance property of classical gradient descent. Crucially, the assumptions required for this analysis are actually satisfied in these applications, unlike in classical results that rely on restrictive Lipschitz smoothness conditions. We further analyze a perturbed variant that efficiently attains second-order stationarity with only logarithmic dependence on dimension, matching similar guarantees of classical gradient methods. Alexander Bodard, Panagiotis Patrinos |
NeurIPS | 2 |
| 2025 | Nonlinearly Preconditioned Gradient Methods: Momentum and Stochastic AnalysisabstractWe study nonlinearly preconditioned gradient methods for smooth nonconvex optimization problems, focusing on sigmoid preconditioners that inherently perform a form of gradient clipping akin to the widely used gradient clipping technique. Building upon this idea, we introduce a novel heavy ball-type algorithm and provide convergence guarantees under a generalized smoothness condition that is less restrictive than traditional Lipschitz smoothness, thus covering a broader class of functions. Additionally, we develop a stochastic variant of the base method and study its convergence properties under different noise assumptions. We compare the proposed algorithms with baseline methods on diverse tasks from machine learning including neural network training. Konstantinos A. Oikonomidis, Jan Quan, Panagiotis Patrinos |
NeurIPS | 3 |
| 2025 | Rethinking PCA Through DualityabstractMotivated by the recently shown connection between self-attention and (kernel) principal component analysis (PCA), we revisit the fundamentals of PCA. Using the difference-of-convex (DC) framework, we present several novel formulations and provide new theoretical insights. In particular, we show the kernelizability and out-of-sample applicability for a PCA-like family of problems. Moreover, we uncover that simultaneous iteration, which is connected to the classical QR algorithm, is an instance of the difference-of-convex algorithm (DCA), offering an optimization perspective on this longstanding method. Further, we describe new algorithms for PCA and empirically compare them with state-of-the-art methods. Lastly, we introduce a kernelizable dual formulation for a robust variant of PCA that minimizes the $l_1$-deviation of the reconstruction errors. Jan Quan, Johan A. K. Suykens, Panagiotis Patrinos |
NeurIPS | 3 |
| 2024 | Unsupervised Neighborhood Propagation Kernel Layers for Semi-supervised Node ClassificationabstractWe present a deep Graph Convolutional Kernel Machine (GCKM) for semi-supervised node classification in graphs. The method is built of two main types of blocks: (i) We introduce unsupervised kernel machine layers propagating the node features in a one-hop neighborhood, using implicit node feature mappings. (ii) We specify a semi-supervised classification kernel machine through the lens of the Fenchel-Young inequality. We derive an effective initialization scheme and efficient end-to-end training algorithm in the dual variables for the full architecture. The main idea underlying GCKM is that, because of the unsupervised core, the final model can achieve higher performance in semi-supervised node classification when few labels are available for training. Experimental results demonstrate the effectiveness of the proposed framework. Sonny Achten, Francesco Tonin, Panagiotis Patrinos, Johan A. K. Suykens |
AAAI | 3 |
| 2024 | Convex Relaxations for Manifold-Valued Markov Random Fields with Approximation Guarantees
Robin Kenis, Emanuel Laude, Panagiotis Patrinos |
ECCV (87) | 3 |
| 2024 | Adaptive Proximal Gradient Methods Are Universal Without ApproximationabstractWe show that adaptive proximal gradient methods for convex problems are not restricted to traditional Lipschitzian assumptions. Our analysis reveals that a class of linesearch-free methods is still convergent under mere local Hölder gradient continuity, covering in particular continuously differentiable semi-algebraic functions. To mitigate the lack of local Lipschitz continuity, popular approaches revolve around $\varepsilon$-oracles and/or linesearch procedures. In contrast, we exploit plain Hölder inequalities not entailing any approximation, all while retaining the linesearch-free nature of adaptive schemes. Furthermore, we prove full sequence convergence without prior knowledge of local Hölder constants nor of the order of Hölder continuity. Numerical experiments make comparisons with baseline methods on diverse tasks from machine learning covering both the locally and the globally Hölder setting. Konstantinos A. Oikonomidis, Emanuel Laude, Puya Latafat, Andreas Themelis, Panagiotis Patrinos |
ICML | 5 |
| 2024 | Learning in Feature Spaces via Coupled Covariances: Asymmetric Kernel SVD and Nyström methodabstractIn contrast with Mercer kernel-based approaches as used e.g. in Kernel Principal Component Analysis (KPCA), it was previously shown that Singular Value Decomposition (SVD) inherently relates to asymmetric kernels and Asymmetric Kernel Singular Value Decomposition (KSVD) has been proposed. However, the existing formulation to KSVD cannot work with infinite-dimensional feature mappings, the variational objective can be unbounded, and needs further numerical evaluation and exploration towards machine learning. In this work, i) we introduce a new asymmetric learning paradigm based on coupled covariance eigenproblem (CCE) through covariance operators, allowing infinite-dimensional feature maps. The solution to CCE is ultimately obtained from the SVD of the induced asymmetric kernel matrix, providing links to KSVD. ii) Starting from the integral equations corresponding to a pair of coupled adjoint eigenfunctions, we formalize the asymmetric Nyström method through a finite sample approximation to speed up training. iii) We provide the first empirical evaluations verifying the practical utility and benefits of KSVD and compare with methods resorting to symmetrization or linear SVD across multiple tasks. Qinghua Tao, Francesco Tonin, Alex Lambert, Yingyi Chen, Panagiotis Patrinos, Johan A. K. Suykens |
ICML | 5 |
| 2024 | Radial basis function neural network training using variable projection and fuzzy means
Despina Karamichailidou, Georgios Gerolymatos, Panagiotis Patrinos, Haralambos Sarimveis, Alex Alexandridis |
Neural Comput. Appl. | 3 |
| 2024 | Deep Kernel Principal Component Analysis for multi-level feature learningabstractPrincipal Component Analysis (PCA) and its nonlinear extension Kernel PCA (KPCA) are widely used across science and industry for data analysis and dimensionality reduction. Modern deep learning tools have achieved great empirical success, but a framework for deep principal component analysis is still lacking. Here we develop a deep kernel PCA methodology (DKPCA) to extract multiple levels of the most informative components of the data. Our scheme can effectively identify new hierarchical variables, called deep principal components, capturing the main characteristics of high-dimensional data through a simple and interpretable numerical optimization. We couple the principal components of multiple KPCA levels, theoretically showing that DKPCA creates both forward and backward dependency across levels, which has not been explored in kernel methods and yet is crucial to extract more informative features. Various experimental evaluations on multiple data types show that DKPCA finds more efficient and disentangled representations with higher explained variance in fewer principal components, compared to the shallow KPCA. We demonstrate that our method allows for effective hierarchical data exploration, with the ability to separate the key generative factors of the input data both for large datasets and when few training samples are available. Overall, DKPCA can facilitate the extraction of useful patterns from high-dimensional data by learning more informative features organized in different levels, giving diversified aspects to explore the variation factors in the data, while maintaining a simple mathematical formulation. Francesco Tonin, Qinghua Tao, Panagiotis Patrinos, Johan A. K. Suykens |
Neural Networks | 3 |
| 2023 | Solving stochastic weak Minty variational inequalities without increasing batch size
Thomas Pethick, Olivier Fercoq, Puya Latafat, Panagiotis Patrinos, Volkan Cevher |
ICLR | 4 |
| 2023 | Extending Kernel PCA through Dualization: Sparsity, Robustness and Fast AlgorithmsabstractThe goal of this paper is to revisit Kernel Principal Component Analysis (KPCA) through dualization of a difference of convex functions. This allows to naturally extend KPCA to multiple objective functions and leads to efficient gradient-based algorithms avoiding the expensive SVD of the Gram matrix. Particularly, we consider objective functions that can be written as Moreau envelopes, demonstrating how to promote robustness and sparsity within the same framework. The proposed method is evaluated on synthetic and realworld benchmarks, showing significant speedup in KPCA training time as well as highlighting the benefits in terms of robustness and sparsity. Francesco Tonin, Alex Lambert, Panagiotis Patrinos, Johan A. K. Suykens |
ICML | 3 |
| 2023 | Safe, learning-based MPC for highway driving under lane-change uncertainty: A distributionally robust approach
Mathijs Schuurmans, Alexander Katriniok, Chris Meissen, H. Eric Tseng, Panagiotis Patrinos |
Artif. Intell. | 5 |
| 2022 | Learning-Based Resource Allocation with Dynamic Data Rate ConstraintsabstractIn this paper, we address the problem of resource allocation (RA) in wireless communication networks, where each user has a dynamic data rate constraint. The objective of RA is to maximize the sum rate (SR) of the users while satisfying the data rate constraints in expectation. For a given set of data rate constraints, a suitable probability distribution for the activation of users is found iteratively with a stochastic gradient descent (SGD) approach to satisfy the data rate constraints in expectation. At each time instant, RA amongst the randomly activated users is performed noniteratively by a centralized deep neural network (DNN). Simulations show that the proposed approach is convergent and not only can consider dynamic data rate constraints accurately, but also that it achieves a SR higher than that of the conventional geometric programming (GP) method. The proposed approach can open up a direction of research for cross-layer RA in the current deep learning-based RA context. Pourya Behmandpoor, Panagiotis Patrinos, Marc Moonen |
ICASSP | 2 |
| 2022 | Escaping limit cycles: Global convergence for constrained nonconvex-nonconcave minimax problems
Thomas Pethick, Puya Latafat, Panagiotis Patrinos, Olivier Fercoq, Volkan Cevher |
ICLR | 3 |
| 2021 | Unsupervised Energy-based Out-of-distribution Detection using Stiefel-Restricted Kernel MachineabstractDetecting out-of-distribution (OOD) samples is an essential requirement for the deployment of machine learning systems in the real world. Until now, research on energy-based OOD detectors has focused on the softmax confidence score from a pre-trained neural network classifier with access to class labels. In contrast, we propose an unsupervised energy-based OOD detector leveraging the Stiefel-Restricted Kernel Machine (St-RKM). Training requires minimizing an objective function with an autoencoder loss term and the RKM energy where the interconnection matrix lies on the Stiefel manifold. Further, we outline multiple energy function definitions based on the RKM framework and discuss their utility. In the experiments on standard datasets, the proposed method improves over the existing energy-based OOD detectors and deep generative models. Through several ablation studies, we further illustrate the merit of each proposed energy function on the OOD detection performance. Francesco Tonin, Arun Pandey, Panagiotis Patrinos, Johan A. K. Suykens |
IJCNN | 3 |
| 2021 | Unsupervised learning of disentangled representations in deep restricted kernel machines with orthogonality constraintsabstractWe introduce Constr-DRKM, a deep kernel method for the unsupervised learning of disentangled data representations. We propose augmenting the original deep restricted kernel machine formulation for kernel PCA by orthogonality constraints on the latent variables to promote disentanglement and to make it possible to carry out optimization without first defining a stabilized objective. After discussing a number of algorithms for end-to-end training, we quantitatively evaluate the proposed method's effectiveness in disentangled feature learning. We demonstrate on four benchmark datasets that this approach performs similarly overall to β-VAE on several disentanglement metrics when few training points are available while being less sensitive to randomness and hyperparameter selection than β-VAE. We also present a deterministic initialization of Constr-DRKM's training algorithm that significantly improves the reproducibility of the results. Finally, we empirically evaluate and discuss the role of the number of layers in the proposed methodology, examining the influence of each principal component in every layer and showing that components in lower layers act as local feature detectors capturing the broad trends of the data distribution, while components in deeper layers use the representation learned by previous layers and more accurately reproduce higher-level features. Francesco Tonin, Panagiotis Patrinos, Johan A. K. Suykens |
Neural Networks | 2 |
| 2020 | Inertial Block Proximal Methods for Non-Convex Non-Smooth OptimizationabstractWe propose inertial versions of block coordinate descent methods for solving non-convex non-smooth composite optimization problems. Our methods possess three main advantages compared to current state-of-the-art accelerated first-order methods: (1) they allow using two different extrapolation points to evaluate the gradients and to add the inertial force (we will empirically show that it is more efficient than using a single extrapolation point), (2) they allow to randomly select the block of variables to update, and (3) they do not require a restarting step. We prove the subsequential convergence of the generated sequence under mild assumptions, prove the global convergence under some additional assumptions, and provide convergence rates. We deploy the proposed methods to solve non-negative matrix factorization (NMF) and show that they compete favorably with the state-of-the-art NMF algorithms. Additional experiments on non-negative approximate canonical polyadic decomposition, also known as nonnegative tensor factorization, are also provided. Hien Le, Nicolas Gillis, Panagiotis Patrinos |
ICML | 3 |
| 2020 | On the Convexity of Bit Depth Allocation for Linear MMSE Estimation in Wireless Sensor NetworksabstractEnergy efficiency is crucial for a wireless sensor network (WSN) since its nodes are generally powered by energy sources of limited capacity, such as batteries. The bit depth used to quantize the sensor signal samples heavily influences energy consumption, as it strongly impacts the amount of information to be transmitted between the sensor nodes. Bit depth allocation problems seek to assign a certain bit depth to each sensor signal such that energy consumption is minimized while respecting a performance constraint. For multi-channel signal estimation tasks these problems are generally non-convex, and they are often solved through simplifying assumptions or through convex relaxation. However, for linear minimum mean squared error (MMSE) estimation, we show how the matrix inversion lemma allows to transform the MMSE constraint into a convex constraint, which can then be interpreted as a constraint on the excess MMSE due to quantization. As a result, as long as the cost function representing energy consumption is convex, this class of bit depth allocation problems is convex, i.e., if the bit depth variable is relaxed to a real-valued variable. This guarantees global optimality up to discretization of the obtained solution. Fernando de la Hucha Arce, Panagiotis Patrinos, Marian Verhelst, Alexander Bertrand |
IEEE Signal Process. Lett. | 2 |
| 2010 | Variable Selection in Nonlinear Modeling Based on RBF Networks and Evolutionary ComputationabstractIn this paper a novel variable selection method based on Radial Basis Function (RBF) neural networks and genetic algorithms is presented. The fuzzy means algorithm is utilized as the training method for the RBF networks, due to its inherent speed, the deterministic approach of selecting the hidden node centers and the fact that it involves only a single tuning parameter. The trade-off between the accuracy and parsimony of the produced model is handled by using Final Prediction Error criterion, based on the RBF training and validation errors, as a fitness function of the proposed genetic algorithm. The tuning parameter required by the fuzzy means algorithm is treated as a free variable by the genetic algorithm. The proposed method was tested in benchmark data sets stemming from the scientific communities of time-series prediction and medicinal chemistry and produced promising results. Panagiotis Patrinos, Alex Alexandridis, Konstantinos Ninos, Haralambos Sarimveis |
Int. J. Neural Syst. | 1 |