Adeel Pervez

dblp:225/4821 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 29% Representation and self-supervised learning · 20% Optimization for machine learning · 17%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 100%

Topics — the 13 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational science and engineering
scientific machine learning
2.532025
Mechanistic PDE Networks for Discovery of Governing Equations · ICML 2025
Scalable Mechanistic Neural Networks · ICLR 2025
Mechanistic Neural Networks for Scientific Machine Learning · ICML 2024
Machine learning › Efficient and distributed learning › model compression
lightweight neural network
0.912025
Scalable Mechanistic Neural Networks · ICLR 2025
Machine learning › Deep learning architectures and training
physics-informed neural network
0.912025
Mechanistic PDE Networks for Discovery of Governing Equations · ICML 2025
Computational science and engineering › partial differential equations
partial differential equation discovery
0.912025
Mechanistic PDE Networks for Discovery of Governing Equations · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.712023
Differentiable Mathematical Programming for Object-Centric Representation Learning · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning
discrete representation learning
0.612022
Stability Regularization for Discrete Representation Learning · ICLR 2022
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE
0.512021
Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders · ICML 2021
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
latent variable discovery
0.512021
Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders · ICML 2021
Machine learning › Generative modeling › variational autoencoder
posterior collapse
0.512021
Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders · ICML 2021
Machine learning › Generative modeling
variational autoencoder
0.512021
Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders · ICML 2021
Machine learning › Optimization for machine learning
gradient estimation
0.412020
Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks · ICML 2020
Machine learning › Deep learning architectures and training
stochastic neural network
0.412020
Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks · ICML 2020
Machine learning › Optimization for machine learning
variance reduction
0.412020
Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks · ICML 2020

Methods — techniques the papers use, named apart from their topics

GPU parallelization · 3.3differentiable programming · 2.4mechanistic neural network · 1.7linear programming solver · 1.5ODE solver · 1.5multigrid solver · 0.9multi-grid solver · 0.9neural conditional poisson networks · 0.7stability regularization · 0.6spectral analysis · 0.5gaussian smoothing · 0.5
YearPublicationVenuePosition
2025 Scalable Mechanistic Neural Networks
abstract
We propose Scalable Mechanistic Neural Network (S-MNN), an enhanced neural network framework designed for scientific machine learning applications involving long temporal sequences. By reformulating the original Mechanistic Neural Network (MNN) (Pervez et al., 2024), we reduce the computational time and space complexities from cubic and quadratic with respect to the sequence length, respectively, to linear. This significant improvement enables efficient modeling of long-term dynamics without sacrificing accuracy or interpretability. Extensive experiments demonstrate that S-MNN matches the original MNN in precision while substantially reducing computational resources. Consequently, S-MNN can drop-in replace the original MNN in applications, providing a practical and efficient tool for integrating mechanistic bottlenecks into neural network models of complex dynamical systems. Source code is available at https://github.com/IST-DASLab/ScalableMNN.
Jiale Chen 0004, Dingling Yao, Adeel Pervez, Dan Alistarh, Francesco Locatello
ICLR3
2025 Mechanistic PDE Networks for Discovery of Governing Equations
abstract
We present Mechanistic PDE Networks -- a model for discovery of governing *partial differential equations* from data. Mechanistic PDE Networks represent spatiotemporal data as space-time dependent *linear* partial differential equations in neural network hidden representations. The represented PDEs are then solved and decoded for specific tasks. The learned PDE representations naturally express the spatiotemporal dynamics in data in neural network hidden space, enabling increased modeling power. Solving the PDE representations in a compute and memory-efficient way, however, is a significant challenge. We develop a native, GPU-capable, parallel, sparse and differentiable multigrid solver specialized for linear partial differential equations that acts as a module in Mechanistic PDE Networks. Leveraging the PDE solver we propose a discovery architecture that can discovers nonlinear PDEs in complex settings, while being robust to noise. We validate PDE discovery on a number of PDEs including reaction-diffusion and Navier-Stokes equations.
Adeel Pervez, Efstratios Gavves, Francesco Locatello
ICML1
2024 Mechanistic Neural Networks for Scientific Machine Learning
abstract
This paper presents *Mechanistic Neural Networks*, a neural network design for machine learning applications in the sciences. It incorporates a new *Mechanistic Block* in standard architectures to explicitly learn governing differential equations as representations, revealing the underlying dynamics of data and enhancing interpretability and efficiency in data modeling. Central to our approach is a novel *Relaxed Linear Programming Solver* (NeuRLP) inspired by a technique that reduces solving linear ODEs to solving linear programs. This integrates well with neural networks and surpasses the limitations of traditional ODE solvers enabling scalable GPU parallel processing. Overall, Mechanistic Neural Networks demonstrate their versatility for scientific machine learning applications, adeptly managing tasks from equation discovery to dynamic systems modeling. We prove their comprehensive capabilities in analyzing and interpreting complex scientific data across various applications, showing significant performance against specialized state-of-the-art methods. Source code is available at https://github.com/alpz/mech-nn.
Adeel Pervez, Francesco Locatello, Stratis Gavves
ICML1
2023 Differentiable Mathematical Programming for Object-Centric Representation Learning
Adeel Pervez, Phillip Lippe, Efstratios Gavves
ICLR1
2023 Scalable Subset Sampling with Neural Conditional Poisson Networks
Adeel Pervez, Phillip Lippe, Efstratios Gavves
ICLR1
2022 Stability Regularization for Discrete Representation Learning
Adeel Pervez, Efstratios Gavves
ICLR1
2021 Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders
abstract
Variational autoencoders with deep hierarchies of stochastic layers have been known to suffer from the problem of posterior collapse, where the top layers fall back to the prior and become independent of input. We suggest that the hierarchical VAE objective explicitly includes the variance of the function parameterizing the mean and variance of the latent Gaussian distribution which itself is often a high variance function. Building on this we generalize VAE neural networks by incorporating a smoothing parameter motivated by Gaussian analysis to reduce higher frequency components and consequently the variance in parameterizing functions and show that this can help to solve the problem of posterior collapse. We further show that under such smoothing the VAE loss exhibits a phase transition, where the top layer KL divergence sharply drops to zero at a critical value of the smoothing parameter that is similar for the same model across datasets. We validate the phenomenon across model configurations and datasets.
Adeel Pervez, Efstratios Gavves
ICML1
2020 Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks
abstract
Stochastic neural networks with discrete random variables are an important class of models for their expressiveness and interpretability. Since direct differentiation and backpropagation is not possible, Monte Carlo gradient estimation techniques are a popular alternative. Efficient stochastic gradient estimators, such Straight-Through and Gumbel-Softmax, work well for shallow stochastic models. Their performance, however, suffers with hierarchical, more complex models. We focus on stochastic networks with Boolean latent variables. To analyze such networks, we introduce the framework of harmonic analysis for Boolean functions to derive an analytic formulation for the bias and variance in the Straight-Through estimator. Exploiting these formulations, we propose \emph{FouST}, a low-bias and low-variance gradient estimation algorithm that is just as efficient. Extensive experiments show that FouST performs favorably compared to state-of-the-art biased estimators and is much faster than unbiased ones.
Adeel Pervez, Taco Cohen, Efstratios Gavves
ICML1