EDBT 2026 Demo / reviewers in the wild / expert
Tanya Marwah
dblp:190/7486
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Deep learning architectures and training · 38% Learning theory · 20% Generative modeling · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Computational science and engineering · 100% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
neural operator |
1.5 | 2 | 2025 | On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025 Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
neural network expressivity |
1.2 | 2 | 2023 | Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023 Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021 |
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving |
1.2 | 2 | 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Graph learning › graph representation learning
edge embedding |
0.9 | 1 | 2025 | Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025 |
Machine learning › Graph learning
graph neural network |
0.9 | 1 | 2025 | Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025 |
Machine learning › Learning theory › computational complexity
time-space tradeoffs |
0.9 | 1 | 2025 | Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025 |
Computational science and engineering
partial differential equations |
0.9 | 1 | 2025 | On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025 |
Computational science and engineering
scientific machine learning |
0.9 | 1 | 2025 | On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025 |
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model |
0.7 | 1 | 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023 |
Machine learning › Deep learning architectures and training › deep generative model
implicit models |
0.7 | 1 | 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 |
Machine learning › Optimization for machine learning
implicit regularization |
0.7 | 1 | 2023 | Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023 |
Machine learning › Deep learning architectures and training › scientific machine learning
neural network approximation of PDEs |
0.7 | 1 | 2023 | Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 1 | 2023 | Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023 |
Computational science and engineering › scientific machine learning
neural operator |
0.7 | 1 | 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 |
Computational science and engineering
partial differential equation solver |
0.7 | 1 | 2023 | Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023 |
Visual content generation and editing
video generation |
0.6 | 2 | 2017 | Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017 Attentive Semantic Video Generation Using Captions · ICCV 2017 |
Machine learning › Learning theory
approximation theory |
0.5 | 1 | 2021 | Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021 |
Visual content generation and editing › video generation
text-to-video generation |
0.3 | 1 | 2017 | Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017 |
Computer vision › Image recognition and object detection
multi-scale inference |
0.3 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Learning theory
curse of dimensionality |
0.2 | 1 | 2023 | Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023 |
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
0.2 | 1 | 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023 |
Mathematical optimization
gradient descent |
0.1 | 1 | 2021 | Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021 |
Computer vision › Vision and language
video captioning |
0.1 | 1 | 2017 | Attentive Semantic Video Generation Using Captions · ICCV 2017 |
Machine learning › Generative modeling
video generation |
0.1 | 1 | 2017 | Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017 |
Methods — techniques the papers use, named apart from their topics
state space model · 1.7multiscale inference scheme · 1.7fourier neural operator · 1.7autoregressive rollout · 1.7implicit differentiation · 1.3gradient descent in hilbert space · 1.3euler-lagrange energy functional · 1.3black-box root solver · 1.3barron norm · 1.3weight tying · 0.7stochastic gradient descent · 0.7network growing proof technique · 0.5hilbert space analysis · 0.5elliptic PDE theory · 0.5variational autoencoder · 0.3recurrent attention mechanism · 0.3latent representation learning · 0.3attentive network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Benefits of Memory for Modeling Time-Dependent PDEsabstractData-driven techniques have emerged as a promising alternative to traditional numerical methods for solving PDEs. For time-dependent PDEs, many approaches are Markovian---the evolution of the trained system only depends on the current state, and not the past states. In this work, we investigate the benefits of using memory for modeling time-dependent PDEs: that is, when past states are explicitly used to predict the future. Motivated by the Mori-Zwanzig theory of model reduction, we theoretically exhibit examples of simple (even linear) PDEs, in which a solution that uses memory is arbitrarily better than a Markovian solution. Additionally, we introduce Memory Neural Operator (MemNO), a neural operator architecture that combines recent state space models (specifically, S4) and Fourier Neural Operators (FNOs) to effectively model memory. We empirically demonstrate that when the PDEs are supplied in low resolution or contain observation noise at train and test time, MemNO significantly outperforms the baselines without memory---with up to $6 \times$ reduction in test error. Furthermore, we show that this benefit is particularly pronounced when the PDE solutions have significant high-frequency Fourier modes (e.g., low-viscosity fluid dynamics) and we construct a challenging benchmark dataset consisting of such PDEs. Ricardo Buitrago Ruiz, Tanya Marwah, Albert Gu, Andrej Risteski |
ICLR | 2 |
| 2025 | Towards characterizing the value of edge embeddings in Graph Neural NetworksabstractGraph neural networks (GNNs) are the dominant approach to solving machine learning problems defined over graphs. Despite much theoretical and empirical work in recent years, our understanding of finer-grained aspects of architectural design for GNNs remains impoverished. In this paper, we consider the benefits of architectures that maintain and update edge embeddings. On the theoretical front, under a suitable computational abstraction for a layer in the model, as well as memory constraints on the embeddings, we show that there are natural tasks on graphical models for which architectures leveraging edge embeddings can be much shallower. Our techniques are inspired by results on time-space tradeoffs in theoretical computer science. Empirically, we show architectures that maintain edge embeddings almost always improve on their node-based counterparts---frequently significantly so in topologies that have "hub" nodes. Dhruv Rohatgi, Tanya Marwah, Zachary C. Lipton, Jianfeng Lu 0001, Ankur Moitra, Andrej Risteski |
ICML | 2 |
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 10 |
| 2023 | Disentangling the Mechanisms Behind Implicit Regularization in SGD
Zachary Novack, Simran Kaur 0001, Tanya Marwah, Zachary C. Lipton |
ICLR | 3 |
| 2023 | Neural Network Approximations of PDEs Beyond Linearity: A Representational PerspectiveabstractA burgeoning line of research has developed deep neural networks capable of approximating the solutions to high dimensional PDEs, opening related lines of theoretical inquiry focused on explaining how it is that these models appear to evade the curse of dimensionality. However, most theoretical analyses thus far have been limited to linear PDEs. In this work, we take a step towards studying the representational power of neural networks for approximating solutions to nonlinear PDEs. We focus on a class of PDEs known as *nonlinear elliptic variational PDEs*, whose solutions minimize an *Euler-Lagrange* energy functional $\mathcal{E}(u) = \int_\Omega L(x, u(x), \nabla u(x)) - f(x) u(x)dx$. We show that if composing a function with Barron norm $b$ with partial derivatives of $L$ produces a function of Barron norm at most $B_L b^p$, the solution to the PDE can be $\epsilon$-approximated in the $L^2$ sense by a function with Barron norm $O\left(\left(dB_L\right)^{\max\{p \log(1/ \epsilon), p^{\log(1/\epsilon)}\}}\right)$. By a classical result due to Barron (1993), this correspondingly bounds the size of a 2-layer neural network needed to approximate the solution. Treating $p, \epsilon, B_L$ as constants, this quantity is polynomial in dimension, thus showing neural networks can evade the curse of dimensionality. Our proof technique involves neurally simulating (preconditioned) gradient in an appropriate Hilbert space, which converges exponentially fast to the solution of the PDE, and such that we can bound the increase of the Barron norm at each iterate. Our results subsume and substantially generalize analogous prior results for linear elliptic PDEs over a unit hypercube. Tanya Marwah, Zachary C. Lipton, Jianfeng Lu 0001, Andrej Risteski |
ICML | 1 |
| 2023 | Deep Equilibrium Based Neural Operators for Steady-State PDEsabstractData-driven machine learning approaches are being increasingly used to solve partial differential equations (PDEs). They have shown particularly striking successes when training an operator, which takes as input a PDE in some family, and outputs its solution. However, the architectural design space, especially given structural knowledge of the PDE family of interest, is still poorly understood. We seek to remedy this gap by studying the benefits of weight-tied neural network architectures for steady-state PDEs. To achieve this, we first demonstrate that the solution of most steady-state PDEs can be expressed as a fixed point of a non-linear operator. Motivated by this observation, we propose FNO-DEQ, a deep equilibrium variant of the FNO architecture that directly solves for the solution of a steady-state PDE as the infinite-depth fixed point of an implicit operator layer using a black-box root solver and differentiates analytically through this fixed point resulting in $\mathcal{O}(1)$ training memory. Our experiments indicate that FNO-DEQ-based architectures outperform FNO-based baselines with $4\times$ the number of parameters in predicting the solution to steady-state PDEs such as Darcy Flow and steady-state incompressible Navier-Stokes. Finally, we show FNO-DEQ is more robust when trained with datasets with more noisy observations than the FNO-based baselines, demonstrating the benefits of using appropriate inductive biases in architectural design for different neural network based PDE solvers. Further, we show a universal approximation result that demonstrates that FNO-DEQ can approximate the solution to any steady-state PDE that can be written as a fixed point equation. Tanya Marwah, Ashwini Pokle, J. Zico Kolter, Zachary C. Lipton, Jianfeng Lu 0001, Andrej Risteski |
NeurIPS | 1 |
| 2021 | Parametric Complexity Bounds for Approximating PDEs with Neural NetworksabstractRecent experiments have shown that deep networks can approximate solutions to high-dimensional PDEs, seemingly escaping the curse of dimensionality. However, questions regarding the theoretical basis for such approximations, including the required network size remain open. In this paper, we investigate the representational power of neural networks for approximating solutions to linear elliptic PDEs with Dirichlet boundary conditions. We prove that when a PDE's coefficients are representable by small neural networks, the parameters required to approximate its solution scale polynomially with the input dimension $d$ and proportionally to the parameter counts of the coefficient networks. To this end, we develop a proof technique that simulates gradient descent (in an appropriate Hilbert space) by growing a neural network architecture whose iterates each participate as sub-networks in their (slightly larger) successors, and converge to the solution of the PDE. Tanya Marwah, Zachary C. Lipton, Andrej Risteski |
NeurIPS | 1 |
| 2017 | Attentive Semantic Video Generation Using CaptionsabstractThis paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term dependencies between video frames and thus generate a video in an incremental manner. Our experiments demonstrate our network architecture’s ability to distinguish between objects, actions and interactions in a video and combine them to generate videos for unseen captions. The network also exhibits the capability to perform spatio-temporal style transfer when asked to generate videos for a sequence of captions. We also show that the network’s ability to learn a latent representation allows it generate videos in an unsupervised manner and perform other tasks such as action recognition. Tanya Marwah, Gaurav Mittal, Vineeth N. Balasubramanian |
ICCV | 1 |
| 2017 | Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive ArchitecturesabstractThis paper introduces a novel approach for generating videos called Synchronized Deep Recurrent Attentive Writer (Sync-DRAW). Sync-DRAW can also perform text-to-video generation which, to the best of our knowledge, makes it the first approach of its kind. It combines a Variational Autoencoder(VAE) with a Recurrent Attention Mechanism in a novel manner to create a temporally dependent sequence of frames that are gradually formed over time. The recurrent attention mechanism in Sync-DRAW attends to each individual frame of the video in sychronization, while the VAE learns a latent distribution for the entire video at the global level. Our experiments with Bouncing MNIST, KTH and UCF-101 suggest that Sync-DRAW is efficient in learning the spatial and temporal information of the videos and generates frames with high structural integrity, and can generate videos from simple captions on these datasets. Gaurav Mittal, Tanya Marwah, Vineeth N. Balasubramanian |
ACM Multimedia | 2 |