Tanya Marwah

dblp:190/7486 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Deep learning architectures and training · 38% Learning theory · 20% Generative modeling · 15%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Computational science and engineering · 100%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural operator
1.522025
On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Machine learning › Deep learning architectures and training
neural network expressivity
1.222023
Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023
Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
1.222023
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.912025
Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025
Machine learning › Graph learning › graph representation learning
edge embedding
0.912025
Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025
Machine learning › Graph learning
graph neural network
0.912025
Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025
Machine learning › Learning theory › computational complexity
time-space tradeoffs
0.912025
Towards characterizing the value of edge embeddings in Graph Neural Networks · ICML 2025
Computational science and engineering
partial differential equations
0.912025
On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025
Computational science and engineering
scientific machine learning
0.912025
On the Benefits of Memory for Modeling Time-Dependent PDEs · ICLR 2025
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model
0.712023
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Machine learning › Learning theory
generalization
0.712023
Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023
Machine learning › Deep learning architectures and training › deep generative model
implicit models
0.712023
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Machine learning › Optimization for machine learning
implicit regularization
0.712023
Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023
Machine learning › Deep learning architectures and training › scientific machine learning
neural network approximation of PDEs
0.712023
Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
Disentangling the Mechanisms Behind Implicit Regularization in SGD · ICLR 2023
Computational science and engineering › scientific machine learning
neural operator
0.712023
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Computational science and engineering
partial differential equation solver
0.712023
Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023
Visual content generation and editing
video generation
0.622017
Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017
Attentive Semantic Video Generation Using Captions · ICCV 2017
Machine learning › Learning theory
approximation theory
0.512021
Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021
Visual content generation and editing › video generation
text-to-video generation
0.312017
Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017
Computer vision › Image recognition and object detection
multi-scale inference
0.312025
Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025
Machine learning › Learning theory
curse of dimensionality
0.212023
Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective · ICML 2023
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
0.212023
Deep Equilibrium Based Neural Operators for Steady-State PDEs · NeurIPS 2023
Mathematical optimization
gradient descent
0.112021
Parametric Complexity Bounds for Approximating PDEs with Neural Networks · NeurIPS 2021
Computer vision › Vision and language
video captioning
0.112017
Attentive Semantic Video Generation Using Captions · ICCV 2017
Machine learning › Generative modeling
video generation
0.112017
Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures · ACM Multimedia 2017

Methods — techniques the papers use, named apart from their topics

state space model · 1.7multiscale inference scheme · 1.7fourier neural operator · 1.7autoregressive rollout · 1.7implicit differentiation · 1.3gradient descent in hilbert space · 1.3euler-lagrange energy functional · 1.3black-box root solver · 1.3barron norm · 1.3weight tying · 0.7stochastic gradient descent · 0.7network growing proof technique · 0.5hilbert space analysis · 0.5elliptic PDE theory · 0.5variational autoencoder · 0.3recurrent attention mechanism · 0.3latent representation learning · 0.3attentive network · 0.3
YearPublicationVenuePosition
2025 On the Benefits of Memory for Modeling Time-Dependent PDEs
abstract
Data-driven techniques have emerged as a promising alternative to traditional numerical methods for solving PDEs. For time-dependent PDEs, many approaches are Markovian---the evolution of the trained system only depends on the current state, and not the past states. In this work, we investigate the benefits of using memory for modeling time-dependent PDEs: that is, when past states are explicitly used to predict the future. Motivated by the Mori-Zwanzig theory of model reduction, we theoretically exhibit examples of simple (even linear) PDEs, in which a solution that uses memory is arbitrarily better than a Markovian solution. Additionally, we introduce Memory Neural Operator (MemNO), a neural operator architecture that combines recent state space models (specifically, S4) and Fourier Neural Operators (FNOs) to effectively model memory. We empirically demonstrate that when the PDEs are supplied in low resolution or contain observation noise at train and test time, MemNO significantly outperforms the baselines without memory---with up to $6 \times$ reduction in test error. Furthermore, we show that this benefit is particularly pronounced when the PDE solutions have significant high-frequency Fourier modes (e.g., low-viscosity fluid dynamics) and we construct a challenging benchmark dataset consisting of such PDEs.
Ricardo Buitrago Ruiz, Tanya Marwah, Albert Gu, Andrej Risteski
ICLR2
2025 Towards characterizing the value of edge embeddings in Graph Neural Networks
abstract
Graph neural networks (GNNs) are the dominant approach to solving machine learning problems defined over graphs. Despite much theoretical and empirical work in recent years, our understanding of finer-grained aspects of architectural design for GNNs remains impoverished. In this paper, we consider the benefits of architectures that maintain and update edge embeddings. On the theoretical front, under a suitable computational abstraction for a layer in the model, as well as memory constraints on the embeddings, we show that there are natural tasks on graphical models for which architectures leveraging edge embeddings can be much shallower. Our techniques are inspired by results on time-space tradeoffs in theoretical computer science. Empirically, we show architectures that maintain edge embeddings almost always improve on their node-based counterparts---frequently significantly so in topologies that have "hub" nodes.
Dhruv Rohatgi, Tanya Marwah, Zachary C. Lipton, Jianfeng Lu 0001, Ankur Moitra, Andrej Risteski
ICML2
2025 Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme
abstract
Conditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability.
Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho
NeurIPS10
2023 Disentangling the Mechanisms Behind Implicit Regularization in SGD
Zachary Novack, Simran Kaur 0001, Tanya Marwah, Zachary C. Lipton
ICLR3
2023 Neural Network Approximations of PDEs Beyond Linearity: A Representational Perspective
abstract
A burgeoning line of research has developed deep neural networks capable of approximating the solutions to high dimensional PDEs, opening related lines of theoretical inquiry focused on explaining how it is that these models appear to evade the curse of dimensionality. However, most theoretical analyses thus far have been limited to linear PDEs. In this work, we take a step towards studying the representational power of neural networks for approximating solutions to nonlinear PDEs. We focus on a class of PDEs known as *nonlinear elliptic variational PDEs*, whose solutions minimize an *Euler-Lagrange* energy functional $\mathcal{E}(u) = \int_\Omega L(x, u(x), \nabla u(x)) - f(x) u(x)dx$. We show that if composing a function with Barron norm $b$ with partial derivatives of $L$ produces a function of Barron norm at most $B_L b^p$, the solution to the PDE can be $\epsilon$-approximated in the $L^2$ sense by a function with Barron norm $O\left(\left(dB_L\right)^{\max\{p \log(1/ \epsilon), p^{\log(1/\epsilon)}\}}\right)$. By a classical result due to Barron (1993), this correspondingly bounds the size of a 2-layer neural network needed to approximate the solution. Treating $p, \epsilon, B_L$ as constants, this quantity is polynomial in dimension, thus showing neural networks can evade the curse of dimensionality. Our proof technique involves neurally simulating (preconditioned) gradient in an appropriate Hilbert space, which converges exponentially fast to the solution of the PDE, and such that we can bound the increase of the Barron norm at each iterate. Our results subsume and substantially generalize analogous prior results for linear elliptic PDEs over a unit hypercube.
Tanya Marwah, Zachary C. Lipton, Jianfeng Lu 0001, Andrej Risteski
ICML1
2023 Deep Equilibrium Based Neural Operators for Steady-State PDEs
abstract
Data-driven machine learning approaches are being increasingly used to solve partial differential equations (PDEs). They have shown particularly striking successes when training an operator, which takes as input a PDE in some family, and outputs its solution. However, the architectural design space, especially given structural knowledge of the PDE family of interest, is still poorly understood. We seek to remedy this gap by studying the benefits of weight-tied neural network architectures for steady-state PDEs. To achieve this, we first demonstrate that the solution of most steady-state PDEs can be expressed as a fixed point of a non-linear operator. Motivated by this observation, we propose FNO-DEQ, a deep equilibrium variant of the FNO architecture that directly solves for the solution of a steady-state PDE as the infinite-depth fixed point of an implicit operator layer using a black-box root solver and differentiates analytically through this fixed point resulting in $\mathcal{O}(1)$ training memory. Our experiments indicate that FNO-DEQ-based architectures outperform FNO-based baselines with $4\times$ the number of parameters in predicting the solution to steady-state PDEs such as Darcy Flow and steady-state incompressible Navier-Stokes. Finally, we show FNO-DEQ is more robust when trained with datasets with more noisy observations than the FNO-based baselines, demonstrating the benefits of using appropriate inductive biases in architectural design for different neural network based PDE solvers. Further, we show a universal approximation result that demonstrates that FNO-DEQ can approximate the solution to any steady-state PDE that can be written as a fixed point equation.
Tanya Marwah, Ashwini Pokle, J. Zico Kolter, Zachary C. Lipton, Jianfeng Lu 0001, Andrej Risteski
NeurIPS1
2021 Parametric Complexity Bounds for Approximating PDEs with Neural Networks
abstract
Recent experiments have shown that deep networks can approximate solutions to high-dimensional PDEs, seemingly escaping the curse of dimensionality. However, questions regarding the theoretical basis for such approximations, including the required network size remain open. In this paper, we investigate the representational power of neural networks for approximating solutions to linear elliptic PDEs with Dirichlet boundary conditions. We prove that when a PDE's coefficients are representable by small neural networks, the parameters required to approximate its solution scale polynomially with the input dimension $d$ and proportionally to the parameter counts of the coefficient networks. To this end, we develop a proof technique that simulates gradient descent (in an appropriate Hilbert space) by growing a neural network architecture whose iterates each participate as sub-networks in their (slightly larger) successors, and converge to the solution of the PDE.
Tanya Marwah, Zachary C. Lipton, Andrej Risteski
NeurIPS1
2017 Attentive Semantic Video Generation Using Captions
abstract
This paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term dependencies between video frames and thus generate a video in an incremental manner. Our experiments demonstrate our network architecture’s ability to distinguish between objects, actions and interactions in a video and combine them to generate videos for unseen captions. The network also exhibits the capability to perform spatio-temporal style transfer when asked to generate videos for a sequence of captions. We also show that the network’s ability to learn a latent representation allows it generate videos in an unsupervised manner and perform other tasks such as action recognition.
Tanya Marwah, Gaurav Mittal, Vineeth N. Balasubramanian
ICCV1
2017 Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive Architectures
abstract
This paper introduces a novel approach for generating videos called Synchronized Deep Recurrent Attentive Writer (Sync-DRAW). Sync-DRAW can also perform text-to-video generation which, to the best of our knowledge, makes it the first approach of its kind. It combines a Variational Autoencoder(VAE) with a Recurrent Attention Mechanism in a novel manner to create a temporally dependent sequence of frames that are gradually formed over time. The recurrent attention mechanism in Sync-DRAW attends to each individual frame of the video in sychronization, while the VAE learns a latent distribution for the entire video at the global level. Our experiments with Bouncing MNIST, KTH and UCF-101 suggest that Sync-DRAW is efficient in learning the spatial and temporal information of the videos and generates frames with high structural integrity, and can generate videos from simple captions on these datasets.
Gaurav Mittal, Tanya Marwah, Vineeth N. Balasubramanian
ACM Multimedia2