Amir Barati Farimani

dblp:206/7414 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
21since 2021 · last 2025
0000-0002-2952-8576ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 20 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Dual Diffusion for Unified Image Generation and Understanding
abstract
Diffusion models have gained tremendous success in text-to-image generation, yet still struggle with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end diffusion model for multi-modal understanding and generation that significantly improves on existing diffusion-based multimodal models, and is the first of its kind to support the full suite of vision-language modeling capabilities. Inspired by the multimodal diffusion transformer (MM-DiT) and recent advances in discrete diffusion language modeling, we leverage a cross-modal maximum likelihood estimation framework that simultaneously trains the conditional likelihoods of both images and text jointly under a single loss function, which is back-propagated through both branches of the diffusion transformer. The resulting model is highly flexible and capable of a wide range of tasks including image generation, captioning, and visual question answering. Our model attained competitive performance compared to recent unified image understanding and generation models, demonstrating the potential of multimodal diffusion modeling as a promising alternative to autoregressive next-token prediction models.
Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger
CVPR4
2025 LLM-SR: Scientific Equation Discovery via Programming with Large Language Models
abstract
Mathematical equations have been unreasonably effective in describing complex natural phenomena across various scientific disciplines. However, discovering such insightful equations from data presents significant challenges due to the necessity of navigating extremely large combinatorial hypothesis spaces. Current methods of equation discovery, commonly known as symbolic regression techniques, largely focus on extracting equations from data alone, often neglecting the domain-specific prior knowledge that scientists typically depend on. They also employ limited representations such as expression trees, constraining the search space and expressiveness of equations. To bridge this gap, we introduce LLM-SR, a novel approach that leverages the extensive scientific knowledge and robust code generation capabilities of Large Language Models (LLMs) to discover scientific equations from data. Specifically, LLM-SR treats equations as programs with mathematical operators and combines LLMs' scientific priors with evolutionary search over equation programs. The LLM iteratively proposes new equation skeleton hypotheses, drawing from its domain knowledge, which are then optimized against data to estimate parameters. We evaluate LLM-SR on four benchmark problems across diverse scientific domains (e.g., physics, biology), which we carefully designed to simulate the discovery process and prevent LLM recitation. Our results demonstrate that LLM-SR discovers physically accurate equations that significantly outperform state-of-the-art symbolic regression baselines, particularly in out-of-domain test settings. We also show that LLM-SR's incorporation of scientific priors enables more efficient equation space exploration than the baselines.
Parshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. Reddy
ICLR4
2025 Text2PDE: Latent Diffusion Models for Accessible Physics Simulation
abstract
Recent advances in deep learning have inspired numerous works on data-driven solutions to partial differential equation (PDE) problems. These neural PDE solvers can often be much faster than their numerical counterparts; however, each presents its unique limitations and generally balances training cost, numerical accuracy, and ease of applicability to different problem setups. To address these limitations, we introduce several methods to apply latent diffusion models to physics simulation. Firstly, we introduce a mesh autoencoder to compress arbitrarily discretized PDE data, allowing for efficient diffusion training across various physics. Furthermore, we investigate full spatiotemporal solution generation to mitigate autoregressive error accumulation. Lastly, we investigate conditioning on initial physical quantities, as well as conditioning solely on a text prompt to introduce text2PDE generation. We show that language can be a compact, interpretable, and accurate modality for generating physics simulations, paving the way for more usable and accessible PDE solvers. Through experiments on both uniform and structured grids, we show that the proposed approach is competitive with current neural PDE solvers in both accuracy and efficiency, with promising scaling behavior up to $\sim$3 billion parameters. By introducing a scalable, accurate, and usable physics simulator, we hope to bring neural PDE solvers closer to practical use.
Anthony Y. Zhou, Michael Schneier, John R. Buchanan Jr., Amir Barati Farimani
ICLR5
2025 LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
abstract
Scientific equation discovery is a fundamental task in the history of scientific progress, enabling the derivation of laws governing natural phenomena. Recently, Large Language Models (LLMs) have gained interest for this task due to their potential to leverage embedded scientific knowledge for hypothesis generation. However, evaluating the true discovery capabilities of these methods remains challenging, as existing benchmarks often rely on common equations that are susceptible to memorization by LLMs, leading to inflated performance metrics that do not reflect actual discovery. In this paper, we introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorization, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Through extensive evaluation of several state-of-the-art methods on LLM-SRBench, using both open and closed LLMs, we find that the best-performing system so far achieves only 31.5% symbolic accuracy. These findings highlight the challenges of scientific equation discovery, positioning LLM-SRBench as a valuable resource for future research.
Parshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani, Khoa D. Doan, Chandan K. Reddy
ICML4
2025 VITaL Pretraining: Visuo-Tactile Pretraining for Tactile and Non-Tactile Manipulation Policies
abstract
Tactile information is a critical tool for dexterous manipulation. As humans, we rely heavily on tactile information to understand objects in our environments and how to interact with them. We use touch not only to perform manipulation tasks but also to learn how to perform these tasks. Therefore, to create robotic agents that can learn to complete manipulation tasks at a human or super-human level of performance, we need to properly incorporate tactile information into both skill execution and skill learning. In this paper, we investigate how we can incorporate tactile information into imitation learning platforms to improve performance on manipulation tasks. We show that incorporating visuo-tactile pretraining improves imitation learning performance, not only for tactile agents (policies that use tactile information at inference), but also for non-tactile agents (policies that do not use tactile information at inference). For these non-tactile agents, pretraining with tactile information significantly improved performance (for example, improving the accuracy on USB plugging from 20% to 85%), reaching a level on par with visuo-tactile agents, and even surpassing them in some cases. For demonstration videos and access to our codebase, see the project website: https://sites.google.com/andrew.cmu.edu/visuo-tactile-pretraining
Selam Gano, Pranav Katragadda, Amir Barati Farimani
ICRA4
2025 Low-Fidelity Visuo-Tactile Pre-Training Improves Vision-Only Manipulation Performance
abstract
Tactile perception is essential for real-world manipulation tasks, yet the high cost and fragility of tactile sensors can limit their practicality. In this work, we explore BeadSight (a low-cost, open-source tactile sensor) alongside a tactile pre-training approach, an alternative method to precise, pre-calibrated sensors. By pre-training with the tactile sensor and then disabling it during downstream tasks, we aim to enhance robustness and reduce costs in manipulation systems. We investigate whether tactile pre-training, even with a low-fidelity sensor like BeadSight, can improve the performance of an imitation learning agent on complex manipulation tasks. Through visuo-tactile pre-training on both similar and dissimilar tasks, we analyze its impact on a longer-horizon downstream task. Our experiments show that visuo-tactile pre-training improved performance on a USB cable plugging task by up to 65% with vision-only inference. Additionally, on a longer-horizon drawer pick-and-place task, pre-training — whether on a similar, dissimilar, or identical task — consistently improved performance, highlighting the potential for a large-scale visuo-tactile pre-trained encoder. Code for this project is available at: https://github.com/selamie/beadsight.
Selam Gano, Amir Barati Farimani
IROS3
2025 Hamiltonian Neural PDE Solvers through Functional Approximation
abstract
Designing neural networks within a Hamiltonian framework offers a principled way to ensure that conservation laws are respected in physical systems. While promising, these capabilities have been largely limited to discrete, analytically solvable systems. In contrast, many physical phenomena are governed by PDEs, which govern infinite-dimensional fields through Hamiltonian functionals and their functional derivatives. Building on prior work, we represent the Hamiltonian functional as a kernel integral parameterized by a neural field, enabling learnable function-to-scalar mappings and the use of automatic differentiation to calculate functional derivatives. This allows for an extension of Hamiltonian mechanics to neural PDE solvers by predicting a functional and learning in the gradient domain. We show that the resulting Hamiltonian Neural Solver (HNS) can be an effective surrogate model through improved stability and conserving energy-like quantities across 1D and 2D PDEs. This ability to respect conservation laws also allows HNS models to better generalize to longer time horizons or unseen initial conditions.
Anthony Y. Zhou, Amir Barati Farimani
NeurIPS2
2024 SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training
abstract
In an era where symbolic mathematical equations are indispensable for modeling complex natural phenomena, scientific inquiry often involves collecting observations and translating them into mathematical expressions. Recently, deep learning has emerged as a powerful tool for extracting insights from data. However, existing models typically specialize in either numeric or symbolic domains, and are usually trained in a supervised manner tailored to specific tasks. This approach neglects the substantial benefits that could arise from a task-agnostic multi-modal understanding between symbolic equations and their numeric counterparts. To bridge the gap, we introduce SNIP, a Symbolic-Numeric Integrated Pre-training model, which employs contrastive learning between symbolic and numeric domains, enhancing their mutual similarities in the embeddings. By performing latent space analysis, we observe that SNIP provides cross-domain insights into the representations, revealing that symbolic supervision enhances the embeddings of numeric data and vice versa. We evaluate SNIP across diverse tasks, including symbolic-to-numeric mathematical property prediction and numeric-to-symbolic equation discovery, commonly known as symbolic regression. Results show that SNIP effectively transfers to various tasks, consistently outperforming fully supervised baselines and competing strongly with established task-specific methods, especially in the low data regime scenarios where available data is limited.
Kazem Meidani, Parshin Shojaee, Chandan K. Reddy, Amir Barati Farimani
ICLR4
2024 SculptBot: Pre-Trained Models for 3D Deformable Object Manipulation
abstract
Deformable object manipulation presents a unique set of challenges in robotic manipulation by exhibiting high degrees of freedom and severe self-occlusion. Choosing state representations for materials that exhibit plastic behavior, like modeling clay or bread dough, is also difficult because they permanently deform under stress and are constantly changing shape. In this work, we investigate each of these challenges using the task of robotic sculpting with a parallel gripper. We propose a system that uses point clouds as the state representation and leverages a pre-trained point cloud reconstruction transformer to learn a latent dynamics model to predict material deformations given a grasp action. We design a novel action sampling algorithm that reasons about geometrical differences between point clouds to further improve the efficiency of model-based planners. All data and experiments are conducted entirely in the real world. Our experiments show the proposed system is able to successfully capture the dynamics of clay, and is able to create a variety of simple shapes. Videos and additional figures are available on our project page at: https://sites.google.com/andrew.cmu.edu/sculptbot
Alison Bartsch, Charlotte Avra, Amir Barati Farimani
ICRA3
2024 SculptDiff: Learning Robotic Clay Sculpting from Humans with Goal Conditioned Diffusion Policy
abstract
Manipulating deformable objects remains a challenge within robotics due to the difficulties of state estimation, long-horizon planning, and predicting how the object will deform given an interaction. These challenges are the most pronounced with 3D deformable objects. We propose SculptDiff, a goal-conditioned diffusion-based imitation learning framework that works with point cloud state observations to directly learn clay sculpting policies for a variety of target shapes. To the best of our knowledge this is the first real-world method that successfully learns manipulation policies for 3D deformable objects. For sculpting videos and access to our dataset and hardware CAD models, see the project website: https://sites.google.com/andrew.cmu.edu/imitation-sculpting/home
Alison Bartsch, Arvind Car, Charlotte Avra, Amir Barati Farimani
IROS4
2024 Fluid viscosity prediction leveraging computer vision and robot interaction
abstract
Accurately determining fluid viscosity is crucial for various industrial and scientific applications. Traditional methods of viscosity measurement, though reliable, often require manual intervention and cannot easily adapt to real-time monitoring. With advancements in machine learning and computer vision, this work explores the feasibility of predicting fluid viscosity by analyzing fluid oscillations captured in video data. The pipeline employs a 3D convolutional autoencoder pretrained in a self-supervised manner to extract and learn features from semantic segmentation masks of oscillating fluids. Then, the latent representations of the input data, produced from the pretrained autoencoder, are processed with a distinct inference head to infer either the fluid category (classification) or the fluid viscosity (regression) in a time-resolved manner. When the latent representations generated by the pre-trained autoencoder are used for classification, the system achieves a 97.1% accuracy across a total of 4140 test datapoints. Similarly, for regression tasks, employing an additional fully-connected network as a regression head allows the pipeline to achieve a mean absolute error of 0.258 cP over 4416 test datapoints. This study represents an innovative contribution to both fluid characterization and the evolving landscape of Artificial Intelligence, demonstrating the potential of deep learning in achieving near real-time viscosity estimation and addressing practical challenges in fluid dynamics through the analysis of video data capturing oscillating fluid dynamics. • Semantic segmentation masks extracted from videos can be used for fluid property prediction, namely viscosity. • Visual features from oscillating fluid videos allow for non-invasive, time-resolved analysis. • Latent vectors generated from the autoencoder contain reduced information about the input video data. • Using binary segmentation masks as learning features poses a challenge in discerning fluid with similar viscosity. • Binary segmentation masks contribute to robustness in lighting conditions during video data collection.
Jong Hoon Park, Gauri Pramod Dalwankar, Alison Bartsch, Amir Barati Farimani
Eng. Appl. Artif. Intell.5
2023 Minimizing Human Assistance: Augmenting a Single Demonstration for Deep Reinforcement Learning
abstract
The use of human demonstrations in reinforcement learning has proven to significantly improve agent performance. However, any requirement for a human to manually ‘teach’ the model is somewhat antithetical to the goals of reinforcement learning. This paper attempts to minimize human involvement in the learning process while retaining the performance advantages by using a single human example collected through a simple-to-use virtual reality simulation to assist with RL training. Our method augments a single demonstration to generate numerous human-like demonstrations that, when combined with Deep Deterministic Policy Gradients and Hindsight Experience Replay (DDPG + HER) significantly improve training time on simple tasks and allows the agent to solve a complex task (block stacking) that DDPG + HER alone cannot solve. The model achieves this significant training advantage using a single human example, requiring less than a minute of human input. Moreover, despite learning from a human example, the agent is not constrained to human-level performance, often learning a policy that is significantly different from the human demonstration.
Alison Bartsch, Amir Barati Farimani
ICRA3
2023 Scalable Transformer for PDE Surrogate Modeling
abstract
Transformer has shown state-of-the-art performance on various applications and has recently emerged as a promising tool for surrogate modeling of partial differential equations (PDEs). Despite the introduction of linear-complexity attention, applying Transformer to problems with a large number of grid points can be numerically unstable and computationally expensive. In this work, we propose Factorized Transformer (FactFormer), which is based on an axial factorized kernel integral. Concretely, we introduce a learnable projection operator that decomposes the input function into multiple sub-functions with one-dimensional domain. These sub-functions are then evaluated and used to compute the instance-based kernel with an axial factorized scheme. We showcase that the proposed model is able to simulate 2D Kolmogorov flow on a $256\times 256$ grid and 3D smoke buoyancy on a $64\times64\times64$ grid with good accuracy and efficiency. The proposed factorized scheme can serve as a computationally efficient low-rank surrogate for the full attention scheme when dealing with multi-dimensional problems.
Dule Shu, Amir Barati Farimani
NeurIPS3
2023 Transformer-based Planning for Symbolic Regression
abstract
Symbolic regression (SR) is a challenging task in machine learning that involves finding a mathematical expression for a function based on its values. Recent advancements in SR have demonstrated the effectiveness of pre-trained transformer models in generating equations as sequences, leveraging large-scale pre-training on synthetic datasets and offering notable advantages in terms of inference time over classical Genetic Programming (GP) methods. However, these models primarily rely on supervised pre-training objectives borrowed from text generation and overlook equation discovery goals like accuracy and complexity. To address this, we propose TPSR, a Transformer-based Planning strategy for Symbolic Regression that incorporates Monte Carlo Tree Search planning algorithm into the transformer decoding process. Unlike conventional decoding strategies, TPSR enables the integration of non-differentiable equation verification feedback, such as fitting accuracy and complexity, as external sources of knowledge into the transformer equation generation process. Extensive experiments on various datasets show that our approach outperforms state-of-the-art methods, enhancing the model's fitting-complexity trade-off, extrapolation abilities, and robustness to noise.
Parshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. Reddy
NeurIPS3
2023 Identification of parametric dynamical systems using integer programming
abstract
Identification of nonlinear dynamical systems using data-driven frameworks facilitates the prediction and control of systems in a range of applications. Identification of a single system from the measurements of the system’s states leads to the discovery of explicit or implicit models that cannot generalize beyond the system for which the data are provided. By learning the effect of parameters in the system, we propose a generalizable model for the Identification of Parametric forms of dynamical systems using Integer Programming (IP2). We first build general libraries of basis functions that take into account both states and parameters. Subsequently, leveraging dimension analysis and the assumption of having integer coefficients in the equations, we show that our framework can identify the exact forms of parametric mechanical dynamical systems like an ideal pendulum or an inverted pendulum on a cart. Moreover, by applying object tracking techniques and taking advantage of a sequential filtering scheme, we can identify the state and energy equations of these dynamical systems from videos of the systems, i.e. pixel space noisy data, rather than state-space measurements. The results show that using integer programming makes the proposed framework significantly (more than 40 times in case of inverted pendulum on a cart) more robust to noise compared to previous optimization models.
Kazem Meidani, Amir Barati Farimani
Expert Syst. Appl.2
2022 TPU-GAN: Learning temporal coherence from dynamic point cloud sequences
Tianqin Li, Amir Barati Farimani
ICLR3
2022 Prototype memory and attention mechanisms for few shot image generation
Tianqin Li, Andrew Luo 0001, Harold Rockwell, Amir Barati Farimani, Tai Sing Lee
ICLR5
2022 Graph neural network-accelerated Lagrangian fluid simulation
abstract
We present a data-driven model for fluid simulation under Lagrangian representation. Our model, Fluid Graph Networks (FGN), uses graphs to represent the fluid field. In FGN, fluid particles are represented as nodes and their interactions are represented as edges. Instead of directly predicting the acceleration or position correction given the current state, FGN decomposes the simulation scheme into separate parts — advection, collision, and pressure projection. For these different predictions tasks, we propose two kinds of graph neural network structures, node-focused networks and edge-focused networks. We show that the learned model can produce accurate results and remain stable in scenarios with different geometries. In addition, FGN is able to retain many important physical properties of incompressible fluids, such as low velocity divergence, and adapt to time step sizes beyond the one used in the training set. FGN is also computationally efficient compared to classical simulation methods as it operates on a smaller neighborhood and does not require iteration at each timestep during the inference.
Amir Barati Farimani
Comput. Graph.2
2022 Online metaheuristic algorithm selection
abstract
The performance of optimization algorithms significantly depends on the landscape of the problems. It is known that there is no single algorithm that outperforms others on problems with different fitness landscapes. One of the issues in metaheuristic algorithms is keeping the balance between exploration and exploitation. The features extracted from analysis of fitness landscapes can be used to select the suitable algorithm for the given problem. However, these features are usually expensive and extracted prior to the optimization process which leads to a single algorithm to be selected. In this work, we propose an intelligent switch mechanism that enjoys an efficient non-convex ratio (ENCR) feature extracted online during the optimization to switch between two choices of algorithms, each favoring a type of landscape in terms of modality. For this work, two case studies including a pair of Harris hawks optimizer (HHO) and differential evolution (DE) and another pair of multiverse optimizer (MVO) and moth-flame optimizer (MFO) are selected among several algorithms to evaluate the performance of this framework. The proposed one-way and two-way switch algorithms take advantage of the merits of the two base algorithms to reach better final solutions and higher convergence rates in the majority of case studies. The overall comparison and ranking of the algorithms, including a random switch baseline, demonstrates the superiority of the intelligent switch mechanisms over the baselines.
Kazem Meidani, Seyedali Mirjalili, Amir Barati Farimani
Expert Syst. Appl.3
2022 Dominant motion identification of multi-particle system using deep learning from video
Yayati Jadhav, Amir Barati Farimani
Neural Comput. Appl.2
2022 Adaptive grey wolf optimizer
Kazem Meidani, AmirPouya Hemmasian, Seyedali Mirjalili, Amir Barati Farimani
Neural Comput. Appl.4