Alireza Rezazadeh

dblp:19/7991 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-2457-9470ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 From Isolated Conversations to Hierarchical Schemas: Dynamic Tree Memory Representation for LLMs
abstract
Recent advancements in large language models have significantly improved their context windows, yet challenges in effective long-term memory management remain. We introduce MemTree, an algorithm that leverages a dynamic, tree-structured memory representation to optimize the organization, retrieval, and integration of information, akin to human cognitive schemas. MemTree organizes memory hierarchically, with each node encapsulating aggregated textual content, corresponding semantic embeddings, and varying abstraction levels across the tree's depths. Our algorithm dynamically adapts this memory structure by computing and comparing semantic embeddings of new and existing information to enrich the model’s context-awareness. This approach allows MemTree to handle complex reasoning and extended interactions more effectively than traditional memory augmentation methods, which often rely on flat lookup tables. Evaluations on benchmarks for multi-turn dialogue understanding and document question answering show that MemTree significantly enhances performance in scenarios that demand structured memory management.
Alireza Rezazadeh, Wei Wei 0019, Yujia Bao
ICLR1
2025 A Parameter-Efficient Tuning Framework for Language-Guided Object Grounding and Robot Grasping
abstract
The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their extensive computation and data demands limit the feasibility of local deployment and customization. To address this, we propose a novel CLIP-based [1] multimodal parameter-efficient tuning (PET) framework designed for three language-guided object grounding and grasping tasks: (1) Referring Expression Segmentation (RES), (2) Referring Grasp Synthesis (RGS), and (3) Referring Grasp Affordance (RGA). Our approach introduces two key innovations: a bi-directional vision-language adapter that aligns multimodal inputs for pixel-level language understanding and a depth fusion branch that incorporates geometric cues to facilitate robot grasping predictions. Experiment results demonstrate superior performance in the RES object grounding task compared with existing CLIP-based full-model tuning or PET approaches. In the RGS and RGA tasks, our model not only effectively interprets object attributes based on simple language descriptions but also shows strong potential for comprehending complex spatial reasoning scenarios, such as multiple identical objects present in the workspace. Project page: https://z.umn.edu/etog-etrg
Houjian Yu, Mingen Li, Alireza Rezazadeh, Yang Yang 0083, Changhyun Choi
ICRA3
2025 InvSlotGNN: Unsupervised Discovery of Viewpoint Invariant Multiobject Representations and Visual Dynamics
abstract
Learning multiobject dynamics purely from visual data is challenging due to the need for robust object representations that can be learned through robot interactions. In previous work (Rezazadeh et al., 2023), we introduced two novel architectures: SlotTransport for discovering object-centric representations from singleview RGB images, referred to as slots, and SlotGNN for predicting scene dynamics from singleview RGB images and robot interactions using the discovered slots. This article introduces InvSlotGNN, a novel framework for learning multiview slot discovery and dynamics that are invariant to the camera viewpoint. First, we demonstrate that SlotTransport can be trained on multiview data such that a single model discovers temporally aligned, object-centric representations from a wide range of different camera angles. These slots bind to objects from various viewpoints, even under occlusion or absence. Next, we introduce InvSlotGNN, an extension of SlotGNN, that learns multiobject dynamics invariant to the camera angle and predicts the future state from observations taken by uncalibrated cameras. InvSlotGNN learns a graph representation of the scene using the slots from SlotTransport and performs relational and spatial reasoning to predict the future state of the scene for arbitrary viewpoints, conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning multiview object-centric features that accurately encode visual and positional information. Furthermore, we highlight the accuracy of InvSlotGNN in downstream robotic tasks, including long-horizon prediction and multiobject rearrangement. Finally, with minimal real data, our framework robustly predicts slots and their dynamics in real-world multiview scenarios.
Alireza Rezazadeh, Houjian Yu, Karthik Desingh, Changhyun Choi
IEEE Trans. Robotics1
2024 SlotGNN: Unsupervised Discovery of Multi-Object Representations and Visual Dynamics
abstract
Learning multi-object dynamics from visual data using unsupervised techniques is challenging due to the need for robust, object representations that can be learned through robot interactions. This paper presents a novel framework with two new architectures: SlotTransport for discovering object representations from RGB images and SlotGNN for predicting their collective dynamics from RGB images and robot interactions. Our SlotTransport architecture is based on slot attention for unsupervised object discovery and uses a feature transport mechanism to maintain temporal alignment in object-centric representations. This enables the discovery of slots that consistently reflect the composition of multi-object scenes. These slots robustly bind to distinct objects, even under heavy occlusion or absence. Our SlotGNN, a novel unsupervised graph-based dynamics model, predicts the future state of multi-object scenes. SlotGNN learns a graph representation of the scene using the discovered slots from SlotTransport and performs relational and spatial reasoning to predict the future appearance of each slot conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning object-centric features that accurately encode both visual and positional information. Further, we highlight the accuracy of SlotGNN in downstream robotic tasks, including challenging multi-object rearrangement and long-horizon prediction. Finally, our unsupervised approach proves effective in the real world. With only minimal additional data, our framework robustly predicts slots and their corresponding dynamics in real-world control tasks. Our project webpage: bit.ly/slotgnn.
Alireza Rezazadeh, Athreyi Badithela, Karthik Desingh, Changhyun Choi
ICRA1
2023 Hierarchical Graph Neural Networks for Proprioceptive 6D Pose Estimation of In-hand Objects
abstract
Robotic manipulation, in particular in-hand object manipulation, often requires an accurate estimate of the object's 6D pose. To improve the accuracy of the estimated pose, state-of-the-art approaches in 6D object pose estimation use observational data from one or more modalities, e.g., RGB images, depth, and tactile readings. However, existing approaches make limited use of the underlying geometric structure of the object captured by these modalities, thereby, increasing their reliance on visual features. This results in poor performance when presented with objects that lack such visual features or when visual features are simply occluded. Furthermore, current approaches do not take advantage of the proprioceptive information embedded in the position of the fingers. To address these limitations, in this paper: (1) we introduce a hierarchical graph neural network architecture for combining multimodal (vision and touch) data that allows for a geometrically informed 6D object pose estimation, (2) we introduce a hierarchical message passing operation that flows the information within and across modalities to learn a graph-based object representation, and (3) we introduce a method that accounts for the proprioceptive information for in-hand object representation. We evaluate our model on a diverse subset of objects from the YCB Object and Model Set, and show that our method substantially outperforms existing state-of-the-art work in accuracy and robustness to occlusion. We also deploy our proposed framework on a real robot and qualitatively demonstrate successful transfer to real settings.
Alireza Rezazadeh, Snehal Dikhale, Soshi Iba, Nawid Jamali
ICRA1
2011 A new artificial bee swarm algorithm for optimization of proton exchange membrane fuel cell model parameters
abstract
An appropriate mathematical model can help researchers to simulate, evaluate, and control a proton exchange membrane fuel cell (PEMFC) stack system. Because a PEMFC is a nonlinear and strongly coupled system, many assumptions and approximations are considered during modeling. Therefore, some differences are found between model results and the real performance of PEMFCs. To increase the precision of the models so that they can describe better the actual performance, optimization of PEMFC model parameters is essential. In this paper, an artificial bee swarm optimization algorithm, called ABSO, is proposed for optimizing the parameters of a steady-state PEMFC stack model suitable for electrical engineering applications. For studying the usefulness of the proposed algorithm, ABSO-based results are compared with the results from a genetic algorithm (GA) and particle swarm optimization (PSO). The results show that the ABSO algorithm outperforms the other algorithms.
Alireza Askarzadeh, Alireza Rezazadeh
J. Zhejiang Univ. Sci. C2
2010 Coordination of PSS and TCSC controller using modified particle swarm optimization algorithm to improve power system dynamic performance
abstract
This paper develops a modified optimization procedure for coordination of a power system stabilizer (PSS) and a thyristor controlled series compensator (TCSC) controller to enhance the power system small signal stability. The new approach employs eigenvalue-based and time-domain simulation based objective functions simultaneously to improve the optimization convergence rate. A modified particle swarm optimization (MPSO) algorithm is used as the optimization algorithm. The results of simulations and eigenvalue analysis for a single machine infinite bus (SMIB) system equipped with the proposed PSS and TCSC controllers confirm that the new approach is effective in enhancing the system stability.
Alireza Rezazadeh, Mostafa Sedighizadeh, Ahmad Hasaninia
J. Zhejiang Univ. Sci. C1