EDBT 2026 Demo / reviewers in the wild / expert
Amin Banitalebi-Dehkordi
dblp:136/5269 · also Amin Banitalebi
· DBLP profile ↗
21ranked-venue papers
9as first author
14since 2021 · last 2024
0000-0001-6407-1762ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Skin Tone Disentanglement in 2D Makeup Transfer With Graph Neural NetworksabstractMakeup transfer involves transferring makeup from a reference image to a target image while maintaining the target’s identity. Existing methods, which use Generative Adversarial Networks, often transfer not just makeup but also the reference image’s skin tone. This limits their use to similar skin tones and introduces bias. Our solution introduces a skin tone-robust makeup embedding achieved by augmenting the reference image with varied skin tones. Using Graph Neural Networks, we establish connections between target, reference, and augmented images to create this robust representation that preserves the target’s skin tone. In a user study, our approach outperformed other methods 66% of the time, showcasing its resilience to skin tone variations. Masoud Mokhtari, Fatemeh Taheri Dezaki, Timo Bolkart, Betty Mohler Tesch, Rahul Suresh, Amin Banitalebi-Dehkordi |
ICASSP | 6 |
| 2023 | ArchBERT: Bi-Modal Understanding of Neural Architectures and Natural LanguagesabstractBuilding multi-modal language models has been a trend in the recent years, where additional modalities such as image, video, speech, etc. are jointly learned along with natural languages (i.e., textual information).Despite the success of these multi-modal language models with different modalities, there is no existing solution for neural network architectures and natural languages.Providing neural architectural information as a new modality allows us to provide fast architecture-2-text and text-2-architecture retrieval/generation services on the cloud with a single inference.Such solution is valuable in terms of helping beginner and intermediate ML users to come up with better neural architectures or AutoML approaches with a simple text query.In this paper, we propose ArchBERT, a bi-modal model for joint learning and understanding of neural architectures and natural languages, which opens up new avenues for research in this area.We also introduce a pre-training strategy named Masked Architecture Modeling (MAM) for a more generalized joint learning.Moreover, we introduce and publicly release two new bi-modal datasets for training and validating our methods.The ArchBERT's performance is verified through a set of numerical experiments on different downstream tasks such as architecture-oriented reasoning, question answering, and captioning (summarization).Datasets, codes, and demos are available as supplementary materials 1 . Saeed Ranjbar Alvar, Behnam Kamranian, Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
CoNLL | 4 |
| 2023 | Asynchronous, Option-Based Multi-Agent Policy Gradient: A Conditional Reasoning ApproachabstractCooperative multi-agent problems often require coordination between agents, which can be achieved through a centralized policy that considers the global state. Multi-agent policy gradient (MAPG) methods are commonly used to learn such policies, but they are often limited to problems with low-level action spaces. In complex problems with large state and action spaces, it is advantageous to extend MAPG methods to use higher-level actions, also known as options, to improve the policy search efficiency. However, multi-robot option executions are often asynchronous, that is, agents may select and complete their options at different time steps. This makes it difficult for MAPG methods to derive a centralized policy and evaluate its gradient, as centralized policy always select new options at the same time. In this work, we propose a novel, conditional reasoning approach to address this problem and demonstrate its effectiveness on representative option-based multi-agent cooperative tasks through empirical validation. Find code and videos at: https://sites.google.com/view/mahrlsupp/ Xubo Lyu, Amin Banitalebi-Dehkordi, Mo Chen 0001, Yong Zhang 0004 |
IROS | 2 |
| 2023 | Exact Combinatorial Optimization with Temporo-Attentional Graph Neural Networks
Mehdi Seyfi, Amin Banitalebi-Dehkordi, Zirui Zhou, Yong Zhang 0004 |
ECML/PKDD (4) | 2 |
| 2023 | A survey on adversarial attacks and defenses for object detection and their applications in autonomous vehicles
Abdollah Amirkhani, Mohammad Parsa Karimi, Amin Banitalebi-Dehkordi |
Vis. Comput. | 3 |
| 2022 | E-LANG: Energy-Based Joint Inferencing of Super and Swift Language ModelsabstractBuilding huge and highly capable language models has been a trend in the past years.Despite their great performance, they incur high computational cost.A common solution is to apply model compression or choose light-weight architectures, which often need a separate fixed-size model for each desirable computational budget, and may lose performance in case of heavy compression.This paper proposes an effective dynamic inference approach, called E-LANG, which distributes the inference between large accurate Supermodels and light-weight Swift models.To this end, a decision making module routes the inputs to Super or Swift models based on the energy characteristics of the representations in the latent space.This method is easily adoptable and architecture agnostic.As such, it can be applied to black-box pre-trained models without a need for architectural manipulations, reassembling of modules, or re-training.Unlike existing methods that are only applicable to encoder-only backbones and classification tasks, our method also works for encoderdecoder structures and sequence-to-sequence tasks such as translation.The E-LANG performance is verified through a set of experiments with T5 and BERT backbones on GLUE, Su-perGLUE, and WMT.In particular, we outperform T5-11B with an average computations speed-up of 3.3× on GLUE and 2.9× on SuperGLUE.We also achieve BERT-based SOTA on GLUE with 3.2× less computations.Code and demo are available here. Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
ACL (1) | 2 |
| 2022 | SemAug: Semantically Meaningful Image Augmentations for Object Detection Through Language Grounding
Morgan Lindsay Heisler, Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
ECCV (36) | 2 |
| 2022 | Deep Reinforcement Learning for Exact Combinatorial Optimization: Learning to BranchabstractBranch-and-bound is a systematic enumerative method for combinatorial optimization, where the performance highly relies on the variable selection strategy. State-of-the-art handcrafted heuristic strategies suffer from relatively slow inference time for each selection, while the current machine learning methods require a significant amount of labeled data. We propose a new approach for solving the data labeling and inference latency issues in combinatorial optimization based on the use of the reinforcement learning (RL) paradigm. We use imitation learning to bootstrap an RL agent and then use Proximal Policy Optimization (PPO) to further explore global optimal actions. Then, a value network is used to run Monte-Carlo tree search (MCTS) to enhance the policy network. We evaluate the performance of our method on four different categories of combinatorial optimization problems and show that our approach performs strongly compared to the state-of-the-art machine learning and heuristics based methods. Tianyu Zhang 0003, Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
ICPR | 2 |
| 2022 | Extending Momentum Contrast With Cross Similarity Consistency RegularizationabstractContrastive self-supervised representation learning methods maximize the similarity between the positive pairs, and at the same time tend to minimize the similarity between the negative pairs. However, in general the interplay between the negative pairs is ignored as they do not put in place special mechanisms to treat negative pairs differently according to their specific differences and similarities. In this paper, we present Extended Momentum Contrast (XMoCo), a self-supervised representation learning method founded upon the legacy of the momentum-encoder unit proposed in the MoCo family configurations. To this end, we introduce a cross consistency regularization loss, with which we extend the transformation consistency to dissimilar images (negative pairs). Under the cross consistency regularization rule, we argue that semantic representations associated with any pair of images (positive or negative) should preserve their cross-similarity under pretext transformations. Moreover, we further regularize the training loss by enforcing a uniform distribution of similarity over the negative pairs across a batch. The proposed regularization can easily be added to existing self-supervised learning algorithms in a plug-and-play fashion. Empirically, we report a competitive performance on the standard Imagenet-1K linear head classification benchmark. In addition, by transferring the learned representations to common downstream tasks, we show that using XMoCo with the prevalently utilized augmentations can lead to improvements in the performance of such tasks. We hope the findings of this paper serve as a motivation for researchers to take into consideration the important interplay among the negative examples in self-supervised learning. Mehdi Seyfi, Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | EBJR: Energy-Based Joint Reasoning for Adaptive Inference
Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
BMVC | 2 |
| 2021 | Repaint: Improving the Generalization of Down-Stream Visual Tasks by Generating Multiple Instances of Training Examples
Amin Banitalebi-Dehkordi, Yong Zhang 0004 |
BMVC | 1 |
| 2021 | Model Composition: Can Multiple Neural Networks Be Combined into a Single Network Using Only Unlabeled Data?
Amin Banitalebi-Dehkordi, Xinyu Kang, Yong Zhang 0004 |
BMVC | 1 |
| 2021 | SimROD: A Simple Adaptation Method for Robust Object DetectionabstractThis paper presents a Simple and effective unsupervised adaptation method for Robust Object Detection (SimROD). To overcome the challenging issues of domain shift and pseudo-label noise, our method integrates a novel domain-centric data augmentation, a gradual self-labeling adaptation procedure, and a teacher-guided fine-tuning mechanism. Using our method, target domain samples can be leveraged to adapt object detection models without changing the model architecture or generating synthetic data. When applied to image corruptions and high-level cross-domain adaptation benchmarks, our method outperforms prior baselines on multiple domain adaptation benchmarks. SimROD achieves new state-of-the-art on standard real-to-synthetic and cross-camera setup benchmarks. On the image corruption benchmark, models adapted with our method achieved a relative robustness improvement of 15-25% AP50 on Pascal-C and 5-6% AP on COCO-C and Cityscapes-C. On the cross-domain benchmark, our method outperformed the best baseline performance by up to 8% and 4% AP50 on Comic and Watercolor respectively.1 Rindranirina Ramamonjison, Amin Banitalebi-Dehkordi, Xinyu Kang, Xiaolong Bai, Yong Zhang 0004 |
ICCV | 2 |
| 2021 | Auto-Split: A General Framework of Collaborative Edge-Cloud AIabstractIn many industry scale applications, large and resource consuming machine learning models reside in powerful cloud servers. At the same time, large amounts of input data are collected at the edge of cloud. The inference results are also communicated to users or passed to downstream tasks at the edge. The edge often consists of a large number of low-power devices. It is a big challenge to design industry products to support sophisticated deep model deployment and conduct model inference in an efficient manner so that the model accuracy remains high and the end-to-end latency is kept low. This paper describes the techniques and engineering practice behind Auto-Split, an edge-cloud collaborative prototype of Huawei Cloud. This patented technology is already validated on selected applications, is on its way for broader systematic edge-cloud application integration, and is being made available for public use as an automated pipeline service for end-to-end cloud-edge collaborative intelligence deployment. To the best of our knowledge, there is no existing industry product that provides the capability of Deep Neural Network (DNN) splitting. Amin Banitalebi-Dehkordi, Naveen Vedula, Jian Pei 0001, Lanjun Wang, Yong Zhang 0004 |
KDD | 1 |
| 2018 | Face recognition using a new compressive sensing-based feature extraction method
Mehdi Banitalebi Dehkordi, Amin Banitalebi-Dehkordi, Jamshid Abouei, Konstantinos N. Plataniotis |
Multim. Tools Appl. | 2 |
| 2018 | Saliency inspired quality assessment of stereoscopic 3D video
Amin Banitalebi-Dehkordi, Panos Nasiopoulos |
Multim. Tools Appl. | 1 |
| 2017 | A learning-based visual saliency prediction model for stereoscopic 3D video (LBVS-3D)
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
Multim. Tools Appl. | 1 |
| 2016 | An efficient human visual system based quality metric for 3D video
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
Multim. Tools Appl. | 1 |
| 2014 | Effect of eye dominance on the perception of stereoscopic 3D videoabstractAsymmetric schemes have widespread applications in the 3D video transmission pipeline. The significance of eye dominance becomes a concern when designing such schemes. In this paper, in order to investigate the effect of eye dominance on the perceptual 3D video quality, a database of representative asymmetric stereoscopic sequences is prepared and the overall 3D quality of these sequences is evaluated through subjective experiments. Experiment results showed that viewers find an asymmetric video more pleasant when the view with higher quality is projected to their dominant eye. Moreover, the eye dominance changes the mean opinion quality score by 16 % at most, a result caused by slight asymmetric video compression. For all other representative types of asymmetry, the statistical difference is much lower and in some cases even negligible. Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
ICIP | 1 |
| 2014 | Compression of high dynamic range video using the HEVC and H.264/AVC standardsabstractThe existing video coding standards such as H.264/AVC and High Efficiency Video Coding (HEVC) have been designed based on the statistical properties of Low Dynamic Range (LDR) videos and are not accustomed to the characteristics of High Dynamic Range (HDR) content. In this study, we investigate the performance of the latest LDR video compression standard, HEVC, as well as the recent widely commercially used video compression standard, H.264/AVC, on HDR content. Subjective evaluations of results on an HDR display show that viewers clearly prefer the videos coded via an HEVC-based encoder to the ones encoded using an H.264/AVC encoder. In particular, HEVC outperforms H.264/AVC by an average of 10.18% in terms of mean opinion score and 25.08% in terms of bit rate savings. Amin Banitalebi-Dehkordi, Mehran Azimi, Mahsa T. Pourazad, Panos Nasiopoulos |
QSHINE | 1 |
| 2013 | 3D video quality metric for mobile applicationsabstractIn this paper, we propose a new full-reference quality metric for mobile 3D content. Our method is modeled around the Human Visual System, fusing the information of both left and right channels, considering color components, the cyclopean views of the two videos and disparity. Our method is assessing the quality of 3D videos displayed on a mobile 3DTV, taking into account the effect of resolution, distance from the viewers' eyes, and dimensions of the mobile display. Performance evaluations showed that our mobile 3D quality metric monitors the degradation of quality caused by several representative types of distortion with 82% correlation with results of subjective tests, an accuracy much better than that of the state-of-the-art mobile 3D quality metric. Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 1 |