Asim Munawar

dblp:05/1134 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 1 since 2021Systems, architecture and hardware · 7 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 38% Language models and text generation · 25% Knowledge representation and reasoning · 18%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 33% Electronic design automation · 33% Embedded and real-time systems · 33%

Topics — the 20 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls · EMNLP 2025
Natural language and speech › Language models and text generation
alignment
0.812024
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback · ICML 2024
Machine learning › Transfer learning and domain adaptation
distribution matching
0.812024
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.812024
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback · ICML 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback · ICML 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
0.812024
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback · ICML 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
abstract meaning representation
0.712023
Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning · ACL (1) 2023
Machine learning › Reinforcement learning
textual reinforcement learning
0.712023
Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › rule learning
differentiable rule learning
0.512021
Neuro-Symbolic Approaches for Text-Based Policy Learning · EMNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning
0.512021
Neuro-Symbolic Approaches for Text-Based Policy Learning · EMNLP (1) 2021
Machine learning › Reinforcement learning › knowledge-based reinforcement learning
neuro-symbolic reinforcement learning
0.512021
Neuro-Symbolic Reinforcement Learning with First-Order Logic · EMNLP (1) 2021
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.412020
Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games · EMNLP (1) 2020
Robotics › Robot manipulation › assembly › assembly task
assembly task execution
0.312018
MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018
Human-robot interaction
human-robot collaboration
0.312018
MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018
Natural language and speech › Language models and text generation › code generation
API call generation
0.212024
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs · ACL (1) 2024
Embedded and real-time systems › automotive embedded systems
automotive e/e architectures
0.212014
Scalable Co-Simulation of Functional Models With Accurate Event Exchange · DAC 2014
Electronic design automation › hardware/software co-design
co-simulation
0.212014
Scalable Co-Simulation of Functional Models With Accurate Event Exchange · DAC 2014
Performance modeling and evaluation
simulation
0.212014
Scalable Co-Simulation of Functional Models With Accurate Event Exchange · DAC 2014
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
first-order logic
0.112021
Neuro-Symbolic Reinforcement Learning with First-Order Logic · EMNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
rule and ontology reasoning
0.112018
MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018

Methods — techniques the papers use, named apart from their topics

benchmark construction · 1.7self-normalized baseline · 0.8instruction tuning · 0.8distributional policy gradient · 0.8bayes' rule · 0.8symbolic rule learning · 0.7symbolic policy learning · 0.5reinforcement learning · 0.5first-order logic · 0.5end-to-end differentiable rule learning · 0.5symbolic planning · 0.3ontology · 0.3natural language understanding · 0.3functional simulation · 0.2event exchange · 0.2
YearPublicationVenuePosition
2025 NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
abstract
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Xin Wang, Luis A. Lastras, Pavan Kapanipathi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kinjal Basu 0002, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Luis A. Lastras, Pavan Kapanipathi
EMNLP8
2024 API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
abstract
Kinjal Basu, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan, Maxwell Crouse, Asim Munawar, Vernon Austel, Sadhana Kumaravel, Vinod Muthusamy, Pavan Kapanipathi, Luis Lastras. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kinjal Basu 0002, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan, Maxwell Crouse, Asim Munawar, Vernon Austel, Sadhana Kumaravel, Vinod Muthusamy, Pavan Kapanipathi, Luis A. Lastras
ACL (1)6
2024 BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
abstract
Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
Gaurav Pandey 0001, Yatin Nandwani, Tahira Naseem, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo
ICML8
2023 Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning
abstract
Subhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander Gray. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Subhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander G. Gray
ACL (1)10
2021 Neuro-Symbolic Approaches for Text-Based Policy Learning
abstract
Text-Based Games (TBGs) have emerged as important testbeds for reinforcement learning (RL) in the natural language domain.Previous methods using LSTM-based action policies are uninterpretable and often overfit the training games showing poor performance to unseen test games.We present SymboLic Action policy for Textual Environments (SLATE), that learns interpretable action policy rules from symbolic abstractions of textual observations for improved generalization.We outline a method for end-to-end differentiable symbolic rule learning and show that such symbolic policies outperform previous stateof-the-art methods in text-based RL for the coin collector environment from 5 -10x fewer training games.Additionally, our method provides human-understandable policy rules that can be readily verified for their logical consistency and can be easily debugged.1
Subhajit Chaudhury, Prithviraj Sen, Masaki Ono, Daiki Kimura, Michiaki Tatsubori, Asim Munawar
EMNLP (1)6
2021 Neuro-Symbolic Reinforcement Learning with First-Order Logic
abstract
Daiki Kimura, Masaki Ono, Subhajit Chaudhury, Ryosuke Kohita, Akifumi Wachi, Don Joven Agravante, Michiaki Tatsubori, Asim Munawar, Alexander Gray. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Daiki Kimura, Masaki Ono, Subhajit Chaudhury, Ryosuke Kohita, Akifumi Wachi, Don Joven Agravante, Michiaki Tatsubori, Asim Munawar, Alexander G. Gray
EMNLP (1)8
2021 Data-Efficient Framework for Real-World Multiple Sound Source 2d Localization
abstract
Deep neural networks have recently led to promising results for the task of multiple sound source localization. Yet, they require a lot of training data to cover a variety of acoustic conditions and micro-phone array layouts. One can leverage acoustic simulators to inexpensively generate labeled training data. However, models trained on synthetic data tend to perform poorly with real-world recordings due to the domain mismatch. Moreover, learning for different microphone array layouts makes the task more complicated due to the infinite number of possible layouts. We propose to use adversarial learning methods to close the gap between synthetic and real do-mains. Our novel ensemble-discrimination method significantly improves the localization performance without requiring any label from the real data. Furthermore, we propose a novel explicit transformation layer to be embedded in the localization architecture. It enables the model to be trained with data from specific microphone array layouts while generalizing well to unseen layouts during inference.
Guillaume Le Moing, Phongtharin Vinayavekhin, Don Joven Agravante, Tadanobu Inoue, Jayakorn Vongkulbhisal, Asim Munawar, Ryuki Tachibana
ICASSP6
2020 Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games
abstract
Subhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Subhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana
EMNLP (1)5
2020 Adversarial Discriminative Attention for Robust Anomaly Detection
abstract
Existing methods for visual anomaly detection predominantly rely on global level pixel comparisons for anomaly score computation without emphasizing on unique local features. However, images from real-world applications are susceptible to unwanted noise and distractions, that might jeopardize the robustness of such anomaly score. To alleviate this problem, we propose a self-supervised masking method that specifically focuses on discriminative parts of images to enable robust anomaly detection. Our experiments reveal that discriminator's class activation map in adversarial training evolves in three stages and finally fixates on the foreground location in the images. Using this property of the activation map, we construct a mask that suppresses spurious signals from the background thus enabling robust anomaly detection by focusing on local discriminative attributes. Additionally, our method can further improve the accuracy by learning a semi-supervised discriminative classifier in cases where a few samples from anomaly classes are available during the training. Experimental evaluations on four different types of datasets demonstrate that our method outperforms previous state-of-the-art methods for each condition and in all domains.
Daiki Kimura, Subhajit Chaudhury, Minori Narita, Asim Munawar, Ryuki Tachibana
WACV4
2019 Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports Videos
abstract
Image-based sports analytics enable automatic retrieval of key events in a game to speed up the analytics process for human experts. However, most existing methods focus on structured television broadcast video datasets with a straight and fixed camera having minimum variability in the capturing pose. In this paper, we study the case of event detection in sports videos for unstructured environments with arbitrary camera angles. The transition from structured to unstructured video analysis produces multiple challenges that we address in our paper. Specifically, we identify and solve two major problems: unsupervised identification of players in an unstructured setting and generalization of the trained models to pose variations due to arbitrary shooting angles. For the first problem, we propose a temporal feature aggregation algorithm using person re-identification features to obtain high player retrieval precision by boosting a weak heuristic scoring method. Additionally, we propose a data augmentation technique, based on multi-modal image translation model, to reduce bias in the appearance of training samples. Experimental evaluations show that our proposed method improves precision for player retrieval from 0.78 to 0.86 for obliquely angled videos. Additionally, we obtain an improvement in F1 score for rally detection in table tennis videos from 0.79 in case of global frame-level features to 0.89 using our proposed player-level features. Please see the supplementary video submission at https://ibm.biz/BdzeZA.
Subhajit Chaudhury, Hiroki Ozaki, Daiki Kimura, Phongtharin Vinayavekhin, Asim Munawar, Ryuki Tachibana, Koji Ito, Yuki Inaba, Minoru Matsumoto, Shuji Kidokoro
ISM5
2019 Injective State-Image Mapping facilitates Visual Adversarial Imitation Learning
abstract
The growing use of virtual autonomous agents in applications like games and entertainment demands better control policies for natural-looking movements and actions. Unlike the conventional approach of hard-coding motion routines, we propose a deep learning method for obtaining control policies by directly mimicking raw video demonstrations. Previous methods in this domain rely on extracting low-dimensional features from expert videos followed by a separate hand-crafted reward estimation step. We propose an imitation learning framework that reduces the dependence on hand-engineered reward functions by jointly learning the feature extraction and reward estimation steps using Generative Adversarial Networks (GANs). Our main contribution in this paper is to show that under injective mapping between low-level joint state (angles and velocities) trajectories and corresponding raw video stream, performing adversarial imitation learning on video demonstrations is equivalent to learning from the state trajectories. Experimental results show that the proposed adversarial learning method from raw videos produces a similar performance to state-of-the-art imitation learning techniques while frequently outperforming existing hand-crafted video imitation methods. Furthermore, we show that our method can learn action policies by imitating video demonstrations on YouTube with similar performance to learned agents from true reward signal. Please see the supplementary video submission at https://ibm.biz/BdzzNA.
Subhajit Chaudhury, Daiki Kimura, Asim Munawar, Ryuki Tachibana
MMSP3
2019 Learning Multiple Sound Source 2D Localization
abstract
In this paper, we propose novel deep learning based algorithms for multiple sound source localization. Specifically, we aim to find the 2D Cartesian coordinates of multiple sound sources in an enclosed environment by using multiple microphone arrays. To this end, we use an encoding-decoding architecture and propose two improvements on it to accomplish the task. In addition, we also propose two novel localization representations which increase the accuracy. Lastly, new metrics are developed relying on resolution-based multiple source association which enables us to evaluate and compare different localization approaches. We tested our method on both synthetic and real world data. The results show that our method improves upon the previous baseline approach for this problem.
Guillaume Le Moing, Phongtharin Vinayavekhin, Tadanobu Inoue, Jayakorn Vongkulbhisal, Asim Munawar, Ryuki Tachibana, Don Joven Agravante
MMSP5
2018 Focusing on What is Relevant: Time-Series Learning and Understanding using Attention
abstract
This paper is a contribution towards interpretability of the deep learning models in different applications of time-series. We propose a temporal attention layer that is capable of selecting the relevant information to perform various tasks, including data completion, key-frame detection and classification. The method uses the whole input sequence to calculate an attention value for each time step. This results in more focused attention values and more plausible visualisation than previous methods. We apply the proposed method to three different tasks. Experimental results show that the proposed network produces comparable results to a state of the art. In addition, the network provides better interpretability of the decision, that is, it generates more significant attention weight to related frames compared to similar techniques attempted in the past.
Phongtharin Vinayavekhin, Subhajit Chaudhury, Asim Munawar, Don Joven Agravante, Giovanni De Magistris, Daiki Kimura, Ryuki Tachibana
ICPR3
2018 MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning
abstract
This paper describes a framework called MaestROBe It is designed to make the robots perform complex tasks with high precision by simple high-level instructions given by natural language or demonstration. To realize this, it handles a hierarchical structure by using the knowledge stored in the forms of ontology and rules for bridging among different levels of instructions. Accordingly, the framework has multiple layers of processing components; perception and actuation control at the low level, symbolic planner and Watson APIs for cognitive capabilities and semantic understanding, and orchestration of these components by a new open source robot middleware called Project Intu at its core. We show how this framework can be used in a complex scenario where multiple actors (human, a communication robot, and an industrial robot) collaborate to perform a common industrial task. Human teaches an assembly task to Pepper (a humanoid robot from SoftBank Robotics) using natural language conversation and demonstration. Our framework helps Pepper perceive the human demonstration and generate a sequence of actions for UR5 (collaborative robot arm from Universal Robots), which ultimately performs the assembly (e.g. insertion) task.
Asim Munawar, Giovanni De Magistris, Tu-Hoa Pham, Daiki Kimura, Michiaki Tatsubori, Takao Moriyama, Ryuki Tachibana, Grady Booch
ICRA1
2017 Deep reinforcement learning for high precision assembly tasks
abstract
The high precision assembly of mechanical parts requires precision that exceeds that of robots. Conventional part-mating methods used in the current manufacturing require numerous parameters to be tediously tuned before deployment. We show how a robot can successfully perform a peg-in-hole task with a tight clearance through training a recurrent neural network with reinforcement learning. In addition to reducing manual effort, the proposed method also shows a better fitting performance with a tighter clearance and robustness against positional and angular errors for the peg-in-hole task. The neural network learns to take the optimal action by observing the sensors of a robot to estimate the system state. The advantages of our proposed method are validated experimentally on a 7-axis articulated robot arm.
Tadanobu Inoue, Giovanni De Magistris, Asim Munawar, Tsuyoshi Yokoya, Ryuki Tachibana
IROS3
2017 Spatio-Temporal Anomaly Detection for Industrial Robots through Prediction in Unsupervised Feature Space
abstract
Spatio-temporal anomaly detection by unsupervised learning have applications in a wide range of practical settings. In this paper we present a surveillance system for industrial robots using a monocular camera. We propose a new unsupervised learning method to train a deep feature extractor from unlabeled images. Without any data augmentation, the algorithm co-learns the network parameters on different pseudo-classes simultaneously to create unbiased feature representation. Combining the learned features with a prediction system, we can detect irregularities in high dimensional data feed (e.g. video of a robot performing pick and place task). The results show how the proposed approach can detect previously unseen anomalies in the robot surveillance video. Although the technique is not designed for classification, we show the use of the learned features in a more traditional classification application for CIFAR-10 dataset.
Asim Munawar, Phongtharin Vinayavekhin, Giovanni De Magistris
WACV1
2016 Low-computation egocentric barcode detector for the blind
abstract
Linear barcodes are the principal labeling system for retail products. Barcode reader apps found on smartphones always assume that the localization and framing of the barcode is performed manually by a sighted human operator. This is problematic for visually-impaired people since they don't know where to position the camera to scan the barcode. To solve this problem we propose a hand-free interface to detect barcode using a wearable camera. The user rotate a query product in front of him/her and is informed when and where the barcode is visible. The challenge is to detect small barcodes at arm's length in a video with potentially large motion blur. In this paper we propose a novel technique for barcode detection using very little computation (adapted to wearable systems), presenting very good robustness to blur and size variations, and able to run on HD video streams. The proposed system perform significantly better than the state-of-the-art methods on existing public datasets, while being much faster. A new and challenging egocentric product video dataset is also provided with this paper.
Clement Creusot, Asim Munawar
ICIP2
2015 Real-time small obstacle detection on highways using compressive RBM road reconstruction
abstract
Small objects on the road can become hazardous obstacles when driving at high speed. Detecting such obstacles is vital to guaranty the safety of self-driving car users, especially on highways. Such tasks cannot be performed using existing active sensors such as radar or LIDAR due to their limited range and resolution at long distances. In this paper we propose a technique to detect anomalous patches on the road from color images using a Restricted Boltzman Machine neural network specifically trained to reconstruct the appearance of the road. The differences between the observed and reconstructed road patches yield a more relevant segmentation of anomalies than classic image processing techniques. We evaluated our technique on texture-based synthetic datasets as well as on real video footage of anomalous objects on highways.
Clement Creusot, Asim Munawar
Intelligent Vehicles Symposium2
2015 Real-Time Barcode Detection in the Wild
abstract
The linear 1D barcode is the main tagging system for billions of products sold each day. Barcodes have many advantages but require a laser scanner for fast and robust scanning. Solutions exist to read barcodes from cell phones but they assume a carefully framed image within the field of view. This undermines the true potential of barcodes in a wide range of scenarios. In this paper we present a real time technique to detect barcodes in the wild from video streams. Our technique outperforms the state-of-the-art passive techniques both in accuracy and speed. Potential commercial applications enabled by such passive scanning system are also discussed in this paper.
Clement Creusot, Asim Munawar
WACV2
2014 Scalable Co-Simulation of Functional Models With Accurate Event Exchange
abstract
Recent trends in the automotive industry have forced OEMs and suppliers to adopt simulation-driven development processes. To meet the demands of the rapidly changing industry, simulation technologies must innovate to provide faster and accurate simulations of increasingly complex models. In this paper we present a set of technologies that were developed to provide a no-compromise, high fidelity, scalable simulation of heterogeneous functional models. Using our new technology we were able to produce more accurate results in orders of magnitude less time than other frequently used techniques in the industry. The proposed approach to co-simulation allows us to perform functional simulation of full-vehicle without sacrificing the solution accuracy or simulation speed.
Asim Munawar, Shuichi Shimizu
DAC1
2013 ISOMER: integrated selection, partitioning, and placement methodology for reconfigurable architectures
abstract
Quality system design on dynamic partially reconfigurable platform needs exploration of a vast and multidimensional design space for (1) selection among implementation variants of hardware accelerators, (2) partitioning the reconfigurable fabric, and (3) their placement on the reconfigurable fabric partitions. This paper presents a novel methodology ISOMER for integrated solution of selection, partitioning and placement for performance optimization. Architecture under consideration is a general purpose processor coupled with reconfigurable fabric that can be partitioned in multi-sized partially reconfigurable bins. Our methodology determines performance-efficient partitioning and usage of reconfigurable fabric. Extensive evaluation illustrates that our methodology is scalable and outperforms state-of-the-art techniques for non-partially reconfigurable architectures.
Rana Muhammad Bilal, Rehan Hafiz, Muhammad Shafique 0001, Saad Shoaib, Asim Munawar, Jörg Henkel
ICCAD5
2011 Advanced genetic algorithm to solve MINLP problems over GPU
abstract
In this paper we propose a many-core implementation of evolutionary computation for GPGPU (General-Purpose Graphic Processing Unit) to solve non-convex Mixed Integer Non-Linear Programming (MINLP) and non-convex Non Linear Programming (NLP) problems using a stochastic algorithm. Stochastic algorithms being random in their behavior are difficult to implement over GPU like architectures. In this paper we not only succeed in implementation of a stochastic algorithm over GPU but show considerable speedups over CPU implementations. The stochastic algorithm considered for this paper is an adaptive resolution approach to genetic algorithm (arGA), developed by the authors of this paper. The technique uses the entropy measure of each variable to adjust the intensity of the genetic search around promising individuals. Performance is further improved by hybridization with adaptive resolution local search (arLS) operator. In this paper, we describe the challenges and design choices involved in parallelization of this algorithm to solve complex MINLPs over a commodity GPU using Compute Unified Device Architecture (CUDA) programming model. Results section shows several numerical tests and performance measurements obtained by running the algorithm over an nVidia Fermi GPU. We show that for difficult problems we can obtain a speedup of up to 20x with double precision and up to 42x with single precision.
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama
IEEE Congress on Evolutionary Computation1
2011 Optimization of parallel Genetic Algorithms for nVidia GPUs
abstract
Led by General Purpose computing over Graphical Processing Units (GPGPUs), the parallel computing area is witnessing a rapid change in dominant parallel systems. A major hurdle in this switch is the Single Instruction Multiple Thread (SIMT) architecture of GPUs which is usually not suitable for the design of legacy parallel algorithms. Genetic Algorithms (GAs) is no exception for that. GAs are commonly parallelized due to the high demanding computational needs. Given the performance of GPGPUs, the need to best exploit them to maximize computing efficiency for parallel GAs is demandingly growing. The goal of this paper is to shed light on the challenges parallel GAs designers/programmers will likely face while trying to achieve this, and to provide some practical advice on how to maximize GPGPU exploitation as a result. To that end, this paper provides a study on adapting legacy parallel GAs on GPGPU systems. The paper exposes the design challenges of nVidia's GPU architecture to the parallel GAs community by: discussing features of GPU, reviewing design issues in GPU relevant to parallel GAs, the design and introduction of new techniques to achieve an efficient implementation for parallel GAs and observing the effect of the pivotal points that both capitalize on the strengths of GPU and limit the deficiencies/overheads of GPUs. The paper demonstrates the performance of designed-for-GPGPU parallel GAs representing the entire spectrum of legacy parallel model of GAs over nVidia Tesla C1060 workstation showing a significant improvement in performance after optimizing and tuning the algorithms for GPU.
Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama
IEEE Congress on Evolutionary Computation2
2011 A Framework for Cloud Embedded Web Services Utilized by Cloud Applications
abstract
Cloud computing is impacting the modern Internet computing and businesses in every aspect. One feature of clouds is the convenience of using the services offered by the cloud. Consequently, most cloud service providers use WS for users and developers to interface with the cloud. However, the current cloud WS are focused into core and fundamental modern computing functionalities. We anticipate as cloud developments tools mature and cloud applications become more popular, there will be an opportunity for designing and implementing applications/services to be embedded in the cloud for use by applications in the cloud. We propose a framework for WS deployment in the cloud to be usable by applications residing in the same cloud. The framework capitalizes on the cloud strong points to offer a higher value to the service consumer inside the cloud. The authoritative nature of clouds would enable more efficient models for WS publishing, indexing and description. Moreover, being hosted in the cloud, WScan build on the high scalability offered by the cloud with a much higher reliability. Finally, scheduling the instances using the WS in bundle with the WS instances could offer a LAN-like connectivity performance driving down the latency to the magnitude of lower microseconds. In this paper, we highlight the challenges and opportunities of cloud applications using cloud embedded Web services. We give a description of the different aspects by illustrating the different components, together with an end-to-end use case to show the applicability of the proposed system.
Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama
SERVICES2
2010 A Bayesian Optimization Algorithm for De Novo ligand design based docking running over GPU
abstract
A principal fragment-based design approach is De Novo ligand design at which small-molecule structures from a database of existing compounds (or compounds that could be made) are docked into the protein binding site following a virtual synthesis scheme. New virtual structures can easily be constructed from combinatorial building blocks. Typically, tens of thousands of orientations are generated for each ligand candidate, therefore global optimization algorithms are usually employed to search the chemical space by generating new molecular structures through probing many different fragments in a combinatorial fashion. We propose using Bayesian Optimization Algorithm (BOA), a meta-heuristic algorithm, in searching the combination of pre-docked fragments through minimizing the energy of ligand-receptor docking. We further introduce the use of GPU (Graphical Processing Unit) to overcome the very long time required in evaluating each possible fragment combination. We show how the GPU utilization enables experimenting larger fragments and target receptors for more complex instances. The experiments resulted in regenerating three drug-like compounds defined in the ZINC database as well as finding a new compound. The Results show how the nVidia's Tesla C1060 GPU was utilized to accelerate the docking process by two orders of magnitude.
Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama
IEEE Congress on Evolutionary Computation2
2010 The design, usage, and performance of GridUFO: A Grid based Unified Framework for Optimization
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama
Future Gener. Comput. Syst.1
2009 Theoretical and Empirical Analysis of a GPU Based Parallel Bayesian Optimization Algorithm
abstract
General purpose computing over graphical processing units (GPGPUs) is a huge shift of paradigm in parallel computing that promises a dramatic increase in performance. But GPGPUs also bring an unprecedented level of complexity in algorithmic design and software development. In this paper we describe the challenges and design choices involved in parallelization of Bayesian optimization algorithm (BOA) to solve complex combinatorial optimization problems over nVidia commodity graphics hardware using compute unified device architecture (CUDA). BOA is a well-known multivariate estimation of distribution algorithm (EDA) that incorporates methods for learning Bayesian network (BN). It then uses BN to sample new promising solutions. Our implementation is fully compatible with modern commodity GPUs and therefore we call it gBOA (BOA on GPU). In the results section, we show several numerical tests and performance measurements obtained by running gBOA over an nVidia Tesla C1060 GPU. We show that in the best case we can obtain a speedup of up to 13x.
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama
PDCAT1
2008 Solving Large Instances of Capacitated Vehicle Routing Problem over Cell BE
abstract
This paper presents a method to solve large instances of capacitated vehicle routing problem (CVRP) using cellular genetic algorithm (cGA) with local search (LS) over cell broadband engine (cell BE) architecture. We propose a unique parallelization model where computationally intensive local search (LS) runs on the available synergistic processing elements (SPEs) in parallel, while the power processor element (PPE) runs the cGA and acts as a controller for all the SPEs. We reproduce the results from earlier work in PPE only implementation of the algorithm, and we show a considerable reduction in execution time for parallel implementation over cell BE. Moreover, we extended it further to solve larger instances of CVRP (compared to the ones present in the CVRP literature), and got acceptable results in a reasonable amount of time.
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama
HPCC1
2008 A Survey: Genetic Algorithms and the Fast Evolving World of Parallel Computing
abstract
This paper gives a survey about the impact of modern parallel/distributed computing paradigms over parallel genetic algorithms (PGAs). Helping the GA community to feel more comfortable with the evolving parallel paradigms, and marking some areas of research for the high-performance computing (HPC) community is the major inspiration behind this survey. In the modern parallel computing paradigms we have considered only two major areas that have evolved very quickly during the past few years, namely, multicore computing and Grid computing. We discuss the challenges involved, and give potential solutions for these challenges. We also propose a hierarchical PGA suitable for Grid environment with multicore computational resources.
Asim Munawar, Mohamed Wahib, Masaharu Munetomo, Kiyoshi Akama
HPCC1
2007 Optimization problem solving framework employing GAs with linkage identification over a grid environment
abstract
This paper is a step towards a general purpose optimization problem-solving framework that can solve a large number of global optimization problems on its own with a min imal input from the user. It relies on competent GAs (Genetic Algorithms) as the solver and depends on Grid computing for the required computational resources. In this paper we will discuss the architecture of the framework in detail. In the results section we will discuss the speedups obtained by using parallel GAs over a Grid computing environment and the effects of Grid overheads on the speedup. Even though there are various advantages of using Grids but in the results section we will focus on the reduction in total execution time due to parallelism.
Asim Munawar, Masaharu Munetomo, Kiyoshi Akama
IEEE Congress on Evolutionary Computation1
2007 MHGrid: Towards an Ideal Optimization Environment for Global Optimization Problems Using Grid Computing
abstract
This paper introduces MHGrid, a framework that exploits meta-heuristics based search methods and grid computing to enable the transparent sharing of heterogeneous and dynamic resources offering a grid based global optimization framework. MHGrid allows a user to solve almost all kinds of global optimization problems in a black box manner with a minimal input from the user, it also allows the user to integrate his own solver into MHGrid. In this paper we will discuss the architecture and motivation of such a system. We will also discuss the challenges/complexities involved in constructing MHGrid.
Mohamed Wahib, Asim Munawar, Masaharu Munetomo, Kiyoshi Akama
PDCAT2