Mehdi Bahrami

dblp:79/3898 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-authorArtificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 AutoDW-TS: Automated Data Wrangling for Time-Series Data
Lei Liu 0061, So Hasegawa, Shailaja Sampat, Mehdi Bahrami, Wei-Peng Chen, Kodai Toyota, Takashi Kato, Takumi Akazaki, Akira Ura, Tatsuya Asai
CIKM4
2024 LLM Diagnostic Toolkit: Evaluating LLMs for Ethical Issues
abstract
The rapid proliferation of large language models (LLMs) has brought with it both opportunities and challenges. While LLMs and more broadly generative AI technologies are capable of providing excellent improvements for various routine and autonomous tasks thereby enabling cost and performance benefit, they are also prone to personal and societal harms such as biases, stereotypes, misinformation, and hallucinations to name a few. These ethical concerns have in turn triggered stakeholders across the world to call in for regulatory measures that ensure safe and beneficial use of generative AI technologies. In parallel, there are also research efforts to alleviate these issues through the development of generative AI bias detection and mitigating strategies. Towards advancing this goal, in this paper, we propose an accessible and end-user-friendly LLM diagnostic toolkit whereby diverse stakeholders such as software engineers, business executives, and consumers can examine a suite of LLMs for uncovering a host of ethical issues including biases and misinformation embedded in LLMs. We also demonstrate that our toolkit can be used to diagnose for issues related to commonsense reasoning capabilities of LLMs. Extensive experiments on challenging tasks and datasets demonstrates the effectiveness of our diagnostic toolkit.
Mehdi Bahrami, Ryosuke Sonoda, Ramya Srinivasan 0002
IJCNN1
2022 An Intelligent Data-Centric Web Crawler Service for API Corpus Construction at Scale
abstract
The number of web APIs is growing rapidly. API adoption is increasing across all industries with executives prioritizing investments in the API economy. Each API provider offers API documentation which includes complex descriptions. In order to collect and understand the applications and operations of diverse APIs, software engineers read lengthy and complicated API documentations. Understanding the variety of API documentations is a labor intensive and error-prone process. In this paper, we introduce a data-centric web crawler service to collect, analyze, and construct a large corpus of API documentations. The generated API Corpus can be used in machine programming (i.e., code generation, code search). The proposed API web-crawler intelligently harvests more than 2.8M API documentation pages where it uses a machine-learning-based approach with an accuracy of 91.32% to select only web API pages (REST). We also conducted an extensive and end-to-end real-world evaluation, where the proposed API web-crawler not only collects a sheer number of API pages, but also successfully validates 1,222 APIs out of 1,521 target APIs with a success rate of 80.34%.
Mehdi Assefi, Mehdi Bahrami, Sarthak Arora, Thiab R. Taha, Hamid R. Arabnia, Khaled Rasheed, Wei-Peng Chen
ICWS2
2022 Automatic Generation of Visualizations for Machine Learning Pipelines
abstract
Visualization is very important for machine learning (ML) pipelines because it can show explorations of the data to inspire data scientists and show explanations of the pipeline to improve understandability. In this paper, we present a novel approach that automatically generates visualizations for ML pipelines by learning visualizations from highly-upvoted Kaggle pipelines. The solution extracts both code and dataset features from these high-quality human-written pipelines and corresponding training datasets, learns the mapping rules from code and dataset features to visualizations using association rule mining (ARM), and finally uses the learned rules to predict visualizations for unseen ML pipelines. The evaluation results show that the proposed solution is feasible and effective to generate visualizations for ML pipelines.
Lei Liu 0061, Wei-Peng Chen, Mehdi Bahrami, Mukul R. Prasad
ASE3
2021 A Systematic Investigation of KB-Text Embedding Alignment at Scale
abstract
Vardaan Pahuja, Yu Gu, Wenhu Chen, Mehdi Bahrami, Lei Liu, Wei-Peng Chen, Yu Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Vardaan Pahuja, Yu Gu 0016, Wenhu Chen, Mehdi Bahrami, Wei-Peng Chen, Yu Su 0001
ACL/IJCNLP (1)4
2020 Automatic Generation of IFTTT Mashup Infrastructures
abstract
In recent years, IF-This-Then-That (IFTTT) services are becoming more and more popular. Many platforms such as Zapier, IFTTT.com, and Workato provide such services, which allow users to create workflows with "triggers" and "actions" by using Web Application Programming Interfaces (APIs). However, the number of IFTTT recipes in the above platforms increases much slower than the growth of Web APIs. This is because human efforts are still largely required to build and deploy IFTTT recipes in the above platforms. To address this problem, in this paper, we present an automation tool to automatically generate the IFTTT mashup infrastructure. The proposed tool provides 5 REST APIs, which can automatically generate triggers, rules, and actions in AWS, and create a workflow XML to describe an IFTTT mashup by connecting the triggers, rules, and actions. This workflow XML is automatically sent to Fujitsu RunMyProcess (RMP) to set up and execute IFTTT mashup. The proposed tool, together with its associated method and procedure, enables an end-to-end solution for automatically creating, deploying, and executing IFTTT mashups in a few seconds, which can greatly reduce the development cycle and cost for new IFTTT mashups.
Mehdi Bahrami, Wei-Peng Chen
ASE2
2020 Deep SAS: A Deep Signature-based API Specification Learning Approach
abstract
The number and variety of Web APIs is growing exponentially. Software engineers need to expend a significant amount of time and effort reading and understanding the accompanying documentation. In addition, system automation may use API to interact with each other. However, this is not always a simple task since the API documentation of a provider can be anything from a single HTML page description through to a complex structure with information spanning several pages. Understanding this wide variety of API documentation structures and styles is therefore a labor intensive and error-prone task for engineers. By providing a machine-learning platform that can extract and standardize API usage information, however, we believe we can accelerate the creation of API-enabled systems by using automation to simplify the task of understanding. In this paper we introduce a novel approach to automating and standardizing usage information about APIs, combining several machine-learning algorithms in order to extract key attributes from API documentation and generate a machine readable Open API Specification (OAS). We develop i) a content-based learning model that identifies the context of a block of extracted API features; ii) a signature-based machine learning model that recognizes a sequence of successful/unsuccessful extracted API endpoints; and iii) a deep mapping model that pinpoints fine-grained mapping of extracted API attributes to OAS objects. Results of our experiments show that the proposed approach successfully works with an accuracy of 99%, 94% and 97% for content-based learning, signature-based learning, and Deep Mapping of API attributes respectively. We then use the models to produce OAS compliant API Specifications for more than 2,585 public APIs, validate them via API calls and finally deploy the validated APIs to the RunMyProcess software automation platform.
Mehdi Bahrami, Mehdi Assefi, Ian Thomas, Wei-Peng Chen, Shridhar Choudhary, Hamid R. Arabnia
SMC1
2019 Accelerating the Digital Transformation of Business and Society Through Composite Business Ecosystems
Shridhar Choudhary, Ian Thomas, Mehdi Bahrami, Motoshi Sumioka
AINA3
2019 WATAPI: Composing Web API Specification from API Documentations through an Intelligent and Interactive Annotation Tool
abstract
The number of web APIs grows rapidly. Each API provider offers API documentations which comes with diversity and complexity of structures. In order to understand a large number of diverse web APIs, we employ deep learning that extracts key objects of a large number of web APIs and produce a unified API specification (Open API Specification). However, the unified API specification is not a simple task due to heterogeneous API documentations and it is required additional human evaluation to adjust incorrect information which produced by machine. In this paper, we introduce WATAPI which represents an automation tool for representing machine learning outcomes and adding a user as a humanin-the-loop to transparently interact with complex machinelearning components to process diverse API documentations. WATAPI allows the user to annotate API documentations and/or adjust the automated annotation process which produced by machine-learning models' predictions. WATAPI is also capable to perform as a semi-automated annotation tool where it records user's interactions. WATAPI is able to automatically apply user's interaction of one API annotation to another API documentations with a similar structure. The user's annotation adjustment provides feedback to machinelearning components for improving the accuracy of extracting Open API Specifications.
Mehdi Bahrami, Wei-Peng Chen
IEEE BigData1
2019 A Deep Learning Based Autonomous Mobile Robotic Assistive Care Giver
abstract
Assistive robotic technology is increasingly employed in many industries including health care. One of the most important features of this assistive technology is its autonomous verbal communication skill. We propose a new theory for autonomous agent based on the five human senses. Then we proceed to address one of the five senses, the speech. Our approach to address and develop an autonomous verbal communication is to apply deep learning to learn about different topics in healthcare. We developed a novel approach where we created a set of question-answer dataset from articles and interviews with physician specialist from U.S. National Public Radio (NPR). We trained a deep learning model which is able to listen to conversations between a patient and a physician to answer to the questions when the physician is not able to answer or it might not answer the question completely. We discuss the corpus on Health Science which shows what NPR can teach to machine. We share the usage of the Corpus to train a deep learning model to be used by pepper which is a humanoid robot that can be implemented in helping provide elderly care in individuals diagnosed with early stages of dementia.
Arshia Khan, Mehdi Bahrami, Yumna Anwar
HealthCom2
2017 Compliance-Aware Provisioning of Containers on Cloud
abstract
Deploying applications in containers has several advantages, such as rapid development, portability across different machines, and simplified maintenance. In a cloud computing environment, container scheduling algorithms coordinate with different aspects of physical systems, such as memory allocation for tasks of different users. The scheduled containers on a host may process sensitive data. For instance, containers may process healthcare information. In that case, diverse cloud environments with different components and subsystems may lead to a potential personal health information leakage and violation of data privacy. In this paper, we introduce a novel compliance-aware analysis model for provisioning containers in the cloud, that provides a HIPAA compliance model. The proposed method dynamically analyzes different requirements of HIPAA complaint containers (HIPAA parameters) and their associated risk values. Based on the risk optimization of the compliance parameters for data security and data privacy of the containers, our proposed method determines scheduling of containers that offer the lowest risk to healthcare data and to the compliance posture of the container. The model describes the resources that are associated with highlevel risks and provides real-time resource recommendation for a container scheduler to decrease the risk of HIPAA compliance violation.
Mehdi Bahrami, Abhishek Malvankar, Karan Kumar Budhraja, Chinmay Kundu, Mukesh Singhal, Ashish Kundu
CLOUD1
2017 Risk-Based Packet Routing for Privacy and Compliance-Preserving SDN
abstract
Software Defined Networking (SDN) is increasingly being used in data centers as well as enterprise networks. In an environment that has strict compliance requirements, such as HIPAA compliance, a critical role for an SDN controller is to route all data packets while considering data privacy preservation and compliance-preservation. In this paper, we address this problem by proposing a routing protocol for SDN which is an efficient risk-based swarm routing protocol. The programmable capability of controllers is exploited in order to minimize privacy and compliance risks in data transmission. The proposed routing protocol is based on the Ant Colony Optimization technique and machine learning, while the data for learning is obtained from OVSDB and the OpenvSwitch Database management protocol. We collect a history of packet transfers for training purposes and learn from the training data to efficiently and intelligently route sensitive data packets while it preserves the target compliance. This routing is obtained by intelligent eviction of rules that are downloaded to the switches. We have implemented the proposed schemes based on an RYU controller.
Karan Kumar Budhraja, Abhishek Malvankar, Mehdi Bahrami, Chinmay Kundu, Ashish Kundu, Mukesh Singhal
CLOUD3
2017 ICN-FC: An Information-Centric Networking based framework for efficient functional chaining
abstract
In this paper, we present ICN-FC, which is an Information-Centric Networking (ICN) based framework for efficient functional chaining (FC). The key enabling techniques for ICN-FC includes naming semantics, Interest & Data processing and an efficient FC forwarding strategy. By using the proposed solutions, a functional chaining request, which consists of the name of raw data and an ordered set of functions, can be executed seamlessly, dynamically and flexibly in the network. In addition, the novel FC forwarding strategy can be used to improve the forwarding efficiency for functional chaining requests. The overall feasibility and efficiency of the proposed solutions are validated by using both experimental prototype and network simulation. The results show that the proposed solutions outperform previous works such as named function networking to support functional chaining applications.
Mehdi Bahrami, Liguang (Ted) Xie, Akira Ito 0004, Sevak Mnatsakanyan, Zilong Ye, Huiping Guo
ICC3
2015 A dynamic cloud computing platform for eHealth systems
abstract
Cloud Computing technology offers new opportunities for outsourcing data, and outsourcing computation to individuals, start-up businesses, and corporations in health care. Although cloud computing paradigm provides interesting, and cost effective opportunities to the users, it is not mature, and using the cloud introduces new obstacles to users. For instance, vendor lock-in issue that causes a healthcare system rely on a cloud vendor infrastructure, and it does not allow the system to easily transit from one vendor to another. Cloud data privacy is another issue and data privacy could be violated due to outsourcing data to a cloud computing system, in particular for a healthcare system that archives and processes sensitive data. In this paper, we present a novel cloud computing platform based on a Service-Oriented cloud architecture. The proposed platform can be ran on the top of heterogeneous cloud computing systems that provides standard, dynamic and customizable services for eHealth systems. The proposed platform allows heterogeneous clouds provide a uniform service interface for eHealth systems that enable users to freely transfer their data and application from one vendor to another with minimal modifications. We implement the proposed platform for an eHealth system that maintains patients' data privacy in the cloud. We consider a data accessibility scenario with implementing two methods, AES and a light-weight data privacy method to protect patients' data privacy on the proposed platform. We assess the performance and the scalability of the implemented platform for a massive electronic medical record. The experimental results show that the proposed platform have not introduce additional overheads when we run data privacy protection methods on the proposed platform.
Mehdi Bahrami, Mukesh Singhal
HealthCom1