Behrooz Omidvar-Tehrani

dblp:200/8321 · DBLP profile ↗
← Back
27ranked-venue papers
12as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 12 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Enhancing Language Model Agents using Diversity of Thoughts
abstract
A popular approach to building agents using Language Models (LMs) involves iteratively prompting the LM, reflecting on its outputs, and updating the input prompts until the desired task is achieved. However, our analysis reveals two key shortcomings in the existing methods: $(i)$ limited exploration of the decision space due to repetitive reflections, which result in redundant inputs, and $(ii)$ an inability to leverage insights from previously solved tasks. To address these issues, we introduce DoT (Diversity of Thoughts), a novel framework that a) explicitly reduces redundant reflections to enhance decision-space exploration, and b) incorporates a task-agnostic memory component to enable knowledge retrieval from previously solved tasks—unlike current approaches that operate in isolation for each task. Through extensive experiments on a suite of programming benchmarks (HumanEval, MBPP, and LeetCodeHardGym) using a variety of LMs, DoT demonstrates up to a $\textbf{10}$% improvement in Pass@1 while maintaining cost-effectiveness. Furthermore, DoT is modular by design. For instance, when the diverse reflection module of DoT is integrated with existing methods like Tree of Thoughts (ToT), we observe a significant $\textbf{13}$% improvement on Game of 24 (one of the main benchmarks of ToT), highlighting the broad applicability and impact of our contributions across various reasoning tasks.
Vijay Lingam, Behrooz Omidvar-Tehrani, Sujay Sanghavi, Linbo Liu, Jun Huan, Anoop Deoras
ICLR2
2025 Model reusability in Reinforcement Learning
abstract
Abstract The ability to reuse trained models in Reinforcement Learning (RL) holds substantial practical value in particular for complex tasks. While model reusability is widely studied for supervised models in data management, to the best of our knowledge, this is the first ever principled study that is proposed for RL. To capture trained policies, we develop a framework based on an expressive and lossless graph data model that accommodates Temporal Difference Learning and Deep-RL based RL algorithms. Our framework is able to capture arbitrary reward functions that can be composed at inference time. The framework comes with theoretical guarantees and shows that it yields the same result as policies trained from scratch. We design a parameterized algorithm that strikes a balance between efficiency and quality w.r.t cumulative reward. Our experiments with two common RL tasks (query refinement and robot movement) corroborate our theory and show the effectiveness and efficiency of our algorithms.
Sepideh Nikookar, Sohrab Namazi Nia, Senjuti Basu Roy, Sihem Amer-Yahia, Behrooz Omidvar-Tehrani
VLDB J.5
2024 Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
abstract
We propose a new method to measure the task-specific accuracy of Retrieval-Augmented Large Language Models (RAG). Evaluation is performed by scoring the RAG on an automatically-generated synthetic exam composed of multiple choice questions based on the corpus of documents associated with the task. Our method is an automated, cost-efficient, interpretable, and robust strategy to select the optimal components for a RAG system. We leverage Item Response Theory (IRT) to estimate the quality of an exam and its informativeness on task-specific accuracy. IRT also provides a natural way to iteratively improve the exam by eliminating the exam questions that are not sufficiently informative about a model’s ability. We demonstrate our approach on four new open-ended Question-Answering tasks based on Arxiv abstracts, StackExchange questions, AWS DevOps troubleshooting guides, and SEC filings. In addition, our experiments reveal more general insights into factors impacting RAG performance like size, retrieval mechanism, prompting and fine-tuning. Most notably, our findings show that choosing the right retrieval algorithms often leads to bigger performance gains than simply using a larger language model.
Gauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent Callot
ICML2
2024 Reasoning and Planning with Large Language Models in Code Development
abstract
Large Language Models (LLMs) are revolutionizing the field of code development by leveraging their deep understanding of code patterns, syntax, and semantics to assist developers in various tasks, from code generation and testing to code understanding and documentation. In this survey, accompanying our proposed lecture-style tutorial for KDD 2024, we explore the multifaceted impact of LLMs on the code development, delving into techniques for generating a high-quality code, creating comprehensive test cases, automatically generating documentation, and engaging in an interactive code reasoning. Throughout the survey, we highlight some crucial components surrounding LLMs, including pre-training, fine-tuning, prompt engineering, iterative refinement, agent planning, and hallucination mitigation. We put forward that such ingredients are essential to harness the full potential of these powerful AI models in revolutionizing software engineering and paving the way for a more efficient, effective, and innovative future in code development.
Hao Ding 0003, Ziwei Fan 0001, Ingo Gühring, Wooseok Ha, Jun Huan, Linbo Liu, Behrooz Omidvar-Tehrani, Shiqi Wang 0002, Hao Zhou 0036
KDD8
2022 What is Your Current Mindset?
abstract
Is recommendation the new search? Recommender systems have shortened the search for information in everyday activities such as following the news, media, and shopping. In this paper, we address the challenges of capturing the situational needs of the user and linking them to the available datasets with the concept of Mindsets. Mindsets are categories such as “I’m hungry” and “Surprise me” designed to lead the users to explicitly state their intent, control the recommended content, save time, get inspired, and gain shortcuts for a satisficing exploration of POI recommendations. In our methodology, we first compiled Mindsets with a card sorting workshop and a formative evaluation. Using the insights gathered from potential end users, we then quantified Mindsets by linking them to POI utility measures using approximated lexicographic multi-objective optimisation. Finally, we ran a summative evaluation of Mindsets and derived guidelines for designing novel categories for recommender systems.
Sruthi Viswanathan, Behrooz Omidvar-Tehrani, Jean-Michel Renders
CHI2
2022 Guided Text-based Item Exploration
abstract
Exploratory Data Analysis (EDA) provides guidance to users to help them refine their needs and find items of interest in large volumes of structured data. In this paper, we develop GUIDES, a framework for guided Text-based Item Exploration (TIE). TIE raises new challenges: (i) the need to abstract and query textual data and (ii) the need to combine queries on both structured and unstructured content. GUIDES represents text dimensions such as sentiment and topics, and introduces new text-based operators that are seamlessly integrated with traditional EDA operators. To train TIE policies, it relies on a multi-reward function that captures different textual dimensions, and extends the Deep Q-Networks (DQN) architecture with multi-objective optimization. Our experiments on Amazon and IMDb, two real-world datasets, demonstrate the necessity of capturing fine-grained text dimensions, the superiority of using both text-based and attribute-based operators over attribute-based operators only, and the need for multi-objective optimization.
Behrooz Omidvar-Tehrani, Aurélien Personnaz, Sihem Amer-Yahia
CIKM1
2020 Designing Ambient Wanderer: Mobile Recommendations for Urban Exploration
abstract
Recommender systems are widely integrated into our everyday activities. These intelligent systems succeed in learning the user's profile to recommend movies, music, news and more. However, for designing context-aware recommendations, new challenges emerge in predicting the situational needs of the user. We prototyped Ambient Wanderer, our personalised and contextualised Point-of-Interest (POI) recommender system and experimented it with new locals, people who have recently relocated to a city. Our key findings include: sudden breakdowns during urban exploration, trust issues with the recommendations from people unlike them, feeling bored as the trigger to POI search, intent to find free activities, information needs on areas-of-interest beyond points-of-interest and the demand to build a new social life. For each of these needs, we present the implications to design mobile recommendations for urban exploration.
Sruthi Viswanathan, Behrooz Omidvar-Tehrani, Adrien Bruyat, Frédéric Roulland, Antonietta Grasso
Conference on Designing Interactive Systems2
2020 Interactive and Explainable Point-of-Interest Recommendation using Look-alike Groups
abstract
Recommending Points-of-Interest (POIs) is surfacing in many location-based applications. The literature contains personalized and socialized POI recommendation approaches which employ historical check-ins and social links to make recommendations. However these systems still lack customizability and contextuality particularly in cold start situations. In this paper, we propose LikeMind, a POI recommendation system which tackles the challenges of cold start, customizability, contextuality, and explainability by exploiting look-alike groups mined in public POI datasets. LikeMind reformulates the problem of POI recommendation, as recommending explainable look-alike groups (and their POIs) which are in line with user's interests. LikeMind frames the task of POI recommendation as an exploratory process where users interact with the system by expressing their favorite POIs, and their interactions impact the way look-alike groups are selected out. Moreover, LikeMind employs "mindsets", which capture actual situation and intent of the user, and enforce the semantics of POI interestingness. In an extensive set of experiments, we show the quality of our approach in recommending relevant look-alike groups and their POIs, in terms of efficiency and effectiveness.
Behrooz Omidvar-Tehrani, Sruthi Viswanathan, Jean-Michel Renders
SIGSPATIAL/GIS1
2020 Visual exploration of rating datasets and user groups
Fabian Colque Zegarra, Juan C. Carbajal Ipenza, Behrooz Omidvar-Tehrani, Viviane Pereira Moreira, Sihem Amer-Yahia, João Luiz Dihl Comba
Future Gener. Comput. Syst.3
2020 Guided Exploration of User Groups
abstract
Finding a set of users of interest serves several applications in behavioral analytics. Often times, identifying users requires to explore the data and gradually choose potential targets. This is a special case of Exploratory Data Analysis (EDA), an iterative and tedious process. In this paper, we formalize and solve the problem of guided exploration of user groups whose purpose is to find target users. We model exploration as an iterative decision-making process, where an agent is shown a set of groups, chooses users from those groups, and selects the best action to move to the next step. To solve our problem, we apply reinforcement learning to discover an efficient exploration strategy from a simulated agent experience, and propose to use the learned strategy to recommend an exploration policy that can be applied to the same task for any dataset. Our framework accepts a wide class of exploration actions and does not need to gather exploration logs. Our experiments show that the agent naturally captures manual exploration by human analysts, and succeeds to learn an interpretable and transferable exploration policy.
Mariia Seleznova, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Eric Simon
Proc. VLDB Endow.2
2020 User Group Analytics Survey and Research Opportunities
abstract
User data can be acquired from various domains and is characterized by a combination of demographics such as age and occupation, and user actions such as rating a movie or recording one's blood pressure. User data is appealing to analysts in their role as data scientists who seek to conduct large-scale population studies, and gain insights on various population segments. It is also appealing to users in their role as information consumers who use the social Web for routine tasks such as finding a book club or choosing a physical activity. User data analytics usually relies on identifying group-level behaviors such as “Asian women who publish regularly in databases”. Group analytics addresses peculiarities of user data such as noise and sparsity to enable insights. In this survey, we discuss different approaches for each component of user group analytics, i.e., discovery, exploration, and visualization. We focus on related work which arises from combining those components. We also discuss challenges and future directions of having an all-in-one system, where all those components are combined. This survey has been presented in the form of two tutorials [1] , [2].
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
IEEE Trans. Knowl. Data Eng.1
2020 Cohort analytics: efficiency and applicability
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Laks V. S. Lakshmanan
VLDB J.1
2019 GroupTravel: Customizing Travel Packages for Groups
abstract
International audience
Sihem Amer-Yahia, Shady Elbassuoni, Behrooz Omidvar-Tehrani, Ria Mae Borromeo, Mehrdad Farokhnejad
EDBT3
2019 Data Pipelines for User Group Analytics
abstract
User data is becoming increasingly available in various domains ranging from the social Web to electronic patient health records (EHRs). User data is characterized by a combination of demographics (e.g., age, gender, life status) and user actions (e.g., posting a tweet, following a diet). Domain experts rely on user data to conduct large-scale population studies. Information consumers, on the other hand, rely on user data for routine tasks such as finding a book club and getting advice from look-alike patients. User data analytics is usually based on identifying group-level behaviors such as "teenage females who watch Titanic" and "old male patients in Paris who suffer from Bronchitis." In this tutorial, we review data pipelines for User Group Analytics (UGA). These pipelines admit raw user data as input and return insights in the form of user groups. We review research on UGA pipelines and discuss approaches and open challenges for discovering, exploring, and visualizing user groups. Throughout the tutorial, we will illustrate examples in two key domains: "the social Web" and "health-care".
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
SIGMOD Conference1
2019 OntoSIDES: Ontology-based student progress monitoring on the national evaluation system of French Medical Schools
Olivier Palombi, Fabrice Jouanot, Nafissetou Nziengam 0002, Behrooz Omidvar-Tehrani, Marie-Christine Rousset, Adam Sanchez
Artif. Intell. Medicine4
2019 COVIZ: A System for Visual Formation and Exploration of Patient Cohorts
abstract
We demonstrate COVIZ, an interactive system to visually form and explore patient cohorts. COVIZ seamlessly integrates visual cohort formation and exploration, making it a single destination for hypothesis generation. COVIZ is easy to use by medical experts and offers many features: (1) It provides the ability to isolate patient demographics (e.g., their age group and location), health markers (e.g., their body mass index), and treatments (e.g., Ventilation for respiratory problems), and hence facilitates cohort formation; (2) It summarizes the evolution of treatments of a cohort into health trajectories, and lets medical experts explore those trajectories; (3) It guides them in examining different facets of a cohort and generating hypotheses for future analysis; (4) Finally, it provides the ability to compare the statistics and health trajectories of multiple cohorts at once. COVIZ relies on QDS, a novel data structure that encodes and indexes various data distributions to enable their efficient retrieval. Additionally, COVIZ visualizes air quality data in the regions where patients live to help with data interpretations. We demonstrate two key scenarios, ecological scenario and case cross-over scenario . A video demonstration of COVIZ is accessible via http://bit.ly/video-coviz.
Cícero A. L. Pahins, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Valérie Siroux, Jean Louis Pépin, Jean-Christian Borel, João Luiz Dihl Comba
Proc. VLDB Endow.2
2019 User group analytics: hypothesis generation and exploratory analysis of user data
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Ria Mae Borromeo
VLDB J.1
2018 User Group Analytics: Discovery, Exploration and Visualization
abstract
User data is becoming increasingly available in various domains from the social Web to patient health records. User data is characterized by a combination of demographics (e.g., age, gender, occupation) and user actions (e.g., rating a movie, following a diet). User data analytics is usually based on identifying group-level behaviors such as "countryside teachers who watch Woody Allen movies." User Group Analytics (UGA) addresses peculiarities of user data such as noise and sparsity. This tutorial reviews research on UGA and discusses different approaches and open challenges for group discovery, exploration, and visualization.
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
CIKM1
2018 Cohort Representation and Exploration
abstract
The abundant availability of health-care data calls for effective analysis methods which help medical experts gain a better understanding of their data. While the focus has been largely on prediction, "representation" and "exploration" of health-care data have received little attention. In this paper, we introduce CORE, a framework for representing and exploring patient cohorts. Obtaining a readable and succinct representation of health data of a cohort is challenging because cohorts often consist of hundreds of patients whose medical actions are of various types and occur at different points in time. We extend the Needleman-Wunsch algorithm for sequence matching to handle temporal sequences, and propose "trajectory families", a customized index to efficiently compare and aggregate patient trajectories into a cohort representation. We define cohort exploration as finding similar cohorts to a given cohort. This problem is challenging because the potential number of similar cohorts is huge. We propose a two-staged approach based on limiting the search space to "contrast cohorts" and then computing their similarity to the given cohort. To speed up cohort similarity computation, we use "event sets" in the same spirit as the double dictionary encoding proposed for keyword search. We run qualitative and quantitative experiments on real data to explore the efficiency and usefulness of CORE. We show that CORE representations reduce time-to-insight from hours to seconds and help medical experts find insights better than state-of-the-art Visual Analytics tools.
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Laks V. S. Lakshmanan
DSAA1
2018 Exploration of User Groups in VEXUS
abstract
We demonstrate VEXUS, an interactive visualization framework for exploring user data to fulfill tasks such as finding a set of experts, forming discussion groups and analyzing collective behaviors. User data is characterized by a combination of demographics like age and occupation, and actions such as rating a movie, writing a paper or following a medical treatment. The ubiquity of user data requires tools that help explorers, be they specialists or novice users, acquire new insights. VEXUS lets explorers interact with user data via visual primitives and builds an exploration profile to recommend the next exploration steps. VEXUS combines state-of-the-art visualization techniques with appropriate indexing of user data to provide fast and relevant exploration.
Sihem Amer-Yahia, Behrooz Omidvar-Tehrani, João Luiz Dihl Comba, Viviane Pereira Moreira, Fabian Colque Zegarra
ICDE2
2017 Online Lattice-Based Abstraction of User Groups
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
DEXA (1)1
2017 Characterizing Driving Context from Driver Behavior
abstract
Because of the increasing availability of spatiotemporal data, a variety of data-analytic applications have become possible. Characterizing driving context, where context may be thought of as a combination of location and time, is a new challenging application. An example of such a characterization is finding the correlation between driving behavior and traffic conditions. This contextual information enables analysts to validate observation-based hypotheses about the driving of an individual. In this paper, we present DriveContext, a novel framework to find the characteristics of a context, by extracting significant driving patterns (e.g., a slow-down), and then identifying the set of potential causes behind patterns (e.g., traffic congestion). Our experimental results confirm the feasibility of the framework in identifying meaningful driving patterns, with improvements in comparison with the state-of-the-art. We also demonstrate how the framework derives interesting characteristics for different contexts, through real-world examples.
Sobhan Moosavi, Behrooz Omidvar-Tehrani, R. Bruce Craig, Arnab Nandi 0001, Rajiv Ramnath
SIGSPATIAL/GIS2
2017 DV8: Interactive Analysis of Aviation Data
abstract
The vast volume of real-time air traffic data, being produced through new digital transmissions of the movement of aircraft throughout the US National Airspace System (NAS), is a rich resource for evaluating the performance of the system. To date, the potential for comprehensively analyzing this data has yet to be tapped, precisely due to a lack of tools that limit fully interactive data visualization. In this paper, we propose DV8, an interactive data visualization framework which provides in immediately visualized aviation-oriented insights, with a focus on evaluating the deviations among flights by route, type, airport, and aircraft performance. By providing scenarios validated by aviation experts, we illustrate different utilities of DV8 in areas such as capacity planning, flight route prediction, and fuel consumption.
Behrooz Omidvar-Tehrani, Arnab Nandi 0001, Dalton Flanagan, Seth Young
ICDE1
2016 Multi-Objective Group Discovery on the Social Web
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Pierre-François Dutot, Denis Trystram
ECML/PKDD (1)1
2015 Interactive User Group Analysis
abstract
User data is becoming increasingly available in multiple domains ranging from phone usage traces to data on the social Web. The analysis of user data is appealing to scientists who work on population studies, recommendations, and large-scale data analytics. We argue for the need for an interactive analysis to understand the multiple facets of user data and address different analytics scenarios. Since user data is often sparse and noisy, we propose to produce labeled groups that describe users with common properties and develop IUGA, an interactive framework based on group discovery primitives to explore the user space. At each step of IUGA, an analyst visualizes group members and may take an action on the group (add/remove members) and choose an operation (exploit/explore) to discover more groups and hence more users. Each discovery operation results in k most relevant and diverse groups. We formulate group exploitation and exploration as optimization problems and devise greedy algorithms to enable efficient group discovery. Finally, we design a principled validation methodology and run extensive experiments that validate the effectiveness of IUGA on large datasets for different user space analysis scenarios.
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Alexandre Termier
CIKM1
2015 Group Recommendation with Temporal Affinities
abstract
International audience
Sihem Amer-Yahia, Behrooz Omidvar-Tehrani, Senjuti Basu Roy, Nafiseh Shabib
EDBT2
2015 PGLCM: efficient parallel mining of closed frequent gradual itemsets
Trong Dinh Thac Do, Alexandre Termier, Anne Laurent, Benjamin Négrevergne, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
Knowl. Inf. Syst.5