VLDB 2026 Research / reviewers in the wild / expert
Ali Payani
dblp:184/3921
· DBLP profile ↗
32ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0003-4054-2958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or HarmabstractAgents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes with significant safety and ethical risks. Unethical behavior by these agents can directly result in serious real-world consequences, including physical harm and financial loss. To efficiently steer the ethical behavior of agents, we frame agent behavior steering as a model editing task, which we term Behavior Editing. Model editing is an emerging area of research that enables precise and efficient modifications to LLMs while preserving their overall capabilities. To systematically study and evaluate this approach, we introduce BehaviorBench, a multi-tier benchmark grounded in psychological moral theories. This benchmark supports both the evaluation and editing of agent behaviors across a variety of scenarios, with each tier introducing more complex and ambiguous scenarios. We first demonstrate that Behavior Editing can dynamically steer agents toward the target behavior within specific scenarios. Moreover, Behavior Editing enables not only scenario-specific local adjustments but also more extensive shifts in an agent’s global moral alignment. We demonstrate that Behavior Editing can be used to promote ethical and benevolent behavior or, conversely, to induce harmful or malicious behavior. Through extensive evaluations of agents built on frontier LLMs, BehaviorBench validates the effectiveness of behavior editing across a wide range of models and scenarios. Our findings offer key insights into a new paradigm for steering agent behavior, highlighting both the promise and perils of Behavior Editing. Baixiang Huang, Zhen Tan 0001, Haoran Wang 0005, Dawei Li 0008, Ali Payani, Huan Liu 0001, Tianlong Chen 0001, Kai Shu |
AAAI | 6 |
| 2026 | Benchmarking LLMs for Political Science: A United Nations PerspectiveabstractLarge Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplored. This paper addresses the gap by focusing on the application of LLMs to the United Nations (UN) decision-making process, where the stakes are particularly high and political decisions can have far-reaching consequences. We introduce a novel dataset comprising publicly available UN Security Council (UNSC) records from 1994 to 2024, including draft resolutions, voting records, and diplomatic speeches. Using this dataset, we propose the United Nations Benchmark (UNBench), the first comprehensive benchmark designed to evaluate LLMs across four interconnected political science tasks: co-penholder judgment, representative voting simulation, draft adoption prediction, and representative statement generation. These tasks span the three stages of the UN decision-making process—drafting, voting, and discussing—and aim to assess LLMs' ability to understand and simulate political dynamics. Our experimental analysis demonstrates the potential and challenges of applying LLMs in this domain, providing insights into their strengths and limitations in political science. To the best of our knowledge, this is the first benchmark to systematically evaluate LLMs in UN decision-making, contributing to the growing intersection of AI and political science. Yueqing Liang, Liangwei Yang, Chen Wang 0018, Congying Xia, Xiongxiao Xu, Haoran Wang 0005, Ali Payani, Kai Shu |
AAAI | 8 |
| 2025 | Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource LanguagesabstractThe development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM's internal representation of various languages against a baseline derived from English, we can assess the model's multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM’s performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources. Zirui Liu 0001, Fan Yang 0023, Ali Payani, Ninghao Liu 0001, Mengnan Du |
AAAI | 5 |
| 2025 | Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World ModelabstractEnhancing the reasoning capabilities of language models (LMs) remains a key challenge, especially for tasks that require complex, multistep decision-making where existing Chain-of-Thought (CoT) approaches struggle with consistency and verification.In this paper, we propose a novel reasoning framework, referred to as Structure-aware Planning with an Accurate World Model (SWAP), that integrates structured knowledge representation with learned planning.Unlike prior methods that rely purely on natural language reasoning, SWAP leverages entailment graphs to encode structured dependencies and enable symbolic verification of intermediate steps.To systematically construct and update the graph, SWAP employs a policy model to propose candidate expansions and a world model to predict structural updates.To improve accuracy, the world model generates multiple alternative updates, and a discriminator re-ranks them based on plausibility.To encourage diverse exploration, we introduce Diversity-based Modelling (DM), which samples candidates from the remaining probability mass after removing previously sampled candidates from the original policy distribution.Additionally, SWAP improves the discrimination accuracy through Contrastive Ranking (CR), which directly compares candidates within prompts and incorporates metaknowledge to improve ranking quality.We evaluate SWAP across diverse reasoning-intensive benchmarks including math reasoning, logical reasoning, and coding tasks.Extensive experiments demonstrate that SWAP significantly improves upon the base models and consistently outperforms existing reasoning methods 1 . Siheng Xiong, Ali Payani, Yuan Yang 0007, Faramarz Fekri |
ACL (1) | 2 |
| 2025 | FABLE: Fairness Attack in Abusive Language Detection
Yueqing Liang, Lu Cheng 0001, Ali Payani, Kai Shu |
IEEE Big Data | 3 |
| 2025 | SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsabstractLarge language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings.This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces.We employ sparse autoencoders (SAEs) to obtain sparse latent representations that aim to disentangle semantic attributes from model activations.Then we train linear classifiers to identify a small subspace of task-relevant dimensions in latent representations.Finally, we learn supervised steering vectors constrained to this subspace, optimized to align with target behaviors.Experiments across sentiment, truthfulness, and politics polarity steering tasks with multiple LLMs demonstrate that our supervised steering vectors achieve higher success rates with minimal degradation in generation quality compared to existing methods.Further analysis reveals that a notably small subspace is sufficient for effective steering, enabling more targeted and interpretable interventions.Our Zirui He, Mingyu Jin, Ali Payani, Yongfeng Zhang 0003, Mengnan Du |
EMNLP | 4 |
| 2025 | AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-ScientistsabstractYifei Li, Hanane Nour Moussa, Ziru Chen, Shijie Chen, Botao Yu, Mingyi Xue, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Chen Wei, Qianheng Zhang, Tianyu Zhang, Song Gao, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yifei Li 0005, Hanane Nour Moussa, Ziru Chen, Botao Yu, Mingyi Xue 0001, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Qianheng Zhang, Song Gao 0001, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun 0001 |
EMNLP | 18 |
| 2025 | Effective Training Data Synthesis for Improving MLLM Chart UnderstandingabstractBeing able to effectively read scientific plots, or chart understanding, is a central part toward building effective agents for science. However, existing multimodal large language models (MLLMs), especially open-source ones, are still falling behind with a typical success rate of 30%-50% on challenging benchmarks. Previous studies on fine-tuning MLLMs with synthetic charts are often restricted by their inadequate similarity to the real charts, which could compromise model training and performance on complex real-world charts. In this study, we show that modularizing chart generation and diversifying visual details improves chart understanding capabilities. In particular, we design a five-step data synthesis pipeline, where we separate data and function creation for single plot generation, condition the generation of later subplots on earlier ones for multi-subplot figures, visually diversify the generated figures, filter out low quality data, and finally generate the question-answer (QA) pairs with GPT-4o. This approach allows us to streamline the generation of fine-tuning datasets and introduce the effective chart dataset (ECD), which contains 10k+ chart images and 300k+ QA pairs, covering 25 topics and featuring 250+ chart type combinations with high visual complexity. We show that ECD consistently improves the performance of various MLLMs on a range of real-world and synthetic test sets. Code, data and models are available at: https://github.com/yuweiyang-anu/ECD. Yunzhong Hou, Zhuowan Li, Gaowen Liu, Ali Payani, Yuan-Sen Ting, Liang Zheng 0001 |
ICCV | 6 |
| 2025 | Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian DistributionabstractProbing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vector of a certain concept in the representation space. However, the single vector identified for a concept varies with both data and training, making it less robust and weakening its effectiveness in real-world applications. To address this challenge, we propose an approach to approximate the subspace representing a specific concept. Built on linear probing classifiers, we extend the concept vectors into Gaussian Concept Subspace (GCS). We demonstrate GCS's effectiveness through measuring its faithfulness and plausibility across multiple LLMs with different sizes and architectures. Additionally, we use representation intervention tasks to showcase its efficacy in real-world applications such as emotion steering. Experimental results indicate that GCS concept vectors have the potential to balance steering performance and maintaining the fluency in natural language generation tasks. Haiyan Zhao 0003, Ali Payani, Fan Yang 0023, Mengnan Du |
ICLR | 4 |
| 2025 | Can Knowledge Editing Really Correct Hallucinations?abstractLarge Language Models (LLMs) suffer from hallucinations, referring to the non-factual information in generated content, despite their superior capacities across tasks. Meanwhile, knowledge editing has been developed as a new popular paradigm to correct erroneous factual knowledge encoded in LLMs with the advantage of avoiding retraining from scratch. However, a common issue of existing evaluation datasets for knowledge editing is that they do not ensure that LLMs actually generate hallucinated answers to the evaluation questions before editing. When LLMs are evaluated on such datasets after being edited by different techniques, it is hard to directly adopt the performance to assess the effectiveness of different knowledge editing methods in correcting hallucinations. Thus, the fundamental question remains insufficiently validated: Can knowledge editing really correct hallucinations in LLMs? We proposed HalluEditBench to holistically benchmark knowledge editing methods in correcting real-world hallucinations. First, we rigorously construct a massive hallucination dataset with 9 domains, 26 topics and more than 6,000 hallucinations. Then, we assess the performance of knowledge editing methods in a holistic way on five dimensions including Efficacy, Generalization, Portability, Locality, and Robustness. Through HalluEditBench, we have provided new insights into the potentials and limitations of different knowledge editing methods in correcting hallucinations, which could inspire future improvements and facilitate progress in the field of knowledge editing. Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani, Kai Shu |
ICLR | 4 |
| 2025 | A Generic Framework for Conformal FairnessabstractConformal Prediction (CP) is a popular method for uncertainty quantification with machine learning models. While conformal prediction provides probabilistic guarantees regarding the coverage of the true label, these guarantees are agnostic to the presence of sensitive attributes within the dataset. In this work, we formalize \textit{Conformal Fairness}, a notion of fairness using conformal predictors, and provide a theoretically well-founded algorithm and associated framework to control for the gaps in coverage between different sensitive groups. Our framework leverages the exchangeability assumption (implicit to CP) rather than the typical IID assumption, allowing us to apply the notion of Conformal Fairness to data types and tasks that are not IID, such as graph data. Experiments were conducted on graph and tabular datasets to demonstrate that the algorithm can control fairness-related gaps in addition to coverage aligned with theoretical expectations. Aditya Vadlamani, Anutam Srinivasan, Pranav Maneriker, Ali Payani, Srinivasan Parthasarathy 0001 |
ICLR | 4 |
| 2025 | InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer InteractionabstractThis paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video.
Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modularity, our agent integrates tool-based and pure vision agents within a highly modular architecture, enabling different models to collaboratively solve decoupled tasks in a step-by-step manner.
Our generality is demonstrated by our ability to evaluate not only pure vision-based real-world benchmarks (i.e., OSWorld), but also more general or tool-intensive benchmarks (e.g., GAIA and SWE-Bench).
Specifically,
we
achieve a $\mathbf{7.27\\%}$ accuracy gain over Claude-Computer-Use on OSWorld.
Codes and evaluation scripts are included in the supplementary material and will be released as open-source. Weitai Kang, Winson Chen, Shan Zuo, Mimi Xie, Ali Payani, Mingyi Hong 0001, Caiwen Ding |
NeurIPS | 8 |
| 2024 | TEILP: Time Prediction over Knowledge Graphs via Logical ReasoningabstractConventional embedding-based models approach event time prediction in temporal knowledge graphs (TKGs) as a ranking problem. However, they often fall short in capturing essential temporal relationships such as order and distance. In this paper, we propose TEILP, a logical reasoning framework that naturaly integrates such temporal elements into knowledge graph predictions. We first convert TKGs into a temporal event knowledge graph (TEKG) which has a more explicit representation of time in term of nodes of the graph. The TEKG equips us to develop a differentiable random walk approach to time prediction. Finally, we introduce conditional probability density functions, associated with the logical rules involving the query interval, using which we arrive at the time prediction. We compare TEILP with state-of-the-art methods on five benchmark datasets. We show that our model achieves a significant improvement over baselines while providing interpretable explanations. In particular, we consider several scenarios where training samples are limited, event types are imbalanced, and forecasting the time of future events based on only past events is desired. In all these cases, TEILP outperforms state-of-the-art methods in terms of robustness. Siheng Xiong, Yuan Yang 0007, Ali Payani, J. Clayton Kerce, Faramarz Fekri |
AAAI | 3 |
| 2024 | When is Tree Search Useful for LLM Planning? It Depends on the DiscriminatorabstractIn this paper, we examine how large language models (LLMs) solve multi-step problems under a language agent framework with three components: a generator, a discriminator, and a planning method.We investigate the practical utility of two advanced planning methods, iterative correction and tree search.We present a comprehensive analysis of how discrimination accuracy affects the overall performance of agents when using these two methods or a simpler method, re-ranking.Experiments on two tasks, text-to-SQL parsing and mathematical reasoning, show that: (1) advanced planning methods demand discriminators with at least 90% accuracy to achieve significant improvements over re-ranking; (2) current LLMs' discrimination abilities have not met the needs of advanced planning methods to achieve such improvements; (3) with LLM-based discriminators, advanced planning methods may not adequately balance accuracy and efficiency.For example, compared to the other two methods, tree search is at least 10-20 times slower but leads to negligible performance gains, which hinders its real-world applications.1 Ziru Chen, Michael White 0001, Raymond J. Mooney, Ali Payani, Yu Su 0001, Huan Sun 0001 |
ACL (1) | 4 |
| 2024 | Large Language Models Can Learn Temporal ReasoningabstractWhile large language models (LLMs) have demonstrated remarkable reasoning capabilities, they are not without their flaws and inaccuracies.Recent studies have introduced various methods to mitigate these limitations.Temporal reasoning (TR), in particular, presents a significant challenge for LLMs due to its reliance on diverse temporal concepts and intricate temporal logic.In this paper, we propose TG-LLM, a novel framework towards languagebased TR.Instead of reasoning over the original context, we adopt a latent representation, temporal graph (TG) that enhances the learning of TR.A synthetic dataset (TGQA), which is fully controllable and requires minimal supervision, is constructed for fine-tuning LLMs on this text-to-TG translation task.We confirmed in experiments that the capability of TG translation learned on our dataset can be transferred to other TR tasks and benchmarks.On top of that, we teach LLM to perform deliberate reasoning over the TGs via Chain-of-Thought (CoT) bootstrapping and graph data augmentation.We observed that those strategies, which maintain a balance between usefulness and diversity, bring more reliable CoTs and final results than the vanilla CoT distillation. 1 * Equal contribution. 1 Code and data are available at https://github.com/ xiongsiheng/TG-LLM.Once upon a time in the quaint town of Weston, a baby boy named John Thompson was brought into the world in the year 1921.Growing up, he had a vibrant spirit and an adventurous soul.…Step 1: Text-to-Temporal Graph Translation True or false: event (John Thompson owned Pearl Network) was longer in duration than event (Sophia Parker was married to John Thompson) ?The duration for each event can be calculated as follows:(John Thompson owned Pearl Network) starts at 1942, ends at 1967 Siheng Xiong, Ali Payani, Ramana Rao Kompella, Faramarz Fekri |
ACL (1) | 2 |
| 2024 | Harnessing the Power of Large Language Models for Natural Language to First-Order Logic TranslationabstractAdvancements in logical reasoning, utilizing LLMs to convert natural language into logical symbolism, combined with the use of external theorem provers, have repositioned the symbolic approach as a central point of interest. The main challenge within this paradigm lies in the LLMs’ capability to accurately translate natural language (NL) statements into first-order-logic (FOL) expressions. Although LLMs have shown notable success, there remains a gap in understanding the limitations and challenges they encounter in NL-FOL translation. This is primarily due to the absence of datasets and evaluation test beds at the required fine-grained level. We present MALLS, a dataset of 28K diverse and verified sentence-level NL-FOL pairs collected from GPT4. We utilize a combined strategy of FOL rule parsing, human annotation, and automatic filtering to ensure quality. We also present LogicLLaMA, a LLaMA2-7B/13B fine-tuned on MALLS for NL-FOL translation, which can be used standalone or to correct previously generated rules by GPT3.5 after being further fine-tuned via a novel reinforcement learning with human feedback (RLHF) framework. We benchmark a wide range of LLMs on MALLS and previous datasets, highlighting weaknesses in them in NL-FOL translation and demonstrating the advantages of MALLS. We also show that LogicLLaMA achieves GPT4-level performance and can generalize to other datasets. Project repo is available at https://github.com/gblackout/LogicLLaMA Yuan Yang 0007, Siheng Xiong, Ali Payani, Ehsan Shareghi, Faramarz Fekri |
ACL (1) | 3 |
| 2024 | Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over QuantityabstractContrastive Language-Image Pre-training (CLIP) on large-scale image-caption datasets learns representations that can achieve remarkable zero-shot generalization. However, such models require a massive amount of pre-training data. Improving the quality of the pre-training data has been shown to be much more effective in improving CLIP’s performance than increasing its volume. Nevertheless, finding small subsets of training data that provably generalize best has remained an open question. In this work, we propose the first theoretically rigorous data selection method for CLIP. We show that subsets that closely preserve the cross-covariance of the images and captions of the full data provably achieve a superior generalization performance.Our extensive experiments on ConceptualCaptions3M and ConceptualCaptions12M demonstrate that subsets found by \textsc{ClipCov} achieve over 2.7x and 1.4x the accuracy of the next best baseline on ImageNet and its shifted versions. Moreover, we show that our subsets obtain 1.5x the average accuracy across 11 downstream datasets, of the next best baseline. The code is available at: \url{https://github.com/BigML-CS-UCLA/clipcov-data-efficient-clip}. Siddharth Joshi 0004, Arnav Jain, Ali Payani, Baharan Mirzasoleiman |
AISTATS | 3 |
| 2024 | Effective Guidance for Model Attention with Simple Yes-no AnnotationsabstractModern deep learning models often make predictions by focusing on irrelevant areas, leading to biased performance and limited generalization. Existing methods aimed at rectifying model attention require explicit labels for irrelevant areas or complex pixel-wise ground truth attention maps. We present Crayon (Correcting Reasoning with Annotations of Yes Or No), offering effective, scalable, and practical solutions to rectify model attention using simple yes-no annotations. Crayon empowers classical and modern model interpretation techniques to identify and guide model reasoning: Crayon-Attention directs classic interpretations based on saliency maps to focus on relevant image regions, while Crayon-Pruning removes irrelevant neurons identified by modern concept-based methods to mitigate their influence. Through extensive experiments with both quantitative and human evaluation, we showcase Crayon’s effectiveness, scalability, and practicality in refining model attention. Crayon achieves state-of-the-art performance, outperforming 12 methods across 3 benchmark datasets, surpassing approaches that require more complex annotations. Seongmin Lee 0007, Ali Payani, Polo Chau |
IEEE Big Data | 2 |
| 2024 | Neural Additive Tensor Decomposition for Sparse TensorsabstractGiven a sparse tensor, how can we accurately capture complex latent structures inherent in the tensor while maintaining the interpretability of those structures? Tensor decomposition is a fundamental technique for analyzing tensors. Classical tensor models provide multi-linear structures that are easy to interpret, but have limitations in capturing complex structures present in real-world sparse tensors. Recent neural tensor models have extended the capabilities of classical tensor models in capturing complex structures within the data. However, this has come at the cost of interpretability: neural tensor models entangle interactions across and within latent structures in a black-box manner, making it difficult to readily understand the discovered structures. Understanding these structures, however, is crucial in applications such as healthcare, which requires transparency in critical decision-making processes. Dawon Ahn, Uday Singh Saini, Evangelos E. Papalexakis, Ali Payani |
CIKM | 4 |
| 2024 | Prompt Mining for Language Models-based Mobility Flow ForecastingabstractWith the advancement of large language models, language model-based forecasting has recently emerged as an innovative approach for predicting mobility flow patterns. The core idea is to use prompts to transform the raw mobility data given as numerical values into natural language sentences so that the language models can be leveraged to generate the description for future observations. However, previous studies have only employed fixed and manually designed templates to transform numerical values into sentences. Since the forecasting performance of language models heavily relies on prompts, using fixed templates for prompting may limit the forecasting capability of language models. In this paper, we propose a novel framework for prompt mining in language model-based mobility forecasting, aiming to explore diverse prompt design strategies. Specifically, the framework includes a prompt generation stage based on the information entropy of prompts and a prompt refinement stage to integrate mechanisms such as the chain of thought. Experimental results on real-world large-scale data demonstrate the superiority of generated prompts from our prompt mining pipeline. Additionally, the comparison of different prompt variants shows that the proposed prompt refinement process is effective. Our study presents a promising direction for further advancing language model-based mobility forecasting. Hao Xue 0001, Tianye Tang, Ali Payani, Flora D. Salim |
SIGSPATIAL/GIS | 3 |
| 2024 | Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsabstractPretrained machine learning models need to be adapted to distribution shifts when deployed in new target environments. When obtaining labeled data from the target distribution is expensive, few-shot adaptation with only a few examples from the target distribution becomes essential. In this work, we propose MixPro, a lightweight and highly data-efficient approach for few-shot adaptation. MixPro first generates a relatively large dataset by mixing (linearly combining) pre-trained embeddings of large source data with those of the few target examples. This process preserves important features of both source and target distributions, while mitigating the specific noise in the small target data. Then, it trains a linear classifier on the mixed embeddings to effectively adapts the model to the target distribution without overfitting the small target data. Theoretically, we demonstrate the advantages of MixPro over previous methods. Our experiments, conducted across various model architectures on 8 datasets featuring different types of distribution shifts, reveal that MixPro can outperform baselines by as much as 7%, with only 2-4 target examples. Yihao Xue, Ali Payani, Yu Yang 0007, Baharan Mirzasoleiman |
ICML | 2 |
| 2024 | A Federated Stochastic Multi-level Compositional Minimax Algorithm for Deep AUC MaximizationabstractAUC maximization is an effective approach to address the imbalanced data classification problem in federated learning. In the past few years, a couple of federated AUC maximization approaches have been developed based on the minimax optimization. However, directly solving a minimax optimization problem to maximize the AUC score cannot achieve satisfactory performance. To address this issue, we propose to maximize AUC via optimizing a federated multi-level compositional minimax problem. Specifically, we develop a novel federated multi-level compositional minimax algorithm with rigorous theoretical guarantees to solve this new learning paradigm in both algorithmic design and theoretical analysis. To the best of our knowledge, this is the first work studying the multi-level minimax optimization problem. Additionally, extensive empirical evaluations confirm the efficacy of our proposed approach. Xinwen Zhang, Ali Payani, Myungjin Lee, Richard Souvenir, Hongchang Gao |
ICML | 2 |
| 2024 | Temporal Inductive Logic Reasoning over Hypergraphs
Yuan Yang 0007, Siheng Xiong, Ali Payani, J. Clayton Kerce, Faramarz Fekri |
IJCAI | 3 |
| 2024 | LightPure: Realtime Adversarial Image Purification for Mobile Devices Using Diffusion ModelsabstractAutonomous mobile systems increasingly rely on deep neural networks for perception and decision-making. While effective, these systems are vulnerable to adversarial machine learning attacks where small perturbations in the input could significantly impact the outcome of the system. Common countermeasures include leveraging adversarial training and/or data or network transformation. Although widely used, the main drawback of these countermeasures is that they require full and invasive access to the classifiers, which are typically proprietary. Additionally, the cost of training or retraining is often prohibitively expensive for large models. To tackle this, purification models have recently been proposed. The aim is to incorporate a "purification" layer before classification, thereby eliminating the necessity to modify the classifier. Despite their effectiveness, state-of-the-art purification methods are compute-intensive, rendering them unsuitable for mobile systems where resources are constrained and large latency is not desired. Hossein Khalili, Vincent Li, Brandan Bright, Ali Payani, Ramana Rao Kompella, Nader Sehatbakhsh |
MobiCom | 5 |
| 2024 | MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical ProblemsabstractRecent advancements in large language models, such as GPT-4, have demonstrated remarkable capabilities in processing standard queries. Despite these advancements, their performance substantially declines in advanced mathematical problems requiring complex, multi-step logical reasoning. To enhance their inferential capabilities, current research has delved into prompting engineering, exemplified by methodologies such as the Tree of Thought and Graph of Thought.
Nonetheless, these existing approaches encounter two significant limitations. Firstly, their effectiveness in tackling complex mathematical problems is somewhat constrained. Secondly, the necessity to design distinct prompts for individual problems hampers their generalizability.
In response to these limitations, this paper introduces the Multi-Agent System for conditional Mining (MACM) prompting method. It not only resolves intricate mathematical problems but also demonstrates strong generalization capabilities across various mathematical contexts.
With the assistance of MACM, the accuracy of GPT-4 Turbo on the most challenging level five mathematical problems in the MATH dataset increase from $\mathbf{54.68\\%} \text{ to } \mathbf{76.73\\%}$. Yi Zhang 0144, Shan Zuo, Ali Payani, Caiwen Ding |
NeurIPS | 4 |
| 2024 | SketchQL: Video Moment Querying with a Visual Query InterfaceabstractLocalizing video moments based on the movement patterns of objects is an important task in video analytics. Existing video analytics systems offer two types of querying interfaces based on natural language and SQL, respectively. However, both types of interfaces have major limitations. SQL-based systems require high query specification time, whereas natural language-based systems require large training datasets to achieve satisfactory retrieval accuracy. To address these limitations, we present SketchQL, a video database management system (VDBMS) for offline, exploratory video moment retrieval that is both easy to use and generalizes well across multiple video moment datasets. To improve ease-of-use, SketchQL features a visual query interface that enables users to sketch complex visual queries through intuitive drag-and-drop actions. To improve generalizability, SketchQL operates on object-tracking primitives that are reliably extracted across various datasets using pre-trained models. We present a learned similarity search algorithm for retrieving video moments closely matching the user's visual query based on object trajectories. SketchQL trains the model on a diverse dataset generated with a novel simulator, that enhances its accuracy across a wide array of datasets and queries. We evaluate SketchQL on four real-world datasets with nine queries, demonstrating its superior usability and retrieval accuracy over state-of-the-art VDBMSs. Renzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu 0002, Joy Arulraj, Kexin Rong 0001 |
Proc. ACM Manag. Data | 3 |
| 2024 | SketchQL Demonstration: Zero-shot Video Moment Querying with SketchesabstractIn this paper, we will present SketchQL, a video database management system (VDBMS) for retrieving video moments with a sketch-based query interface. This novel interface allows users to specify object trajectory events with simple mouse drag-and-drop operations. Users can use trajectories of single objects as building blocks to compose complex events. Using a pre-trained model that encodes trajectory similarity, SketchQL achieves zero-shot video moments retrieval by performing similarity searches over the video to identify clips that are the most similar to the visual query. In this demonstration, we introduce the graphic user interface of SketchQL and detail its functionalities and interaction mechanisms. We also demonstrate the end-to-end usage of SketchQL from query composition to video moments retrieval using real-world scenarios. Renzhi Wu, Pramod Chunduri, Dristi J. Shah, Ashmitha Julius Aravind, Ali Payani, Xu Chu 0002, Joy Arulraj, Kexin Rong 0001 |
Proc. VLDB Endow. | 5 |
| 2023 | LogicDP: Creating Labels for Graph Data via Inductive Logic Programming
Yuan Yang 0007, Faramarz Fekri, J. Clayton Kerce, Ali Payani |
ICLR | 4 |
| 2021 | Visual Question Answering based on Formal LogicabstractVisual question answering (VQA) has been gaining a lot of traction in the machine learning community in the recent years due to the challenges posed in understanding information coming from multiple modalities (i.e., images, language). In VQA, a series of questions are posed based on a set of images and the task at hand is to arrive at the answer. To achieve this, we take a symbolic reasoning based approach using the framework of formal logic. The image and the questions are converted into symbolic representations on which explicit reasoning is performed. We propose a formal logic framework where (i) images are converted to logical background facts with the help of scene graphs, (ii) the questions are translated to first-order predicate logic clauses using a transformer based deep learning model, and (iii) perform satisfiability checks, by using the background knowledge and the grounding of predicate clauses, to obtain the answer. Our proposed method is highly interpretable and each step in the pipeline can be easily analyzed by a human. We validate our approach on the CLEVR and the GQA dataset. We achieve near perfect accuracy of 99.6% on the CLEVR dataset comparable to the state of art models, showcasing that formal logic is a viable tool to tackle visual question answering. Our model is also data efficient, achieving 99.1% accuracy on CLEVR dataset when trained on just 10% of the training data. Muralikrishnna G. Sethuraman, Ali Payani, Faramarz Fekri, J. Clayton Kerce |
ICMLA | 2 |
| 2021 | Predicting Mobile Users Traffic and Access-Time Behavior Using Recurrent Neural NetworksabstractPredicting mobile users' web-access behavior can have substantial impacts on resource allocation and cost reduction for wireless networks. Therefore, we propose a machine learning platform to forecast the web traffic and access time of mobile users. Based on the observation, the traffic patterns exhibit complex dependency on time, location, and popularity of webpages. Thus, a Recurrent Neural Network (RNN) with Long Short Term Memory (LSTM) is developed based on distinct engineered features to learn and predict users' web browsing activities. Then, to forecast the future access-time of mobile users, we propose a Self-exciting Memory Neural Network (SMNN). The access activities are modeled as self-exciting point processes, and intensities are adopted for prediction. Moreover, we extend the proposed predicting framework to cell towers. To cope with the diversity of traffic at the cell tower, we resort to clustering methods to group the similar users of each tower. Then, we develop an LSTM model for each cluster separately to predict the web domain traffic activities for the cell tower. Finally, we show that our proposed models outperform the baseline prediction models based on cellular networks dataset. We also show that for the cell tower access prediction task, the clustering method can significantly improve the prediction accuracy. Abdulrahman Alamoudi, Mingliu Liu, Ali Payani, Faramarz Fekri, Deshi Li |
WCNC | 3 |
| 2017 | Seismic Data Compression Using Online Double-Sparse Dictionary Learning SchemesabstractSeismic data (traces) usually demonstrate high correlation. We propose a scheme based on online dictionary learning, which explores the resemblance among local seismic traces to facilitate compression for communication. In order to alleviate the transmission overhead caused by the slow convergence of online dictionary scheme, sparse constraints and a sliding window mechanism are applied to the incremental components of the dictionaries, which significantly improve the performance of online dictionary learning scheme in the sense of communication cost. Entao Liu, Ali Payani, Faramarz Fekri |
DCC | 2 |
| 2017 | Learning dictionary for efficient signal compressionabstractWe consider the problem of learning dictionaries for data compression. Different from ordinary learning methods, the objective is to design a dictionary such that the signal has a low entropy representation in the basis of the dictionary, rather than giving a sparse or low-energy representation. To achieve this goal, we need to consider the effect of quantization on the rate-distortion curve as well as an estimation of the distributions of the coefficients. Based on this probability estimation, the coefficients are computed, quantized and then entropy-coded. As such, we have developed algorithms for different classes of dictionaries; orthonormal, union of orthonormals and general dictionaries with unit-norm atoms, to iteratively learn the dictionary and the distribution models of the coefficients. A mixture of Gaussians is adopted to estimate the probability and is updated using the expectation maximization algorithm together with the dictionary learning. Simulation results on the real seismic data show the effectiveness of the proposed algorithm compared to ordinary dictionary learning methods. Afshin Abdi, Ali Payani, Faramarz Fekri |
ICASSP | 2 |