Ali Emami

dblp:75/10772 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA
abstract
As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness.Meanwhile, using LLMs themselves as evaluators without external grounding remains unreliable for objective tasks, as they systematically over-accept incorrect answers, fabricate supporting rationales, and degrade sharply on questions that fall outside their training data.We propose Search-AuGmented Evaluation (SAGE), a framework to assess LLM outputs without fixed groundtruth answers.Unlike conventional metrics that compare to static references or depend solely on LLM-as-a-judge knowledge, SAGE acts as an agent that actively retrieves and synthesizes external evidence.It iteratively generates web queries, collects information, summarizes findings, and refines subsequent searches through reflection.By reducing dependence on static reference-driven evaluation protocols, SAGE offers a scalable and adaptive alternative for evaluating the factuality of LLMs.Experimental results on multiple free-form QA benchmarks show that SAGE achieves substantial to perfect agreement with human evaluations.
Sher Badshah, Ali Emami, Hassan Sajjad 0001
ACL (1)2
2026 Reasoning Traces Shape Outputs but Models Won't Say So
abstract
Can we trust the reasoning traces that large reasoning models (LRMs) produce?We investigate whether these traces faithfully reflect what drives model outputs, and whether models will honestly report their influence.We introduce THOUGHT INJECTION, a method that injects synthetic reasoning snippets into a model's trace, then measures whether the model follows the injected reasoning and acknowledges doing so.Across 45,000 samples from three LRMs, we find that injected hints reliably alter outputs, confirming that reasoning traces causally shape model behavior.However, when asked to explain their changed answers, models overwhelmingly refuse to disclose the influence: overall non-disclosure exceeds 90% for extreme hints across 30,000 follow-up samples.Instead of acknowledging the injected reasoning, models fabricate aligned-appearing but unrelated explanations.Activation analysis reveals that sycophancy-and deceptionrelated directions are strongly activated during these fabrications, suggesting systematic patterns rather than incidental failures.Our findings reveal a gap between the reasoning LRMs follow and the reasoning they report, raising concern that aligned-appearing explanations may not be equivalent to genuine alignment.
Yijie Hao, Lingjie Chen, Ali Emami, Joyce C. Ho
ACL (1)3
2026 Common to Whom? Regional Cultural Commonsense and LLM Bias in India
abstract
Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino, Jad Kabbara, Ali Emami. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Queiroz Dividino, Jad Kabbara, Ali Emami
ACL (1)6
2025 NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers
abstract
Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Connections, a collection of 358 simple word classification puzzles derived from the New York Times Connections game. This benchmark is designed to penalize quick, intuitive “System 1” thinking, isolating fundamental reasoning skills. We evaluated six recent LLMs, a simple machine learning heuristic, and humans across three configurations: single-attempt, multiple attempts without hints, and multiple attempts with contextual hints. Our findings reveal a significant performance gap: even top-performing LLMs like GPT-4 fall short of human performance by nearly 30%. Notably, advanced prompting techniques such as Chain-of-Thought and Self-Consistency show diminishing returns as task difficulty increases. NYT-Connections uniquely combines linguistic isolation, resistance to intuitive shortcuts, and regular updates to mitigate data leakage, offering a novel tool for assessing LLM reasoning capabilities.
Angel Yahir Loredo Lopez, Tyler McDonald, Ali Emami
COLING3
2025 Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index
abstract
As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a novel metric that combines accuracy scores with token consumption, adjusted by a user-specified cost concern level to reflect different resource constraints. Our study examines 6 advanced prompting techniques, including Chain-of-Thought, Self-Consistency, and Tree of Thoughts, across 10 widely-used language models and 4 diverse datasets. We demonstrate that approaches such as Self-Consistency often provide statistically insignificant gains while becoming cost-prohibitive. For example, on high-performing models like Claude 3.5 Sonnet, the EPI of simpler techniques like Chain-of-Thought (0.72) surpasses more complex methods like Self-Consistency (0.64) at slight cost concern levels. Our findings suggest a reevaluation of complex prompting strategies in resource-constrained scenarios, potentially reshaping future research priorities and improving cost-effectiveness for end-users.
Tyler McDonald, Anthony Colosimo, Ali Emami
COLING4
2025 We Politely Insist: Your LLM Must Learn the Persian Art of Taarof
abstract
Large language models (LLMs) struggle to navigate culturally specific communication norms, limiting their effectiveness in global contexts.We focus on Persian taarof, a social norm in Iranian interactions, which is a sophisticated system of ritual politeness that emphasizes deference, modesty, and indirectness, yet remains absent from existing cultural benchmarks.We introduce TAAROFBENCH, the first benchmark for evaluating LLM understanding of taarof, comprising 450 role-play scenarios covering 12 common social interaction topics, validated by native speakers.Our evaluation of five frontier LLMs reveals substantial gaps in cultural competence, with accuracy rates 40-48% below native speakers when taarof is culturally appropriate.Performance varies between interaction topics, improves with Persian-language prompts, and exhibits gender-based asymmetries.We also show that responses rated "polite" by standard metrics often violate taarof norms, indicating the limitations of Western politeness frameworks.Through supervised fine-tuning and Direct Preference Optimization, we achieve 21.8% and 42.3% improvement in model alignment with cultural expectations.Our human study with 33 participants (11 native Persian, 11 heritage, and 11 non-Iranian speakers) forms baselines in varying degrees of familiarity with Persian norms.This work lays the foundation for developing diverse and culturally aware LLMs, enabling applications that better navigate complex social interactions.1
Nikta Gohari Sadr, Sahar Heidariasl, Karine Megerdoomian, Laleh Seyyed-Kalantari, Ali Emami
EMNLP5
2025 Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
abstract
As Large Language Models (LLMs) increasingly integrate into everyday workflows, where users shape outcomes through multi-turn collaboration, a critical question emerges: do users with different personality traits systematically prefer certain LLMs over others?We conducted a study with 32 participants evenly distributed across four Keirsey personality types, evaluating their interactions with GPT-4 and Claude 3.5 across four collaborative tasks: data analysis, creative writing, information retrieval, and writing assistance.Results revealed significant personality-driven preferences: Rationals strongly preferred GPT-4, particularly for goaloriented tasks, while idealists favored Claude 3.5, especially for creative and analytical tasks.Other personality types showed task-dependent preferences.Sentiment analysis of qualitative feedback confirmed these patterns.Notably, aggregate helpfulness ratings were similar across models, showing how personality-based analysis reveals LLM differences that traditional evaluations miss.
Sarfaroz Yunusov, Kaige Chen, Kazi Nishat Anwar, Ali Emami
EMNLP4
2025 Fine-Tuned LLMs are "Time Capsules" for Tracking Societal Bias Through Books
abstract
Sangmitra Madhusudan, Robert Morabito, Skye Reid, Nikta Gohari Sadr, Ali Emami. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sangmitra Madhusudan, Robert Morabito, Skye Reid, Nikta Gohari Sadr, Ali Emami
NAACL (Long Papers)5
2024 Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
abstract
As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models.We introduce the concept of Confidence-Probability Alignment, that connects an LLM's internal confidence, quantified by token probabilities, to the confidence conveyed in the model's response when explicitly asked about its certainty.Using various datasets and prompting techniques that encourage model introspection, we probe the alignment between models' internal and expressed confidence.These techniques encompass using structured evaluation scales to rate confidence, including answer options when prompting, and eliciting the model's confidence level for outputs it does not recognize as its own.Notably, among the models analyzed, OpenAI's GPT-4 showed the strongest confidence-probability alignment, with an average Spearman's ρ of 0.42, across a wide range of tasks.Our work contributes to the ongoing efforts to facilitate risk assessment in the application of LLMs and to further our understanding of model trustworthiness. 1
Robert Morabito, Sanzhar Umbet, Jad Kabbara, Ali Emami
ACL (1)5
2024 Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
abstract
Research on Large Language Models (LLMs) has often neglected subtle biases that, although less apparent, can significantly influence the models' outputs toward particular social narratives.This study addresses two such biases within LLMs: representative bias, which denotes a tendency of LLMs to generate outputs that mirror the experiences of certain identity groups, and affinity bias, reflecting the models' evaluative preferences for specific narratives or viewpoints.We introduce two novel metrics to measure these biases: the Representative Bias Score (RBS) and the Affinity Bias Score (ABS), and present the Creativity-Oriented Generation Suite (CoGS), a collection of open-ended tasks such as short story writing and poetry composition, designed with customized rubrics to detect these subtle biases.Our analysis uncovers marked representative biases in prominent LLMs, with a preference for identities associated with being white, straight, and men.Furthermore, our investigation of affinity bias reveals distinctive evaluative patterns within each model, akin to 'bias fingerprints'.This trend is also seen in human evaluators, highlighting a complex interplay between human and machine bias perceptions. 1
Sarfaroz Yunusov, Ali Emami
ACL (1)3
2024 Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge
abstract
Large Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning.However, applying this reasoning to multimodal domains, where understanding text and images together is essential, remains a substantial challenge.To address this, we introduce WINOVIS, a novel dataset specifically designed to probe text-to-image models on pronoun disambiguation within multimodal contexts.Utilizing GPT-4 for prompt generation and Diffusion Attentive Attribution Maps (DAAM) for heatmap analysis, we propose a novel evaluation framework that isolates the models' ability in pronoun disambiguation from other visual processing challenges.Evaluation of successive model versions reveals that, despite incremental advancements, Stable Diffusion 2.0 achieves a precision of 56.7% on WINOVIS, showing minimal improvement from past iterations and only marginally surpassing random guessing.Further error analysis identifies important areas for future research aimed at advancing text-to-image models in their ability to interpret and interact with the complex visual world.
Brendan Park, Madeline Janecek, Naser Ezzati-Jivan, Ali Emami
ACL (1)5
2024 EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
abstract
While Large Language Models (LLMs) excel at the Winograd Schema Challenge (WSC), a coreference resolution task testing common-sense reasoning through pronoun disambiguation, they struggle with instances that feature minor alterations or rewording. To address this, we introduce EvoGrad, an open-source platform that harnesses a human-in-the-loop approach to create a dynamic dataset tailored to such altered WSC instances. Leveraging ChatGPT’s capabilities, we expand our task instances from 182 to 3691, setting a new benchmark for diverse common-sense reasoning datasets. Additionally, we introduce the error depth metric, assessing model stability in dynamic tasks. Our results emphasize the challenge posed by EvoGrad: Even the best performing LLM, GPT-3.5, achieves an accuracy of 65.0% with an average error depth of 7.2, a stark contrast to human performance of 92.8% accuracy without perturbation errors. This highlights ongoing model limitations and the value of dynamic datasets in uncovering them.
Jing Han Sun, Ali Emami
LREC/COLING2
2024 WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
abstract
The Winograd Schema Challenge (WSC) serves as a prominent benchmark for evaluating machine understanding.While Large Language Models (LLMs) excel at answering WSC questions, their ability to generate such questions remains less explored.In this work, we propose Tree-of-Experts (ToE), a novel prompting method which enhances the generation of WSC instances (50% valid cases vs. 10% in recent methods).Using this approach, we introduce WSC+, a novel dataset comprising 3,026 LLM-generated sentences.Notably, we extend the WSC framework by incorporating new 'ambiguous' and 'offensive' categories, providing a deeper insight into model overconfidence and bias.Our analysis reveals nuances in generation-evaluation consistency, suggesting that LLMs may not always outperform in evaluating their own generated questions when compared to those crafted by other models.On WSC+, GPT-4, the top-performing LLM, achieves an accuracy of 68.7%, significantly below the human benchmark of 95.1%.
Pardis Sadat Zahraei, Ali Emami
EACL (1)2
2024 STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
abstract
Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing.However, many current methodologies evaluate scenarios in isolation, without considering the broader context or the spectrum of potential biases within each situation.To address this, we introduce the Sensitivity Testing on Offensive Progressions (STOP) dataset, which includes 450 offensive progressions containing 2,700 unique sentences of varying severity that progressively escalate from less to more explicitly offensive.Covering a broad spectrum of 9 demographics and 46 sub-demographics, STOP ensures inclusivity and comprehensive coverage.We evaluate several leading closed-and open-source models, including GPT-4, Mixtral, and Llama 3. Our findings reveal that even the best-performing models detect bias inconsistently, with success rates ranging from 19.3% to 69.8%.We also demonstrate how aligning models with human judgments on STOP can improve model answer rates on sensitive tasks such as BBQ, StereoSet, and CrowS-Pairs by up to 191%, while maintaining or even improving performance.STOP presents a novel framework for assessing the complex nature of biases in LLMs, which will enable more effective bias mitigation strategies and facilitates the creation of fairer language models.
Robert Morabito, Sangmitra Madhusudan, Tyler McDonald, Ali Emami
EMNLP4
2024 MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models
abstract
This study explores the effectiveness of Large Language Models (LLMs) in creating personalized "mirror stories" that reflect and resonate with individual readers' identities, addressing the significant lack of diversity in literature.We present MIRRORSTORIES, a corpus of 1,500 personalized short stories generated by integrating elements such as name, gender, age, ethnicity, reader interest, and story moral.We demonstrate that LLMs can effectively incorporate diverse identity elements into narratives, with human evaluators identifying personalized elements in the stories with high accuracy.Through a comprehensive evaluation involving 26 diverse human judges, we compare the effectiveness of MIRRORSTORIES against generic narratives.We find that personalized LLMgenerated stories not only outscore generic human-written and LLM-generated ones across all metrics of engagement (with average ratings of 4.22 versus 3.37 on a 5-point scale), but also achieve higher textual diversity while preserving the intended moral.We also provide analyses that include bias assessments and a study on the potential for integrating images into personalized stories. 1
Sarfaroz Yunusov, Hamza Sidat, Ali Emami
EMNLP3
2021 ADEPT: An Adjective-Dependent Plausibility Task
abstract
Ali Emami, Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ali Emami, Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung
ACL/IJCNLP (1)1
2021 EBST: An Evolutionary Multi-Objective Optimization Based Tool for Discovering Potential Biomarkers in Ovarian Cancer
abstract
Ovarian cancer is the deadliest gynecologic malignancy, mainly due to limitations in early diagnosis. With advances in high-throughput technologies, research interest in identifying novel and customized tumor biomarkers for early detection and diagnosis is rapidly growing. Here we introduce a new tool called EBST to select microRNAs with biomarker potency in ovarian cancer. This tool has pre-processing options and Its core is the use of Modified Multi Objective Imperialist Competitive Algorithm and six objective functions based on the classifier performance/structure evaluation, clustering error and mRMR filter. In this paper, we used the FDR filter in the pre-processing stage and considered five objective functions, four of which relate to thel1-SVM classifier performance and one to the average mRMR ranking. The proposed method has identified 11 microRNAs includinghsa-miR-6784-5p, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-6756-5p, hsa-miR-1307-3p, hsa-miR-4697-5p, hsa-miR-3663-3p, hsa-miR-328-5p, hsa-miR-1228-3p, hsa-miR-6821-5p, hsa-miR-1268a. Data classification by the proposed model showed 100 percent sensitivity, 99.38 percent specificity, 99.69 percent accuracy and 99.39 percent positive predictive value. In comparison with routine state-of-the-art methods, superiority of our method was confirmed. The biological evaluation of selected microRNAs using bioinformatics tools and published articles confirms their role in cancer signaling pathways. The tool and its MATLAB code are freely available athttps://github.com/hanif-y.
Hanif Yaghoobi, Esmaeil Babaei, Bashdar Mahmud Hussen, Ali Emami
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 An Analysis of Dataset Overlap on Winograd-Style Tasks
abstract
The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR).Model performance on the WSC has quickly progressed from chance-level to near-human using neural language models trained on massive corpora.In this paper, we analyze the effects of varying degrees of overlap between these training corpora and the test instances in WSC-style tasks.We find that a large number of test instances overlap considerably with the corpora on which state-of-the-art models are (pre)trained, and that a significant drop in classification accuracy occurs when we evaluate models on instances with minimal overlap.Based on these results, we develop the KNOWREF-60K dataset, which consists of over 60k pronoun disambiguation problems scraped from web data.KNOWREF-60K is the largest corpus to date for WSC-style common-sense reasoning and exhibits a significantly lower proportion of overlaps with current pretraining corpora.
Ali Emami, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung
COLING1
2020 ReDMark: Framework for residual diffusion watermarking based on deep networks
Mahdi Ahmadi, Alireza Norouzi, Nader Karimi, Shadrokh Samavi, Ali Emami
Expert Syst. Appl.5
2020 BlessMark: a blind diagnostically-lossless watermarking framework for medical applications based on deep neural networks
Hamidreza Zarrabi, Ali Emami, Pejman Khadivi, Nader Karimi, Shadrokh Samavi
Multim. Tools Appl.2
2019 The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution
abstract
We introduce a new benchmark for coreference resolution and NLI, KnowRef, that targets common-sense understanding and world knowledge. Previous coreference resolution tasks can largely be solved by exploiting the number and gender of the antecedents, or have been handcrafted and do not reflect the diversity of naturally occurring text. We present a corpus of over 8,000 annotated text passages with ambiguous pronominal anaphora. These instances are both challenging and realistic. We show that various coreference systems, whether rule-based, feature-rich, or neural, perform significantly worse on the task than humans, who display high inter-annotator agreement. To explain this performance gap, we show empirically that state-of-the art models often fail to capture context, instead relying on the gender or number of candidate antecedents to make a decision. We then use problem-specific insights to propose a data-augmentation trick called antecedent switching to alleviate this tendency in models. Finally, we show that antecedent switching yields promising results on other tasks as well: we use it to achieve state-of-the-art results on the GAP coreference task.
Ali Emami, Paul Trichelair, Adam Trischler, Kaheer Suleman, Hannes Schulz, Jackie Chi Kit Cheung
ACL (1)1
2019 How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG
abstract
Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung
EMNLP/IJCNLP (1)2
2019 Exploiting Uncertainty of Deep Neural Networks for Improving Segmentation Accuracy in MRI Images
abstract
Deep neural networks have shown great achievements in solving complex problems. However, there are fundamental challenges which limit their real world applications. Lack of a measurable criterion for estimating uncertainty of the network predictions is one of these challenges. However, we can compute the variance of the network output by applying spatial transformations, distortions or noise injection to network inputs and interpret these variances as uncertainty of the network predictions. In other words, as long as the deformations do not conceptually alter target of interest, we expect the network to produce the same result. Hence, any outputs changes can be a sign of uncertainty in the network predictions. In order to estimate the prediction uncertainty of deep convolutional neural networks we use simple random transformations. By exploiting the network uncertainty, we improve the overall performance of the system. For a real use case, we apply the proposed method to segment left ventricle in MRI cardiac images. Experimental results demonstrate state-of- the-art performance and highlight the potential capabilities of simple ideas in conjunction with deep neural networks.
Alireza Norouzi, Ali Emami, Kayvan Najarian, Nader Karimi, Shadrokh Samavi, S. Mohamad R. Soroushmehr
ICASSP2
2018 A Knowledge Hunting Framework for Common Sense Reasoning
abstract
We introduce an automatic system that achieves state-of-the-art results on the Winograd Schema Challenge (WSC), a common sense reasoning task that requires diverse, complex forms of inference and knowledge.Our method uses a knowledge hunting module to gather text from the web, which serves as evidence for candidate problem resolutions.Given an input problem, our system generates relevant queries to send to a search engine, then extracts and classifies knowledge from the returned results and weighs them to make a resolution.Our approach improves F1 performance on the full WSC by 0.21 over the previous best and represents the first system to exceed 0.5 F1.We further demonstrate that the approach is competitive on the Choice of Plausible Alternatives (COPA) task, which suggests that it is generally applicable.
Ali Emami, Noelia De La Cruz, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung
EMNLP1
2017 Modeling Glucagon Action in Patients With Type 1 Diabetes
abstract
The dual-hormone artificial pancreas is an emerging technology to treat type 1 diabetes (T1D). It consists of a glucose sensor, infusion pumps, and a dosing algorithm that directs hormonal delivery. Preclinical optimization of dosing algorithms using computer simulations has the potential to accelerate the pace of development for this technology. However, current simulation environments consider glucose regulation models that either do not include glucagon action submodels or include submodels that were proposed without comparison to other candidate models. We consider here nine candidate models of glucagon action featuring a number of possible characteristics: insulin-independent glucagon action, insulin/glucagon ratio effect on hepatic glucose production, insulin-dependent suppression of glucagon action, and the effect of rate of change of glucagon. To assess the models, we use measurements of plasma insulin, plasma glucagon, and endogenous glucose production collected from experiments involving eight subjects with T1D who receive four subcutaneous glucagon boluses. We estimate each model's parameters using a Bayesian approach, and the models are contrasted based on the deviance information criterion. The model achieving the best fit features insulin-dependent suppression of glucagon action and incorporates effects of both glucagon levels and its rate of change.
Ali Emami, Joseph El Youssef, Remi Rabasa-Lhoret, Joelle Pineau, Jessica Castle, Ahmad Haidar
IEEE J. Biomed. Health Informatics1
2015 Novelty detection in human tracking based on spatiotemporal oriented energies
Ali Emami, Mehrtash Harandi, Farhad Dadgostar, Brian C. Lovell
Pattern Recognit.1
2012 Role of Spatiotemporal Oriented Energy Features for Robust Visual Tracking in Video Surveillance
abstract
We propose an effective approach to take advantage of the rich description provided by Spatiotemporal Oriented Energy features for the purpose of robust tracking. There are two core components in our system. The first one is a compound measure of 'Coherent Motion' and 'Identity Motion Signature' which is introduced based on motion dynamics of the targets. This measure is used for robust optimisation in occluded situations as well as an adaptive template updating scheme. The second component is a state machine which detects various states of the targets based on statistical analysis of their 'Motion Signature'. Empirical evaluations demonstrate improvement in performance of the tracking system along with the role of each component.
Ali Emami, Farhad Dadgostar, Abbas Bigdeli, Brian C. Lovell
AVSS1