EDBT 2026 Demo / reviewers in the wild / expert
Alireza Fathi
dblp:70/3898
· DBLP profile ↗
39ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0002-2909-144XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object PlacementabstractScene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to 3D scene generation is hindered by their limited grounding on 3D geometry. In this paper, we investigate how to best work with MLLMs in an object placement task. Towards this goal, we introduce a novel framework, FirePlace, that applies existing MLLMs in (1) 3D geometric reasoning and the extraction of relevant geometric details from the 3D scene, (2) constructing and solving geometric constraints on the extracted low-level geometry, and (3) pruning for final placements that conform to common sense. By combining geometric reasoning with real-world understanding of MLLMs, our method can propose object placements that satisfy both geometric constraints as well as high-level semantic common-sense considerations. Our experiments show that these capabilities allow our method to place objects more effectively in complex scenes with intricate geometry, surpassing the quality of prior work. Ian Huang, Yanan Bao, Karen Truong, Howard Zhou, Cordelia Schmid, Leonidas J. Guibas, Alireza Fathi |
CVPR | 7 |
| 2025 | Visual Lexicon: Rich Image Features in Language SpaceabstractWe present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are often challenging to convey in natural language. Unlike traditional methods that prioritize either high-level semantics (e.g., CLIP) or pixel-level reconstruction (e.g., VAE), ViLex simultaneously captures rich semantic content and fine visual details, enabling high-quality image generation and comprehensive visual scene understanding. Through a self-supervised learning pipeline, ViLex generates tokens optimized for reconstructing input images using a frozen text-to-image (T2I) diffusion model, preserving the detailed information necessary for high-fidelity semantic-level reconstruction. As an image embedding in the language space, ViLex tokens leverage the compositionality of natural languages, allowing them to be used independently as "text tokens" or combined with natural language tokens to prompt pretrained T2I models with both visual and textual inputs, mirroring how we interact with vision-language models (VLMs). Experiments demonstrate that ViLex achieves higher fidelity in image reconstruction compared to text embeddings—even with a single ViLex token. Moreover, ViLex successfully performs various DreamBooth tasks in a zero-shot, unsupervised manner without fine-tuning T2I models. Additionally, ViLex serves as a powerful vision encoder, consistently improving vision-language model performance across 15 benchmarks relative to a strong SigLIP baseline. Xudong Wang 0007, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid |
CVPR | 3 |
| 2025 | Language-Guided Image Tokenization for GenerationabstractImage tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally have limited compression rates, making high-resolution image generation computationally expensive. To address this challenge, we propose to leverage language for efficient image tokenization, and we call our method Text-Conditioned Image Tokenization (TexTok). TexTok is a simple yet effective tokenization framework that leverages language to provide a compact, high-level semantic representation. By conditioning the tokenization process on descriptive text captions, TexTok simplifies semantic learning, allowing more learning capacity and token space to be allocated to capture fine-grained visual details, leading to enhanced reconstruction quality and higher compression rates. Compared to the conventional tokenizer without text conditioning, TexTok achieves average reconstruction FID improvements of 29.2% and 48.1% on ImageNet-256 and -512 benchmarks respectively, across varying numbers of tokens. These tokenization improvements consistently translate to 16.3% and 34.3% average improvements in generation FID. By simply replacing the tokenizer in Diffusion Transformer (DiT) with TexTok, our system can achieve a 93.5× inference speedup while still outperforming the original DiT using only 32 tokens on ImageNet-512. TexTok with a vanilla DiT generator achieves state-of-the-art FID scores of 1.46 and 1.62 on ImageNet-256 and -512 respectively. Furthermore, we demonstrate TexTok’s superiority on the text-to-image generation task, effectively utilizing the off-the-shelf text captions in tokenization. Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross, Cordelia Schmid, Dina Katabi, Xiuye Gu |
CVPR | 3 |
| 2025 | Temporal Chain of Thought: Long-Video Understanding by Thinking in FramesabstractDespite recent advances in Vision-Language Models (VLMs), long-video understanding remains a challenging problem. Although state-of-the-art long-context VLMs can process around 1000 input frames, they still struggle to effectively leverage this sequence length, and succumb to irrelevant distractors within the context window. We present Dynamic Context Aggregation, an inference strategy for video question-answering that curates the model's input context. We use the VLM itself to iteratively identify and extract the most relevant frames from the video, which are then used for answering. We demonstrate how leveraging more computation at inference-time to select the most relevant context leads to improvements in accuracy, in agreement with recent work on inference-time scaling of LLMs. Moreover, we achieve state-of-the-art results on 4 diverse video question-answering datasets, showing consistent improvements with 3 different VLMs. In particular, our method shines on longer videos which would not otherwise fit in the model's context window: On longer videos of more than 1 hour on LVBench, our approach using a context window of 32K outperforms the same VLM using standard inference with a 700K context window by 2.8 points. Anurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid |
NeurIPS | 4 |
| 2024 | A Generative Approach for Wikipedia-Scale Visual Entity RecognitionabstractIn this paper, we address web-scale visual entity recognition, specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way of approaching a problem of such scale is using dual-encoder models (e.g. CLIP), where all the entity names and query images are embedded into a unified space, paving the way for an approximate kNN search. Alternatively, it is also possible to re-purpose a captioning model to directly generate the entity names for a given image. In contrast, we introduce a novel Generative Entity Recognition (GER)framework, which given an input image learns to auto-regressively decode a semantic and discriminative “code” identifying the target entity. Our experiments demonstrate the efficacy of this GER paradigm, showcasing state-of-the-art performance on the challenging OVEN benchmark. GER surpasses strong captioning, dual-encoder, visual matching and hier-archical classification baselines, affirming its advantage in tackling the complexities of web-scale recognition. Mathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia Schmid |
CVPR | 3 |
| 2024 | Retrieval-Enhanced Contrastive Vision-Text ModelsabstractContrastive image-text models such as CLIP form the building blocks of many state-of-the-art systems. While they excel at recognizing common generic concepts, they still struggle on fine-grained entities which are rare, or even absent from the pre-training dataset. Hence, a key ingredient to their success has been the use of large-scale curated pre-training data aiming at expanding the set of concepts that they can memorize during the pre-training stage. In this work, we explore an alternative to encoding fine-grained knowledge directly into the model's parameters: we instead train the model to retrieve this knowledge from an external memory. Specifically, we propose to equip existing vision-text models with the ability to refine their embedding with cross-modal retrieved information from a memory at inference time, which greatly improves their zero-shot predictions. Remarkably, we show that this can be done with a light-weight, single-layer, fusion transformer on top of a frozen CLIP. Our experiments validate that our retrieval-enhanced contrastive (RECO) training improves CLIP performance substantially on several challenging fine-grained tasks: for example +10.9 on Stanford Cars, +10.2 on CUB-2011 and +7.3 on the recent OVEN benchmark, where we even outperform the fine-tuned models on unseen classes. Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid |
ICLR | 3 |
| 2024 | SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender CodeabstractThis paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hundred 3D assets. This process requires complex spatial planning and arrangement. We tackle these challenges through a combination of advanced abstraction, strategic planning, and library learning. SceneCraft first models a scene graph as a blueprint, detailing the spatial relationships among assets in the scene. SceneCraft then writes Python scripts based on this graph, translating relationships into numerical constraints for asset layout. Next, SceneCraft leverages the perceptual strengths of vision-language foundation models like GPT-V to analyze rendered images and iteratively refine the scene. On top of this process, SceneCraft features a library learning mechanism that compiles common script functions into a reusable library, facilitating continuous self-improvement without expensive LLM parameter tuning. Our evaluation demonstrates that SceneCraft surpasses existing LLM-based agents in rendering complex scenes, as shown by its adherence to constraints and favorable human assessments. We also showcase the broader application potential of SceneCraft by reconstructing detailed 3D scenes from the Sintel movie and guiding a video generative model with generated scenes as intermediary control signal. Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A. Ross, Cordelia Schmid, Alireza Fathi |
ICML | 8 |
| 2024 | Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachabstractWeb-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this paper, we propose a novel methodology to curate such a dataset, leveraging a multimodal large language model (LLM) for label verification, metadata generation, and rationale explanation. Instead of relying on the multimodal LLM to directly annotate data, which we found to be suboptimal, we prompt it to reason about potential candidate entity labels by accessing additional contextually relevant information (such as Wikipedia), resulting in more accurate annotations. We further use the multimodal LLM to enrich the dataset by generating question-answer pairs and a grounded fine-grained textual description (referred to as "rationale") that explains the connection between images and their assigned entities. Experiments demonstrate that models trained on this automatically curated data achieve state-of-the-art performance on web-scale visual entity recognition tasks (e.g. +6.9% improvement in OVEN entity task), underscoring the importance of high-quality training data in this domain. Mathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet Iscen |
NeurIPS | 2 |
| 2023 | Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryabstractIn this paper, we propose an end-to-end Retrieval-Augmented Visual Language Model (REVEAL) that learns to encode world knowledge into a large-scale memory, and to retrieve from it to answer knowledge-intensive queries. Reveal consists of four key components: the memory, the encoder, the retriever and the generator. The large-scale memory encodes various sources of multimodal world knowledge (e.g. image-text pairs, question answering pairs, knowledge graph triplets, etc.) via a unified encoder. The retriever finds the most relevant knowledge entries in the memory, and the generator fuses the retrieved knowledge with the input query to produce the output. A key novelty in our approach is that the memory, encoder, retriever and generator are all pre-trained end-to-end on a massive amount of data. Furthermore, our approach can use a diverse set of multimodal knowledge sources, which is shown to result in significant gains. We show that Reveal achieves state-of-the-art results on visual question answering and image captioning. The project page of this work is reveal. github. io. Ziniu Hu, Ahmet Iscen, Chen Sun 0002, Kai-Wei Chang 0001, Yizhou Sun, Cordelia Schmid, David A. Ross, Alireza Fathi |
CVPR | 9 |
| 2023 | Improving Image Recognition by Retrieving from Web-Scale Image-Text DataabstractRetrieval augmented models are becoming increasingly popular for computer vision tasks after their recent success in NLP problems. The goal is to enhance the recognition capabilities of the model by retrieving similar examples for the visual input from an external memory set. In this work, we introduce an attention-based memory module, which learns the importance of each retrieved example from the memory. Compared to existing approaches, our method removes the influence of the irrelevant retrieved examples, and retains those that are beneficial to the input query. We also thoroughly study various ways of constructing the memory dataset. Our experiments show the benefit of using a massive-scale memory dataset of 1B image-text pairs, and demonstrate the performance of different memory representations. We evaluate our method in three different classification tasks, namely long-tailed recognition, learning with noisy labels, and fine-grained classification, and show that it achieves state-of-the-art accuracies in ImageNet-LT, Places-LT and Webvision datasets. Ahmet Iscen, Alireza Fathi, Cordelia Schmid |
CVPR | 2 |
| 2023 | AVIS: Autonomous Visual Information Seeking with Large Language Model AgentabstractIn this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task. This data is then used to design a system comprised of three components: an LLM-powered planner that dynamically determines which tool to use next, an LLM-powered reasoner that analyzes and extracts key information from the tool outputs, and a working memory component that retains the acquired information throughout the process. The collected user behavior serves as a guide for our system in two key ways. First, we create a transition graph by analyzing the sequence of decisions made by users. This graph delineates distinct states and confines the set of actions available at each state. Second, we use examples of user decision-making to provide our LLM-powered planner and reasoner with relevant contextual instances, enhancing their capacity to make informed decisions. We show that AVIS achieves state-of-the-art results on knowledge-based visual question answering benchmarks such as Infoseek and OK-VQA. Ziniu Hu, Ahmet Iscen, Chen Sun 0002, Kai-Wei Chang 0001, Yizhou Sun, David A. Ross, Cordelia Schmid, Alireza Fathi |
NeurIPS | 8 |
| 2022 | A Memory Transformer Network for Incremental Learning
Ahmet Iscen, Thomas Bird, Mathilde Caron, Alireza Fathi, Cordelia Schmid |
BMVC | 4 |
| 2022 | Panoptic Neural Fields: A Semantic Object-Aware Neural Scene RepresentationabstractWe present Panoptic Neural Fields (PNF), an object-aware neural scene representation that decomposes a scene into a set of objects (things) and background (stuff). Each object is represented by an oriented 3D bounding box and a multi-layer perceptron (MLP) that takes position, direction, and time and outputs density and radiance. The background stuff is represented by a similar MLP that additionally outputs semantic labels. Each object MLPs are instance-specific and thus can be smaller and faster than previous object-aware approaches, while still leveraging category-specific priors incorporated via meta-learned initialization. Our model builds a panoptic radiance field representation of any scene from just color images. We use off-the-shelf algorithms to predict camera poses, object tracks, and 2D image semantic segmentations. Then we jointly optimize the MLP weights and bounding box parameters using analysis-by-synthesis with self-supervision from color images and pseudo-supervision from predicted semantic segmentations. During experiments with real-world dynamic scenes, we find that our model can be used effectively for several tasks like novel view synthesis, 2D panoptic segmentation, 3D scene editing, and multiview depth prediction. Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J. Guibas, Andrea Tagliasacchi, Frank Dellaert, Thomas A. Funkhouser |
CVPR | 4 |
| 2022 | PreTraM: Self-supervised Pre-training via Connecting Trajectory and Map
Chenfeng Xu, Chen Tang 0001, Lingfeng Sun, Kurt Keutzer, Masayoshi Tomizuka, Alireza Fathi |
ECCV (39) | 7 |
| 2020 | 3D-MPA: Multi-Proposal Aggregation for 3D Semantic Instance SegmentationabstractWe present 3D-MPA, a method for instance segmentation on 3D point clouds. Given an input point cloud, we propose an object-centric approach where each point votes for its object center. We sample object proposals from the predicted object centers. Then, we learn proposal features from grouped point features that voted for the same object center. A graph convolutional network introduces inter-proposal relations, providing higher-level feature learning in addition to the lower-level point features. Each proposal comprises a semantic label, a set of associated points over which we define a foreground-background mask, an objectness score and aggregation features. Previous works usually perform non-maximum-suppression (NMS) over proposals to obtain the final object detections or semantic instances. However, NMS can discard potentially correct predictions. Instead, our approach keeps all proposals and groups them together based on the learned aggregation features. We show that grouping proposals improves over NMS and outperforms previous state-of-the-art methods on the tasks of 3D object detection and semantic instance segmentation on the ScanNetV2 benchmark and the S3DIS dataset. Francis Engelmann, Martin Bokeloh, Alireza Fathi, Bastian Leibe, Matthias Nießner |
CVPR | 3 |
| 2020 | DOPS: Learning to Detect 3D Objects and Predict Their 3D ShapesabstractWe propose DOPS, a fast single-stage 3D object detection method for LIDAR data. Previous methods often make domain-specific design decisions, for example projecting points into a bird-eye view image in autonomous driving scenarios. In contrast, we propose a general-purpose method that works on both indoor and outdoor scenes. The core novelty of our method is a fast, single-pass architecture that both detects objects in 3D and estimates their shapes. 3D bounding box parameters are estimated in one pass for every point, aggregated through graph convolutions, and fed into a branch of the network that predicts latent codes representing the shape of each detected object. The latent shape space and shape decoder are learned on a synthetic dataset and then used as supervision for the end-to-end training of the 3D object detection pipeline. Thus our model is able to extract shapes without access to ground-truth shape information in the target dataset. During experiments, we find that our proposed method achieves state-of-the-art results by~5% on object detection in ScanNet scenes, and it gets top results by 3.4% in the Waymo Open Dataset, while reproducing the shapes of detected cars. Mahyar Najibi, Guangda Lai, Abhijit Kundu, Zhichao Lu, Vivek Rathod, Thomas A. Funkhouser, Caroline Pantofaru, David A. Ross, Larry Davis 0001, Alireza Fathi |
CVPR | 10 |
| 2020 | An LSTM Approach to Temporal 3D Object Detection in LiDAR Point Clouds
Wanyue Zhang, Abhijit Kundu, Caroline Pantofaru, David A. Ross, Thomas A. Funkhouser, Alireza Fathi |
ECCV (18) | 7 |
| 2020 | Virtual Multi-view Fusion for 3D Semantic Segmentation
Abhijit Kundu, Xiaoqi Yin, Alireza Fathi, David A. Ross, Brian Brewington, Thomas A. Funkhouser, Caroline Pantofaru |
ECCV (24) | 3 |
| 2020 | Pillar-Based Object Detection for Autonomous Driving
Yue Wang 0041, Alireza Fathi, Abhijit Kundu, David A. Ross, Caroline Pantofaru, Thomas A. Funkhouser, Justin Solomon 0001 |
ECCV (22) | 2 |
| 2019 | The Devil is in the Decoder: Classification, Regression and GANs
Zbigniew Wojna, Vittorio Ferrari, Sergio Guadarrama, Nathan Silberman, Liang-Chieh Chen, Alireza Fathi, Jasper R. R. Uijlings |
Int. J. Comput. Vis. | 6 |
| 2018 | Instance Embedding Transfer to Unsupervised Video Object SegmentationabstractWe propose a method for unsupervised video object segmentation by transferring the knowledge encapsulated in image-based instance embedding networks. The instance embedding network produces an embedding vector for each pixel that enables identifying all pixels belonging to the same object. Though trained on static images, the instance embeddings are stable over consecutive video frames, which allows us to link objects together over time. Thus, we adapt the instance networks trained on static images to video object segmentation and incorporate the embeddings with objectness and optical flow features, without model retraining or online fine-tuning. The proposed method outperforms state-of-the-art unsupervised segmentation methods in the DAVIS dataset and the FBMS dataset. Siyang Li 0002, Bryan Seybold, Alexey Vorobyov, Alireza Fathi, Qin Huang 0006, C.-C. Jay Kuo |
CVPR | 4 |
| 2018 | Tracking Emerges by Colorizing Videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, Kevin Murphy 0002 |
ECCV (13) | 3 |
| 2017 | The Devil is in the Decoder
Zbigniew Wojna, Jasper R. R. Uijlings, Sergio Guadarrama, Nathan Silberman, Liang-Chieh Chen, Alireza Fathi, Vittorio Ferrari |
BMVC | 6 |
| 2017 | Speed/Accuracy Trade-Offs for Modern Convolutional Object DetectorsabstractThe goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detection systems. A number of successful systems have been proposed in recent years, but apples-toapples comparisons are difficult due to different base feature extractors (e.g., VGG, Residual Networks), different default image resolutions, as well as different hardware and software platforms. We present a unified implementation of the Faster R-CNN [30], R-FCN [6] and SSD [25] systems, which we view as meta-architectures and trace out the speed/accuracy trade-off curve created by using alternative feature extractors and varying other critical parameters such as image size within each of these meta-architectures. On one extreme end of this spectrum where speed and memory are critical, we present a detector that achieves real time speeds and can be deployed on a mobile device. On the opposite end in which accuracy is critical, we present a detector that achieves state-of-the-art performance measured on the COCO detection task. Jonathan Huang, Vivek Rathod, Chen Sun 0002, Menglong Zhu, Anoop Korattikara Balan, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song 0009, Sergio Guadarrama, Kevin Murphy 0002 |
CVPR | 6 |
| 2017 | Comparisons of several variants of continuous quantum-inspired evolutionary algorithmsabstractIn this study, an extensive numerical analysis is carried out to investigate the effects of different quantum-based operators on the performance of continuous quantum-inspired evolutionary algorithms (QEAs). In this context, different variants of quantum-inspired evolutionary operators are adopted for numerical simulations. Furthermore, some novel chaos-enhanced QEAs are proposed and their performances are evaluated through the numerical comparative study. Based on evaluating the accuracy, robustness, convergence, scalability and sensitivity to initialisation of the rival methods, it is indicated that the algorithmic structure of QEAs is prone to being combined with chaotic maps. The results demonstrate that chaotically implemented QEAs can effectively explore/exploit the solution spaces of different landscapes and dimensionality, and finally, converge to acceptable regions within the solution domain. Ahmad Mozaffari, Mahdi Emami, Nasser L. Azad, Alireza Fathi |
J. Exp. Theor. Artif. Intell. | 4 |
| 2016 | Mixed continuous/binary quantum-inspired learning system with non-negative least square optimisation for automated design of regularised ensemble extreme learning machinesabstractIn this paper, a hybrid quantum-inspired evolutionary algorithm (QIEA) is proposed to automatically design regularised ensemble extreme learning machines (EELMs). Quantum evolutionary computing is a relatively recent spot-lighted concept which takes advantage from both the evolutionary and quantum computing laws. In general, QIEAs have been proven to be really powerful for optimising complex engineering tasks. The fascinating trait of observation operator in QIEA enables us to transform the quantum bits to both the binary and continuous spaces. Here, the authors present a mix continuous/binary version of QIEA, to find out whether it is suited for designing regularised EELMs. Indeed, the design process of EELM is conducted at two different levels, i.e. hyper and low levels. At the low level, some novel criteria are presented in the form of penalty functions to enable the optimiser searching for parsimonious, compact and accurate regularised extreme learning machines, as individual components of the ensemble. At the hyper-level, the non-negative least square error optimisation technique is utilised to deterministically find the most eligible components for designing the ensemble. Through extensive numerical experiments, the authors demonstrate that the proposed method is really efficient for the automated design of EELM identifiers. Ahmad Mozaffari, Nasser L. Azad, Mahdi Emami, Alireza Fathi |
J. Exp. Theor. Artif. Intell. | 4 |
| 2015 | On the efficacy of chaos-enhanced heuristic walks with nature-based controllers for robust and accurate intelligent search, part A: an experimental analysisabstractIn this investigation, the authors intend to find out whether chaos-embedded heuristic walks with nature-based controllers are capable of increasing the performance of the standard random walks of evolutionary algorithms. Indeed, the research motivation emanates in the pursuit of addressing an increasing interest in using chaos sequences for modifying the exploitation capability of evolutionary optimisers. Here, the authors propose a nature-based controller to dynamically balance the exploration/exploitation capability of the evolutionary optimiser. To test the validity of the proposed method, a number of well-known Gaussian distribution-based evolutionary operators are considered, and their random parameters are changed with discrete-time chaotic sequences. Besides, to obtain authentic and reliable results regarding the accuracy and robustness of the rival techniques, the authors went through the existing literature and extracted 23 scalable, multimodal, constraint and unconstraint numerical benchmark problems and one engineering problem, i.e. optimal control of shape memory alloy actuators. The convergence rate and computational time of the resulting self-controlled chaos-based evolutionary approaches are analysed to find out whether they are capable of stabilising the chromosomes within a logical time. The results of the numerical experiments indicate that the self-controlled chaos-enhanced heuristic walks significantly increase the global searching capability of the evolutionary optimisers and have an aptitude to show a robust performance. Ahmad Mozaffari, Mahdi Emami, Nasser L. Azad, Alireza Fathi |
J. Exp. Theor. Artif. Intell. | 4 |
| 2015 | An experimentally derived hybrid intelligent tool for analysing and optimising the clad height and melt-pool depth in laser solid freeform fabrication processabstractIn this investigation, a comparative experimental study is conducted to obtain an efficient hybrid intelligent framework for analysing and optimising the operating parameters of the laser solid freeform fabrication (LSFF) process. Here, the experimental studies are conducted in two different stages. In the first stage, different concepts of machine learning systems are taken into account to find a simple yet accurate intelligent model for identifying the LSFF process. To do so, multi-layered neural network with different types of analytical, gradient-based and heuristic learning strategies, i.e. extreme learning machine, back-propagation and steepest descend gradient-based learning and Nelder–Mead simplex heuristic, respectively, are adopted and applied to the LSFF process. In the second stage, different types of swarm- and evolutionary-based metaheuristics, i.e. differential evolutionary algorithm, particle swarm optimisation, the great salmon run, firefly algorithm, bee algorithm, are used to simultaneously find the optimal values of melt-pool depth and clad height during the LSFF process. The statistical results of the simulation indicate that the conducted experiments can result in a fast, robust and accurate hybrid intelligent system which can easily cope with the nonlinearities and uncertainties of the resulting optimisation problem. Ahmad Mozaffari, Alireza Fathi |
J. Exp. Theor. Artif. Intell. | 2 |
| 2014 | Reasoning about Object Affordances in a Knowledge Base Representation
Yuke Zhu, Alireza Fathi, Li Fei-Fei 0001 |
ECCV (2) | 2 |
| 2013 | Modeling Actions through State ChangesabstractIn this paper we present a model of action based on the change in the state of the environment. Many actions involve similar dynamics and hand-object relationships, but differ in their purpose and meaning. The key to differentiating these actions is the ability to identify how they change the state of objects and materials in the environment. We propose a weakly supervised method for learning the object and material states that are necessary for recognizing daily actions. Once these state detectors are learned, we can apply them to input videos and pool their outputs to detect actions. We further demonstrate that our method can be used to segment discrete actions from a continuous video of an activity. Our results outperform state-of-the-art action recognition and activity segmentation results. Alireza Fathi, James M. Rehg |
CVPR | 1 |
| 2013 | Learning to Predict Gaze in Egocentric VideoabstractWe present a model for gaze prediction in egocentric video by leveraging the implicit cues that exist in camera wearer's behaviors. Specifically, we compute the camera wearer's head motion and hand location from the video and combine them to estimate where the eyes look. We further model the dynamic behavior of the gaze, in particular fixations, as latent variables to improve the gaze prediction. Our gaze prediction results outperform the state-of-the-art algorithms by a large margin on publicly available egocentric vision datasets. In addition, we demonstrate that we get a significant performance boost in recognizing daily actions and segmenting foreground objects by plugging in our gaze predictions into state-of-the-art methods. Yin Li 0003, Alireza Fathi, James M. Rehg |
ICCV | 2 |
| 2012 | Social interactions: A first-person perspectiveabstractThis paper presents a method for the detection and recognition of social interactions in a day-long first-person video of u social event, like a trip to an amusement park. The location and orientation of faces are estimated and used to compute the line of sight for each face. The context provided by all the faces in a frame is used to convert the lines of sight into locations in space to which individuals attend. Further, individuals are assigned roles based on their patterns of attention. The rotes and locations of individuals are analyzed over time to detect and recognize the types of social interactions. In addition to patterns of face locations and attention, the head movements of the first-person can provide additional useful cues as to their attentional focus. We demonstrate encouraging results on detection and recognition of social interactions in first-person videos captured from multiple days of experience in amusement parks. Alireza Fathi, Jessica K. Hodgins, James M. Rehg |
CVPR | 1 |
| 2012 | Learning to Recognize Daily Actions Using Gaze
Alireza Fathi, Yin Li 0003, James M. Rehg |
ECCV (1) | 1 |
| 2012 | Detecting eye contact using wearable eye-tracking glassesabstractWe describe a system for detecting moments of eye contact between an adult and a child, based on a single pair of gaze-tracking glasses which are worn by the adult. Our method utilizes commercial gaze tracking technology to determine the adult's point of gaze, and combines this with computer vision analysis of video of the child's face to determine their gaze direction. Eye contact is then detected as the event of simultaneous, mutual looking at faces by the dyad. We report encouraging findings from an initial implementation and evaluation of this approach. Zhefan Ye, Yin Li 0003, Alireza Fathi, Yi Han 0005, Agata Rozga, Gregory D. Abowd, James M. Rehg |
UbiComp | 3 |
| 2011 | Combining Self Training and Active Learning for Video SegmentationabstractPresented at the 22nd British Machine Vision Conference (BMVC 2011), 29 August-2 September 2011, University of Dundee, Scotland, UK. Alireza Fathi, Maria-Florina Balcan, Xiaofeng Ren, James M. Rehg |
BMVC | 1 |
| 2011 | Learning to recognize objects in egocentric activitiesabstractThis paper addresses the problem of learning object models from egocentric video of household activities, using extremely weak supervision. For each activity sequence, we know only the names of the objects which are present within it, and have no other knowledge regarding the appearance or location of objects. The key to our approach is a robust, unsupervised bottom up segmentation method, which exploits the structure of the egocentric domain to partition each frame into hand, object, and background categories. By using Multiple Instance Learning to match object instances across sequences, we discover and localize object occurrences. Object representations are refined through transduction and object-level classifiers are trained. We demonstrate encouraging results in detecting novel object instances using models produced by weakly-supervised learning. Alireza Fathi, Xiaofeng Ren, James M. Rehg |
CVPR | 1 |
| 2011 | Understanding egocentric activitiesabstractWe present a method to analyze daily activities, such as meal preparation, using video from an egocentric camera. Our method performs inference about activities, actions, hands, and objects. Daily activities are a challenging domain for activity recognition which are well-suited to an egocentric approach. In contrast to previous activity recognition methods, our approach does not require pre-trained detectors for objects and hands. Instead we demonstrate the ability to learn a hierarchical model of an activity by exploiting the consistent appearance of objects, hands, and actions that results from the egocentric context. We show that joint modeling of activities, actions, and objects leads to superior performance in comparison to the case where they are considered independently. We introduce a novel representation of actions based on object-hand interactions and experimentally demonstrate the superior performance of our representation in comparison to standard activity representations such as bag of words. Alireza Fathi, Ali Farhadi, James M. Rehg |
ICCV | 1 |
| 2008 | Action recognition by learning mid-level motion featuresabstractThis paper presents a method for human action recognition based on patterns of motion. Previous approaches to action recognition use either local features describing small patches or large-scale features describing the entire human figure. We develop a method constructing mid-level motion features which are built from low-level optical flow information. These features are focused on local regions of the image sequence and are created using a variant of AdaBoost. These features are tuned to discriminate between different classes of action, and are efficient to compute at run-time. A battery of classifiers based on these mid-level features is created and used to classify input sequences. State-of-the-art results are presented on a variety of standard datasets. Alireza Fathi, Greg Mori |
CVPR | 1 |
| 2007 | Human Pose Estimation using Motion ExemplarsabstractWe present a motion exemplar approach for finding body configuration in monocular videos. A motion correlation technique is employed to measure the motion similarity at various space-time locations between the input video and stored video templates. These observations are used to predict the conditional state, distributions of exemplars and joint positions. Exemplar sequence selection and joint position estimation are then solved with approximate inference using Gibbs sampling and gradient ascent. The presented approach is able to find joint positions accurately for people with textured clothing. Results are presented on a dataset containing slow, fast and incline walk videos of various people from different view angles. The results demonstrate an overall improvement compared to previous methods. Alireza Fathi, Greg Mori |
ICCV | 1 |