Khaled S. Refaat

dblp:42/2444 · DBLP profile ↗
← Back
15ranked-venue papers
9as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 CausalAgents: A Robustness Benchmark for Motion Forecasting
abstract
As machine learning models become increasingly prevalent in motion forecasting for autonomous vehicles (AVs), it is critical to ensure that model predictions are safe and reliable. In this paper, we examine the robustness of motion forecasting to non-causal perturbations. We construct a new benchmark for evaluating and improving model robustness by applying perturbations to existing data. Specifically, we conduct an extensive labeling effort to identify causal agents, or agents whose presence influences human drivers’ behavior, in the Waymo Open Motion Dataset (WOMD), and we use these labels to perturb the data by deleting non-causal agents from the scene. We evaluate a diverse set of state-of-the-art deep-learning models on our proposed benchmark and find that all evaluated models exhibit large shifts under non-causal perturbation: we observe a surprising 25-38% relative change in minADE as compared to the original. In addition, we investigate techniques to improve model robustness, including increasing the training dataset size and using targeted data augmentations that randomly drop non-causal agents throughout training. Finally, we release the causal agent labels as an extension to WOMD and the robustness benchmarks to aid the community in building more reliable and safe deep-learning models for motion forecasting1.
Rebecca Roelofs, Benjamin Caine, Khaled S. Refaat, Benjamin Sapp, Scott Ettinger, Wei Chai
ICRA4
2023 MotionLM: Multi-Agent Motion Forecasting as Language Modeling
abstract
Reliable forecasting of the future behavior of road agents is a critical component to safe planning in autonomous vehicles. Here, we represent continuous trajectories as sequences of discrete motion tokens and cast multi-agent motion prediction as a language modeling task over this domain. Our model, MotionLM, provides several advantages: First, it does not require anchors or explicit latent variable optimization to learn multimodal distributions. Instead, we leverage a single standard language modeling objective, maximizing the average log probability over sequence tokens. Second, our approach bypasses post-hoc interaction heuristics where individual agent trajectory generation is conducted prior to interactive scoring. Instead, MotionLM produces joint distributions over interactive agent futures in a single autoregressive decoding process. In addition, the model’s sequential factorization enables temporally causal conditional rollouts. The proposed approach establishes new state-of-the-art performance for multi-agent motion prediction on the Waymo Open Motion Dataset, ranking 1ston the interactive challenge leaderboard.
Ari Seff, Brian Cera, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S. Refaat, Rami Al-Rfou, Benjamin Sapp
ICCV7
2023 Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints
abstract
Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory pre-diction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study.
Jiachen Li 0001, Xinwei Shi, Jonathan Stroud, Zhishuai Zhang, Junhua Mao, Jeonhyung Kang, Khaled S. Refaat, Weilong Yang, Eugene Ie
ICRA9
2023 Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
abstract
Motion forecasting for autonomous driving is a challenging task because complex driving scenarios involve a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road geometry, lane connectivity, time-varying traffic light state, and history of a dynamic set of agents and their interactions into an effective encoding. To model this diverse set of input features, many approaches proposed to design an equally complex system with a diverse set of modality specific modules. This results in systems that are difficult to scale, extend, or tune in rigorous ways to trade off quality and efficiency. In this paper, we present Wayformer, a family of simple and homogeneous attention based architectures for motion forecasting. Wayformer offers a compact model description consisting of an attention based scene encoder and a decoder. In the scene encoder we study the choice of early, late and hierarchical fusion of input modalities. For each fusion type we explore strategies to trade off efficiency and quality via factorized attention or latent query attention. We show that early fusion, despite its simplicity, is not only modality agnostic but also achieves state-of-the-art results on both Waymo Open Motion Dataset (WOMD) and Argoverse leaderboards, demonstrating the effectiveness of our design philosophy.
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou, Kratarth Goel, Khaled S. Refaat, Benjamin Sapp
ICRA5
2022 MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction
abstract
Predicting the future behavior of road users is one of the most challenging and important problems in autonomous driving. Applying deep learning to this problem requires fusing heterogeneous world state in the form of rich perception signals and map information, and inferring highly multi-modal distributions over possible futures. In this paper, we present MultiPath++, a future prediction model that achieves state-of-the-art performance on popular benchmarks. MultiPath++ improves the MultiPath architecture [34] by revisiting many design choices. The first key design difference is a departure from dense image-based encoding of the input world state in favor of a sparse encoding of heterogeneous scene elements: MultiPath++ consumes compact and efficient polylines to describe road features, and raw agent state information directly (e.g., position, velocity, acceleration). We propose a context-aware fusion of these elements and develop a reusable multi-context gating fusion component. Second, we reconsider the choice of pre-defined static anchors, and develop a way to learn latent anchor embeddings end-to-end in the model. Lastly, we explore ensembling and output aggregation techniques—common in other ML domains—and find effective variants for our probabilistic multimodal output representation. We perform an extensive ablation on these design choices, and show that our proposed model achieves state-of-the-art performance on the Argoverse Motion Forecasting Competition [10] and the Waymo Open Dataset Motion Prediction Challenge [13].
Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava, Khaled S. Refaat, Nigamaa Nayakanti, Andre Cornman, Bertrand Douillard, Chi-Pang Lam, Dragomir Anguelov, Benjamin Sapp
ICRA4
2019 Agent Prioritization for Autonomous Navigation
abstract
In autonomous navigation, a planning system reasons about other agents to plan a safe and plausible trajectory. Before planning starts, agents are typically processed with computationally intensive models for recognition, tracking, motion estimation and prediction. With limited computational resources and a large number of agents to process in real time, it becomes important to efficiently rank agents according to their impact on the decision making process. This allows spending more time processing the most important agents. We propose a system to rank agents around an autonomous vehicle (AV) in real time. We automatically generate a ranking data set by running the planner in simulation on real-world logged data, where we can afford to run more accurate and expensive models on all the agents. The causes of various planner actions are logged and used for assigning ground truth importance scores. The generated data set can be used to learn ranking models. In particular, we show the utility of combining learned features, via a convolutional neural network, with engineered features designed to capture domain knowledge. We show the benefits of various design choices experimentally. When tested on real AVs, our system demonstrates the capability of understanding complex driving situations.
Khaled S. Refaat, Natalia Ponomareva 0001, Stéphane Ross
IROS1
2015 Data Compression for Learning MRF Parameters
Khaled S. Refaat, Adnan Darwiche
IJCAI1
2015 An Upper Bound on the Global Optimum in Parameter Estimation
Khaled S. Refaat, Adnan Darwiche
UAI1
2014 Decomposing Parameter Estimation Problems
Khaled S. Refaat, Arthur Choi, Adnan Darwiche
NIPS1
2013 EDML for Learning Parameters in Directed and Undirected Graphical Models
abstract
EDML is a recently proposed algorithm for learning parameters in Bayesian networks. It was originally derived in terms of approximate inference on a meta-network, which underlies the Bayesian approach to parameter estimation. While this initial derivation helped discover EDML in the first place and provided a concrete context for identifying some of its properties (e.g., in contrast to EM), the formal setting was somewhat tedious in the number of concepts it drew on. In this paper, we propose a greatly simplified perspective on EDML, which casts it as a general approach to continuous optimization. The new perspective has several advantages. First, it makes immediate some results that were non-trivial to prove initially. Second, it facilitates the design of EDML algorithms for new graphical models, leading to a new algorithm for learning parameters in Markov networks. We derive this algorithm in this paper, and show, empirically, that it can sometimes learn better estimates from complete data, several times faster than commonly used optimization methods, such as conjugate gradient and L-BFGS.
Khaled S. Refaat, Arthur Choi, Adnan Darwiche
NIPS1
2012 New Advances and Theoretical Insights into EDML
Khaled S. Refaat, Arthur Choi, Adnan Darwiche
UAI1
2011 EDML: A Method for Learning Parameters in Bayesian Networks
Arthur Choi, Khaled S. Refaat, Adnan Darwiche
UAI2
2010 Efficient Stochastic Analysis of Real-Time Systems via Random Sampling
abstract
This paper provides a stochastic approach to the analysis of real-time systems under preemptive priority-driven scheduling. The main idea is to simplify the execution time distributions via random sampling to decrease complexity. This beneficial effect is counterbalanced by an increase in pessimism. However, the proposed analysis is significantly less pessimistic than the classical worst-case deterministic analysis. In addition, it could be tuned according to the memory and time availability. Thus, the proposed method provides, for the first time, a relation between pessimism and computational resources. The testing results show the effectiveness of the sampling approach in terms of practicality and optimism.
Khaled S. Refaat, Pierre-Emmanuel Hladik
ECRTS1
2009 Hand-Drawn Shape Recognition Using the SVM'ed Kernel
Khaled S. Refaat, Amir F. Atiya
ICANN (2)1
2008 A new approach for context-independent handwritten offline diagram recognition using support vector machines
abstract
Structured diagrams are very prevalent in many document types. Most people who need to create such diagrams use structured graphics editors such as Microsoft Visio. Structured graphics editors are extremely powerful and expressive but they can be cumbersome to use. We have shown through extensive timing experiments that structured diagrams drawn by hand will take only about 10% of the time it takes to draw one using a tool like Visio. This indicates the value of automated recognition of hand-written diagrams. Recently, applications have been developed that use online systems running on pen-input PCs that allow users to create structured diagrams by drawing the diagram on the PC tablet. The progress of offline diagram recognition is still minimal. The objective of this paper is to propose a context-independent off-line diagram recognition system. Our approach utilizes support vector machines for recognition and line primitive extraction by interpretation of line continuation for segmentation.
Khaled S. Refaat, Wael N. Helmy, AbdelRahman H. Ali, Mohamed S. AbdelGhany, Amir F. Atiya
IJCNN1