EDBT 2026 Demo / reviewers in the wild / expert
Yi-Ting Chen 0001
dblp:12/5268-1
· DBLP profile ↗
28ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-7906-0828ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 6 since 2021Systems, architecture and hardware · 8 · 4 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
Nakul Agarwal, Yi-Ting Chen 0001, Behzad Dariush |
IV | 2 |
| 2025 | Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian SplattingabstractBird’s-eye view (BEV) perception has gained significant attention because it provides a unified representation to fuse multiple view images and enables a wide range of downstream autonomous driving tasks, such as forecasting and planning. Recent state-of-the-art models utilize projection-based methods which formulate BEV perception as query learning to bypass explicit depth estimation. While we observe promising advancements in this paradigm, they still fall short of real-world applications because of the lack of uncertainty modeling and expensive computational requirement. In this work, we introduce GaussianLSS, an uncertainty-aware BEV perception framework that revisits the unprojection-based method, specifically the Lift-Splat-Shoot (LSS) paradigm, and enhances it with depth uncertainty modeling. Our GaussianLSS represents spatial dispersion by learning a soft depth mean and computing the variance of the depth distribution, which implicitly captures object extents. We then transform the depth distribution into 3D Gaussians and rasterize them to construct uncertainty-aware BEV features. We evaluate GaussianLSS on the nuScenes dataset, achieving state-of-the-art performance compared to unprojection-based methods. In particular, it provides significant advantages in speed, running 2x faster, and in memory efficiency, using 0.3x less memory compared to projection-based methods, while achieving competitive performance with only a 0.7% IoU difference. See our project page for more details: https://hcis-lab.github.io/GaussianLSS/. Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting Chen 0001 |
CVPR | 3 |
| 2025 | Reinvigorating Structured Knowledge and Ontologies for Trustworthy and Beneficial AI and RoboticsabstractThis position paper contends that as artificial intelligence (AI) and robotics advance at a rapid pace—transforming industries, healthcare, transportation, and more—the development of structured knowledge and formal ontologies struggles to keep up. This imbalance poses substantial risks, as increasingly powerful and autonomous systems can make decisions that are both opaque and difficult to verify. While AI and robotics research benefit from significant funding and widespread attention, ontology and structured knowledge efforts often remain under-resourced, under-integrated, and often sidelined—leaving a growing semantic gap in system design and governance. To address this gap, we argue that formal ontology research must be reinvigorated to keep pace with the accelerating demands of AI and robotics—not merely as a support function but as a core contributor to trustworthy and beneficial AI/robotics system design. This requires renewed investment in the academic foundations of ontology engineering, reintegration of semantic modeling into AI/robotics development workflows, and the lowering of practical barriers that have hindered broader adoption. Specifically, we advocate for ontology-augmented verification methods that incorporate semantic constraints into behavioral validation; for the institutionalization of semantic auditing practices that allow ontologies to serve as transparent, inspectable system referents; and for the adoption of FAIR publishing standards that treat ontologies as first-class research outputs. These directions, we argue, are essential to bridge the gap between AI/robotics and ontology research and to ensure that autonomous AI/robotics systems remain not only performant but also transparent, accountable, and ultimately beneficial to humanity. Chun-Yien Chang, Yi-Ting Chen 0001, Ying-Ping Chen |
FOIS | 2 |
| 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningabstractUnderstanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling the temporal order of actions, but has not explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by Large Language Models (LLMs) as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining unseen "What if" scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. We conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, action phase classification, frame retrieval, multi-instance retrieval, and action recognition. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks. Chi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 0001, David Crandall, Yi-Hsuan Tsai |
ICCV | 4 |
| 2025 | Potential Fields as Scene Affordance for Behavior Change-Based Visual Risk Object IdentificationabstractWe study behavior change-based visual risk object identification (Visual-ROI), a critical framework designed to detect potential hazards for intelligent driving systems. Existing methods often show significant limitations in spatial accuracy and temporal consistency, stemming from an incomplete understanding of scene affordance. For example, these methods frequently misidentify vehicles that do not impact the ego vehicle as risk objects. Furthermore, existing behavior change-based methods are inefficient because they implement causal inference in the perspective image space. We propose a new framework with a Bird's Eye View (BEV) representation to overcome the above challenges. Specifically, we utilize potential fields as scene affordance, involving repulsive forces derived from road infrastructure and traffic participants, along with attractive forces sourced from target destinations. In this work, we compute potential fields by assigning different energy levels according to the semantic labels obtained from BEV semantic segmentation. We conduct thorough experiments and ablation studies, comparing the proposed method with various state-of-the-art algorithms on both synthetic and real-world datasets. Our results show a notable increase in spatial accuracy and temporal consistency, with enhancements of 20.3% and 11.6% on the RiskBench dataset, respectively. Additionally, we can improve computational efficiency by 88%. We achieve improvements of 5.4% in spatial accuracy and 7.2% in temporal consistency on the nuScenes dataset. For more qualitative results, please visit our project webpage: project webpage. Pang-Yuan Pao, Shu-Wei Lu, Ze-Yan Lu, Yi-Ting Chen 0001 |
ICRA | 4 |
| 2025 | ATARS: An Aerial Traffic Atomic Activity Recognition and Temporal Segmentation DatasetabstractTraffic Atomic Activity, which describes traffic patterns for topological intersection dynamics, is a crucial topic for the advancement of intelligent driving systems. However, existing atomic activity datasets are collected from an egocentric view, which cannot support the scenarios where traffic activities in an entire intersection are required. Moreover, existing datasets only provide video-level atomic activity annotations, which require exhausting efforts to manually trim the videos for recognition and limit their applications to untrimmed videos. To bridge this gap, we introduce the Aerial Traffic Atomic Activity Recognition and Segmentation (ATARS) dataset, the first aerial dataset designed for multilabel atomic activity analysis. We offer atomic activity labels for each frame, which accurately record the intervals for traffic activities. Moreover, we propose a novel task, Multi-label Temporal Atomic Activity Recognition, enabling the study of accurate temporal localization for atomic activity and easing the burden of manual video trimming for recognition. We conduct extensive experiments to evaluate existing state-of-theart models on both atomic activity recognition and temporal atomic activity segmentation. The results highlight the unique challenges of our ATARS dataset, such as recognizing extremely small objects’ activities. We further provide a comprehensive discussion analyzing these challenges and offer valuable insights for future direction to improve recognition of atomic activity in an aerial view. Our source code and dataset are available at https://github.com/magecliff96/ATARS/. Zihao Chen 0003, Hsuanyu Wu, Chi-Hsi Kung, Yi-Ting Chen 0001, Yan-Tsung Peng |
IROS | 4 |
| 2024 | Action-Slot: Visual Action-Centric Representations for Multi-Label Atomic Activity Recognition in Traffic ScenesabstractIn this paper, we study multi-label atomic activity recognition. Despite the notable progress in action recognition, it is still challenging to recognize atomic activities due to a deficiency in holistic understanding of both multiple road users' motions and their contextual information. In this paper, we introduce Action-slot, a slot attention-based approach that learns visual action-centric representations, capturing both motion and contextual information. Our key idea is to design action slots that are capable of paying attention to regions where atomic activities occur, without the need for explicit perception guidance. To further enhance slot attention, we introduce a background slot that competes with action slots, aiding the training process in avoiding un-necessary focus on background regions devoid of activities. Yet, the imbalanced class distribution in the existing dataset hampers the assessment of rare activities. To address the limitation, we collect a synthetic dataset called TACO, which is four times larger than OATS and features a balanced distribution of atomic activities. To validate the effectiveness of our method, we conduct comprehensive experiments and ablation studies against various action recognition baselines. We also show that the performance of multi-label atomic activity recognition on real-world datasets can be improved by pretraining representations on TACO. Our source code, dataset, and visualization videos are available at https://hcis-lab.github.io/Action-slot/. Chi-Hsi Kung, Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting Chen 0001 |
CVPR | 4 |
| 2024 | RiskBench: A Scenario-based Benchmark for Risk IdentificationabstractIntelligent driving systems aim to achieve a zero-collision mobility experience, requiring interdisciplinary efforts to enhance safety performance. This work focuses on risk identification, the process of identifying and analyzing risks stemming from dynamic traffic participants and unexpected events. While significant advances have been made in the community, the current evaluation of different risk identification algorithms uses independent datasets, leading to difficulty in direct comparison and hindering collective progress toward safety performance enhancement. To address this limitation, we introduce RiskBench, a large-scale scenario-based benchmark for risk identification. We design a scenario taxonomy and augmentation pipeline to enable a systematic collection of ground truth risks under diverse scenarios. We assess the ability of ten algorithms to (1) detect and locate risks, (2) anticipate risks, and (3) facilitate decision-making. We conduct extensive experiments and summarize future research on risk identification. Our aim is to encourage collaborative endeavors in achieving a society with zero collisions. We have made our dataset and benchmark toolkit publicly at this project webpage. Chi-Hsi Kung, Chieh-Chi Yang, Pang-Yuan Pao, Shu-Wei Lu, Pin-Lun Chen, Hsin-Cheng Lu, Yi-Ting Chen 0001 |
ICRA | 7 |
| 2024 | Reliability Engineering in a Time of Rapidly Converging TechnologiesabstractThe convergence of technologies is happening across various aspects, such as communication, computing, medicine, and transportation. The smartphone is a perfect example of convergence, packing features, such as a camera, GPS, artificial intelligence, and Internet connectivity into one sleek device. Autonomous driving is another good example. In a time of rapidly converging technologies, reliability engineering must take into account the potential for cyber threats, the need for cyber trust, the importance of cyber security, and the criticality of cyber resilience. In this way, reliability engineers can ensure the confidentiality, integrity, and availability of computer systems and networks in the face of evolving threats and changing technologies. In this article, we introduce the challenges and current progress of reliability engineering in emerging technologies, including practices and applications of cyber trust and security, AI-empowered autonomous driving systems, modern mobile networks, blockchains and distributed ledger technologies, prognostic and health management, integrated circuit and hardware, and enterprise cybersecurity and threat hunting. Shiuh-Pyng Shieh, Jeffrey M. Voas, Phillip A. Laplante, Jason W. Rupe, Christian K. Hansen, Yu-Sung Wu, Yi-Ting Chen 0001, Chi-Yu Li 0001, Kai-Chiang Wu |
IEEE Trans. Reliab. | 7 |
| 2023 | Ordered Atomic Activity for Fine-grained Interactive Traffic Scenario UnderstandingabstractWe introduce a novel representation called Ordered Atomic Activity for interactive scenario understanding. The representation decomposes each scenario into a set of ordered atomic activities, where each activity consists of an action and the corresponding actors involved and the order denotes the temporal development of the scenario. This design also helps in identifying important interactive relationships, such as yielding. The action is a high-level semantic motion pattern that is grounded in the surrounding road topology, which we decompose into zones and corners with unique IDs. For example, a group of pedestrians crossing in front is denoted as C1 → C4: P+, as depicted in Figure 1. We collect a new large-scale dataset called OATS1(Ordered Atomic Activities in interactive Traffic Scenarios), comprising 1026 video clips (~ 20s) captured at intersections in San Francisco Bay Area. Each clip is labeled with the proposed language, resulting in 59 activity categories and 6512 annotated activity instances. We propose three fine-grained scenario understanding tasks, i.e., multilabel atomic activity recognition, activity order prediction, and interactive scenario retrieval. We also propose a Graph Convolutional Network based framework that models both appearance and motion of traffic participants to tackle the above tasks, that performs favorably against state-of-the-art methods. However, we find that the methods cannot achieve satisfactory performance, indicating rising opportunities for the community to develop new algorithms for these tasks towards better interactive scenario understanding. Nakul Agarwal, Yi-Ting Chen 0001 |
ICCV | 2 |
| 2023 | Content Estimation Through Tactile Interactions with Deformable ContainersabstractPouring snacks and moving containers with beverages are challenging for a service robot. To obtain accurate content properties for planning robotic motion, tactile sensing can provide information about the pressure distribution of the contact surface, which is not obvious by visual observation. In this work, we focus on estimating the content properties of various content materials in distinct deformable containers through tactile interactions. We propose a learning-based model that can estimate content properties by using the tactile data collected by slightly squeezing a container with the content of interest. We analyzed an uncalibrated tactile sensor and collected a dataset consisting of 1125 sets of tactile sequences, which are combinations of five types of deformable containers and eleven types of content materials in different content heights. Experiments were conducted on content estimation with known contents and containers, unknown contents, and unknown containers. For unknown contents, our model can still achieve 8.5% height relative error and 79.7% state of matter accuracy. Furthermore, we analyzed that the tactile features of contents with similar content properties are close in the latent snace to show the effectiveness of our model. Yu-En Liu, Chun-Yu Chai, Yi-Ting Chen 0001, Shiao-Li Tsao |
IROS | 3 |
| 2023 | DROID: Driver-Centric Risk Object IdentificationabstractIdentification of high-risk driving situations is generally approached through collision risk estimation or accident pattern recognition. In this work, we approach the problem from the perspective of subjective risk. We operationalize subjective risk assessment by predicting driver behavior changes and identifying the cause of changes. To this end, we introduce a new task called driver-centric risk object identification (DROID), which uses egocentric video to identify object(s) influencing a driver's behavior, given only the driver's response as the supervision signal. We formulate the task as a cause-effect problem and present a novel two-stage DROID framework, taking inspiration from models of situation awareness and causal inference. A subset of data constructed from the Honda Research Institute Driving Dataset (HDD) is used to evaluate DROID. We demonstrate state-of-the-art DROID performance, even compared with strong baseline models using this dataset. Additionally, we conduct extensive ablative studies to justify our design choices. Moreover, we demonstrate the applicability of DROID for risk assessment. Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Semi-supervised 3D Object Detection via Temporal Graph Neural Networksabstract3D object detection plays an important role in autonomous driving and other robotics applications. However, these detectors usually require training on large amounts of annotated data that is expensive and time-consuming to collect. Instead, we propose leveraging large amounts of unlabeled point cloud videos by semi-supervised learning of 3D object detectors via temporal graph neural networks. Our insight is that temporal smoothing can create more accurate detection results on unlabeled data, and these smoothed detections can then be used to retrain the detector. We learn to perform this temporal reasoning with a graph neural network, where edges represent the relationship between candidate detections in different time frames. After semi-supervised learning, our method achieves state-of-the-art detection performance on the challenging nuScenes [3] and H3D [19] benchmarks, compared to baselines trained on the same amount of labeled data. Project and code are released at https://www.jianrenw.com/SOD-TGNN/. Jianren Wang, Haiming Gang, Siddarth Ancha, Yi-Ting Chen 0001, David Held |
3DV | 4 |
| 2021 | Bird's Eye View Segmentation Using Lifted 2D Semantic Features
Isht Dwivedi, Srikanth Malla, Yi-Ting Chen 0001, Behzad Dariush |
BMVC | 3 |
| 2020 | Unsupervised Domain Adaptation for Spatio-Temporal Action Localization
Nakul Agarwal, Yi-Ting Chen 0001, Behzad Dariush, Ming-Hsuan Yang 0001 |
BMVC | 2 |
| 2020 | Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional NetworksabstractTo enable intelligent automated driving systems, a promising strategy is to understand how human drives and interacts with road users in complicated driving situations. In this paper, we propose a 3D-aware egocentric spatial-temporal interaction framework for automated driving applications. Graph convolution networks (GCN) is devised for interaction modeling. We introduce three novel concepts into GCN. First, we decompose egocentric interactions into ego-thing and ego- stuff interaction, modeled by two GCNs. In both GCNs, ego nodes are introduced to encode the interaction between thing objects (e.g., car and pedestrian), and interaction between stuff objects (e.g., lane marking and traffic light). Second, objects' 3D locations are explicitly incorporated into GCN to better model egocentric interactions. Third, to implement ego-stuff interaction in GCN, we propose a MaskAlign operation to extract features for irregular objects.We validate the proposed framework on tactical driver behavior recognition. Extensive experiments are conducted using Honda Research Institute Driving Dataset, the largest dataset with diverse tactical driver behavior annotations. Our framework demonstrates substantial performance boost over baselines on the two experimental settings by 3.9% and 6.0%, respectively. Furthermore, we visualize the learned affinity matrices, which encode ego-thing and ego-stuff interactions, to showcase the proposed framework can capture interactions effectively. Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001 |
ICRA | 4 |
| 2020 | Who Make Drivers Stop? Towards Driver-centric Risk Assessment: Risk Object Identification via Causal InferenceabstractA significant amount of people die in road accidents due to driver errors. To reduce fatalities, developing intelligent driving systems assisting drivers to identify potential risks is in an urgent need. Risky situations are generally defined based on collision prediction in the existing works. However, collision is only a source of potential risks, and a more generic definition is required. In this work, we propose a novel driver-centric definition of risk, i.e., objects influencing drivers' behavior are risky. A new task called risk object identification is introduced. We formulate the task as the cause-effect problem and present a novel two-stage risk object identification framework based on causal inference with the proposed object-level manipulable driving model. We demonstrate favorable performance on risk object identification compared with strong baselines on the Honda Research Institute Driving Dataset (HDD). Our framework achieves a substantial average performance boost over a strong baseline by 7.5%. Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001 |
IROS | 3 |
| 2020 | Uncertainty-aware Self-supervised 3D Data Associationabstract3D object trackers usually require training on large amounts of annotated data that is expensive and time-consuming to collect. Instead, we propose leveraging vast unlabeled datasets by self-supervised metric learning of 3D object trackers, with a focus on data association. Large scale annotations for unlabeled data are cheaply obtained by automatic object detection and association across frames. We show how these self-supervised annotations can be used in a principled manner to learn point-cloud embeddings that are effective for 3D tracking. We estimate and incorporate uncertainty in self-supervised tracking to learn more robust embeddings, without needing any labeled data. We design embeddings to differentiate objects across frames, and learn them using uncertainty-aware self-supervised training. Finally, we demonstrate their ability to perform accurate data association across frames, towards effective and accurate 3D tracking. Project videos and code are at https://jianrenw.github.io/Self-Supervised-3D-Data-Association/. Jianren Wang, Siddharth Ancha, Yi-Ting Chen 0001, David Held |
IROS | 3 |
| 2020 | Boosting Standard Classification Architectures Through a Ranking RegularizerabstractWe employ triplet loss as a feature embedding regularizer to boost classification performance. Standard architectures, like ResNet and Inception, are extended to support both losses with minimal hyper-parameter tuning. This promotes generality while fine-tuning pretrained networks. Triplet loss is a powerful surrogate for recently proposed embedding regularizers. Yet, it is avoided due to large batch-size requirement and high computational cost. Through our experiments, we re-assess these assumptions.During inference, our network supports both classification and embedding tasks without any computational overhead. Quantitative evaluation highlights a steady improvement on five fine-grained recognition datasets. Further evaluation on an imbalanced video dataset achieves significant improvement. Triplet loss brings feature embedding capabilities like nearest neighbor to classification models. Code available at http://bit.ly/2LNYEqL. Ahmed Taha 0001, Yi-Ting Chen 0001, Teruhisa Misu, Abhinav Shrivastava, Larry Davis 0001 |
WACV | 2 |
| 2019 | Grounding Human-To-Vehicle Advice for Self-Driving VehiclesabstractRecent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human actions, but they lack semantic understanding of image contents. This makes them brittle and potentially unsafe in situations that do not match training data. Here, we propose to address this issue by augmenting training data with natural language advice from a human. Advice includes guidance about what to do and where to attend. We present the first step toward advice giving, where we train an end-to-end vehicle controller that accepts advice. The controller adapts the way it attends to the scene (visual attention) and the control (steering and speed). Attention mechanisms tie controller behavior to salient objects in the advice. We evaluate our model on a novel advisable driving dataset with manually annotated human-to-vehicle advice called Honda Research Institute-Advice Dataset (HAD). We show that taking advice improves the performance of the end-to-end network, while the network cues on a variety of visual features that are provided by advice. The dataset is available at https://usa.honda-ri.com/HAD. Jinkyu Kim 0001, Teruhisa Misu, Yi-Ting Chen 0001, Ashish Tawari, John F. Canny |
CVPR | 3 |
| 2019 | Temporal Recurrent Networks for Online Action DetectionabstractMost work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions as soon as each video frame arrives, based only on current and historical observations. In this paper, we propose a novel framework, the Temporal Recurrent Network (TRN), to model greater temporal context of each frame by simultaneously performing online action detection and anticipation of the immediate future. At each moment in time, our approach makes use of both accumulated historical evidence and predicted future information to better recognize the action that is currently occurring, and integrates both of these into a unified end-to-end architecture. We evaluate our approach on two popular online action detection datasets, HDD and TVSeries, as well as another widely used dataset, THUMOS'14. The results show that TRN significantly outperforms the state-of-the-art. Mingfei Gao, Yi-Ting Chen 0001, Larry Davis 0001, David Crandall |
ICCV | 3 |
| 2019 | The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenesabstract3D multi-object detection and tracking are crucial for traffic scene understanding. However, the community pays less attention to these areas due to the lack of a standardized benchmark dataset to advance the field. Moreover, existing datasets (e.g., KITTI [1]) do not provide sufficient data and labels to tackle challenging scenes where highly interactive and occluded traffic participants are present. To address the issues, we present the Honda Research Institute 3D Dataset (H3D), a large-scale full-surround 3D multi-object detection and tracking dataset collected using a 3D LiDAR scanner. H3D comprises of 160 crowded and highly interactive traffic scenes with a total of 1 million labeled instances in 27,721 frames. With unique dataset size, rich annotations, and complex scenes, H3D is gathered to stimulate research on full-surround 3D multi-object detection and tracking. To effectively and efficiently annotate a large-scale 3D point cloud dataset, we propose a labeling methodology to speed up the overall annotation cycle. A standardized benchmark is created to evaluate full-surround 3D multi-object detection and tracking algorithms. 3D object detection and tracking algorithms are trained and tested on H3D. Finally, sources of errors are discussed for the development of future algorithms. Abhishek Patil, Srikanth Malla, Haiming Gang, Yi-Ting Chen 0001 |
ICRA | 4 |
| 2018 | Toward Driving Scene Understanding: A Dataset for Learning Driver Behavior and Causal ReasoningabstractDriving Scene understanding is a key ingredient for intelligent transportation systems. To achieve systems that can operate in a complex physical and social environment, they need to understand and learn how humans drive and interact with traffic scenes. We present the Honda Research Institute Driving Dataset (HDD), a challenging dataset to enable research on learning driver behavior in real-life environments. The dataset includes 104 hours of real human driving in the San Francisco Bay Area collected using an instrumented vehicle equipped with different sensors. We provide a detailed analysis of HDD with a comparison to other driving datasets. A novel annotation methodology is introduced to enable research on driver behavior understanding from untrimmed data sequences. As the first step, baseline algorithms for driver behavior detection are trained and tested to demonstrate the feasibility of the proposed task. Vasili Ramanishka, Yi-Ting Chen 0001, Teruhisa Misu, Kate Saenko |
CVPR | 2 |
| 2018 | A 3D Dynamic Scene Analysis Framework for Development of Intelligent Transportation SystemsabstractHolistic driving scene understanding is a critical step toward intelligent transportation systems. It involves different levels of analysis, interpretation, reasoning and decision making. In this paper, we propose a 3D dynamic scene analysis framework as the first step toward driving scene understanding. Specifically, given a sequence of synchronized 2D and 3D sensory data, the framework systematically integrates different perception modules to obtain 3D position, orientation, velocity and category of traffic participants and the ego car in a reconstructed 3D semantically labeled traffic scene. We implement this framework and demonstrate the effectiveness in challenging urban driving scenarios. The proposed framework builds a foundation for higher level driving scene understanding problems such as intention and motion prediction of surrounding entities, ego motion planning, and decision making. Chien-Yi Wang, Athma Narayanan, Abhishek Patil, Yi-Ting Chen 0001 |
Intelligent Vehicles Symposium | 5 |
| 2018 | Probabilistic Prediction from Planning Perspective: Problem Formulation, Representation Simplification and Evaluation MetricabstractAccurate probabilistic prediction for intention and motion of road users is a key prerequisite to achieve safe and high-quality decision-making and motion planning for autonomous driving. Typically, the performance of probabilistic predictions was only evaluated by learning metrics for approximation to the motion distribution in the dataset. However, as a module supporting decision and planning, probabilistic prediction should also be evaluated from decision and planning perspective. Moreover, the evaluation of probabilistic prediction highly relies on the problem formulation variation and motion representation simplification, which lacks a formal foundation in a comprehensive framework. To address such concerns, we provide a systematic and unified framework for the analysis of three under-explored aspects of probabilistic prediction: problem formulation, representation simplification and evaluation metric. More importantly, we address the omitted but crucial problems in the three aspects from decision and planning perspective. In addition to a review of learning metrics, metrics to be considered from planning perspective are highlighted, such as planning consequence of inaccurate and erroneous prediction, as well as violations of predicted motions to planning constraints. We address practical formulation variations of prediction problems, such as decision-maker view and blind view for viewpoint, as well as reactive prediction for interaction, so that decision and planning can be facilitated. Arnaud de La Fortelle, Yi-Ting Chen 0001, Ching-Yao Chan, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2017 | Monocular localization in urban environments using road markingsabstractLocalization is an essential problem in autonomous navigation of self-driving cars. We present a monocular vision based approach for localization in urban environments using road markings. We utilize road markings as landmarks instead of traditional visual features (e.g. SIFT) to tackle the localization problem because road markings are more robust against changes in perspective, illumination, and across time. Specifically, we employ Chamfer matching to register edges of road markings against a lightweight 3D map where road markings are represented as a set of sparse points. By only matching geometry of road markings, our localization algorithm further gains robustness against photometric appearance changes in the environment. We take vehicle odometry and epipolar geometry constraints into account and formulate a non-linear optimization problem to estimate the 6 DoF camera pose. We evaluate the proposed method on data collected in the real world. Experimental results show that our method achieves sub-meter localization errors in areas with sufficient road markings. Yi-Ting Chen 0001, Bernd Heisele |
Intelligent Vehicles Symposium | 3 |
| 2015 | Multi-instance object segmentation with occlusion handlingabstractWe present a multi-instance object segmentation algorithm to tackle occlusions. As an object is split into two parts by an occluder, it is nearly impossible to group the two separate regions into an instance by purely bottomup schemes. To address this problem, we propose to incorporate top-down category specific reasoning and shape prediction through exemplars into an intuitive energy minimization framework. We perform extensive evaluations of our method on the challenging PASCAL VOC 2012 segmentation set. The proposed algorithm achieves favorable results on the joint detection and segmentation task against the state-of-the-art method both quantitatively and qualitatively. Yi-Ting Chen 0001, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2015 | Extracting Image Regions by Structured Edge PredictionabstractWe present two approaches to extract regions from structured edge detection. While the state-of-the-art algorithm based on globalized probability of boundary (gPb) generates a hierarchical region tree, it entails significant computational load. In this work, we exploit an efficient algorithm for structured edge prediction to extract regions. To generate high quality regions, we develop a novel algorithm to link the structured edge and gPb hierarchical image segmentation framework with steerable filters. The extracted regions are grouped by the proposed hierarchical grouping method to generate object proposals for effective detection and recognition problems. We demonstrate the effectiveness of our region generation for image segmentation on the BSDS500 database, and region generation for object proposals on the PASCAL VOC 2007 benchmark database. Experimental results show that the proposed algorithm achieves the comparable or superior quality to the state-of-the-art methods. Yi-Ting Chen 0001, Jimei Yang, Ming-Hsuan Yang 0001 |
WACV | 1 |