EDBT 2026 Demo / reviewers in the wild / expert
Yutong Ban
dblp:188/7582
· DBLP profile ↗
18ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-5396-9251ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hypergraph-Transformer (HGT) for Interaction Event Prediction in Laparoscopic and Robotic SurgeryabstractUnderstanding and anticipating events and actions is critical for intraoperative assistance and decision-making during minimally invasive surgery. We propose a predictive neural network that is capable of understanding and predicting critical interaction aspects of surgical workflow based on endoscopic, intracorporeal video data, while flexibly leveraging surgical knowledge graphs. The approach incorporates a hypergraph-transformer (HGT) structure that encodes expert knowledge into the network design and predicts the hidden embedding of the graph. We verify our approach on established surgical datasets and applications, including the prediction of action-triplets, and the achievement of the Critical View of Safety (CVS), which is a critical safety measure. Moreover, we address specific, safety-related forecasts of surgical processes, such as predicting the clipping of the cystic duct or artery without prior achievement of the CVS. Our results demonstrate improvement in prediction of interactive event when incorporating with our approach compared to unstructured alternatives. Lianhao Yin, Yutong Ban, Jennifer A. Eckhoff, Ozanan R. Meireles, Daniela Rus, Guy Rosman |
ICRA | 2 |
| 2025 | Tracking-Aware Deformation Field Estimation for Non-rigid 3D Reconstruction in Robotic SurgeriesabstractMinimally invasive procedures have been advanced rapidly by the robotic laparoscopic surgery. The latter greatly assists surgeons in sophisticated and precise operations with reduced invasiveness. Nevertheless, it is still safety critical to be aware of even the least tissue deformation during instrument-tissue interactions, especially in 3D space. To address this, recent works rely on NeRF to render 2D videos from different perspectives and eliminate occlusions. However, most of the methods fail to predict the accurate 3D shapes and associated deformation estimates robustly. Differently, we propose Tracking-Aware Deformation Field (TADF), a novel framework which reconstructs the 3D mesh along with the 3D tissue deformation simultaneously. It first tracks the key points of soft tissue by a foundation vision model, providing an accurate 2D deformation field. Then, the 2D deformation field is smoothly incorporated with a neural implicit reconstruction network to obtain tissue deformation in the 3D space. Finally, we experimentally demonstrate that the proposed method provides more accurate deformation estimation compared with other 3D neural reconstruction methods in two public datasets. Our demo is available at https://kasumigaoka-utaha.github.io/TADF-web/. Our code is available at https://github.com/Zing110/TADF. Zeqing Wang, Yutong Ban |
IROS | 4 |
| 2025 | Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement LearningabstractSymmetry is pervasive in robotics and has been widely exploited to improve sample efficiency in deep reinforcement learning (DRL). However, existing approaches primarily focus on spatial symmetries—such as reflection, rotation, and translation—while largely neglecting temporal symmetries. To address this gap, we explore time reversal symmetry, a form of temporal symmetry commonly found in robotics tasks such as door opening and closing. We propose Time Reversal symmetry enhanced Deep Reinforcement Learning (TR-DRL), a framework that combines trajectory reversal augmentation and time reversal guided reward shaping to efficiently solve temporally symmetric tasks. Our method generates reversed transitions from fully reversible transitions, identified by a proposed dynamics-consistent filter, to augment the training data. For partially reversible transitions, we apply reward shaping to guide learning, according to successful trajectories from the reversed task. Extensive experiments on the Robosuite and MetaWorld benchmarks demonstrate that TR-DRL is effective in both single-task and multi-task settings, achieving higher sample efficiency and stronger final performance compared to baseline methods. Yunpeng Jiang, Jianshu Hu, Paul Weng, Yutong Ban |
NeurIPS | 4 |
| 2025 | State-novelty guided action persistence in deep reinforcement learning
Jianshu Hu, Paul Weng, Yutong Ban |
Mach. Learn. | 3 |
| 2024 | INViT: A Generalizable Routing Problem Solver with Invariant Nested View TransformerabstractRecently, deep reinforcement learning has shown promising results for learning fast heuristics to solve routing problems. Meanwhile, most of the solvers suffer from generalizing to an unseen distribution or distributions with different scales. To address this issue, we propose a novel architecture, called Invariant Nested View Transformer (INViT), which is designed to enforce a nested design together with invariant views inside the encoders to promote the generalizability of the learned solver. It applies a modified policy gradient algorithm enhanced with data augmentations. We demonstrate that the proposed INViT achieves a dominant generalization performance on both TSP and CVRP problems with various distributions and different problem scales. Our source code and datasets are available in supplementary materials. Zhihao Song, Paul Weng, Yutong Ban |
ICML | 4 |
| 2024 | Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation ModelsabstractAs autonomous driving technology matures, end-to-end methodologies have emerged as a leading strategy, promising seamless integration from perception to control via deep learning. However, existing systems grapple with challenges such as unexpected open set environments and the complexity of black-box models. At the same time, the evolution of deep learning introduces larger, multimodal foundational models, offering multi-modal visual and textual understanding. In this paper, we harness these multimodal foundation models to enhance the robustness and adaptability of autonomous driving systems. We introduce a method to extract nuanced spatial features from transformers and the incorporation of latent space simulation for improved training and policy debugging. We use pixel/patch-aligned feature descriptors to expand foundational model capabilities to create an end-to-end multimodal driving model, demonstrating unparalleled results in diverse tests. Our solution combines language with visual perception and achieves significantly greater robustness on out-of-distribution situations. Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao 0003, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, Daniela Rus |
ICRA | 4 |
| 2024 | Concept Graph Neural Networks for Surgical Video UnderstandingabstractAnalysis of relations between objects and comprehension of abstract concepts in the surgical video is important in AI-augmented surgery. However, building models that integrate our knowledge and understanding of surgery remains a challenging endeavor. In this paper, we propose a novel way to integrate conceptual knowledge into temporal analysis tasks using temporal concept graph networks. In the proposed networks, a knowledge graph is incorporated into the temporal video analysis of surgical notions, learning the meaning of concepts and relations as they apply to the data. We demonstrate results in surgical video data for tasks such as verification of the critical view of safety, estimation of the Parkland grading scale as well as recognizing instrument-action-tissue triplets. The results show that our method improves the recognition and detection of complex benchmarks as well as enables other analytic applications of interest. Yutong Ban, Jennifer A. Eckhoff, Thomas M. Ward, Daniel A. Hashimoto, Ozanan R. Meireles, Daniela Rus, Guy Rosman |
IEEE Trans. Medical Imaging | 1 |
| 2023 | On the Forward Invariance of Neural ODEsabstractWe propose a new method to ensure neural ordinary differential equations (ODEs) satisfy output specifications by using invariance set propagation. Our approach uses a class of control barrier functions to transform output specifications into constraints on the parameters and inputs of the learning system. This setup allows us to achieve output specification guarantees simply by changing the constrained parameters/inputs both during training and inference. Moreover, we demonstrate that our invariance set propagation through data-controlled neural ODEs not only maintains generalization performance but also creates an additional degree of robustness by enabling causal manipulation of the system’s parameters/inputs. We test our method on a series of representation learning tasks, including modeling physical dynamics and convexity portraits, as well as safe collision avoidance for autonomous vehicles. Wei Xiao 0003, Tsun-Hsuan Wang, Ramin M. Hasani, Mathias Lechner, Yutong Ban, Chuang Gan 0001, Daniela Rus |
ICML | 5 |
| 2023 | Infrastructure-based End-to-End Learning and Prevention of Driver FailureabstractIntelligent intersection managers can improve safety by detecting dangerous drivers or failure modes in autonomous vehicles, warning oncoming vehicles as they approach an intersection. In this work, we present FailureNet, a recurrent neural network trained end-to-end on trajectories of both nominal and reckless drivers in a scaled miniature city. FailureNet observes the poses of vehicles as they approach an intersection and detects whether a failure is present in the autonomy stack, warning cross-traffic of potentially dangerous drivers. FailureNet can accurately identify control failures, upstream perception errors, and speeding drivers, distinguishing them from nominal driving. The network is trained and deployed with autonomous vehicles in the MiniCity. Compared to speed or frequency-based predictors, FailureNet's recurrent neural network structure provides improved predictive power, yielding upwards of 84% accuracy when deployed on hardware. Noam Buckman, Shiva Sreeram, Mathias Lechner, Yutong Ban, Ramin M. Hasani, Sertac Karaman, Daniela Rus |
ICRA | 4 |
| 2023 | TransCenter: Transformers With Dense Representations for Multiple-Object TrackingabstractTransformers have proven superior performance for a wide variety of tasks since they were introduced. In recent years, they have drawn attention from the vision community in tasks such as image classification and object detection. Despite this wave, an accurate and efficient multiple-object tracking (MOT) method based on transformers is yet to be designed. We argue that the direct application of a transformer architecture with quadratic complexity and insufficient noise-initialized sparse queries - is not optimal for MOT. We propose TransCenter, a transformer-based MOT architecture with dense representations for accurately tracking all the objects while keeping a reasonable runtime. Methodologically, we propose the use of image-related dense detection queries and efficient sparse tracking queries produced by our carefully designed query learning networks (QLN). On one hand, the dense image-related detection queries allow us to infer targets' locations globally and robustly through dense heatmap outputs. On the other hand, the set of sparse tracking queries efficiently interacts with image features in our TransCenter Decoder to associate object positions through time. As a result, TransCenterexhibits remarkable performance improvements and outperforms by a large margin the current state-of-the-art methods in two standard MOT benchmarks with two tracking settings (public/private). TransCenter is also proven efficient and accurate by an extensive ablation study and, comparisons to more naive alternatives and concurrent works. The code is made publicly available at https://github.com/yihongxu/transcenter. Yutong Ban, Guillaume Delorme 0002, Chuang Gan 0001, Daniela Rus, Xavier Alameda-Pineda |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | A Deep Concept Graph Network for Interaction-Aware Trajectory PredictionabstractTemporal patterns (how vehicles behave in our observed past) underline our reasoning of how people drive on the road, and can explain why we make certain predictions about interactions among road agents. In this paper we propose the ConceptNet trajectory predictor - a novel prediction framework that is able to incorporate agent interactions as explicit edges in a temporal knowledge graph. We demonstrate the sample efficiency and the overall accuracy of the proposed approach, and show that using the graphical structure to explicitly model interactions enables better detection of agent interactions and improved trajectory predictions on a large real-world driving dataset. Yutong Ban, Xiao Li 0025, Guy Rosman, Igor Gilitschenski, Ozanan R. Meireles, Sertac Karaman, Daniela Rus |
ICRA | 1 |
| 2021 | Aggregating Long-Term Context for Learning Laparoscopic and Robot-Assisted Surgical WorkflowsabstractAnalyzing surgical workflow is crucial for surgical assistance robots to understand surgeries. With the understanding of the complete surgical workflow, the robots are able to assist the surgeons in intra-operative events, such as by giving a warning when the surgeon is entering specific keys or high-risk phases. Deep learning techniques have recently been widely applied to recognizing surgical workflows. Many of the existing temporal neural network models are limited in their capability to handle long-term dependencies in the data, instead, relying upon the strong performance of the underlying per-frame visual models. We propose a new temporal network structure that leverages task-specific network representation to collect long-term sufficient statistics that are propagated by a sufficient statistics model (SSM). We implement our approach within an LSTM backbone for the task of surgical phase recognition and explore several choices for propagated statistics. We demonstrate superior results over existing and novel state-of-the-art segmentation techniques on two laparoscopic cholecystectomy datasets: the publicly available Cholec80 dataset and MGH100, a novel dataset with more challenging and clinically meaningful segment labels. Yutong Ban, Guy Rosman, Thomas M. Ward, Daniel A. Hashimoto, Taisei Kondo, Hidekazu Iwaki, Ozanan R. Meireles, Daniela Rus |
ICRA | 1 |
| 2021 | Variational Bayesian Inference for Audio-Visual Tracking of Multiple SpeakersabstractIn this article, we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature and roles of these two modalities in order to accurately estimate smooth trajectories of the tracked persons, to deal with the partial or total absence of one of the modalities over short periods of time, and to estimate the acoustic status-either speaking or silent-of each tracked person over time. We propose to cast the problem at hand into a generative audio-visual fusion (or association) model formulated as a latent-variable temporal graphical model. This may well be viewed as the problem of maximizing the posterior joint distribution of a set of continuous and discrete latent variables given the past and current observations, which is intractable. We propose a variational inference model which amounts to approximate the joint distribution with a factorized distribution. The solution takes the form of a closed-form expectation maximization procedure. We describe in detail the inference algorithm, we evaluate its performance and we compare it with several baseline methods. These experiments show that the proposed audio-visual tracker performs well in informal meetings involving a time-varying number of people. Yutong Ban, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | How to Train Your Deep Multi-Object TrackerabstractThe recent trend in vision-based multi-object tracking (MOT) is heading towards leveraging the representational power of deep learning to jointly learn to detect and track objects. However, existing methods train only certain sub-modules using loss functions that often do not correlate with established tracking evaluation measures such as Multi-Object Tracking Accuracy (MOTA) and Precision (MOTP). As these measures are not differentiable, the choice of appropriate loss functions for end-to-end training of multi-object tracking methods is still an open research problem. In this paper, we bridge this gap by proposing a differentiable proxy of MOTA and MOTP, which we combine in a loss function suitable for end-to-end training of deep multi-object trackers. As a key ingredient, we propose a Deep Hungarian Net (DHN) module that approximates the Hungarian matching algorithm. DHN allows estimating the correspondence between object tracks and ground truth objects to compute differentiable proxies of MOTA and MOTP, which are in turn used to optimize deep trackers directly. We experimentally demonstrate that the proposed differentiable framework improves the performance of existing multi-object trackers, and we establish a new state of the art on the MOTChallenge benchmark. Our code is publicly available from https://github.com/yihongXU/deepMOT. Aljosa Osep, Yutong Ban, Radu Horaud, Laura Leal-Taixé, Xavier Alameda-Pineda |
CVPR | 3 |
| 2019 | Audio-Visual Variational Fusion for Multi-Person Tracking with RobotsabstractRobust multi-person tracking with robots opens the door to analysing engagement and social signals in real-world environments. Multi-person scenarios are charaterised by (i) a time-varying number of people, (ii) intermittent auditory (\eg speech turns) and visual cues (\eg person appearing/disappearing) and (iii) impact of the robot actions in perception. The various sensors (cameras and microphones) available for perception, provide a rich flow of information of intermittent and complementary nature. How to jointly exploit these cues to tackle the multi-person tracking problem with an autonomous system has been an intense research line of the Perception Team in the past few years. In this demo we want to present our, now mature, achievements in the field, and demonstrate two robotic systems able to track multiple persons using auditory and visual cues, when they are available. We will bring the two robots and the necessary computing resources with us, as well as the required presentation materials to discuss the models, methods and tools supporting this technology with the attendants. Xavier Alameda-Pineda, Soraya Arias, Yutong Ban, Guillaume Delorme 0002, Laurent Girin, Radu Horaud, Xiaofei Li 0001, Bastien Mourgue, Guillaume Sarrazin |
ACM Multimedia | 3 |
| 2019 | Tracking Multiple Audio Sources With the von Mises Distribution and Variational EMabstractIn this letter, we address the problem of simultaneously tracking several moving audio sources, namely the problem of estimating source trajectories from a sequence of observed features. We propose to use the von Mises distribution to model audio-source directions of arrival with circular random variables. This leads to a Bayesian filtering formulation, which is intractable because of the combinatorial explosion of associating observed variables with latent variables, over time. We propose a variational approximation of the filtering distribution. We infer a variational expectation-maximization algorithm that is both computationally tractable and time efficient. We propose an audio-source birth method that favors smooth source trajectories and which is used both to initialize the number of active sources and to detect new sources. We perform experiments with the recently released LOCATA dataset comprising two moving sources and a moving microphone array mounted onto a robot. Yutong Ban, Xavier Alameda-Pineda, Christine Evers, Radu Horaud |
IEEE Signal Process. Lett. | 1 |
| 2018 | Accounting for Room Acoustics in Audio-Visual Multi-Speaker TrackingabstractMultiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reverberation. We investigate how to combine state-of-the-art audio-source localization techniques with Bayesian multi-person tracking. Our experiments demonstrate that the performance of the proposed system is not affected by changes in the acoustic environment. Yutong Ban, Xiaofei Li 0001, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
ICASSP | 1 |
| 2017 | Tracking a varying number of people with a visually-controlled robotic headabstractMulti-person tracking with a robotic platform is one of the cornerstones of human-robot interaction. Challenges arise from occlusions, appearance changes and a time-varying number of people. Furthermore, the final system is constrained by the hardware platform: low computational capacity and limited field-of-view. In this paper, we propose a novel method to simultaneously track a time-varying number of persons in three-dimensions and perform visual servoing. The complementary nature of the tracking and visual servoing enables the system to: (i) track several persons while compensating for large ego-movements and (ii) visually control the robot to keep a selected person of interest within the field of view. We propose a variational Bayesian formulation allowing us to effectively solve the inference problem through the use of closed-form solutions. Importantly, this leads to a computationally efficient procedure that runs at 10 FPS. The experiments on the NAO-MPVS dataset confirm the importance of using visual servoing for tracking multiple persons. Yutong Ban, Xavier Alameda-Pineda, Fabien Badeig, Sileye O. Ba, Radu Horaud |
IROS | 1 |