Pia Bideau

dblp:179/2629 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-8145-1732ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Salience-SGG: Enhancing Unbiased Scene Graph Generation with Iterative Salience Estimation
abstract
Scene Graph Generation (SGG) suffers from a long-tailed distribution, where a few predicate classes dominate while many others are underrepresented, leading to biased models that underperform on rare relations. Unbiased-SGG methods address this by implementing debiasing strategies, but often at the cost of spatial understanding—resulting in over-reliance on semantic priors. We introduce Salience-SGG, a novel framework featuring an Iterative Salience Decoder (ISD) that emphasizes triplets with salient spatial structures. To support this, we propose semantic-agnostic salience labels guiding ISD. Evaluations on Visual Genome, Open Images V6, and GQA-200 show that Salience-SGG achieves state-of-the-art performance and improves existing Unbiased-SGG methods in their spatial understanding as demonstrated by the Pairwise Localization Average Precision. Code is available at: https://github.com/runfeng-q/Salience-SGG.
Runfeng Qu, Ole Hall, Pia Bideau, Julie Ouerfelli-Ethier, Martin Rolfs, Klaus Obermayer, Olaf Hellwich
WACV3
2026 Watching Swarm Dynamics from Above: A Framework for Advanced Object Tracking in Drone Videos
abstract
Abstract Easily accessible technologies, such as drones equipped with diverse onboard sensors, have greatly expanded opportunities to study animal behavior in natural environments. However, analyzing large volumes of unlabeled video data, often spanning hours, remains a significant challenge for machine learning, particularly in computer vision. Existing approaches typically process only a small number of frames, and accurate georeferencing of tracked positions is still largely unresolved, particularly in dynamic environments where static landmarks cannot be established. In this work, we focus on long-term tracking of animal behavior in real-world geographic coordinates. To address this challenge, we utilize classical probabilistic methods for state estimation, such as particle filtering. Particle filters offer a useful algorithmic structure for recursively adding new incoming information and thus ensuring time consistency. By incorporating recent developments in semantic object segmentation, we enable continuous tracking of rapidly evolving object formations, even in scenarios with limited data availability. We propose a novel approach for tracking schools of fish in the open ocean from drone videos. Our framework not only performs classical object tracking in image coordinates, instead it additionally tracks the position and spatial expansion of the fish school in geographic coordinates by fusing video data and the drone’s on board sensor information (GPS and IMU). No landmarks with known geographic coordinates are required, making the proposed method adaptable to unstructured, dynamic environments like the open ocean, where static landmarks are unavailable. With this, the presented framework enables researchers to study the collective behavior of fish schools within their social and environmental context.
Pia Bideau, Duc Pham, Félicie Dhellemmes, Matthew Hansen, Jens Krause
Int. J. Comput. Vis.1
2025 Active Event Alignment for Monocular Distance Estimation
abstract
Event cameras provide a natural and data efficient representation of visual information, motivating novel computational strategies towards extracting visual information. Inspired by the biological vision system, we propose a behavior driven approach for object-wise distance estimation from event camera data. This behavior-driven method mimics how biological systems, like the human eye, stabilize their view based on ob-ject distance: distant objects require minimal compensatory rotation to stay in focus, while nearby objects demand greater adjustments to maintain alignment. This adaptive strategy leverages natural stabilization behaviors to estimate relative distances effectively. Unlike traditional vision algorithms that estimate depth across the entire image, our approach targets local depth estimation within a specific region of interest. By aligning events within a small region, we estimate the angular velocity required to stabilize the image motion. We demonstrate that, under certain assumptions, the compensatory rotational flow is inversely proportional to the object's distance. The proposed approach achieves new state-of-the-art accuracy in distance estimation - a performance gain of 16% on EVIM02. EVIM02 event sequences comprise complex camera motion and substantial variance in depth of static real world scenes. Code: https://github.com/pbideau/AAEDepth.
Nan Cai, Pia Bideau
WACV2
2024 Model AI Assignments 2024
abstract
The Model AI Assignments session seeks to gather and dis- seminate the best assignment designs of the Artificial In- telligence (AI) Education community. Recognizing that as- signments form the core of student learning experience, we here present abstracts of five AI assignments from the 2024 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment spec- ifications and supporting resources may be found at http://modelai.gettysburg.edu.
Todd W. Neller, Pia Bideau, David Bierbach, Wolfgang Hönig, Nir Lipovetzky, Christian J. Muise, Lino Coria, Claire Wong, Stephanie Rosenthal
AAAI2
2024 The Right Spin: Learning Object Motion from Rotation-Compensated Flow Fields
abstract
Abstract A good understanding of geometrical concepts as well as a broad familiarity with objects lead to excellent human perception of moving objects. The human ability to detect and segment moving objects works in the presence of multiple objects, complex background geometry, motion of the observer and even camouflage. How we perceive moving objects so reliably is a longstanding research question in computer vision and borrows findings from related areas such as psychology, cognitive science and physics. One approach to the problem is to teach a deep network to model all of these effects. This is in contrast with the strategy used by human vision, where cognitive processes and body design are tightly coupled and each is responsible for certain aspects of correctly identifying moving objects. Similarly, from the computer vision perspective there is evidence that classical, geometry-based techniques are better suited to the “motion-based” parts of the problem, while deep networks are more suitable for modeling appearance. In this work, we argue that the coupling of camera rotation and camera translation can create complex motion fields that are difficult for a deep network to untangle directly. We present a novel probabilistic model to estimate the camera’s rotation given the motion field. We then rectify the flow field to obtain a rotation-compensated motion field for subsequent segmentation. This strategy of first estimating camera motion, and then allowing a network to learn the remaining parts of the problem, yields improved results on the widely used DAVIS benchmark as well as the more recent motion segmentation data set MoCA (Moving Camouflaged Animals).
Pia Bideau, Erik G. Learned-Miller, Cordelia Schmid, Karteek Alahari
Int. J. Comput. Vis.1
2022 Action-Based Contrastive Learning for Trajectory Prediction
Marah Halawa, Olaf Hellwich, Pia Bideau
ECCV (39)3
2021 The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data
abstract
Event cameras, inspired by biological vision systems, provide a natural and data efficient representation of visual information. Visual information is acquired in the form of events that are triggered by local brightness changes. However, because most brightness changes are triggered by relative motion of the camera and the scene, the events recorded at a single sensor location seldom correspond to the same world point. To extract meaningful information from event cameras, it is helpful to register events that were triggered by the same underlying world point. In this work we propose a new model of event data that captures its natural spatio-temporal structure. We start by developing a model for aligned event data. That is, we develop a model for the data as though it has been perfectly registered already. In particular, we model the aligned data as a spatio-temporal Poisson point process. Based on this model, we develop a maximum likelihood approach to registering events that are not yet aligned. That is, we find transformations of the observed events that make them as likely as possible under our model. In particular we extract the camera rotation that leads to the best event alignment. We show new state of the art accuracy for rotational velocity estimation on the DAVIS 240C dataset [20]. In addition, our method is also faster and has lower computational complexity than several competing methods. Code: https://github.com/pbideau/Event-ST-PPP
Erik G. Learned-Miller, Daniel Sheldon, Guillermo Gallego 0002, Pia Bideau
ICCV5
2020 C-14: assured timestamps for drone videos
abstract
Inexpensive and highly capable unmanned aerial vehicles (aka drones) have enabled people to contribute high-quality videos at a global scale. However, a key challenge exists for accepting videos from untrusted sources: establishing when a particular video was taken. Once a video has been received or posted publicly, it is evident that the video was created before that time, but there are no current methods for establishing how long it was made before that time.
Zhipeng Tang, Fabien Delattre, Pia Bideau, Mark D. Corner, Erik G. Learned-Miller
MobiCom3
2018 The Best of Both Worlds: Combining CNNs and Geometric Constraints for Hierarchical Motion Segmentation
abstract
Traditional methods of motion segmentation use powerful geometric constraints to understand motion, but fail to leverage the semantics of high-level image understanding. Modern CNN methods of motion analysis, on the other hand, excel at identifying well-known structures, but may not precisely characterize well-known geometric constraints. In this work, we build a new statistical model of rigid motion flow based on classical perspective projection constraints. We then combine piecewise rigid motions into complex deformable and articulated objects, guided by semantic segmentation from CNNs and a second "object-level" statistical model. This combination of classical geometric knowledge combined with the pattern recognition abilities of CNNs yields excellent performance on a wide range of motion segmentation benchmarks, from complex geometric scenes to camouflaged animals.
Pia Bideau, Aruni Roy Chowdhury, Rakesh R. Menon, Erik G. Learned-Miller
CVPR1
2016 It's Moving! A Probabilistic Model for Causal Motion Segmentation in Moving Camera Videos
Pia Bideau, Erik G. Learned-Miller
ECCV (8)1