Christoph Mertz

dblp:53/204 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-7540-5211ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Systems, architecture and hardware · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neuro-Geometric Zero-Shot Anomaly Detection for Lab Automation
Kashish Gandhi, Christoph Mertz
ICPR (14)4
2025 ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Anurag Ghosh, Robert Tamburo, Khiem Vuong, Juan R. Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan
ICCV9
2023 Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection
abstract
Real-time efficient perception is critical for autonomous navigation and city scale sensing. Orthogonal to architectural improvements, streaming perception approaches have exploited adaptive sampling improving real-time detection performance. In this work, we propose a learnable geometry-guided prior that incorporates rough geometry of the 3D scene (a ground plane and a plane above) to resample images for efficient object detection. This significantly improves small and far-away object detection performance while also being more efficient both in terms of latency and memory. For autonomous navigation, using the same detector and scale, our approach improves detection rate by +4.1 APsor +39% and in real-time performance by +5.3 sAPs or +63% for small objects over state-of-the-art (SOTA). For fixed traffic cameras, our approach detects small objects at image scales other methods cannot. At the same scale, our approach improves detection of small objects by 195% (+12.5 APS) over naive-downsampling and 63% (+4.2 APS) over SOTA.
Anurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. Narasimhan
CVPR3
2023 Toward Map Updates with Crosswalk Change Detection Using a Monocular Bus Camera
abstract
Detecting when road maps change is useful for autonomous vehicles to drive safely and legally, for city planners to make more educated decisions, and web maps to better serve consumers. Many public vehicles drive around the city on a regular basis and collect road data for security and safety purposes through dash cams, yet few cities and companies have considered this as a data source for city monitoring. We present an automatic method and system for crosswalk change detection at city intersections using a monocular camera on a city bus and analyze longitudinal results over the course of a year. Using images recorded by a bus two years ago as reference, multiple city intersections are reconstructed, fitted for ground planes, and labelled for crosswalks. Subsequent images from the bus are imported and processed to detect if changes have occurred since intersections were first seen by first localizing current images with respect to the reference images, detecting for crosswalks, and computing detection overlaps in the bird’s-eye-view. Our method makes improvements upon baseline methods by checking for crosswalk visibility and localization errors, is able to generate results typically seen by using more expensive LiDAR sensors, and has been successfully deployed live for one month.
Tom Bu, Christoph Mertz, John M. Dolan
IV2
2023 CrackFormer Network for Pavement Crack Segmentation
abstract
In this paper, we rethink our earlier work on self-attention based crack segmentation, and propose an upgraded CrackFormer network (CrackFormer-II) for pavement crack segmentation, instead of only for fine-grained crack-detection tasks. This work embeds novel Transformer encoder modules into a SegNet-like encoder-decoder structure, where the basic module is composed of novel Transformer encoder blocks with effective relative positional embedding and long range interactions to extract efficient contextual information from feature-channels. Further, fusion modules of scaling-attention are proposed to integrate the results of each respective encoder and decoder block to highlight semantic features and suppress non-semantic ones. Moreover, we update the Transformer encoder blocks enhanced by the local feed-forward layer and skip-connections, and optimize the channel configurations to compress the model parameters. Compared with the original CrackFormer, the CrackFormer-II is trained and evaluated on more general crack datasets. It achieves higher accuracy than the original CrackFormer, and the state-of-the-art (SOTA) method with$6.7 \times $fewer FLOPs and$6.2 \times $fewer parameters, and its practical inference speed is comparable to most classical CNN models. The experimental results show that it achieves the F-measures on Optimal Dataset Scale (ODS) of 0.912, 0.908, 0.914 and 0.869, respectively, on the four benchmarks. Codes are available athttps://github.com/LouisNUST/CrackFormer-II.
Huajun Liu, Jing Yang 0053, Xiangyu Miao, Christoph Mertz, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.4
2022 Multimodal Object Detection via Probabilistic Ensembling
Jinghao Shi, Zelin Ye, Christoph Mertz, Deva Ramanan, Shu Kong
ECCV (9)4
2021 CrackFormer: Transformer Network for Fine-Grained Crack Detection
abstract
Cracks are irregular line structures that are of interest in many computer vision applications. Crack detection (e.g., from pavement images) is a challenging task due to intensity in-homogeneity, topology complexity, low contrast and noisy background. The overall crack detection accuracy can be significantly affected by the detection performance on fine-grained cracks. In this work, we propose a Crack Transformer network (CrackFormer) for fine-grained crack detection. The CrackFormer is composed of novel attention modules in a SegNet-like encoder-decoder architecture. Specifically, it consists of novel self-attention modules with 1x1 convolutional kernels for efficient contextual information extraction across feature-channels, and efficient positional embedding to capture large receptive field contextual information for long range interactions. It also introduces new scaling-attention modules to combine outputs from the corresponding encoder and decoder blocks to suppress non-semantic features and sharpen semantic ones. The CrackFormer is trained and evaluated on three classical crack datasets. The experimental results show that the CrackFormer achieves the Optimal Dataset Scale (ODS) values of 0.871, 0.877 and 0.881, respectively, on the three datasets and outperforms the state-of-the-art methods.
Huajun Liu, Xiangyu Miao, Christoph Mertz, Cheng-Zhong Xu 0001, Hui Kong 0001
ICCV3
2021 Linear Inverse Problem for Depth Completion with RGB Image and Sparse LIDAR Fusion
abstract
Comprehensive depth information from surrounding scenes is important for perception in autonomous driving and robots. Sparse LIDAR sensors give a low-density point cloud of the environment, but are more affordable than their high-density counterparts. In this paper, we propose a novel sensor fusion architecture for sparse LIDAR depth completion. Instead of the traditional end-to-end neural network-based algorithm, we formulate depth completion as a Linear Inverse Problem (LIP) with a multi-modal proximal operator. This sensor fusion architecture allows a better signal prior and finds the unique optimal solution to the LIP. Instead of learning a unified network for the sparse input which treats pixels evenly, the proposed architecture guarantees both the data consistency and smoothness of the predicted depth map. To demonstrate the performance of our algorithm, we benchmark on the simulation dataset TartanAir, and the real indoor NYUdepthv2 and real outdoor KITTI datasets. Our proposed method outperforms previous methods and uses fewer parameters in both indoor and outdoor datasets.
Christoph Mertz, John M. Dolan
ICRA2
2020 Depth Completion via Inductive Fusion of Planar LIDAR and Monocular Camera
abstract
Modern high-definition LIDAR is expensive for commercial autonomous driving vehicles and small indoor robots. An affordable solution to this problem is fusion of planar LIDAR with RGB images to provide a similar level of perception capability. Even though state-of-the-art methods provide approaches to predict depth information from limited sensor input, they are usually a simple concatenation of sparse LIDAR features and dense RGB features through an end-to-end fusion architecture. In this paper, we introduce an inductive late-fusion block which better fuses different sensor modalities inspired by a probability model. The proposed demonstration and aggregation network propagates the mixed context and depth features to the prediction network and serves as a prior knowledge of the depth completion. This late-fusion block uses the dense context features to guide the depth prediction based on demonstrations by sparse depth features. In addition to evaluating the proposed method on benchmark depth completion datasets including NYUDepthV2 and KITTI, we also test the proposed method on a simulated planar LIDAR dataset. Our method shows promising results compared to previous approaches on both the benchmark datasets and simulated dataset with various 3D densities.
Chiyu Dong, Christoph Mertz, John M. Dolan
IROS3
2019 DeepDA: LSTM-based Deep Data Association Network for Multi-Targets Tracking in Clutter
Huajun Liu, Christoph Mertz
FUSION3
2018 PCN: Point Completion Network
abstract
Shape completion, the problem of estimating the complete geometry of objects from partial observations, lies at the core of many vision and robotics applications. In this work, we propose Point Completion Network (PCN), a novel learning-based approach for shape completion. Unlike existing shape completion methods, PCN directly operates on raw point clouds without any structural assumption (e.g. symmetry) or annotation (e.g. semantic class) about the underlying shape. It features a decoder design that enables the generation of fine-grained completions while maintaining a small number of parameters. Our experiments show that PCN produces dense, complete point clouds with realistic structures in the missing regions on inputs with various levels of incompleteness and noise, including cars from LiDAR scans in the KITTI dataset.
Tejas Khot, David Held, Christoph Mertz, Martial Hebert
3DV4
2014 Visual sensing for developing autonomous behavior in snake robots
abstract
Snake robots are uniquely qualified to investigate a large variety of settings including archaeological sites, natural disaster zones, and nuclear power plants. For these applications, modular snake robots have been tele-operated to perform specific tasks using images returned to it from an onboard camera in the robots head. In order to give the operator an even richer view of the environment and to enable the robot to perform autonomous tasks we developed a structured light sensor that can make three-dimensional maps of the environment. This paper presents a sensor that is uniquely qualified to meet the severe constraints in size, power and computational footprint of snake robots. Using range data, in the form of 3D pointclouds, we show that it is possible to pair high-level planning with mid-level control to accomplish complex tasks without operator intervention.
Hugo Ponte, Max Queenan, Chaohui Gong, Christoph Mertz, Matthew J. Travers, Florian Enner, Martial Hebert, Howie Choset
ICRA4
2014 Vision for road inspection
abstract
Road surface inspection in cities is for the most part, a task performed manually. Being a subjective and labor intensive process, it is an ideal candidate for automation. We propose a solution based on computer vision and data-driven methods to detect distress on the road surface. Our method works on images collected from a camera mounted on the windshield of a vehicle. We use an automatic procedure to select images suitable for inspection based on lighting and weather conditions. From the selected data we segment the ground plane and use texture, color and location information to detect the presence of pavement distress. We describe an over-segmentation algorithm that identifies coherent image regions not just in terms of color, but also texture. We also discuss the problem of learning from unreliable human-annotations and propose using a weakly supervised learning algorithm (Multiple Instance Learning) to train a classifier. We present results from experiments comparing the performance of this approach against multiple individual human labelers, with the ground-truth labels obtained from an ensemble of other human labelers. Finally, we show results of pavement distress scores computed using our method over a subset of a citywide road network.
Srivatsan Varadharajan, Sobhagya Jose, Karan Sharma, Lars Wander, Christoph Mertz
WACV5
2009 Planning-based prediction for pedestrians
abstract
We present a novel approach for determining robot movements that efficiently accomplish the robot's tasks while not hindering the movements of people within the environment. Our approach models the goal-directed trajectories of pedestrians using maximum entropy inverse optimal control. The advantage of this modeling approach is the generality of its learned cost function to changes in the environment and to entirely different environments. We employ the predictions of this model of pedestrian trajectories in a novel incremental planner and quantitatively show the improvement in hindrance-sensitive robot trajectory planning provided by our approach.
Brian D. Ziebart, Nathan D. Ratliff, Garratt Gallagher, Christoph Mertz, Kevin M. Peterson, J. Andrew Bagnell, Martial Hebert, Anind K. Dey, Siddhartha S. Srinivasa
IROS4
2003 Safe Robot Driving in Cluttered Environments
Charles E. Thorpe, Justin Carlson, David Duggins, Jay Gowdy, Robert A. MacLachlan, Christoph Mertz, Arne Suppé, Bob Wang
ISRR6