VLDB 2026 Research / reviewers in the wild / expert
Christoph Mertz
dblp:53/204
· DBLP profile ↗
15ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-7540-5211ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Systems, architecture and hardware · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neuro-Geometric Zero-Shot Anomaly Detection for Lab Automation
Kashish Gandhi, Christoph Mertz |
ICPR (14) | 4 |
| 2025 | ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Anurag Ghosh, Robert Tamburo, Khiem Vuong, Juan R. Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan |
ICCV | 9 |
| 2023 | Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object DetectionabstractReal-time efficient perception is critical for autonomous navigation and city scale sensing. Orthogonal to architectural improvements, streaming perception approaches have exploited adaptive sampling improving real-time detection performance. In this work, we propose a learnable geometry-guided prior that incorporates rough geometry of the 3D scene (a ground plane and a plane above) to resample images for efficient object detection. This significantly improves small and far-away object detection performance while also being more efficient both in terms of latency and memory. For autonomous navigation, using the same detector and scale, our approach improves detection rate by +4.1 APsor +39% and in real-time performance by +5.3 sAPs or +63% for small objects over state-of-the-art (SOTA). For fixed traffic cameras, our approach detects small objects at image scales other methods cannot. At the same scale, our approach improves detection of small objects by 195% (+12.5 APS) over naive-downsampling and 63% (+4.2 APS) over SOTA. Anurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. Narasimhan |
CVPR | 3 |
| 2023 | Toward Map Updates with Crosswalk Change Detection Using a Monocular Bus CameraabstractDetecting when road maps change is useful for autonomous vehicles to drive safely and legally, for city planners to make more educated decisions, and web maps to better serve consumers. Many public vehicles drive around the city on a regular basis and collect road data for security and safety purposes through dash cams, yet few cities and companies have considered this as a data source for city monitoring. We present an automatic method and system for crosswalk change detection at city intersections using a monocular camera on a city bus and analyze longitudinal results over the course of a year. Using images recorded by a bus two years ago as reference, multiple city intersections are reconstructed, fitted for ground planes, and labelled for crosswalks. Subsequent images from the bus are imported and processed to detect if changes have occurred since intersections were first seen by first localizing current images with respect to the reference images, detecting for crosswalks, and computing detection overlaps in the bird’s-eye-view. Our method makes improvements upon baseline methods by checking for crosswalk visibility and localization errors, is able to generate results typically seen by using more expensive LiDAR sensors, and has been successfully deployed live for one month. Tom Bu, Christoph Mertz, John M. Dolan |
IV | 2 |
| 2023 | CrackFormer Network for Pavement Crack SegmentationabstractIn this paper, we rethink our earlier work on self-attention based crack segmentation, and propose an upgraded CrackFormer network (CrackFormer-II) for pavement crack segmentation, instead of only for fine-grained crack-detection tasks. This work embeds novel Transformer encoder modules into a SegNet-like encoder-decoder structure, where the basic module is composed of novel Transformer encoder blocks with effective relative positional embedding and long range interactions to extract efficient contextual information from feature-channels. Further, fusion modules of scaling-attention are proposed to integrate the results of each respective encoder and decoder block to highlight semantic features and suppress non-semantic ones. Moreover, we update the Transformer encoder blocks enhanced by the local feed-forward layer and skip-connections, and optimize the channel configurations to compress the model parameters. Compared with the original CrackFormer, the CrackFormer-II is trained and evaluated on more general crack datasets. It achieves higher accuracy than the original CrackFormer, and the state-of-the-art (SOTA) method with$6.7 \times $fewer FLOPs and$6.2 \times $fewer parameters, and its practical inference speed is comparable to most classical CNN models. The experimental results show that it achieves the F-measures on Optimal Dataset Scale (ODS) of 0.912, 0.908, 0.914 and 0.869, respectively, on the four benchmarks. Codes are available athttps://github.com/LouisNUST/CrackFormer-II. Huajun Liu, Jing Yang 0053, Xiangyu Miao, Christoph Mertz, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Multimodal Object Detection via Probabilistic Ensembling
Jinghao Shi, Zelin Ye, Christoph Mertz, Deva Ramanan, Shu Kong |
ECCV (9) | 4 |
| 2021 | CrackFormer: Transformer Network for Fine-Grained Crack DetectionabstractCracks are irregular line structures that are of interest in many computer vision applications. Crack detection (e.g., from pavement images) is a challenging task due to intensity in-homogeneity, topology complexity, low contrast and noisy background. The overall crack detection accuracy can be significantly affected by the detection performance on fine-grained cracks. In this work, we propose a Crack Transformer network (CrackFormer) for fine-grained crack detection. The CrackFormer is composed of novel attention modules in a SegNet-like encoder-decoder architecture. Specifically, it consists of novel self-attention modules with 1x1 convolutional kernels for efficient contextual information extraction across feature-channels, and efficient positional embedding to capture large receptive field contextual information for long range interactions. It also introduces new scaling-attention modules to combine outputs from the corresponding encoder and decoder blocks to suppress non-semantic features and sharpen semantic ones. The CrackFormer is trained and evaluated on three classical crack datasets. The experimental results show that the CrackFormer achieves the Optimal Dataset Scale (ODS) values of 0.871, 0.877 and 0.881, respectively, on the three datasets and outperforms the state-of-the-art methods. Huajun Liu, Xiangyu Miao, Christoph Mertz, Cheng-Zhong Xu 0001, Hui Kong 0001 |
ICCV | 3 |
| 2021 | Linear Inverse Problem for Depth Completion with RGB Image and Sparse LIDAR FusionabstractComprehensive depth information from surrounding scenes is important for perception in autonomous driving and robots. Sparse LIDAR sensors give a low-density point cloud of the environment, but are more affordable than their high-density counterparts. In this paper, we propose a novel sensor fusion architecture for sparse LIDAR depth completion. Instead of the traditional end-to-end neural network-based algorithm, we formulate depth completion as a Linear Inverse Problem (LIP) with a multi-modal proximal operator. This sensor fusion architecture allows a better signal prior and finds the unique optimal solution to the LIP. Instead of learning a unified network for the sparse input which treats pixels evenly, the proposed architecture guarantees both the data consistency and smoothness of the predicted depth map. To demonstrate the performance of our algorithm, we benchmark on the simulation dataset TartanAir, and the real indoor NYUdepthv2 and real outdoor KITTI datasets. Our proposed method outperforms previous methods and uses fewer parameters in both indoor and outdoor datasets. Christoph Mertz, John M. Dolan |
ICRA | 2 |
| 2020 | Depth Completion via Inductive Fusion of Planar LIDAR and Monocular CameraabstractModern high-definition LIDAR is expensive for commercial autonomous driving vehicles and small indoor robots. An affordable solution to this problem is fusion of planar LIDAR with RGB images to provide a similar level of perception capability. Even though state-of-the-art methods provide approaches to predict depth information from limited sensor input, they are usually a simple concatenation of sparse LIDAR features and dense RGB features through an end-to-end fusion architecture. In this paper, we introduce an inductive late-fusion block which better fuses different sensor modalities inspired by a probability model. The proposed demonstration and aggregation network propagates the mixed context and depth features to the prediction network and serves as a prior knowledge of the depth completion. This late-fusion block uses the dense context features to guide the depth prediction based on demonstrations by sparse depth features. In addition to evaluating the proposed method on benchmark depth completion datasets including NYUDepthV2 and KITTI, we also test the proposed method on a simulated planar LIDAR dataset. Our method shows promising results compared to previous approaches on both the benchmark datasets and simulated dataset with various 3D densities. Chiyu Dong, Christoph Mertz, John M. Dolan |
IROS | 3 |
| 2019 | DeepDA: LSTM-based Deep Data Association Network for Multi-Targets Tracking in Clutter
Huajun Liu, Christoph Mertz |
FUSION | 3 |
| 2018 | PCN: Point Completion NetworkabstractShape completion, the problem of estimating the complete geometry of objects from partial observations, lies at the core of many vision and robotics applications. In this work, we propose Point Completion Network (PCN), a novel learning-based approach for shape completion. Unlike existing shape completion methods, PCN directly operates on raw point clouds without any structural assumption (e.g. symmetry) or annotation (e.g. semantic class) about the underlying shape. It features a decoder design that enables the generation of fine-grained completions while maintaining a small number of parameters. Our experiments show that PCN produces dense, complete point clouds with realistic structures in the missing regions on inputs with various levels of incompleteness and noise, including cars from LiDAR scans in the KITTI dataset. Tejas Khot, David Held, Christoph Mertz, Martial Hebert |
3DV | 4 |
| 2014 | Visual sensing for developing autonomous behavior in snake robotsabstractSnake robots are uniquely qualified to investigate a large variety of settings including archaeological sites, natural disaster zones, and nuclear power plants. For these applications, modular snake robots have been tele-operated to perform specific tasks using images returned to it from an onboard camera in the robots head. In order to give the operator an even richer view of the environment and to enable the robot to perform autonomous tasks we developed a structured light sensor that can make three-dimensional maps of the environment. This paper presents a sensor that is uniquely qualified to meet the severe constraints in size, power and computational footprint of snake robots. Using range data, in the form of 3D pointclouds, we show that it is possible to pair high-level planning with mid-level control to accomplish complex tasks without operator intervention. Hugo Ponte, Max Queenan, Chaohui Gong, Christoph Mertz, Matthew J. Travers, Florian Enner, Martial Hebert, Howie Choset |
ICRA | 4 |
| 2014 | Vision for road inspectionabstractRoad surface inspection in cities is for the most part, a task performed manually. Being a subjective and labor intensive process, it is an ideal candidate for automation. We propose a solution based on computer vision and data-driven methods to detect distress on the road surface. Our method works on images collected from a camera mounted on the windshield of a vehicle. We use an automatic procedure to select images suitable for inspection based on lighting and weather conditions. From the selected data we segment the ground plane and use texture, color and location information to detect the presence of pavement distress. We describe an over-segmentation algorithm that identifies coherent image regions not just in terms of color, but also texture. We also discuss the problem of learning from unreliable human-annotations and propose using a weakly supervised learning algorithm (Multiple Instance Learning) to train a classifier. We present results from experiments comparing the performance of this approach against multiple individual human labelers, with the ground-truth labels obtained from an ensemble of other human labelers. Finally, we show results of pavement distress scores computed using our method over a subset of a citywide road network. Srivatsan Varadharajan, Sobhagya Jose, Karan Sharma, Lars Wander, Christoph Mertz |
WACV | 5 |
| 2009 | Planning-based prediction for pedestriansabstractWe present a novel approach for determining robot movements that efficiently accomplish the robot's tasks while not hindering the movements of people within the environment. Our approach models the goal-directed trajectories of pedestrians using maximum entropy inverse optimal control. The advantage of this modeling approach is the generality of its learned cost function to changes in the environment and to entirely different environments. We employ the predictions of this model of pedestrian trajectories in a novel incremental planner and quantitatively show the improvement in hindrance-sensitive robot trajectory planning provided by our approach. Brian D. Ziebart, Nathan D. Ratliff, Garratt Gallagher, Christoph Mertz, Kevin M. Peterson, J. Andrew Bagnell, Martial Hebert, Anind K. Dey, Siddhartha S. Srinivasa |
IROS | 4 |
| 2003 | Safe Robot Driving in Cluttered Environments
Charles E. Thorpe, Justin Carlson, David Duggins, Jay Gowdy, Robert A. MacLachlan, Christoph Mertz, Arne Suppé, Bob Wang |
ISRR | 6 |