Chengtang Yao

dblp:250/9040 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-5987-8510ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 Diving into the Fusion of Monocular Priors for Generalized Stereo Matching
abstract
The matching formulation makes it naturally hard for the stereo matching to handle ill-posed regions like occlusions and non-Lambertian surfaces. Fusing monocular priors has been proven helpful for ill-posed matching, but the biased monocular prior learned from small stereo datasets constrains the generalization. Recently, stereo matching has progressed by leveraging the unbiased monocular prior from the vision foundation model (VFM) to improve the generalization in ill-posed regions. We dive into the fusion process and observe three main problems limiting the fusion of the VFM monocular prior. The first problem is the misalignment between affine-invariant relative monocular depth and absolute depth of disparity. Besides, when we use the monocular feature in an iterative update structure, the over-confidence in the disparity update leads to local optima results. A direct fusion of a monocular depth map could alleviate the local optima problem, but noisy disparity results computed at the first several iterations will misguide the fusion. In this paper, we propose a binary local ordering map to guide the fusion, which converts the depth map into a binary relative format, unifying the relative and absolute depth representation. The computed local ordering map is also used to re-weight the initial disparity update, resolving the local optima and noisy problem. In addition, we formulate the final direct fusion of monocular depth to the disparity as a registration problem, where a pixel-wise linear regression module can globally and adaptively align them. Our method fully exploits the monocular prior to support stereo matching results effectively and efficiently. We significantly improve the performance from the experiments when generalizing from SceneFlow to Middlebury and Booster datasets while barely reducing the efficiency.
Chengtang Yao, Lidong Yu, Zhidan Liu 0005, Jiaxi Zeng, Yuwei Wu 0001, Yunde Jia
ICCV1
2025 3D Visual Illusion Depth Estimation
abstract
3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, making a flat artwork or object look three-dimensional in the human visual system. In this paper, we reveal that the machine visual system is also seriously fooled by 3D visual illusions, including monocular and binocular depth estimation. In order to explore and analyze the impact of 3D visual illusion on depth estimation, we collect a large dataset containing almost 3k scenes and 200k images to train and evaluate SOTA monocular and binocular depth estimation methods. We also propose a 3D visual illusion depth estimation framework that uses common sense from the vision language model to adaptively fuse depth from binocular disparity and monocular depth. Experiments show that SOTA monocular, binocular, and multi-view depth estimation approaches are all fooled by various 3D visual illusions, while our method achieves SOTA performance.
Chengtang Yao, Zhidan Liu 0005, Jiaxi Zeng, Lidong Yu, Yuwei Wu 0001, Yunde Jia
NeurIPS1
2025 Inter-Scale Similarity Guided Cost Aggregation for Stereo Matching
abstract
Stereo matching aims to estimate 3D geometry by computing disparity from a rectified image pair. Most deep learning based stereo matching methods aggregate multi-scale cost volumes computed by downsampling and achieve good performance. However, their effectiveness in fine-grained areas is limited by significant detail loss during downsampling and the use of fixed weights in upsampling. In this paper, we propose an inter-scale similarity-guided cost aggregation method that dynamically upsamples the cost volumes according to the content of images for stereo matching. The method consists of two modules: inter-scale similarity measurement and stereo-content-aware cost aggregation. Specifically, we use inter-scale similarity measurement to generate similarity guidance from feature maps in adjacent scales. The guidance, generated from both reference and target images, is then used to aggregate the cost volumes from low-resolution to high-resolution via stereo-content-aware cost aggregation. We further split the 3D aggregation into 1D disparity and 2D spatial aggregation to reduce the computational cost. Experimental results on various benchmarks (e.g., SceneFlow, KITTI, Middlebury and ETH3D-two-view) show that our method achieves consistent performance gain on multiple models (e.g., PSM-Net, HSM-Net, CF-Net, FastAcv, and FactAcvPlus). The code can be found athttps://github.com/Pengxiang-Li/issga-stereo.
Pengxiang Li 0002, Chengtang Yao, Yunde Jia, Yuwei Wu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Temporally Consistent Stereo Matching
Jiaxi Zeng, Chengtang Yao, Yuwei Wu 0001, Yunde Jia
ECCV (31)2
2023 Sparse Point Guided 3D Lane Detection
abstract
3D lane detection usually builds a dense correspondence between the front-view space and the BEV space to estimate lane points in the 3D space. 3D lanes only occupy a small ratio of the dense correspondence, while most correspondence belongs to the redundant background. This sparsity phenomenon bottlenecks valuable computation and raises the computation cost of building a high-resolution correspondence for accurate results. In this paper, we propose a sparse point-guided 3D lane detection, focusing on points related to 3D lanes. Our method runs in a coarse-to-fine manner, including coarse-level lane detection and iterative fine-level sparse point refinements. In coarse-level lane detection, we build a dense but efficient correspondence between the front view and BEV space at a very low resolution to compute coarse lanes. Then in fine-level sparse point refinement, we sample sparse points around coarse lanes to extract local features from the high-resolution front-view feature map. The high-resolution local information brought by sparse points refines 3D lanes in the BEV space hierarchically from low resolution to high resolution. The sparse point guides a more effective information flow and greatly promotes the SOTA result by 3 points on the overall F1-score and 6 points on several hard situations while reducing almost half memory cost and speeding up 2 times.
Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia
ICCV1
2023 Parameterized Cost Volume for Stereo Matching
abstract
Stereo matching becomes computationally challenging when dealing with a large disparity range. Prior methods mainly alleviate the computation through dynamic cost volume by focusing on a local disparity space, but it requires many iterations to get close to the ground truth due to the lack of a global view. We find that the dynamic cost volume approximately encodes the disparity space as a single Gaussian distribution with a fixed and small variance at each iteration, which results in an inadequate global view over disparity space and a small update step at every iteration. In this paper, we propose a parameterized cost volume to encode the entire disparity space using multi-Gaussian distribution. The disparity distribution of each pixel is parameterized by weights, means, and variances. The means and variances are used to sample disparity candidates for cost computation, while the weights and means are used to calculate the disparity output. The above parameters are computed through a JS-divergence-based optimization, which is realized as a gradient descent update in a feed-forward differential module. Experiments show that our method speeds up the runtime of RAFT-Stereo by 4 ~ 15 times, achieving real-time performance and comparable accuracy. The code is available at https://github.com/jiaxiZeng/Parameterized-Cost-Volume-for-Stereo-Matching.
Jiaxi Zeng, Chengtang Yao, Lidong Yu, Yuwei Wu 0001, Yunde Jia
ICCV2
2022 FoggyStereo: Stereo Matching with Fog Volume Representation
abstract
Stereo matching in foggy scenes is challenging as the scattering effect of fog blurs the image and makes the matching ambiguous. Prior methods deem the fog as noise and discard it before matching. Different from them, we propose to explore depth hints from fog and improve stereo matching via these hints. The exploration of depth hints is designed from the perspective of rendering. The rendering is conducted by reversing the atmospheric scattering process and removing the fog within a selected depth range. The quality of the rendered image reflects the correctness of the selected depth, as the closer it is to the real depth, the clearer the rendered image is. We introduce a fog volume representation to collect these depth hints from the fog. We construct the fog volume by stacking images rendered with depths computed from disparity candidates that are also used to build the cost volume. We fuse the fog volume with cost volume to rectify the ambiguous matching caused by fog. Experiments show that our fog volume representation significantly promotes the SOTA result on foggy scenes by 10% ~ 30% while maintaining a comparable performance in clear scenes.
Chengtang Yao, Lidong Yu
CVPR1
2021 A Decomposition Model for Stereo Matching
abstract
In this paper, we present a decomposition model for stereo matching to solve the problem of excessive growth in computational cost (time and memory cost) as the resolution increases. In order to reduce the huge cost of stereo matching at the original resolution, our model only runs dense matching at a very low resolution and uses sparse matching at different higher resolutions to recover the disparity of lost details scale-by-scale. After the decomposition of stereo matching, our model iteratively fuses the sparse and dense disparity maps from adjacent scales with an occlusion-aware mask. A refinement network is also applied to improving the fusion result. Compared with high-performance methods like PSMNet and GANet, our method achieves 10−100× speed increase while obtaining comparable disparity estimation results.
Chengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li 0002, Yuwei Wu 0001
CVPR1
2020 Face Spoofing Detection Using Relativity Representation on Riemannian Manifold
abstract
Face recognition and verification systems are susceptible to spoofing attacks using photographs, videos or masks. Most existing methods focus on spoofing detection in Euclidean space, and ignore the features' manifold structure and interrelationships, thus limiting their capabilities of discrimination and generalization. In this paper, we propose a relativity representation on Riemannian manifold for face spoofing detection. The relativity representation improves generalization capability while ensuring discriminability, at both levels of feature description and classification score. The feature-level relativity representation generalizes information by modeling interrelationships among basic features, and would not depend too much on characteristics of a particular dataset. The score-level relativity representation makes decisions relatively, not absolutely, according to interrelationships (via Riemannian metric) and competitions (via example reweighting) among data samples on Riemannian manifold. The discriminability is ensured by the high-order nature of the feature-level relativity representation as well as Riemannian reweighted discriminative learning of the score-level relativity representation. Moreover, we integrate an attack-sensitive SVM classifier in Euclidean space to improve spoofing detection. Experiments demonstrate the effectiveness of our method on both intra-dataset and cross-dataset testing.
Chengtang Yao, Yunde Jia, Huijun Di, Yuwei Wu 0001
IEEE Trans. Inf. Forensics Secur.1
2019 Diffusion-based kernel matrix model for face liveness detection
Changyong Yu, Chengtang Yao, Mingtao Pei, Yunde Jia
Image Vis. Comput.2