Wenbing Tao

dblp:73/188 · DBLP profile ↗
← Back
107ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0003-3284-864XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 66 · 5 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 5 first-author · 26 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Unsupervised asymmetry-agnostic stereo matching with selective adaptive patch correlation
Ximeng Li 0007, Wenbing Tao
Expert Syst. Appl.3
2026 Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy
Yu-Xin Zhang 0004, Jie Gui, Baosheng Yu, Xiaofeng Cong, Xin Gong 0001, Wenbing Tao, Dacheng Tao
Int. J. Comput. Vis.6
2026 SparseLiSplat: LiDAR meets neural Gaussian splatting for novel view synthesis from sparse input
Chen Zhang 0043, Mingyang Liang, Xu Ge, Jixiang Ma, Wenbing Tao
Neurocomputing8
2026 MGC-Net: Learning feature matching with multi-geometry cooperation
Luxia Ai, Kun Sun 0002, Chen Zhang 0043, Nanjun Yuan, Qun Jiang, Wenbing Tao
Knowl. Based Syst.6
2026 OIF-PCR++: Point Cloud Registration via Progressive Distillation of Conditional Positional Encoding
abstract
Transformer architecture has shown significant potential in various visual tasks, including point cloud registration. Positional encoding, as an order-aware module, plays a crucial role in Transformer framework. In this paper, we propose OIF-PCR++, a conditional positional encoding (CPE) method for point cloud registration. The core CPE module utilizes length and vector encoding at different stages, conditioned on the relative pose states between the point clouds to be registered. As a result, it progressively alleviates feature ambiguity through the incorporation of geometric cues. Building upon CPE, we introduce an iterative positional encoding optimization pipeline comprising two stages: 1) We find one correspondence via a differentiable optimal transport layer, and use it to encode length information into point cloud features, enhancing spatial consistency across different reference frames. 2) We apply a progressive direction alignment strategy to achieve rough alignment between paired point clouds, and then gradually incorporate direction information with the aid of this alignment, further enhancing feature distinctiveness and reducing feature ambiguity. Through this iterative optimization process, length and direction information are effectively integrated to achieve consistent and distinctive positional encoding, enabling the learning of discriminative point cloud features. Additionally, we present an inlier propagation mechanism that harmoniously integrates consistent geometric information for positional encoding. The proposed method is highly efficient, introducing marginal computational overhead while significantly improving feature distinguishability. Extensive experiments demonstrate superior performance over state-of-the-art methods on indoor, outdoor, object-level, and multi-way benchmarks, as well as strong generalization to complex real-world scenarios.
Fan Yang 0088, Zhi Chen 0011, Nanjun Yuan, Lin Guo 0019, Wenbing Tao
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 One-Stage Absolute Human Mesh Recovery
abstract
The reconstruction of realistic and precise human meshes in world coordinates is facilitated by considering scene information. Challenges related to accuracy, robustness, and computation time are faced by existing absolute human mesh recovery methods. In this paper, a one-stage model for absolute human mesh recovery with superior reconstruction precision and inference speed is presented. The proposed one-stage model is composed of two parallel branches to achieve root position estimation and human mesh regression. To effectively connect the two branches, a scene-image information aggregation module is designed. The accuracy of the estimated human meshes is improved and the end-to-end training of the whole model is facilitated by this module. Experiments are conducted on three diverse datasets, and a GMPJPE decrease of 72.3 mm/27.32% and an MPJPE reduction of 25.6 mm/27.26% are achieved by the proposed method with the lowest inference time compared to previous SOTA methods.
Xinyao Liao, Wanjuan Su, Chen Zhang 0043, Ximeng Li 0007, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.5
2025 Cross-View Referring Multi-Object Tracking
abstract
Referring Multi-Object Tracking (RMOT) is an important topic in the current tracking field. Its task form is to guide the tracker to track objects that match the language description. Current research mainly focuses on referring multi-object tracking under single-view, which refers to a view sequence or multiple unrelated view sequences. However, in the single-view, some appearances of objects are easily invisible, resulting in incorrect matching of objects with the language description. In this work, we propose a new task, called Cross-view Referring Multi-Object Tracking (CRMOT). It introduces the cross-view to obtain the appearances of objects from multiple views, avoiding the problem of the invisible appearances of objects in RMOT task. CRMOT is a more challenging task of accurately tracking the objects that match the language description and maintaining the identity consistency of objects in each cross-view. To advance CRMOT task, we construct a cross-view referring multi-object tracking benchmark based on CAMPUS and DIVOTrack datasets, named CRTrack. Specifically, it provides 13 different scenes and 221 language descriptions. Furthermore, we propose an end-to-end cross-view referring multi-object tracking method, named CRTracker. Extensive experiments on the CRTrack benchmark verify the effectiveness of our method.
En Yu, Wenbing Tao
AAAI3
2025 High-Fidelity Lightweight Mesh Reconstruction from Point Clouds
abstract
Recently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To solve the problem, we propose an adaptive meshing method to extract resolution-adaptive meshes based on surface curvature, enabling the recovery of high-fidelity lightweight meshes. Specifically, we first use point-based representation to perceive implicit surfaces and calculate surface curvature. A vertex generator is designed to produce curvature-adaptive vertices with any specified number on the implicit surface, preserving the overall structure and high-curvature features. Then we develop a Delaunay meshing algorithm to generate meshes from vertices, ensuring geometric fidelity and correct topology. In addition, to obtain accurate SDFs for adaptive meshing and achieve better lightweight reconstruction, we design a hybrid representation combining feature grid and feature tri-plane for better detail capture. Experiments demonstrate that our method can generate high-quality lightweight meshes from point clouds. Compared with methods from various categories, our approach achieves superior results, especially in capturing more details with fewer elements.
Chen Zhang 0043, Ximeng Li 0007, Xinyao Liao, Wanjuan Su, Wenbing Tao
CVPR6
2025 Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
Enyu Liu, En Yu, Wenbing Tao
ICCV4
2025 OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer
abstract
Open-vocabulary multiple object tracking aims to generalize trackers to unseen categories during training, enabling their application across a variety of real-world scenarios. However, the existing open-vocabulary tracker is constrained by its framework structure, isolated frame-level perception, and insufficient modal interactions, which hinder its performance in open-vocabulary classification and tracking. In this paper, we propose OVTR (End-to-End Open-Vocabulary Multiple Object Tracking with TRansformer), the first end-to-end open-vocabulary tracker that models motion, appearance, and category simultaneously. To achieve stable classification and continuous tracking, we design the CIP (Category Information Propagation) strategy, which establishes multiple high-level category information priors for subsequent frames. Additionally, we introduce a dual-branch structure for generalization capability and deep multimodal interaction, and incorporate protective strategies in the decoder to enhance performance. Experimental results show that our method surpasses previous trackers on the open-vocabulary MOT benchmark while also achieving faster inference speeds and significantly reducing preprocessing requirements. Moreover, the experiment transferring the model to another dataset demonstrates its strong adaptability.
En Yu, Wenbing Tao
ICLR4
2025 Unhackable Temporal Reward for Scalable Video MLLMs
abstract
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the “anti-scaling law”, where more data and larger models lead to worse performance. This study unmasks the culprit: “temporal hacking”, a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.
En Yu, Kangheng Lin, Yana Wei, Zining Zhu 0004, Jianjian Sun, Zheng Ge, Xiangyu Zhang 0005, Jingyu Wang 0001, Wenbing Tao
ICLR11
2025 FIELD: Fast Information-driven Autonomous Exploration using Larger Perception Distance
abstract
Autonomous exploration is a critical challenge for various unmanned aerial vehicle (UAV) applications. Existing methods often suffer from low exploration rates due to limitations such as inefficient global coverage and inadequate sensor data utilization. In this paper, we introduce FIELD, a Fast Information-driven aerial robot Exploration planner using Larger perception Distance. FIELD leverages a larger perception distance to identify high-information-gain viewpoints while maintaining mapping precision and utilizing more sensor data to guide the exploration process. Then, the method incorporates a history-aware coverage path to determine a consistent and reasonable sequence for visiting frontier viewpoints. Local viewpoints are refined to find the optimal combination of these viewpoints. We compare our method with state-of-the-art frontier-based approaches in benchmark environments. Our method significantly improves exploration efficiency by 13% to 17%.
Yuefeng Zhang, Fan Yang 0088, Nanjun Yuan, Wenbing Tao
IROS4
2025 Perception-R1: Pioneering Perception Policy with Reinforcement Learning
abstract
Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual perplexity is a major factor in determining the effectiveness of RL. We also observe that reward design plays a crucial role in further approaching the upper limit of model perception. To leverage these findings, we propose Perception-R1, a scalable RL framework using GRPO during MLLM post-training. With a standard Qwen2-VL-2B-Instruct, Perception-R1 achieves +4.2% on RefCOCO+, +17.9% on PixMo-Count, +4.2% on PageOCR, and notably, 31.9% AP on COCO2017 val for the first time, establishing a strong baseline for perception policy learning.
En Yu, Kangheng Lin, Jisheng Yin, Yana Wei, Yuang Peng, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Jingyu Wang 0001, Wenbing Tao
NeurIPS14
2025 Context-Aware Multi-view Stereo Network for Efficient Edge-Preserving Depth Estimation
Wanjuan Su, Wenbing Tao
Int. J. Comput. Vis.2
2025 Learning Meshing from Delaunay Triangulation for 3D Shape Representation
Chen Zhang 0043, Wenbing Tao
Int. J. Comput. Vis.2
2025 A hybrid algorithm with inlier-guided Hough voting for point cloud registration
Qun Jiang, Zhi Chen 0011, Fan Yang 0088, Lin Guo 0019, Luxia Ai, Wenbing Tao
Neurocomputing6
2025 Learning hierarchical image feature for efficient image rectification
Nanjun Yuan, Fan Yang 0088, Yuefeng Zhang, Luxia Ai, Wenbing Tao
Neurocomputing5
2025 FE-GS: 3D feature-embedded Gaussian splatting with geometric regularizations for high-fidelity rendering
Yining Peng, Chen Zhang 0043, Wanjuan Su, Wenbing Tao
Knowl. Based Syst.4
2025 InstaHMR: Instance-Aware One-Stage Multi-Person Human Mesh Recovery
abstract
Human mesh recovery aims to estimate all human meshes within a given image. In this article, we propose an Instance-aware Multi-person 3D Human Mesh Recovery (InstaHMR) network based on the one-stage framework. Compared to former one-stage methods, instance-aware single person feature is exploited to represent more accurate human mesh. Specifically, we propose the Contextual Instance Guidance (CIG) module which generates instance-aware single person feature by leveraging spatial and channel attention operations. In this way, it preserves more instance-specific information compared to the pixel-level feature used in some existing one-stage methods. Besides, we further introduce two auxiliary losses for better mesh recovery, namely the Human Triplet Planes (HTP) loss and the T-pose Shape (TS) loss. The HTP loss encourages the model to capture subtle differences in human joint positions, while the TS loss facilitates the learning of abstract shape parameters. By incorporating these advancements, our model achieves state-of-the-art results on four multi-person datasets.
Xinyao Liao, Chen Zhang 0043, Jianyao Xu, Wanjuan Su, Zhi Chen 0011, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.6
2025 PSDF: Prior-Driven Neural Implicit Surface Learning for Multi-View Reconstruction
abstract
Surface reconstruction has traditionally relied on the Multi-View Stereo (MVS)-based pipeline, which often suffers from noisy and incomplete geometry. This is due to that although MVS has been proven to be an effective way to recover the geometry of the scenes, especially for locally detailed areas with rich textures, it struggles to deal with areas with low texture and large variations of illumination where the photometric consistency is unreliable. Recently, Neural Implicit Surface Reconstruction (NISR) combines surface rendering and volume rendering techniques and bypasses the MVS as an intermediate step, which has emerged as a promising alternative to overcome the limitations of traditional pipelines. While NISR has shown impressive results on simple scenes, it remains challenging to recover delicate geometry from uncontrolled real-world scenes which is caused by its underconstrained optimization. To this end, the framework PSDF is proposed which resorts to external geometric priors from a pretrained MVS network and internal geometric priors inherent in the NISR model to facilitate high-quality neural implicit surface learning. Specifically, the visibility-aware feature consistency loss and depth prior-assisted sampling based on external geometric priors are introduced. These proposals provide powerfully geometric consistency constraints and aid in locating surface intersection points, thereby significantly improving the accuracy and delicate reconstruction of NISR. Meanwhile, the internal prior-guided importance rendering is presented to enhance the fidelity of the reconstructed surface mesh by mitigating the biased rendering issue in NISR. Extensive experiments on Tanks and Temples datasets show that PSDF achieves state-of-the-art performance on complex uncontrolled scenes.
Wanjuan Su, Chen Zhang 0043, Qingshan Xu 0001, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.4
2025 PG-NeuS: Robust and Efficient Point Guidance for Multi-View Neural Surface Reconstruction
abstract
Recently, learning multi-view neural surface reconstruction with the supervision of point clouds or depth maps has been a promising way. However, due to weak perception and underutilization of prior information, current methods still struggle with the challenges of limited accuracy and excessive time complexity. In addition, prior data perturbation is also an important yet rarely considered issue, often resulting in distorted geometry. To address these challenges, we propose a novel point-guided method named PG-NeuS, which achieves accurate and efficient reconstruction while robustly coping with point noise. Specifically, the aleatoric uncertainty of the point cloud is modeled to capture the noise distribution, estimating the reliability of each point and enhancing robustness against noise. Moreover, a Neural Projection module is proposed to connect points and images, adding geometric constraints to the implicit surface and achieving more precise point guidance. To better compensate for geometric bias between volume rendering and point modeling, we additionally design a Bias network that leverages the geometric information in high-fidelity points to enhance detail representation. Benefiting from the effective point guidance, the proposed PG-NeuS achieves an 11x speed increase and a 33.3% accuracy improvement compared to NeuS on DTU, even with a lightweight network. Extensive experiments show that our method yields high-quality surfaces with high efficiency, especially for fine-grained details and smooth regions, outperforming the state-of-the-art methods. Moreover, it exhibits strong robustness to noisy data and sparse data.
Chen Zhang 0043, Wanjuan Su, Qingshan Xu 0001, Xinyao Liao, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.5
2024 IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo Matching
abstract
Recently, there has been a growing interest in 3D CNN-based stereo matching methods due to their remarkable accuracy. However, the high complexity of 3D convolution makes it challenging to strike a balance between accuracy and speed. Notably, explicit 3D volumes contain considerable redundancy. In this study, we delve into more compact 2D implicit network to eliminate redundancy and boost real-time performance. However, simply replacing explicit 3D networks with 2D implicit networks causes issues that can lead to performance degradation, including the loss of structural information, the quality decline of inter-image information, as well as the inaccurate regression caused by low-level features. To address these issues, we first integrate intra-image information to fuse with inter-image information, facilitating propagation guided by structural cues. Subsequently, we introduce the Fast Multi-scale Score Volume (FMSV) and Confidence Based Filtering (CBF) to efficiently acquire accurate multi-scale, noise-free inter-image information. Furthermore, combined with the Residual Context-aware Upsampler (RCU), our Intra-Inter Fusing network is meticulously designed to enhance information transmission on both feature-level and disparity-level, thereby enabling accurate and robust regression. Experimental results affirm the superiority of our network in terms of both speed and accuracy compared to all other fast methods.
Ximeng Li 0007, Chen Zhang 0043, Wanjuan Su, Wenbing Tao
AAAI4
2024 Delving into the Trajectory Long-tail Distribution for Muti-object Tracking
abstract
Multiple Object Tracking (MOT) is a critical area within computer vision, with a broad spectrum of practical im-plementations. Current research has primarily focused on the development of tracking algorithms and enhancement of post-processing techniques. Yet, there has been a lack of thorough examination concerning the nature of tracking data it self. In this study, we pioneer an exploration into the distribution patterns of tracking data and iden-tify a pronounced long-tail distribution issue within existing MOT datasets. We note a significant imbalance in the distribution of trajectory lengths across different pedestri-ans, a phenomenon we refer to as “pedestrians trajectory long-tail distribution”. Addressing this challenge, we intro-duce a bespoke strategy designed to mitigate the effects of this skewed distribution. Specifically, we propose two data augmentation strategies, including Stationary Camera View Data Augmentation (SVA) and Dynamic Camera View Data Augmentation (DVA), designed for viewpoint states and the Group Softmax (GS) module for Re-ID. SVA is to backtrack and predict the pedestrian trajectory of tail classes, and DVA is to use diffusion model to change the background of the scene. GS divides the pedestrians into unrelated groups and performs softmax operation on each group individually. Our proposed strategies can be integrated into numerous existing tracking systems, and extensive experimentation validates the efficacy of our method in reducing the influ-ence of long-tail distribution on multi-object tracking per-formance. The code is available at https://github.com/chen-si-jia/Trajectory-Long-tail-Distribution-for-MOT.
En Yu, Wenbing Tao
CVPR4
2024 Few-Shot NeRF by Adaptive Rendering Loss Regularization
Qingshan Xu 0001, Xuanyu Yi, Jianyao Xu, Wenbing Tao, Yew-Soon Ong, Hanwang Zhang
ECCV (66)4
2024 Merlin: Empowering Multimodal LLMs with Foresight Minds
En Yu, Yana Wei, Dongming Wu 0005, Lingyu Kong, Tiancai Wang, Zheng Ge, Xiangyu Zhang 0005, Wenbing Tao
ECCV (4)11
2024 A Comprehensive Survey and Taxonomy on Point Cloud Registration Based on Deep Learning
Yu-Xin Zhang 0004, Jie Gui, Xiaofeng Cong, Xin Gong 0001, Wenbing Tao
IJCAI5
2024 QTrack: Embracing Quality Clues for Robust 3D Multi-Object Tracking
abstract
3D Multi-Object Tracking (MOT) has achieved tremendous achievement thanks to the rapid development of 3D object detection and 2D MOT. Recent advanced works generally employ a series of object attributes, e.g., position, size, velocity, and appearance, to provide the clues for the association in 3D MOT. However, these cues may not be reliable due to some visual noise, such as occlusion and blur, leading to tracking performance bottlenecks. To reveal the dilemma, we conduct extensive empirical analysis to expose the key bottleneck of each clue and how they correlate with each other. The analysis results motivate us to efficiently absorb the merits among all cues and adaptively produce an optimal tracking manner. Specifically, we present Location and Velocity Quality Learning, which efficiently guides the network to estimate the quality of predicted object attributes. Based on these quality estimations, we propose a quality-aware object association (QOA) strategy to leverage the quality score as an important reference factor for achieving robust association. Despite its simplicity, extensive experiments indicate that the proposed strategy significantly boosts tracking performance by 2.2% AMOTA and our method outperforms all existing state-of-the-art works on nuScenes by a large margin. Moreover, QTrack achieves 51.1%, 54.8% and 56.6% AMOTA tracking performance on the nuScenes test sets with BEVDepth, VideoBEV, and StreamPETR models respectively, which significantly reduces the performance gap between the pure camera and LiDAR-based trackers.
En Yu, Xiaoping Li 0005, Wenbing Tao
IROS5
2024 Learning compact and overlap-biased interactions for point cloud registration
Lin Guo 0019, Zhi Chen 0011, Senmao Cheng, Fan Yang 0088, Wenbing Tao
Neurocomputing5
2024 DualGroup for 3D instance and panoptic segmentation
Lin Zhao 0012, Wenbing Tao
Pattern Recognit. Lett.4
2024 LIF-Seg: LiDAR and Camera Image Fusion for 3D LiDAR Semantic Segmentation
abstract
Camera and 3D LiDAR sensors have become indispensable devices in modern autonomous driving vehicles. Camera provides fine-grained texture and color information in 2D space, while LiDAR captures more precise and farther-away distance measurements of the surrounding environments. The complementary information from these two sensors makes the fusion of two modalities a desired option. However, two primary challenges in the fusion of camera and LiDAR hinder its performance, i.e., how to effectively fuse the information from these two modalities and how to precisely align them (suffering from the weak spatiotemporal synchronization problem). This article proposes a coarse-to-fine LiDAR and camera fusion-based network, named LIF-Seg, for LiDAR segmentation. For the first challenge, unlike these previous works fusing the point cloud and image information in a one-to-one manner, the proposed method introduces a simple but effective early-fusion strategy to fully utilize the contextual information of images. Second, to tackle the weak spatiotemporal synchronization problem, an offset rectification approach is designed to align the features of the two modalities. The cooperation of these two components leads to the success of the effective camera-LiDAR fusion. Experimental results on the nuScenes dataset show the superiority of LIF-Seg over existing methods by a large margin. Ablation studies and analyses further illustrate that the LIF-Seg can effectively address the weak spatiotemporal synchronization problem.
Lin Zhao 0012, Hui Zhou 0005, Xinge Zhu, Xiao Song 0002, Hongsheng Li 0001, Wenbing Tao
IEEE Trans. Multim.6
2023 Efficient Edge-Preserving Multi-View Stereo Network for Depth Estimation
abstract
Over the years, learning-based multi-view stereo methods have achieved great success based on their coarse-to-fine depth estimation frameworks. However, 3D CNN-based cost volume regularization inevitably leads to over-smoothing problems at object boundaries due to its smooth properties. Moreover, discrete and sparse depth hypothesis sampling exacerbates the difficulty in recovering the depth of thin structures and object boundaries. To this end, we present an Efficient edge-Preserving multi-view stereo Network (EPNet) for practical depth estimation. To keep delicate estimation at details, a Hierarchical Edge-Preserving Residual learning (HEPR) module is proposed to progressively rectify the upsampling errors and help refine multi-scale depth estimation. After that, a Cross-view Photometric Consistency (CPC) is proposed to enhance the gradient flow for detailed structures, which further boosts the estimation accuracy. Last, we design a lightweight cascade framework and inject the above two strategies into it to achieve better efficiency and performance trade-offs. Extensive experiments show that our method achieves state-of-the-art performance with fast inference speed and low memory usage. Notably, our method tops the first place on challenging Tanks and Temples advanced dataset and ETH3D high-res benchmark among all published learning-based methods. Code will be available at https://github.com/susuwj/EPNet.
Wanjuan Su, Wenbing Tao
AAAI2
2023 Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language Representation
abstract
Although existing multi-object tracking (MOT) algorithms have obtained competitive performance on various benchmarks, almost all of them train and validate models on the same domain. The domain generalization problem of MOT is hardly studied. To bridge this gap, we first draw the observation that the high-level information contained in natural language is domain invariant to different tracking domains. Based on this observation, we propose to introduce natural language representation into visual MOT models for boosting the domain generalization ability. However, it is infeasible to label every tracking target with a textual description. To tackle this problem, we design two modules, namely visual context prompting (VCP) and visual-language mixing (VLM). Specifically, VCP generates visual prompts based on the input frames. VLM joints the information in the generated visual prompts and the textual prompts from a pre-defined Trackbook to obtain instance-level pseudo textual description, which is domain invariant to different tracking scenes. Through training models on MOT17 and validating them on MOT20, we observe that the pseudo textual descriptions generated by our proposed modules improve the generalization performance of query-based trackers by large margins.
En Yu, Zhuoling Li, Shoudong Han, Wenbing Tao
AAAI7
2023 DMNet: Delaunay Meshing Network for 3D Shape Representation
abstract
Recently, there has been a growing interest in learningbased explicit methods due to their ability to respect the original input and preserve details. However, the connectivity on complex structures is still difficult to infer due to the limited local shape perception, resulting in artifacts and non-watertight triangles. In this paper, we present a novel learning-based method with Delaunay triangulation to achieve high-precision reconstruction. We model the Delaunay triangulation as a dual graph, extract local geometric information from the points, and embed it into the structural representation of Delaunay triangulation in an organic way, benefiting fine-grained details reconstruction. To encourage neighborhood information interaction of edges and nodes in the graph, we introduce a local graph iteration algorithm, which is a variant of graph neural network. Moreover, a geometric constraint loss further improves the classification of tetrahedrons. Benefiting from our fully local network, a scaling strategy is designed to enable large-scale reconstruction. Experiments show that our method yields watertight and high-quality meshes. Especially for some thin structures and sharp edges, our method shows better performance than the current state-of-the-art methods. Furthermore, it has a strong adaptability to point clouds of different densities.
Chen Zhang 0043, Ganzhangqin Yuan, Wenbing Tao
ICCV3
2023 3D hand pose and shape estimation from monocular RGB via efficient 2D cues
abstract
Estimating 3D hand shape from a single-view RGB image is important for many applications. However, the diversity of hand shapes and postures, depth ambiguity, and occlusion may result in pose errors and noisy hand meshes. Making full use of 2D cues such as 2D pose can effectively improve the quality of 3D human hand shape estimation. In this paper, we use 2D joint heatmaps to obtain spatial details for robust pose estimation. We also introduce a depth-independent 2D mesh to avoid depth ambiguity in mesh regression for efficient hand-image alignment. Our method has four cascaded stages: 2D cue extraction, pose feature encoding, initial reconstruction, and reconstruction refinement. Specifically, we first encode the image to determine semantic features during 2D cue extraction; this is also used to predict hand joints and for segmentation. Then, during the pose feature encoding stage, we use a hand joints encoder to learn spatial information from the joint heatmaps. Next, a coarse 3D hand mesh and 2D mesh are obtained in the initial reconstruction step; a mesh squeeze-and-excitation block is used to fuse different hand features to enhance perception of 3D hand structures. Finally, a global mesh refinement stage learns non-local relations between vertices of the hand mesh from the predicted 2D mesh, to predict an offset hand mesh to fine-tune the reconstruction results. Quantitative and qualitative results on the FreiHAND benchmark dataset demonstrate that our approach achieves state-of-the-art performance.
Fenghao Zhang, Lin Zhao 0012, Shengling Li, Wanjuan Su, Liman Liu, Wenbing Tao
Comput. Vis. Media6
2023 Edge-Aware Spatial Propagation Network for Multi-view Depth Estimation
Qingshan Xu 0001, Wanjuan Su, Wenbing Tao
Neural Process. Lett.4
2023 SC$^{2}$2-PCR++: Rethinking the Generation and Selection for Efficient and Robust Point Cloud Registration
abstract
Outlier removal is a critical part of feature-based point cloud registration. In this paper, we revisit the model generation and selection of the classic RANSAC approach for fast and robust point cloud registration. For the model generation, we propose a second-order spatial compatibility (SC$^{2}$) measure to compute the similarity between correspondences. It takes into account global compatibility instead of local consistency, allowing for more distinctive clustering between inliers and outliers at an early stage. The proposed measure promises to find a certain number of outlier-free consensus sets using fewer samplings, making the model generation more efficient. For the model selection, we propose a new Feature and Spatial consistency constrained Truncated Chamfer Distance (FS-TCD) metric for evaluating the generated models. It considers the alignment quality, the feature matching properness, and the spatial consistency constraint simultaneously, enabling the correct model to be selected even when the inlier rate of the putative correspondence set is extremely low. Extensive experiments are carried out to investigate the performance of our method. In addition, we also experimentally prove that the proposed SC$^{2}$measure and the FS-TCD metric are general and can be easily plugged into deep learning based frameworks. The code will be available athttps://github.com/ZhiChen902/SC2-PCR-plusplus.
Zhi Chen 0011, Kun Sun 0002, Fan Yang 0088, Lin Guo 0019, Wenbing Tao
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Multi-Scale Geometric Consistency Guided and Planar Prior Assisted Multi-View Stereo
abstract
In this paper, we propose some efficient multi-view stereo methods for accurate and complete depth map estimation. We first present our basic methods with Adaptive Checkerboard sampling and Multi-Hypothesis joint view selection (ACMH & ACMH+). Based on our basic models, we develop two frameworks to deal with the depth estimation of ambiguous regions (especially low-textured areas) from two different perspectives: multi-scale information fusion and planar geometric clue assistance. For the former one, we propose a multi-scale geometric consistency guidance framework (ACMM) to obtain the reliable depth estimates for low-textured areas at coarser scales and guarantee that they can be propagated to finer scales. For the latter one, we propose a planar prior assisted framework (ACMP). We utilize a probabilistic graphical model to contribute a novel multi-view aggregated matching cost. At last, by taking advantage of the above frameworks, we further design a multi-scale geometric consistency guided and planar prior assisted multi-view stereo (ACMMP). This greatly enhances the discrimination of ambiguous regions and helps their depth sensing. Experiments on extensive datasets show our methods achieve state-of-the-art performance, recovering the depth estimation not only in low-textured areas but also in details. Related codes are available at https://github.com/GhiXu.
Qingshan Xu 0001, Weihang Kong, Wenbing Tao, Marc Pollefeys
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 JSNet++: Dynamic Filters and Pointwise Correlation for 3D Point Cloud Instance and Semantic Segmentation
abstract
In this paper, we propose a novel joint instance and semantic segmentation approach, called JSNet++, to address the instance and semantic segmentation tasks of 3D point clouds simultaneously. We first introduce a basic joint segmentation framework (JSNet). It fuses features from different layers of the backbone network to obtain more discriminative features and makes the two tasks take advantage of each other with a joint instance and semantic segmentation (JISS) module. Specifically, the JISS transforms semantic features into instance embedding space, and then the transformed features are fused with instance features to facilitate instance segmentation. Meanwhile, the JISS module also makes semantic segmentation benefit from instance segmentation by aggregating instance features to semantic feature space. To further reduce the memory consumption of JSNet, we design a dynamic filters for convolution (DFConv) on point clouds. Specifically, we exploit the geometry and density information to generate the dynamic filters, which are used to perform depthwise convolution with the input features. Afterwards, we unify the spatial correlation and channel correlation into a module to fully explore the pointwise correlation in point clouds, and we develop an improved JISS module (JISS*) by using the pointwise correlation module to further improve the accuracy of segmentation. Finally, based on the JSNet, DFConv and JISS*, we propose a new joint segmentation network, termed JSNet++. Experimental results on the benchmarks S3DIS and ScanNet v2 datasets demonstrate the effectiveness of our approach, and our method achieves significant performance improvements over baseline on both instance and semantic segmentation.
Lin Zhao 0012, Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.2
2023 SSRNet: Scalable 3D Surface Reconstruction Network
abstract
Learning-based surface reconstruction methods have received considerable attention in recent years due to their excellent expressiveness. However, existing learning-based methods lack scalability in processing large-scale point clouds. This paper proposes a novel scalable learning-based 3D surface reconstruction method based on octree, called SSRNet. SSRNet works in a scalable reconstruction pipeline, which divides oriented point clouds into different local parts and then processes them in parallel. Accommodating this scalable design pattern, SSRNet constructs local geometric features for octree vertices. Such features comprise the relation between the vertices and the implicit surface, ensuring geometric perception. Focusing on local geometric information also enables the network to avoid the overfitting problem and generalize well on different datasets. Finally, as a learning-based method, SSRNet can process large-scale point clouds in a short time. And to further solve the efficiency problem, we provide a lightweight and efficient version that is about five times faster while maintaining reconstruction performance. Experiments show that our methods achieve state-of-the-art performance with outstanding efficiency.
Ganzhangqin Yuan, Qiancheng Fu, Zhenxing Mi, Wenbing Tao
IEEE Trans. Vis. Comput. Graph.5
2023 LGP-MVS: combined local and global planar priors guidance for indoor multi-view stereo
Weihang Kong, Qingshan Xu 0001, Wanjuan Su, Wenbing Tao
Vis. Comput.5
2022 DeTarNet: Decoupling Translation and Rotation by Siamese Network for Point Cloud Registration
abstract
Point cloud registration is a fundamental step for many tasks. In this paper, we propose a neural network named DetarNet to decouple the translation t and rotation R, so as to overcome the performance degradation due to their mutual interference in point cloud registration. First, a Siamese Network based Progressive and Coherent Feature Drift (PCFD) module is proposed to align the source and target points in high-dimensional feature space, and accurately recover translation from the alignment process. Then we propose a Consensus Encoding Unit (CEU) to construct more distinguishable features for a set of putative correspondences. After that, a Spatial and Channel Attention (SCA) block is adopted to build a classification network for finding good correspondences. Finally, the rotation is obtained by Singular Value Decomposition (SVD). In this way, the proposed network decouples the estimation of translation and rotation, resulting in better performance for both of them. Experimental results demonstrate that the proposed DetarNet improves registration performance on both indoor and outdoor scenes. Our code will be available in https://github.com/ZhiChen902/DetarNet.
Zhi Chen 0011, Fan Yang 0088, Wenbing Tao
AAAI3
2022 SC2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration
abstract
In this paper, we present a second order spatial compat-ibility (SC2) measure based method for efficient and robust point cloud registration (PCR), called SC2-PCR 1. Firstly, we propose a second order spatial compatibility (SC2) mea-sure to compute the similarity between correspondences. It considers the global compatibility instead of local consis-tency, allowing for more distinctive clustering between in-liers and outliers at early stage. Based on this measure, our registration pipeline employs a global spectral technique to find some reliable seeds from the initial correspondences. Then we design a two-stage strategy to expand each seed to a consensus set based on the SC2measure matrix. Finally, we feed each consensus set to a weighted SVD algorithm to generate a candidate rigid transformation and select the best model as the final result. Our method can guarantee to find a certain number of outlier-free consensus sets using fewer samplings, making the model estimation more ef-ficient and robust. In addition, the proposed SC2measure is general and can be easily plugged into deep learning based frameworks. Extensive experiments are carried out to in-vestigate the performance of our method.
Zhi Chen 0011, Kun Sun 0002, Fan Yang 0088, Wenbing Tao
CVPR4
2022 Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view Reconstruction
abstract
Recently, neural implicit surfaces learning by volume rendering has become popular for multi-view reconstruction. However, one key challenge remains: existing approaches lack explicit multi-view geometry constraints, hence usually fail to generate geometry-consistent surface reconstruction. To address this challenge, we propose geometry-consistent neural implicit surfaces learning for multi-view reconstruction. We theoretically analyze that there exists a gap between the volume rendering integral and point-based signed distance function (SDF) modeling. To bridge this gap, we directly locate the zero-level set of SDF networks and explicitly perform multi-view geometry optimization by leveraging the sparse geometry from structure from motion (SFM) and photometric consistency in multi-view stereo. This makes our SDF optimization unbiased and allows the multi-view geometry constraints to focus on the true surface optimization. Extensive experiments show that our proposed method achieves high-quality surface reconstruction in both complex thin structures and large smooth regions, thus outperforming the state-of-the-arts by a large margin.
Qiancheng Fu, Qingshan Xu 0001, Yew-Soon Ong, Wenbing Tao
NeurIPS4
2022 One-Inlier is First: Towards Efficient Position Encoding for Point Cloud Registration
abstract
Transformer architecture has shown great potential for many visual tasks, including point cloud registration. As an order-aware module, position encoding plays an important role in Transformer architecture applied to point cloud registration task. In this paper, we propose OIF-PCR, a one-inlier based position encoding method for point cloud registration network. Specifically, we first find one correspondence by a differentiable optimal transport layer, and use it to normalize each point for position encoding. It can eliminate the challenges brought by the different reference frames of two point clouds, and mitigate the feature ambiguity by learning the spatial consistency. Then, we propose a joint approach for establishing correspondence and position encoding, presenting an iterative optimization process. Finally, we design a progressive way for point cloud alignment and feature learning to gradually optimize the rigid transformation. The proposed position encoding is very efficient, requiring only a small addition of memory and computing overhead. Extensive experiments demonstrate the proposed method can achieve competitive performance with the state-of-the-art methods in both indoor and outdoor scenes.
Fan Yang 0088, Lin Guo 0019, Zhi Chen 0011, Wenbing Tao
NeurIPS4
2022 Sparse prior guided deep multi-view stereo
Yuhang Qi, Wanjuan Su, Qingshan Xu 0001, Wenbing Tao
Comput. Graph.4
2022 Learning Inverse Depth Regression for Pixelwise Visibility-Aware Multi-View Stereo Networks
Qingshan Xu 0001, Wanjuan Su, Yuhang Qi, Wenbing Tao, Marc Pollefeys
Int. J. Comput. Vis.4
2022 Multi-scale receptive field fusion network for lightweight image super-resolution
Lin Zhao 0012, Wenbing Tao
Neurocomputing4
2022 Robust consensus-aware network for 3D point registration
Fan Yang 0088, Zhi Chen 0011, Kun Sun 0002, Liman Liu, Wenbing Tao
Neurocomputing5
2022 Bridging the gap between one-to-many and one-to-one label assignment via NMS-aware alignment module
Lin Zhao 0012, Liman Liu, Wenbing Tao
Neurocomputing5
2022 Uncertainty Guided Multi-View Stereo Network for Depth Estimation
abstract
Deep learning has greatly promoted the development of multi-view stereo in recent years. However, how to measure the reliability of the estimated depth map for practical applications and make reasonable depth hypothesis sampling for the cost volume building in the coarse-to-fine architecture are still unresolved crucial problems. To this end, an Uncertainty Guided multi-view Network (UGNet) is proposed in this paper. In order to enable the network to perceive the uncertainty, an uncertainty-aware loss function is introduced, which not only can infer uncertainty implicitly in an unsupervised manner but also can reduce the bad impact of high uncertainty regions and the erroneous labels in the training set during training. Moreover, an uncertainty-based depth hypothesis sampling strategy is further proposed to adaptively determine the depth search range of each pixel for finer stages, which helps to generate more rational depth intervals compared with other methods and build more compact cost volumes without redundancy. Experimental results on DTU dataset, BlendedMVS dataset, Tanks and Temples dataset and ETH3D high-res benchmark show that our method achieves promising reconstruction results compared with other state-of-the-art methods.
Wanjuan Su, Qingshan Xu 0001, Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.3
2021 Cascade Network with Guided Loss and Hybrid Attention for Finding Good Correspondences
abstract
Finding good correspondences is a critical prerequisite in many feature based tasks. Given a putative correspondence set of an image pair, we propose a neural network which finds correct correspondences by a binary-class classifier and estimates relative pose through classified correspondences. First, we analyze that due to the imbalance in the number of correct and wrong correspondences, the loss function has a great impact on the classification results. Thus, we propose a new Guided Loss that can directly use evaluation criterion (Fn-measure) as guidance to dynamically adjust the objective function during training. We theoretically prove that the perfect negative correlation between the Guided Loss and Fn-measure, so that the network is always trained towards the direction of increasing Fn-measure to maximize it. We then propose a hybrid attention block to extract feature, which integrates the Bayesian attentive context normalization (BACN) and channel-wise attention (CA). BACN can mine the prior information to better exploit global context and CA can capture complex channel context to enhance the channel awareness of the network. Finally, based on our Guided Loss and hybrid attention block, a cascade network is designed to gradually optimize the result for more superior performance. Experiments have shown that our network achieves the state-of-the-art performance on benchmark datasets. Our code will be available in https://github.com/wenbingtao/GLHA.
Zhi Chen 0011, Fan Yang 0088, Wenbing Tao
AAAI3
2021 DeepDT: Learning Geometry From Delaunay Triangulation for Surface Reconstruction
abstract
In this paper, a novel learning-based network, named DeepDT, is proposed to reconstruct the surface from Delaunay triangulation of point cloud. DeepDT learns to predict inside/outside labels of Delaunay tetrahedrons directly from a point cloud and corresponding Delaunay triangulation. The local geometry features are first extracted from the input point cloud and aggregated into a graph deriving from the Delaunay triangulation. Then a graph filtering is applied on the aggregated features in order to add structural regularization to the label prediction of tetrahedrons. Due to the complicated spatial relations between tetrahedrons and the triangles, it is impossible to directly generate ground truth labels of tetrahedrons from ground truth surface. Therefore, we propose a multi-label supervision strategy which votes for the label of a tetrahedron with labels of sampling locations inside it. The proposed DeepDT can maintain abundant geometry details without generating overly complex surfaces, especially for inner surfaces of open scenes. Meanwhile, the generalization ability and time consumption of the proposed method is acceptable and competitive compared with the state-of-the-art methods. Experiments demonstrate the superior performance of the proposed DeepDT.
Zhenxing Mi, Wenbing Tao
AAAI3
2021 Efficient convex optimization-based texture mapping for large-scale 3D scene reconstruction
Jing Yuan 0001, Wenbing Tao, Bo Tao 0001, Liman Liu
Inf. Sci.3
2021 IoU-uniform R-CNN: Breaking through the limitations of RPN
Zihao Xie, Liman Liu, Bo Tao 0001, Wenbing Tao
Pattern Recognit.5
2020 Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost Volume
abstract
Deep learning has shown to be effective for depth inference in multi-view stereo (MVS). However, the scalability and accuracy still remain an open problem in this domain. This can be attributed to the memory-consuming cost volume representation and inappropriate depth inference. Inspired by the group-wise correlation in stereo matching, we propose an average group-wise correlation similarity measure to construct a lightweight cost volume. This can not only reduce the memory consumption but also reduce the computational burden in the cost volume filtering. Based on our effective cost volume representation, we propose a cascade 3D U-Net module to regularize the cost volume to further boost the performance. Unlike the previous methods that treat multi-view depth inference as a depth regression problem or an inverse depth classification problem, we recast multi-view depth inference as an inverse depth regression task. This allows our network to achieve sub-pixel estimation and be applicable to large-scale scenes. Through extensive experiments on DTU dataset and Tanks and Temples dataset, we show that our proposed network with Correlation cost volume and Inverse DEpth Regression (CIDER1), achieves state-of-the-art results, demonstrating its superior performance on scalability and accuracy.
Qingshan Xu 0001, Wenbing Tao
AAAI2
2020 Planar Prior Assisted PatchMatch Multi-View Stereo
abstract
The completeness of 3D models is still a challenging problem in multi-view stereo (MVS) due to the unreliable photometric consistency in low-textured areas. Since low-textured areas usually exhibit strong planarity, planar models are advantageous to the depth estimation of low-textured areas. On the other hand, PatchMatch multi-view stereo is very efficient for its sampling and propagation scheme. By taking advantage of planar models and PatchMatch multi-view stereo, we propose a planar prior assisted PatchMatch multi-view stereo framework in this paper. In detail, we utilize a probabilistic graphical model to embed planar models into PatchMatch multi-view stereo and contribute a novel multi-view aggregated matching cost. This novel cost takes both photometric consistency and planar compatibility into consideration, making it suited for the depth estimation of both non-planar and planar regions. Experimental results demonstrate that our method can efficiently recover the depth information of extremely low-textured areas, thus obtaining high complete 3D models and achieving state-of-the-art performance.
Qingshan Xu 0001, Wenbing Tao
AAAI2
2020 JSNet: Joint Instance and Semantic Segmentation of 3D Point Clouds
abstract
In this paper, we propose a novel joint instance and semantic segmentation approach, which is called JSNet, in order to address the instance and semantic segmentation of 3D point clouds simultaneously. Firstly, we build an effective backbone network to extract robust features from the raw point clouds. Secondly, to obtain more discriminative features, a point cloud feature fusion module is proposed to fuse the different layer features of the backbone network. Furthermore, a joint instance semantic segmentation module is developed to transform semantic features into instance embedding space, and then the transformed features are further fused with instance features to facilitate instance segmentation. Meanwhile, this module also aggregates instance features into semantic feature space to promote semantic segmentation. Finally, the instance predictions are generated by applying a simple mean-shift clustering on instance embeddings. As a result, we evaluate the proposed JSNet on a large-scale 3D indoor point cloud dataset S3DIS and a part dataset ShapeNet, and compare it with existing approaches. Experimental results demonstrate our approach outperforms the state-of-the-art method in 3D instance segmentation with a significant improvement in 3D semantic prediction and our method is also beneficial for part segmentation. The source code for this work is available at https://github.com/dlinzhao/JSNet.
Lin Zhao 0012, Wenbing Tao
AAAI2
2020 SSRNet: Scalable 3D Surface Reconstruction Network
abstract
Existing learning-based surface reconstruction methods from point clouds are still facing challenges in terms of scalability and preservation of details on large-scale point clouds. In this paper, we propose the SSRNet, a novel scalable learning-based method for surface reconstruction. The proposed SSRNet constructs local geometry-aware features for octree vertices and designs a scalable reconstruction pipeline, which not only greatly enhances the predication accuracy of the relative position between the vertices and the implicit surface facilitating the surface reconstruction quality, but also allows dividing the point cloud and octree vertices and processing different parts in parallel for superior scalability on large-scale point clouds with millions of points. Moreover, SSRNet demonstrates outstanding generalization capability and only needs several surface data for training, much less than other learning-based reconstruction methods, which can effectively avoid overfitting. The trained model of SSRNet on one dataset can be directly used on other datasets with superior performance. Finally, the time consumption with SSRNet on a large-scale point cloud is acceptable and competitive. To our knowledge, the proposed SSRNet is the first to really bring a convincing solution to the scalability issue of the learning-based surface reconstruction methods, and is an important step to make learning-based methods competitive with respect to geometry processing methods on real-world and challenging data. Experiments show that our method achieves a breakthrough in scalability and quality compared with state-of-the-art learning-based methods.
Zhenxing Mi, Wenbing Tao
CVPR3
2020 Localization-aware channel pruning for object detection
Zihao Xie, Lin Zhao 0012, Bo Tao 0001, Liman Liu, Wenbing Tao
Neurocomputing6
2020 Guide to Match: Multi-Layer Feature Matching With a Hybrid Gaussian Mixture Model
abstract
As a fundamental yet challenging task in computer vision, finding correspondences between two sets of feature points has received extensive attention. Among all the proposed methods, the Gaussian Mixture Model (GMM) based algorithms show their great power in formulating such problems. However, they are vulnerable to large portion of outliers in the extracted feature points. In this paper, a new Hybrid Gaussian Mixture Model (HGMM) combined with a multi-layer matching framework is proposed. Different from existing GMM based methods, HGMM uses a set of seed correspondences to guide the matching procedure. To automatically find seed correspondences, the feature points are divided into multiple layers according to their matching potential. With the help of Locality Sensitive Hashing, this can be done economically and efficiently. Correspondences found in lower layers which contain few outliers will be used as hard constraint when matching features in higher layers where a large portion of outliers exist. Extensive experiments show that the proposed method is efficient and more robust to outliers when images have large viewpoint difference or small scene overlap.
Kun Sun 0002, Wenbing Tao
IEEE Trans. Multim.2
2019 Multi-Scale Geometric Consistency Guided Multi-View Stereo
abstract
In this paper, we propose an efficient multi-scale geometric consistency guided multi-view stereo method for accurate and complete depth map estimation. We first present our basic multi-view stereo method with Adaptive Checkerboard sampling and Multi-Hypothesis joint view selection (ACMH). It leverages structured region information to sample better candidate hypotheses for propagation and infer the aggregation view subset at each pixel. For the depth estimation of low-textured areas, we further propose to combine ACMH with multi-scale geometric consistency guidance (ACMM) to obtain the reliable depth estimates for low-textured areas at coarser scales and guarantee that they can be propagated to finer scales. To correct the erroneous estimates propagated from the coarser scales, we present a novel detail restorer. Experiments on extensive datasets show our method achieves state-of-the-art performance, recovering the depth estimation not only in low-textured areas but also in details.
Qingshan Xu 0001, Wenbing Tao
CVPR2
2019 Semantic Object and Plane SLAM for RGB-D Cameras
Longyu Zheng, Wenbing Tao
PRCV (3)2
2019 A center-driven image set partition algorithm for efficient structure from motion
Kun Sun 0002, Wenbing Tao
Inf. Sci.2
2019 Efficient large-scale geometric verification for structure from motion
Qingshan Xu 0001, Wenbing Tao, Delie Ming
Pattern Recognit. Lett.3
2019 Learning Linear Regression via Single-Convolutional Layer for Visual Object Tracking
abstract
Learning a large-scale regression model has proven to be one of the most successful approaches for visual tracking as in recent correlation filter (CF)- based trackers. Different from the conventional CF-based algorithms in which the regression model is solved based on circulant training samples, we propose learning linear regression models via a single-convolutional layer with the gradient descent (GD) technique. In our convolution-based approach, the samples are cropped from an image in a sliding-window manner rather than being circularly shifted from one base sample. As a result, the abundant background context in the images can be fully exploited to learn a robust tracker. The proposed tracker is based on two independent regression models: a holistic regression model and a texture regression model. The holistic regression model is trained based on the entire object patch to predict the object location, whereas the texture regression model is trained based on the local object textures. The foreground map outputted by the texture regression model is not only helpful to boost the location prediction in the case of large variations, but is also an important clue for estimating the object size. With the foreground map outputted by the texture regression model, we are able to estimate the object size by optimizing a novel objective function based on object-background contrast. Our extensive experiments on four popular visual tracking datasets OTB-50, OTB-100, VOT-2016, and TempleColor have proved that the proposed algorithm achieves outstanding performance and outperforms most CF-based trackers.
Kai Chen 0023, Wenbing Tao
IEEE Trans. Multim.2
2018 Point Cloud Noise and Outlier Removal with Locally Adaptive Scale
Zhenxing Mi, Wenbing Tao
PRCV (3)2
2018 Iterative image segmentation with feature driven heuristic four-color labeling
Kunqian Li, Wenbing Tao, Xiaobai Liu, Liman Liu
Pattern Recognit.2
2018 A Constrained Radial Agglomerative Clustering Algorithm for Efficient Structure From Motion
abstract
Building three-dimensional models effectively and accurately is an important issue. In this letter, a new image set partition method for efficient structure from motion (SfM) from a set of unevenly distributed images is proposed. Given the largest connected component in the image matching graph, we first reconstruct a base model from a set of images with large overlap and sufficient feature correspondences. Then, a novel constrained radial agglomerative clustering algorithm is proposed to divide the remaining images, so that each image cluster could be independently added to the base model in parallel. Finally, all the partial models are merged into a complete scene. Experiment results show that the proposed method works better than the popular normalized-cuts-based SfM method.
Kun Sun 0002, Wenbing Tao
IEEE Signal Process. Lett.2
2018 Once for All: A Two-Flow Convolutional Neural Network for Visual Tracking
abstract
The main challenges of visual object tracking arise from the arbitrary appearance of the objects that need to be tracked. Most existing algorithms try to solve this problem by training a new model to regenerate or classify each tracked object. As a result, the model needs to be initialized and retrained for each new object. In this paper, we propose to track different objects in an object-independent approach with a novel two-flow convolutional neural network (YCNN). The YCNN takes two inputs (one is an object image patch, the other is a larger searching image patch), then outputs a response map which predicts how likely and where the object would appear in the search patch. Unlike the object-specific approaches, the YCNN is actually trained to measure the similarity between the two image patches. Thus, this model will not be limited to any specific object. Furthermore, the network is end-to-end trained to extract both shallow and deep dedicated convolutional features for visual tracking. And once properly trained, the YCNN can be used to track all kinds of objects without further training and updating. As a result, our algorithm is able to run at a very high speed of 45 frames-per-second. The effectiveness of the proposed algorithm can also be proved by the experiments on two popular data sets: OTB-100 and VOT-2014.
Kai Chen 0023, Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.2
2018 Convolutional Regression for Visual Tracking
abstract
Recently, discriminatively learned correlation filters (DCF) has attracted much attention in visual object tracking community. The success of DCF is potentially attributed to the fact that a large number of samples are utilized to train the ridge regression model and predict the location of an object. To solve the regression problem in an efficient way, these samples are all generated by circularly shifting from a searching patch. However, these synthetic samples also induce some negative effects which weaken the robustness of DCF based trackers. In this paper, we propose a new approach to learn the regression model for visual tracking with single convolutional layer. Instead of learning the linear regression model in a closed form, we try to solve the regression problem by optimizing a onechannel- output convolution layer with gradient descent (GD). In particular, the kernel size of the convolution layer is set to the size of the object. Contrary to DCF, it is possible to incorporate all "real" samples clipped from the whole image. A critical issue of the GD approach is that most of the convolutional samples are negative and the contribution of positive samples will be suppressed. To address this problem, we propose a novel objective function to eliminate easy negatives and enhance positives. We perform extensive experiments on four widely-used datasets: OTB-100, OTB-50, TempleColor, and VOT-2016. The results show that the proposed algorithm achieves outstanding performance and outperforms most of the existing DCF based algorithms.
Kai Chen 0023, Wenbing Tao
IEEE Trans. Image Process.2
2017 Visual object tracking via enhanced structural correlation filter
Kai Chen 0023, Wenbing Tao, Shoudong Han
Inf. Sci.2
2017 Image Matching via Feature Fusion and Coherent Constraint
abstract
The Gaussian mixture model (GMM)-based methods have achieved great success in point set registration. However, they cannot be directly applied to image matching, because the features extracted from two images usually contain a large portion of outliers. In this letter, we propose a new method to extend the powerful GMM to the field of image feature points matching. The algorithm consists of two main steps. In the first step, points extracted from the images are mapped into a new subspace, in which feature similarity information is fused to get the new representation of the points. The second step performs an improved progressive process with the GMM to find correspondences satisfying the coherent constraint. In this way, finding correspondences among large outliers is feasible and the iteration converges faster. Experimental results on benchmark data sets show that the proposed method can find more correct matches with high accuracy.
Kun Sun 0002, Liman Liu, Wenbing Tao
IEEE Geosci. Remote. Sens. Lett.3
2017 Color-texture cosegmentation based on nonlinear compact multi-scale structure tensor and TV-flow
Shoudong Han, Wenbing Tao
Signal Process.3
2016 Complementary saliency driven co-segmentation with region searching and Hierarchical constraint
Liman Liu, Wenbing Tao, Haihua Liu
Inf. Sci.2
2016 Progressive match expansion via coherent subspace constraint
Kun Sun 0002, Liman Liu, Wenbing Tao
Inf. Sci.3
2016 Pedestrian detection aided by fusion of binocular information
Zhiguo Zhang 0005, Wenbing Tao, Kun Sun 0002
Pattern Recognit.2
2016 Multivideo Object Cosegmentation for Irrelevant Frames Involved Videos
abstract
Even though there have been a large amount of previous work on video segmentation techniques, it is still a challenging task to extract the video objects accurately without interactions, especially for those videos which contain irrelevant frames (frames containing no common targets). In this essay, a novel multivideo object cosegmentation method is raised to cosegment common or similar objects of relevant frames in different videos, which includes three steps: 1) object proposal generation and clustering within each video; 2) weighted graph construction and common objects selection; and 3) irrelevant frames detection and pixel-level segmentation refinement. We apply our method on challenging datasets and exhaustive comparison experiments demonstrate the effectiveness of the proposed method.
Kunqian Li, Wenbing Tao
IEEE Signal Process. Lett.3
2016 Pedestrian Detection in Binocular Stereo Sequence Based on Appearance Consistency
abstract
Pedestrian detection is an important yet challenging task. In this paper, a pedestrian codetection framework for binocular stereo sequences is proposed. Binocular vision and consecutive frames can provide more information than a single image. That information can be used to improve detection performance by reducing the number of candidate detection windows with low confidence or enhancing the high-confidence candidates. To design the framework, we follow the intuition that a pedestrian has consistent appearance when observed from the same or different viewpoints. First, a baseline detector is used in both stereo images with a conservative threshold in order to expand the set of detection candidates. Before the detection process, an assisted stixel world model is computed for both left and right frames. Thus, the search range of the detector is greatly reduced with the help of the stixel world. Second, adjacency-constrained patch matching is proposed to build correspondence between two candidates in both intra and inter sequences of binocular vision. Finally, we establish a mechanism to update the score of the detection aided by the corresponding candidates. The experimental results show that our framework significantly improves the performance of the evaluated baseline pedestrian approaches.
Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.2
2016 Unsupervised Co-Segmentation for Indefinite Number of Common Foreground Objects
abstract
Co-segmentation addresses the problem of simultaneously extracting the common targets appeared in multiple images. Multiple common targets involved object co-segmentation problem, which is very common in reality, has been a new research hotspot recently. In this paper, an unsupervised object co-segmentation method for indefinite number of common targets is proposed. This method overcomes the inherent limitation of traditional proposal selection-based methods for multiple common targets involved images while retaining their original advantages for objects extracting. For each image, the proposed multi-search strategy extracts each target individually and an adaptive decision criterion is raised to give each candidate a reliable judgment automatically, i.e., target or non-target. The comparison experiments conducted on public data sets iCoseg, MSRC, and a more challenging data set Coseg-INCT demonstrate the superior performance of the proposed method.
Kunqian Li, Wenbing Tao
IEEE Trans. Image Process.3
2015 Feature Guided Biased Gaussian Mixture Model for image matching
Kun Sun 0002, Wenbing Tao, Yuan Yan Tang
Inf. Sci.3
2015 SaCoseg: Object Cosegmentation by Shape Conformability
abstract
In this paper, an object cosegmentation method based on shape conformability is proposed. Different from the previous object cosegmentation methods which are based on the region feature similarity of the common objects in image set, our proposed SaCoseg cosegmentation algorithm focuses on the shape consistency of the foreground objects in image set. In the proposed method, given an image set where the implied foreground objects may be varied in appearance but share similar shape structures, the implied common shape pattern in the image set can be automatically mined and regarded as the shape prior of those unsatisfactorily segmented images. The SaCoseg algorithm mainly consists of four steps: 1) the initial Grabcut segmentation; 2) the shape mapping by coherent point drift registration; 3) the common shape pattern discovery by affinity propagation clustering; and 4) the refinement by Grabcut with common shape constraint. To testify our proposed algorithm and establish a benchmark for future work, we built the CoShape data set to evaluate the shape-based cosegmentation. The experiments on CoShape data set and the comparison with some related cosegmentation algorithms demonstrate the good performance of the proposed SaCoseg algorithm.
Wenbing Tao, Kunqian Li, Kun Sun 0002
IEEE Trans. Image Process.1
2015 Robust Point Sets Matching by Fusing Feature and Spatial Information Using Nonuniform Gaussian Mixture Models
abstract
Most of the traditional methods that handle the point sets matching between two images are based on local feature descriptors and the succedent mismatch eliminating strategies, which usually suffers from the sparsity of the initial match set because some correct ambiguous associations are easily filtered out by the ratio test of SIFT matching due to their second ranking in feature similarity. In this paper, we propose a nonuniform Gaussian mixture model (NGMM) for point sets matching between a pair of images which combines feature with position information of the local feature points extracted from the image pair to achieve point sets matching in a GMM framework. The proposed point set matching using an NGMM is able to change the correspondence assignments throughout the matching process and has the potential to match up even ambiguous matches correctly. The proposed NGMM framework can be either used to directly find matches between two point sets obtained from two images or applied to remove outliers in a match set. When finding matches, NGMM tries to learn a nonrigid transformation between the two point sets and provide a probability for every found match to measure the reliability of the match. Then, a probability threshold can be used to get the final robust match set. When removing outliers, NGMM requires that the vector field formed by the correct matches to be coherent and the matches contradicting the coherent vector field will be regarded as mismatches to be removed. A number of comparison and evaluation experiments reveal the good performance of the proposed NGMM framework in both finding matches and discarding mismatches.
Wenbing Tao, Kun Sun 0002
IEEE Trans. Image Process.1
2015 Adaptive Optimal Shape Prior for Easy Interactive Object Segmentation
abstract
For interactive segmentation approaches, object segmentation in complicated background is cumbersome, and usually needs tedious interactions to refine the incomplete segmentations . In this paper, an adaptive optimal shape prior is proposed for easy interactive object segmentation. Different from the traditional shape priors which only provide loose constraint, our adaptive shape prior gives more accurate and individualized constraint by exploiting the shape information of incomplete segmentation. Moreover, by combining the non-rigid shape registration and a local shape consistency evaluation system presented in this paper, such adaptive optimal shape prior could be achieved automatically. Both of these contributions greatly lighten the burden on users and make interactive segmentation much easier. The comparison experiments on the newly-built TypShape dataset with the related algorithms have demonstrated good performance of the proposed algorithm.
Kunqian Li, Wenbing Tao
IEEE Trans. Multim.2
2014 Asymmetrical Gauss Mixture Models for Point Sets Matching
abstract
The probabilistic methods based on Symmetrical Gauss Mixture Model(SGMM)[4, 13, 8] have achieved great success in point sets registration, but are seldom used to find the correspondences between two images due to the complexity of the non-rigid transformation and too many outliers. In this paper we propose an Asymmetrical GMM(AGMM) for point sets matching between a pair of images. Different from the previous SGMM, the AGMM gives each Gauss component a different weight which is related to the feature similarity between the data point and model point, which leads to two effective algorithms: the Single Gauss Model for Mismatch Rejection(SGMR) algorithm and the AGMM algorithm for point sets matching. The SGMR algorithm iteratively filters mismatches by estimating a non-rigid transformation between two images based on the spatial coherence of point sets. The AGMM algorithm combines the feature information with position information of the SIFT feature points extracted from the images to achieve point sets matching so that much more correct correspondences with high precision can be found. A number of comparison and evaluation experiments reveal the excellent performance of the proposed SGMR algorithm and AGMM algorithm.
Wenbing Tao, Kun Sun 0002
CVPR1
2014 Integration of the saliency-based seed extraction and random walks for image segmentation
Chanchan Qin, Yicong Zhou, Wenbing Tao, Zhiguo Cao 0001
Neurocomputing4
2014 Spatial adjacent bag of features with multiple superpixels for object segmentation and classification
Wenbing Tao, Yicong Zhou, Liman Liu, Kunqian Li, Kun Sun 0002, Zhiguo Zhang 0005
Inf. Sci.1
2014 Unsupervised multiphase color-texture image segmentation based on variational formulation and multilayer graph
Tianjiang Wang, Wenbing Tao, Guangpu Shao, Qi Feng 0003
Image Vis. Comput.4
2014 Epipolar geometry estimation for wide baseline stereo by Clustering Pairing Consensus
Dazhi Zhang, Yongtao Wang, Wenbing Tao, Chengyi Xiong
Pattern Recognit. Lett.3
2014 Automatic image segmentation using salient key point extraction and star shape prior
Xiangli Liao, Yicong Zhou, Kunqian Li, Wenbing Tao, Qiuju Guo, Liman Liu
Signal Process.5
2014 An Ordered-Patch-Based Image Classification Approach on the Image Grassmannian Manifold
abstract
This paper presents an ordered-patch-based image classification framework integrating the image Grassmannian manifold to address handwritten digit recognition, face recognition, and scene recognition problems. Typical image classification methods explore image appearances without considering the spatial causality among distinctive domains in an image. To address the issue, we introduce an ordered-patch-based image representation and use the autoregressive moving average (ARMA) model to characterize the representation. First, each image is encoded as a sequence of ordered patches, integrating both the local appearance information and spatial relationships of the image. Second, the sequence of these ordered patches is described by an ARMA model, which can be further identified as a point on the image Grassmannian manifold. Then, image classification can be conducted on such a manifold under this manifold representation. Furthermore, an appropriate Grassmannian kernel for support vector machine classification is developed based on a distance metric of the image Grassmannian manifold. Finally, the experiments are conducted on several image data sets to demonstrate that the proposed algorithm outperforms other existing image classification methods.
Chunyan Xu, Tianjiang Wang, Junbin Gao, Shougang Cao, Wenbing Tao, Fang Liu 0011
IEEE Trans. Neural Networks Learn. Syst.5
2013 Multilayer graph cuts based unsupervised color-texture image segmentation using multivariate mixed student's t-distribution and regional credibility merging
Shoudong Han, Tianjiang Wang, Wenbing Tao, Xue-Cheng Tai
Pattern Recognit.4
2012 Iterative Narrowband-Based Graph Cuts Optimization for Geodesic Active Contours With Region Forces (GACWRF)
abstract
In this paper, an iterative narrow-band-based graph cuts (INBBGC) method is proposed to optimize the geodesic active contours with region forces (GACWRF) model for interactive object segmentation. Based on cut metric on graphs proposed by Boykov and Kolmogorov, an NBBGC method is devised to compute the local minimization of GAC. An extension to an iterative manner, namely, INBBGC, is developed for less sensitivity to the initial curve. The INBBGC method is similar to graph-cuts-based active contour (GCBAC) presented by Xu , and their differences have been analyzed and discussed. We then integrate the region force into GAC. An improved INBBGC (IINBBGC) method is proposed to optimize the GACWRF model, thus can effectively deal with the concave region and complicated real-world images segmentation. Two region force models such as mean and probability models are studied. Therefore, the GCBAC method can be regarded as the special case of our proposed IINBBGC method without region force. Our proposed algorithm has been also analyzed to be similar to the Grabcut method when the Gaussian mixture model region force is adopted, and the band region is extended to the whole image. Thus, our proposed IINBBGC method can be regarded as narrow-band-based Grabcut method or GCBAC with region force method. We apply our proposed IINBBGC algorithm on synthetic and real-world images to emphasize its performance, compared with other segmentation methods, such as GCBAC and Grabcut methods.
Wenbing Tao
IEEE Trans. Image Process.1
2011 Multiple piecewise constant with geodesic active contours (MPC-GAC) framework for interactive image segmentation using graph cut optimization
Wenbing Tao, Xue-Cheng Tai
Image Vis. Comput.1
2011 Texture segmentation using independent-scale component-wise Riemannian-covariance Gaussian mixture model in KL measure based multi-scale nonlinear structure tensor space
Shoudong Han, Wenbing Tao, Xianglin Wu
Pattern Recognit.2
2011 Image segmentation by iterative optimization of multiphase multiple piecewise constant model and Four-Color relabeling
Liman Liu, Wenbing Tao
Pattern Recognit.2
2011 A variational model and graph cuts optimization for interactive foreground extraction
Liman Liu, Wenbing Tao, Jin-Wen Tian
Signal Process.2
2011 Integrating Spatio-Temporal Context With Multiview Representation for Object Recognition in Visual Surveillance
abstract
We present in this paper an integrated solution to rapidly recognizing dynamic objects in surveillance videos by exploring various contextual information. This solution consists of three components. The first one is a multi-view object representation. It contains a set of deformable object templates, each of which comprises an ensemble of active features for an object category in a specific view/pose. The template can be efficiently learned via a small set of roughly aligned positive samples without negative samples. The second component is a unified spatio-temporal context model, which integrates two types of contextual information in a Bayesian way. One is the spatial context, including main surface property (constraints on object type and density) and camera geometric parameters (constraints on object size at a specific location). The other is the temporal context, containing the pixel-level and instance-level consistency models, used to generate the foreground probability map and local object trajectory prediction. We also combine the above spatial and temporal contextual information to estimate the object pose in scene and use it as a strong prior for inference. The third component is a robust sampling-based inference procedure. Taking the spatio-temporal contextual knowledge as the prior model and deformable template matching as the likelihood model, we formulate the problem of object category recognition as a maximum-a-posteriori problem. The probabilistic inference can be achieved by a simple Markov chain Mento Carlo sampler, owing to the informative spatio-temporal context model which is able to greatly reduce the computation complexity and the category ambiguities. The system performance and benefit gain from the spatio-temporal contextual information are quantitatively evaluated on several challenging datasets and the comparison results clearly demonstrate that our proposed algorithm outperforms other state-of-the-art algorithms.
Xiaobai Liu, Liang Lin 0004, Shuicheng Yan, Hai Jin 0001, Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.5
2010 Interactively multiphase image segmentation based on variational formulation and graph cuts
Wenbing Tao, Feng Chang, Liman Liu, Hai Jin 0001, Tianjiang Wang
Pattern Recognit.1
2010 Fast image segmentation based on multilevel banded closed-form method
Shoudong Han, Wenbing Tao, Xianglin Wu, Xue-Cheng Tai, Tianjiang Wang
Pattern Recognit. Lett.2
2009 Image Segmentation Based on GrabCut Framework Integrating Multiscale Nonlinear Structure Tensor
abstract
In this paper, we propose an interactive color natural image segmentation method. The method integrates color feature with multiscale nonlinear structure tensor texture (MSNST) feature and then uses GrabCut method to obtain the segmentations. The MSNST feature is used to describe the texture feature of an image and integrated into GrabCut framework to overcome the problem of the scale difference of textured images. In addition, we extend the Gaussian Mixture Model (GMM) to MSNST feature and GMM based on MSNST is constructed to describe the energy function so that the texture feature can be suitably integrated into GrabCut framework and fused with the color feature to achieve the more superior image segmentation performance than the original GrabCut method. For easier implementation and more efficient computation, the symmetric KL divergence is chosen to produce the estimates of the tensor statistics instead of the Riemannian structure of the space of tensor. The Conjugate norm was employed using Locality Preserving Projections (LPP) technique as the distance measure in the color space for more discriminating power. An adaptive fusing strategy is presented to effectively adjust the mixing factor so that the color and MSNST texture features are efficiently integrated to achieve more robust segmentation performance. Last, an iteration convergence criterion is proposed to reduce the time of the iteration of GrabCut algorithm dramatically with satisfied segmentation accuracy. Experiments using synthesis texture images and real natural scene images demonstrate the superior performance of our proposed method.
Shoudong Han, Wenbing Tao, Xue-Cheng Tai, Xianglin Wu
IEEE Trans. Image Process.2
2008 Layered shape matching and registration: Stochastic sampling with hierarchical graph representation
abstract
To automatically register foreground target in cluttered images, we present a novel hierarchical graph representation and a stochastic computing strategy in Bayesian framework. The graph representation, which contains point-(image primitives), seedgraph-, and subgraph- three levels, are built up following the primal sketch theory to capture geometric, topological, and spatial information both in local and global scale. We use two types of bottom-up algorithms for searching matching candidates to generate the point-level and seedgraph-level representations respectively. Then, the Swendsen-Wang Cuts and Gibbs sampling methods are performed for global optimal solution to generate the final subgraph-level representation, where a mixture bending function and a set of topological operators are defined for matching measurement. Experiments with comparison are demonstrated on standard dataset with outperforming results. Results show that our method can work well even with clutter noise and complex background.
Xiaobai Liu, Liang Lin 0004, Hai Jin 0001, Wenbing Tao
ICPR5
2008 Image Thresholding Using Graph Cuts
abstract
A novel thresholding algorithm is presented in this paper to improve image segmentation performance at a low computational cost. The proposed algorithm uses a normalized graph-cut measure as thresholding principle to distinguish an object from the background. The weight matrices used in evaluating the graph cuts are based on the gray levels of the image, rather than the commonly used image pixels. For most images, the number of gray levels is much smaller than the number of pixels. Therefore, the proposed algorithm requires much smaller storage space and lower computational complexity than other image segmentation algorithms based on graph cuts. This fact makes the proposed algorithm attractive in various real-time vision applications such as automatic target recognition. Several examples are presented, assessing the superior performance of the proposed thresholding algorithm compared with the existing ones. Numerical results also show that the normalized-cut measure is a better thresholding principle compared with other graph-cut measures, such as average-cut and average-association ones.
Wenbing Tao, Hai Jin 0001, Yimin Zhang 0001, Liman Liu
IEEE Trans. Syst. Man Cybern. Part A1
2007 A New Image Thresholding Method Based on Graph Cuts
abstract
A novel thresholding algorithm is presented to achieve improved image segmentation performance at low computational cost in this paper. The proposed algorithm uses a normalized graph cut measure as the thresholding principle to distinguish an object from the background. The weight matrices used in evaluating the graph cuts are based on the gray levels of an image, rather than the commonly used image pixels. Therefore, the proposed algorithm occupies much smaller storage space and requires much lower computational costs and implementation complexity than other image segmentation algorithms based on graph cuts. This fact makes the proposed algorithm attractive in various real-time vision applications such as automatic target recognition (ATR). A large number of examples are presented to show the superior performance of the proposed thresholding algorithm compared to existing thresholding algorithms.
Wenbing Tao, Hai Jin 0001, Liman Liu
ICASSP (1)1
2007 An automatic method for generating affine moment invariants
DeRen Li, Wenbing Tao
Pattern Recognit. Lett.3
2007 Object segmentation using ant colony optimization algorithm and fuzzy entropy
Wenbing Tao, Hai Jin 0001, Liman Liu
Pattern Recognit. Lett.1
2007 Color Image Segmentation Based on Mean Shift and Normalized Cuts
abstract
In this correspondence, we develop a novel approach that provides effective and robust segmentation of color images. By incorporating the advantages of the mean shift (MS) segmentation and the normalized cut (Ncut) partitioning methods, the proposed method requires low computational complexity and is therefore very feasible for real-time image segmentation processing. It preprocesses an image by using the MS algorithm to form segmented regions that preserve the desirable discontinuity characteristics of the image. The segmented regions are then represented by using the graph structures, and the Ncut method is applied to perform globally optimized clustering. Because the number of the segmented regions is much smaller than that of the image pixels, the proposed method allows a low-dimensional image clustering with significant reduction of the complexity compared to conventional graph-partitioning methods that are directly applied to the image pixels. In addition, the image clustering using the segmented regions, instead of the image pixels, also reduces the sensitivity to noise and results in enhanced image segmentation performance. Furthermore, to avoid some inappropriate partitioning when considering every region as only one graph node, we develop an improved segmentation strategy using multiple child nodes for each region. The superiority of the proposed method is examined and demonstrated through a large number of experiments using color natural scene images.
Wenbing Tao, Hai Jin 0001, Yimin Zhang 0001
IEEE Trans. Syst. Man Cybern. Part B1
2003 Image segmentation by three-level thresholding based on maximum fuzzy entropy and genetic algorithm
Wenbing Tao, Jin-Wen Tian, Jian Liu 0011
Pattern Recognit. Lett.1