EDBT 2026 Demo / reviewers in the wild / expert
Mingtao Feng
dblp:184/6596
· DBLP profile ↗
63ranked-venue papers
13as first author
59since 2021 · last 2026
0000-0003-0384-3743ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 11 first-author · 33 since 2021Artificial intelligence and machine learning · 29 · 8 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From perception to cognition: Unifying multi-object 3D visual grounding and dense captioning in monocular images
Keyu Guo, Yongle Huang, Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Huansheng Song |
Expert Syst. Appl. | 6 |
| 2026 | ArgusNet: Understanding 3D scenes more like humans
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001 |
Neurocomputing | 6 |
| 2026 | Diffusion-Driven Self-Supervised Learning for Shape Reconstruction and Pose EstimationabstractFully-supervised category-level pose estimation aims to determine the 6-DoF poses of unseen instances from known categories, requiring expensive manual labeling costs. Recently, various self-supervised category-level pose estimation methods have been proposed to reduce the requirement of the annotated datasets. However, most methods rely on synthetic data or 3D CAD model, and they are typically limited to addressing single-object pose problems without considering multi-objective tasks or shape reconstruction. To overcome these challenges and limitations, we introduce a diffusion-driven self-supervised network for multi-object shape reconstruction and categorical pose estimation, only leveraging the shape priors. Specifically, to capture the SE(3)-equivariant pose features and 3D scale-invariant shape information, we present a Prior-Aware Pyramid 3D Point Transformer. This module adopts a point convolutional layer with radial-kernels for pose-aware learning and a 3D scale-invariant graph convolution layer for object-level shape representation. Furthermore, we introduce a Pretrain-to-Refine Self-Supervised Training Paradigm to train our network. It enables proposed network to capture the associations between shape priors and observations, addressing the challenge of intra-class shape variations by utilising the diffusion mechanism. Extensive experiments conducted on four public datasets and a self-built dataset demonstrate that our method significantly outperforms state-of-the-art self-supervised category-level baselines and even surpasses some fully-supervised instance-level and category-level methods. The project page is released at Self-SRPE. Yaonan Wang 0001, Mingtao Feng, Chao Ding 0006, Zheng Shou 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Learning coherent matrixized representation in latent space for volumetric 4D generation
Qitong Yang, Mingtao Feng, Shijie Sun 0001, Weisheng Dong, Yaonan Wang 0001, Mian M. Ajmal |
Pattern Recognit. | 2 |
| 2026 | Description helps: Semantic and texture consistency constraints for SAR-to-optical translation
Fanhao Zhou, Mingtao Feng, Weisheng Dong |
Pattern Recognit. | 2 |
| 2026 | Soft-Masked Transformer for Point Cloud Processing With Skip Attention-Based UpsamplingabstractPoint cloud processing methods leverage local and global point features to cater to downstream tasks, yet they often overlook the task-level context inherent in point clouds during the encoding stage. We argue that integrating task-level information into the encoding stage significantly enhances performance. To that end, we propose SMTransformer which incorporates task-level information into a vector-based transformer by utilizing a soft mask generated from task-level queries and keys to learn the attention weights. Additionally, to facilitate effective communication between features from the encoding and decoding layers in high-level tasks such as segmentation, we introduce a skip-attention-based up-sampling block. This block dynamically fuses features from various resolution points across the encoding and decoding layers. To mitigate the increase in network parameters and training time resulting from the complexity of the aforementioned blocks, we propose a novel shared point position encoding strategy. This strategy allows various transformer blocks to share the same position information over the same resolution points, thereby reducing network parameters and training time without compromising accuracy. Experimental comparisons with existing methods on multiple datasets demonstrate the efficacy of SMTransformer and skip-attention-based up-sampling for semantic segmentation task. In particular, we achieve state-of-the-art semantic segmentation results of 73.9% mIoU on S3DIS Area 5 and 62.4% mIoU on SWAN dataset. Note to Practitioners—Point cloud processing underpins automation tasks such as robotic perception, navigation, and inspection, where accurate 3D understanding is essential. Existing methods often prioritize vision benchmarks while overlooking automation needs like efficiency on limited hardware and robustness in real-world environments. The proposed SMTransformer embeds task-level guidance into feature learning and employs skip-attention up-sampling to improve segmentation accuracy with practical efficiency. It is well-suited for robotic manipulation, autonomous driving, and inspection applications. Current limitations include reliance on GPUs and sensitivity to extreme density variations. Future work will target edge-device deployment and multi-task extensions. Yong He 0012, Hongshan Yu, Chaoxu Mu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Ajmal Mian |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Shape and Prototype-Guided Diffusion Model for Grape Amodal Completion in Vineyard Phenotyping SystemsabstractOcclusions often lead to underestimated grape phenotypes, thereby impeding accurate vineyard management and yield prediction. In grape de-occlusion tasks, existing amodal completion methods perform poorly due to complex cluster structures and fine-grained local textures. To this end, we propose a shape and prototype guided diffusion model for high-fidelity amodal completion, by developing a variational shape embedding learning strategy and a dynamic prototype prior extraction module. We model the shape embedding as a mixture of von Mises–Fisher distributions, indicating the potential complete shape of occluded grape clusters. Subsequently, we formulate prototype priors as discrete embeddings that are consistent with patch features, which represent local texture characteristics of grape berries. The shape and prototype embeddings are integrated into the reverse diffusion process via cross-attention mechanisms, stabilizing structural predictions and mitigating common artifacts such as deformation and adhesion. We construct and release a grape amodal completion dataset collected from real-world vineyard environments. Experimental results on the grape dataset and public KINS dataset demonstrate the superiority of our method in terms of perceptual fidelity and phenotypic accuracy, highlighting its effectiveness for vision-based phenotyping in practical vineyard applications. Yihan Wang 0006, Jianqiao Luo, Bailin Li, Mingtao Feng, Ajmal Mian |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Spatial Multimodal Knowledge-Driven 3D Scene Graph Prediction With Vision-Language ModelabstractIn-depth understanding of 3D environments not only involves locating and recognizing individual objects but also requires inferring the relationships and interactions among them. However, most existing methods heavily rely on scene-specific contents, which leads to poor performance due to the noisy, cluttered, and partial nature of real-world 3D scenes. In this work, we find that the inherently hierarchical structures of 3D environments, derived from support relationships, aid in the automatic association of semantic and spatial arrangements of objects and provide rich geometric and topological information independent of specific scenarios. To this end, we propose a 3D scene graph generation model that leverages the hierarchical structures of 3D environments as spatial multimodal knowledge to enhance 3D scene graph generation. Specifically, we first devise a cross-modal tuning approach, where a visually-prompted vision language model is learned to infer the support relationships between objects in a low-resource way. Subsequently, we build a hierarchical visual graph and hierarchical symbolic knowledge graph using the fine-tuned vision language model to extract contextualized visual contents and relevant textual facts, respectively. Finally, we progressively accumulate 3D spatial multimodal knowledge about the hierarchical structures by correlating contextualized visual contents and textual facts using a novel graph reasoning network. In addition, to better evaluate the performance of 3D scene graph generation models, we propose a new benchmark 3DSSG-M by reorganizing the widely-used 3D scene graph generation dataset 3DSSG. This reorganization balances the predicate distribution of 3DSSG and reduces the influence of frequency bias. Extensive results and ablations attest to the effectiveness of the hierarchical structures in 3D environments and demonstrate the superiority of our proposed method over current state-of-the-art competitors. Haoran Hou, Mingtao Feng, Yulan Guo, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Uncertainty-Adaptive Volume for Unsupervised Homography EstimationabstractEstimating homography from an image pair is crucial for image alignment, and unsupervised methods that optimize feature reprojection error between target and warped source images have gained attention for their promising performance. In real-world scenes with multiple planes, such as moving objects, outlier rejection strategies are essential to mitigate the influence of non-dominant planes. Existing methods address this by learning a mask based on reprojection error, where high errors indicate non-dominant planes misaligned by homography. However, this error-fitting mask often overextends to the dominant plane, limiting the use of valid image regions for accurate estimation. This paper proposes a novel unsupervised method to compactly exclude non-dominant planes by introducing an uncertainty-adaptive cost volume for homography estimation. We first model uncertainty by assuming image features follow a Gaussian distribution derived from a prior Normal Inverse-Gamma distribution. The network-learned distribution parameters disentangle aleatoric uncertainty, distinguishing data-dependent errors within the total reprojection error. This uncertainty reflects inherent observation noise in image data, effectively indicating non-dominant planes. We then integrate this aleatoric uncertainty into the concatenation volume across image feature maps, creating an adaptive volume that filters out unreliable matching costs associated with non-dominant planes. This adaptive volume simplifies learning homography from the rich, redundant content in the concatenation volume, enabling more efficient and accurate estimation. Experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance both qualitatively and quantitatively. Jianqiao Luo, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Xuebing Liu, Yang Mo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DiffCom: Decoupled Sparse Priors Guided Diffusion Compression for Point CloudsabstractWhile conventional lossy compression methods predominantly depend on autoencoders to map point clouds into latent representations, they often neglect the intrinsic redundancy within these latent points. To address this limitation, this paper presents a diffusion-based architecture steered by sparse priors, designed to minimize latent redundancy while securing superior reconstruction fidelity, particularly in low-bitrate scenarios. A key feature of the framework is an efficient dual-density data flow that alleviates the stringent size constraints imposed on latent points. By integrating a Probabilistic Attention-based Conditional Denoiser (PACD), the method effectively encapsulates critical reconstruction details within sparse priors, which are hierarchically decoupled into intra- and inter-point components. Specifically, separate encoders are utilized to transform the source point cloud into latent points and decoupled sparse priors, respectively. To dynamically exploit geometric and semantic information, an attention-driven latent denoiser, conditioned on these decoupled priors, is applied across the encoding and decoding layers. Furthermore, inter-point distributions are incorporated into the arithmetic codec to refine local context modeling for sparse points, with the final point cloud recovered via a point decoder. Comprehensive experiments conducted on ShapeNet and standard MPEG PCC datasets demonstrate that the proposed method outperforms state-of-the-art techniques, achieving a superior rate-distortion trade-off. Xiaoge Zhang 0003, Mingtao Feng, Mehwish Nasim, Saeed Anwar, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Staged Modulation Diffusion Policy With Complementary Visual Fusion for Robotic Workpiece AssemblyabstractHigh-precision robotic assembly remains a challenge in intelligent manufacturing. Existing vision-based approaches often learn predefined trajectories from large image corpora yet underutilize task-relevant visual cues during execution, limiting deployment. We present a diffusion-based end-to-end assembly policy that performs redundancy-aware cross-scale fusion and provides stage-dependent conditioning for action generation. Specifically, we introduce a bidirectional-attention complementary visual fusion module that aligns cross-scale observations and produces scale-consistent features with reduced redundancy. We then introduce a dual-stream adaptive modulation module that enables time-state correlated routing and progressively shifts emphasis from scene-level stabilization to contact-level refinement within the denoising process. Coupling complementary visual fusion with adaptive-modulation-conditioned denoising, we develop a staged modulation diffusion policy for end-to-end action generation, producing temporally coherent and geometrically accurate actions. Experiments on five real-world tasks demonstrate consistent improvements over representative baselines, with average gains of 41.10% in overall task success rate and 63.16% in precise assembly rate on three representative assembly benchmarks. Yaonan Wang 0001, Mingtao Feng, Renjie Ding, Hui Zhang 0023 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | Second-Order Robust Iterative Pose Optimization for Fine-Grained Cross-View LocalizationabstractFine-grained cross-view localization seeks to estimate precise camera poses by matching ground images with GPS-tagged aerial imagery. Existing methods typically employ first-order iterative optimization to progressively update the camera pose based on cross-view feature correspondences. However, they rely on local features and neglect global and complementary contextual information, making them prone to local optima and slow convergence under large initial errors or strong disturbances. To overcome these limitations, we propose a second-order robust iterative pose estimation framework for fine-grained cross-view localization. Firstly, we devise a second-order deep iterative optimization module to capture complementary forward and backward motion cues, leading to a bidirectional correlation volume. A motion aggregator uses the volume to approximate the dynamics of second-order iterators, substantially facilitating convergence and robustness. In addition, a bidirectional motion-aware robust regularization module mitigates geometric distortions and outlier interference by leveraging bidirectional motion cues to generate fine-grained confidence maps, adaptively suppressing unreliable regions and enhancing the stability of iterative optimization and pose estimation accuracy. Extensive experiments demonstrate that the proposed framework achieves faster convergence and higher pose estimation accuracy than state-of-the-art methods, particularly under large initial errors and challenging conditions. Mingtao Feng, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Image Process. | 1 |
| 2026 | Monocular Multi-Object 3D Visual Language TrackingabstractVisual Language Tracking (VLT) enables machines to perform tracking in real world through human-like language descriptions. However, existing VLT methods are limited to 2D spatial tracking or single-object 3D tracking and do not support multi-object 3D tracking within monocular video. This limitation arises because advancements in 3D multi-object tracking have predominantly relied on sensor-based data (e.g., point clouds, depth sensors) that lacks corresponding language descriptions. Moreover, natural language descriptions in existing VLT literature often suffer from redundancy, impeding the efficient and precise localization of multiple objects. We present the first technique to extend VLT to multi-object 3D tracking using monocular video. We introduce a comprehensive framework that includes (i) a Monocular Multi-object 3D Visual Language Tracking (MoMo-3DVLT) task, (ii) a large-scale dataset, MoMo-3DRoVLT, tailored for this task, and (iii) a custom neural model. Our dataset, generated with the aid of Large Language Models (LLMs) and manual verification, contains 8,216 video sequences annotated with both 2D and 3D bounding boxes, with each sequence accompanied by three freely generated, human-level textual descriptions. We propose MoMo-3DVLTracker, the first neural model specifically designed for MoMo-3DVLT. This model integrates a multimodal feature extractor, a visual language encoder-decoder, and modules for detection and tracking, setting a strong baseline for MoMo-3DVLT. Beyond existing paradigms, it introduces a task-specific structural coupling that integrates a differentiable linked-memory mechanism with depth-guided and language-conditioned reasoning for robust monocular 3D multi-object tracking. Experimental results demonstrate that our approach outperforms existing methods on the MoMo-3DRoVLT dataset. Our dataset and code are available at https://github.com/hongkai-wei/MoMo-3DVLT. Hongkai Wei, Haixiang Hu, Shijie Sun 0001, Mingtao Feng, Keyu Guo, Yongle Huang, Naveed Akhtar |
IEEE Trans. Image Process. | 6 |
| 2026 | Self-Expert Imitation With Purifying Latent Feature for Generalization in Visual Reinforcement LearningabstractThe generalization ability of visual reinforcement learning, which allows the policy trained in the source domain to guide agents in similar unknown target environments, is one of the cores applied to visual navigation and autonomous driving. Recently, methods such as data augmentation techniques, self-supervised learning methods, and the generative adversarial network were employed to enhance the generalization capability of policy neural networks in visual reinforcement learning. However, current state-of-the-art methods, after utilizing domain-general latent features to train the RL policy, result in the loss of certain state-specific features, leading to diminished policy performance following generalization. To tackle these challenges, we designed a technical framework called self-expert imitation with purifying latent features, which enables the trained policy to effectively guide agents in scenarios similar to the training environment, without compromising the performance of the policy-guided agent in task completion. Additionally, a novel method was developed for separating domain-general and domain-specific latent vectors based on a variational autoencoder, enabling the domain-general component to exhibit strong and stable zero-shot generalization performance in unseen visually similar domains. Extensive experiments on the CarRacing game demonstrated that our approach achieves strong and stable generalization performance in unseen environments, without compromising the performance of the policy in guiding agents to complete tasks. Lin Chen 0034, Yang Mo, Yaonan Wang 0001, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Zhen Zhou 0003, Sifei Wang, Danwei Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-LocalizationabstractVisual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing. This makes it challenging to automatically control the relative importance of the two tasks, potentially compromising the retrieval task in favor of the offset regression. Consequently, the model may encounter conflicting dominating gradients during joint training. To address this, we propose to model the semantic ambiguity during the offset regression process by integrating associated uncertainty scores, represented as 2D Gaussian distributions, to mitigate negative transfer effects within the joint tasks. We further introduce an uncertainty-aware similarity metric to enhance similarity assessment between query and reference images, accounting for their semantic ambiguities. This metric propagates uncertainty scores into the retrieval task, focusing on certain samples and learning discriminative feature embeddings, allowing the model to adaptively handle conflicting dominating gradients during joint training. Extensive experiments demonstrate that our method improves the overall performance of the joint tasks, achieving state-of-the-art results on the VIGOR and CVACT datasets. Mingtao Feng, Fenghao Tian, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
AAAI | 1 |
| 2025 | Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionabstractTiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture to enhance the extinguished regions. We for the first time exploit the regions to be enhanced from the perspective of pixel-wise amount of information. Specifically, we model the entire image pixels feature information by minimizing Information Entropy loss, generating an information map to attentively highlight weak activated regions in an unsupervised way. To effectively assist the above phase with more attention to tiny objects, we next introduce the Position Gaussian Distribution Map, explicitly modeled using a Gaussian Mixture distribution, where each Gaussian component's parameters depend on the position and size of object instance labels, serving as supervision for further feature enhancement. Taking the information map as prior knowledge guidance, we construct a multi-scale position gaussian distribution map prediction module, simultaneously modulating the information map and distribution map to focus on tiny objects during training. Extensive experiments on three public tiny object datasets demonstrate the superiority of our method over current state-of-the-art competitors. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
CVPR | 2 |
| 2025 | Beyond Human Perception: Understanding Multi-Object World from Monocular ViewabstractLanguage and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG Keyu Guo, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001, Naveed Akhtar, Ajmal Mian |
CVPR | 5 |
| 2025 | Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingabstractVisual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal with 3D tracking in the confines of monocular video. Unfortunately, advances in 3D tracking mainly rely on expensive sensor inputs, e.g., point clouds, depth measurements, radar. Absence of language counterpart for the outputs of these mildly democratized sensors in the literature also hinders VLT expansion to 3D tracking. Addressing that, we make the first attempt towards extending VLT to 3D tracking based on monocular video. We present a comprehensive framework, introducing (i) the Monocular-Video-based 3D Visual Language Tracking (Mono3DVLT) task, (ii) a large-scale dataset for the task, called Mono3DVLT-V2X, and (iii) a customized neural model for the task. Our dataset is carefully curated, leveraging a Large Langauge Model (LLM) followed by human verification, composing natural language descriptions for 79,158 video sequences aiming at single object tracking, providing 2D and 3D bounding box annotations. Our neural model, termed Mono3DVLT-MT, is the first targeted approach for the Mono3DVLT task. Comprising the pipeline of multi-modal feature extractor, visual-language encoder, tracking decoder and a tracking head, our model sets a strong baseline for the task on Mono3DVLT-V2X. Experimental results show that our method significantly outperforms existing techniques on the Mono3DVLT-V2X dataset. Our dataset and code are available in https://github.com/hongkai-wei/Mono3DVLT. Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Hongli Hu, Huansheng Song, Naveed Akhtar, Ajmal Mian |
CVPR | 4 |
| 2025 | Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesabstractRecent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impact of distributional shifts. In this paper, we observe that classification errors arising from distribution shifts tend to cluster near the true values, suggesting that misclassifications commonly occur in semantically similar, neighboring categories. Furthermore, robust advanced vision foundation models maintain larger inter-class distances while preserving semantic consistency, making them less vulnerable to such shifts. Building on these findings, we propose a new method called GFN (Gain From Neighbors), which uses gradient priors from neighboring classes to perturb input images and incorporates an inter-class distance-weighted loss to improve class separation. This approach encourages the model to learn more resilient features from data prone to errors, enhancing its robustness against shifts in diverse settings. In extensive experiments across various model architectures and benchmark datasets, GFN consistently demonstrated superior performance. For instance, compared to the current state-of-the-art TAPADL method, our approach achieved a higher corruption robustness of 41.4% on ImageNet-C (+2.3%), without requiring additional parameters and using only minimal data. Mingtao Feng, Weisheng Dong, Xin Li 0005, Guangming Shi |
CVPR | 2 |
| 2025 | Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generationabstract3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework designed to efficiently and controllably generate high-resolution 3D models from text prompts or images. Our key insights are three-fold: 1) Hierarchical Gaussian Mixture Model Splatting: We propose a hybrid hierarchical representation to extract fixed number of fine-grained Gaussians with multiscale details from textured object, also establish part-level representation of Gaussians primitives. 2) Mamba with adaptive tree topology: We present a diffusion mamba with tree-topology to adaptively generate Gaussians with disordered spatial structures, without the need for complex preprocessing and maintain linear complexity generation. 3) Controllable Generation: Building on the HGMM tree, we introduce a cascaded diffusion framework combining controllable implicit latent generation, which progressively generates condition-driven latents, and explicit splatting generation, which transforms latents into high-quality Gaussian primitives. Extensive experiments demonstrate the high fidelity and efficiency of our approach. Qitong Yang, Mingtao Feng, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
CVPR | 2 |
| 2025 | Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization
Mingtao Feng, Longlong Mei, Jianqiao Luo, Fenghao Tian, Jie Feng 0003, Weisheng Dong, Yaonan Wang 0001 |
ICCV | 1 |
| 2025 | Generalizing to New Area: Self-Distillation Curriculum Learning for Fine-Grained Cross View LocalizationabstractFine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision. Fenghao Tian, Mingtao Feng, Jianqiao Luo, Longlong Mei, Weisheng Dong, Yaonan Wang 0001 |
ACM Multimedia | 2 |
| 2025 | Disambiguating Holistic Language for 3D Visual Grounding via Neural-Symbolic Reasoning
Yunze Wu, Yufan Zhu, Mingtao Feng, Weisheng Dong |
PRCV (17) | 5 |
| 2025 | Bootstrapping vision-language transformer for monocular 3D visual groundingabstractAbstract In the task of 3D visual grounding using monocular RGB images, it is a challenging problem to perceive visual features and accurately predict the localization of 3D objects based on given geometric and appearances descriptions. Traditional text‐guided attention‐based methods have achieved better results than baselines, but it is argued that there is still potential for improvement in the area of multi‐modal fusion. Thus, Mono3DVG‐TRv2, an end‐to‐end transformer‐based architecture that employs a visual‐text multi‐modal encoder for the alignment and fusion of multi‐modal features, incorporating an enhanced transformer module proven in 2D detection, is introduced. The depth features predicted by the multi‐modal features and the visual‐text features are associated with the learnable queries in the decoder, facilitating more efficient and effective acquisition of geometric information in intricate scenes. Following a comprehensive comparison and ablation study on the Mono3DRefer dataset, this method achieves state‐of‐the‐art performance, markedly surpassing the prior approach. The code will be released at https://github.com/Jade-Ray/Mono3DVGv2 . Shijie Sun 0001, Huansheng Song, Mingtao Feng, Chengzhong Wu |
IET Image Process. | 5 |
| 2025 | Self-supervised contrastive learning for heterophilic graph with latent recurring pattern embedding
Juan Song, Mingtao Feng |
Knowl. Based Syst. | 3 |
| 2025 | Hyperrectangle Embedding for Debiased 3D Scene Graph Prediction From RGB Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment and is regarded as a prerequisite for long-term autonomous robotic operations. A practical research problem here is to predict the 3D scene graph from sequentially captured data. However, existing methods neglect the polysemy of semantic roles that coarse feature vectors are insufficient to represent entities in different relationship semantics. This extremely limits their capability to predict relationships. We propose an approach to tackle the aforementioned challenge by introducing a novel representation, the hyperrectangle embedding, which represents entity using distinctive geometry for more effective scene understanding, rather than learning within vector-based feature with blindly increasing dimensions. By incorporating an entity within two affine-transformed embeddings, each representing either the subject or object and characterized by separate learnable transformations, we achieve the polysemy of semantic roles. The intersections of affine-transformed hyperrectangle embeddings represent the bidirectional relationship between two entities. We identify bias and reliability as two challenges impeding the model learning process. In response to the bias, that arises from long-tailed distributions in the data, we propose a history-guided debiasing strategy that utilizes a confusion history block comprised of previous hyperrectangle embeddings. This strategy mitigates inherent biases by extracting pertinent information and facilitating knowledge transfer from dominant categories to rare ones. To enhance the reliability of predictions, we introduce predictive uncertainty into the 3D scene graph prediction task. We develop a post-hoc reliability enhancement strategy to identify potentially unreliable predictions and subsequently enhance the model's predictive accuracy. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming existing state-of-the-art. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | STR: Spatial-Temporal RetNet for Distributed Multi-Robot NavigationabstractThe core of multi-robot collision avoidance is to guide robots to avoid collisions with other robots and obstacles in a dynamic multi-robot environment, which has recently gained increasing interest among the main challenges of robotics. However, the current multi-robot navigation policy neural network exhibits weak position encoding capabilities for spatial environmental features in mapping environment states and robot actions, as well as an inability to recurrently infer information on dynamic environmental features in the temporal dimension, leading to insufficient safety and effectiveness in guiding robot motion. In this paper, we propose a novel spatial-temporal RetNet (STR) that encodes reciprocal collision avoidance states between robots in both spatial and temporal dimensions, aiming to enhance the safety and effectiveness of the policy neural network in guiding robots to accomplish specified tasks. The spatial state encoder module is developed based on parallel RetNet structure, which enhances the ability of the neural network in multi-robot navigation policies to extract reciprocal collision avoidance states between robots in spatial dimensions and overcomes the weak position encoding capability of advanced transformer-based multi-robot navigation policy neural networks. A temporal state encoder is designed by introducing the recurrent RetNet structure. This enhances the multi-robot navigation policy neural network’s ability to encode features in the temporal dimension of multi-robot movements and overcomes the transformer-based multi-robot navigation policy neural network’s inability to recurrently infer information in the time dimension. Simulation experiments were designed to demonstrate that the safety and effectiveness of our proposed method outperform the previous state-of-the-art approaches in guiding the robot to complete the task. Physical experiments illustrate that our policy can be effectively applied to real-world systemsNote to Practitioners—Multi-robot navigation has a wide range of real-world applications, such as multi-robot formation flying for search and rescue, autonomous warehouse operations, and robots navigating through human crowds. This paper introduces a novel Spatial-Temporal RetNet (STR) framework aimed at enhancing safety and effectiveness in multi-robot collision avoidance. STR addresses the limitations of existing methods by improving the neural network’s ability to extract reciprocal collision avoidance states in both spatial and temporal dimensions. The spatial state encoder strengthens the extraction of spatial features, while the temporal state encoder improves the handling of time-dependent information. Simulation and physical experiments demonstrate that STR enhances robot navigation in dynamic environments, making it suitable for real-world applications such as multi-robot coordination. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Yuanzhe Wang, Yang Mo, Wei He 0001, Hesheng Wang 0001, Danwei Wang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | History-Enhanced 3D Scene Graph Reasoning From RGB-D Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment, and is considered a prerequisite for long-term autonomous robotic operations. However, building rich representations from RGB-D sequences remains a challenging problem. Existing methods ignore the semantic gap between linguistic and geometric feature spaces or neglect the importance of historical context in incrementally captured data. This limits the learning of visual-textual correspondence and the capability of relationship prediction. To address these problems, we propose a history-enhanced 3D scene graph reasoning framework that incrementally builds a consistent 3D semantic scene graph from an RGB-D image sequence. Specifically, we first introduce a cross-domain unified feature representation module to describe the object instances and their relationships distinctly. Next, we build a one-hot candidate matrix-enabled recurrent mechanism to reason the 3D scene graph, combining the perceived global and local history information. Finally, we design history-aware supervised semantics contrastive learning to optimize the scene-specific global history features. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be available athttps://github.com/cbyan1003/HE-3DSGR. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | RAMPGrasp: Retentive Attention-Based Multiscale Perception Grasp Detection NetworkabstractIn robotic grasp detection, challenges such as uncertainty in object type, size, and placement within the scene diminish grasping accuracy. However, the inability to effectively locate the graspable area and incomplete feature extraction for grasp detection are two key factors that hinder grasp detection accuracy and are not considered in current methods. This paper presents a novel retentive attention-based multiscale perception grasp detection network (RAMPGrasp) to address this constraint. First, we introduce retentive attention in the feature extraction module, which significantly improves the efficiency of attention score computation for long sequences in visual tasks. Second, we propose a multiscale spatial pyramid attention module, which can effectively adjust the importance of multiscale feature sequences and feature channels, while enhancing the correlation of multiscale features. Third, we design the prediction module as a coarse-to-fine framework, improving feature representation for grasp detection by considering the distribution trend of grasp poses. As a result, RAMPGrasp achieves state-of-the-art grasp detection accuracy, with 98.4% and 95.6% on the Cornell and Jacquard datasets, respectively. Jianan Huang 0002, Xuebing Liu, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Lin Chen 0034, Danwei Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Distilling Hierarchical Knowledge From Multimodal Fusion for Unimodal Image SegmentationabstractThe application of multimodal image fusion has become increasingly widespread across various fields in the era of deep learning. Existing fusion methods integrate infrared and visible images to provide complementary content and enhance the robustness of complex real-world scenes for high-level visual tasks, such as semantic segmentation and object detection. In return, high-level visual tasks facilitate the fusion of infrared and visible by providing mid-level semantic information. However, such frameworks rely heavily on multimodal data and require strict registration of images from different modalities before fusion, seriously limiting their practical applications due to the common realistic situations of missing modalities or misregistration. To move beyond this limitation, we propose a novel hierarchical knowledge distillation (HKD) framework tailored for unimodal image segmentation with the guidance of multi-modality. This framework aims to retain as much diverse information from multimodal image fusion as possible, thereby enhancing downstream high-level visual tasks when only the unimodal images are available during the inference phase. Our proposed method is two-stage, and we construct a robust multimodal fusion and segmentation interaction network in the first stage as a powerful teacher model. In the second stage, we design a hierarchical distillation method to transfer the fused and segmented multi-layer knowledge from the multimodal teacher model to the unimodal student model. Extensive experimental results on two public datasets, i.e., MFNet and FMB, demonstrate that the proposed hierarchical knowledge distillation framework can effectively transfuse multimodal knowledge into the unimodal student model for image enhancement and segmentation under incomplete multimodal conditions, and achieves considerably competitive results compared to multimodal image fusion and segmentation models. Weisheng Dong, Shuaibo Wang, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Discriminative Correspondence Estimation for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a fundamental task for estimating the rigid transformation matrix between two point clouds, and is regarded as a prerequisite for downstream vision tasks. Recent works have sought to address the registration problem using the obtainable RGB-D sequence, rather than relying solely on point clouds, which may not always be available. However, most existing unsupervised RGB-D point cloud registration works struggle to obtain fine-grained, robust, discriminative correspondences due to the simple concatenation of multimodal features and the increase in vector dimensions. These methods typically follow a common paradigm: extracting features from the input data, estimating correspondences, and obtaining the transformation matrix through geometric fitting. In this work, we design a generative feature extraction module to fully leverage multimodal information, and seek a novel perspective for correspondence estimation which expands the points in the source and target point clouds into hyperrectangle-based embeddings and considers their inner relationships, based on intersections in n-dimensional space, as the basis for estimating correspondences. Each hyperrectangle-based embedding is built upon the natural and discriminative semantics from the proposed generative feature extraction module, which involves a diffusion branch, a geometric branch, and point-pixel fusion. We harness the capability of the generative model to fully leverage the information from both complementary modalities in RGB-D frames. Furthermore, this distinctive geometry space allows for efficient calculation of intersection volumes and model conditional probabilistics for estimating correspondences. Extensive experiments on the 3DMatch and ScanNet datasets show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be released at:https://github.com/cbyan1003/DCE. Chenbo Yan, Mingtao Feng, Yulan Guo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power SceneabstractIn the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | A State Space Model for Multiobject Full 3-D Information Estimation From RGB-D ImagesabstractVisual understanding of 3-D objects is essential for robotic manipulation, autonomous navigation, and augmented reality. However, existing methods struggle to perform this task efficiently and accurately in an end-to-end manner. We propose a single-shot method based on the state space model (SSM) to predict the full 3-D information (pose, size, shape) of multiple 3-D objects from a single RGB-D image in an end-to-end manner. Our method first encodes long-range semantic information from RGB and depth images separately and then combines them into an integrated latent representation that is processed by a modified SSM to infer the full 3-D information in two separate task heads within a unified model. A heatmap/detection head predicts object centers, and a 3-D information head predicts a matrix detailing the pose, size and latent code of shape for each detected object. We also propose a shape autoencoder based on the SSM, which learns canonical shape codes derived from a large database of 3-D point cloud shapes. The end-to-end framework, modified SSM block and SSM-based shape autoencoder form major contributions of this work. Our design includes different scan strategies tailored to different input data representations, such as RGB-D images and point clouds. Extensive evaluations on the REAL275, CAMERA25, and Wild6D datasets show that our method achieves state-of-the-art performance. On the large-scale Wild6D dataset, our model significantly outperforms the nearest competitor, achieving 2.6% and 5.1% improvements on the IOU-50 and 5°10 cm metrics, respectively. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Jian Liu 0014, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Cybern. | 4 |
| 2025 | Uncertainty Guided Deep Lucas-Kanade Homography for Multimodal Image AlignmentabstractHomography estimation for multimodal images poses a considerable challenge in computer vision because of content disparities and the diverse feature points captured by different sensors. Existing methods typically extract feature maps using neural networks and apply the Lucas-Kanade (LK) algorithm, which is based on the brightness constancy assumption, to solve the homography matrix. However, applying this assumption across all pixel features in multimodal images can lead to inaccuracies, as these images often contain noise, such as homogeneous regions or considerable appearance variations, which can corrupt the network’s training. To address this problem, we propose an uncertainty-guided deep LK (UG-DLK) framework that integrates uncertainty predictions to enhance the network’s iterative learning process. Specifically, we employ a probabilistic approach where the network predicts the distribution of the feature map rather than fixed values. By designing an uncertainty neighborhood estimator, we unfold the cost volume along the channels into 2-D slices, allowing the model to focus on neighborhood information at specific locations, effectively reducing the interference from spatial neighborhoods in the estimation of feature uncertainty. Through uncertainty modeling, the network can accurately identify scenes and objects that comply with the brightness constancy constraint, leading to more robust learning outcomes. Additionally, we introduce a novel loss function that incorporates feature uncertainty, leading to a smoother optimization landscape near the true homography parameters and reducing convergence oscillations. Our method, which is evaluated on benchmark datasets such as Google Maps, Google Earth, MSCOCO, and DPDN, demonstrates state-of-the-art performance, confirming the robustness and adaptability of our model across various scenarios. Zhen Zhou 0003, Jianqiao Luo, Qing Zhu 0003, Yaonan Wang 0001, Hang Zhong, Mingtao Feng, Lin Chen 0034 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Locally Aware Visual State Space for Small Defect Segmentation in Complex Component ImagesabstractSegmenting small defects within large imaging fields remains challenging in industrial scenarios due to the difficulty in distinguishing defects from complex component backgrounds and identifying defects comprising only a few pixels in high-resolution images. To address these issues, we propose a novel dual-branch feature extraction architecture, the locally aware visual state space block, which captures global contextual information while maintaining locally aware perception. In addition, we introduce the parallel quad-directional scanning fusion module to extract multiscale information, aggregating high-level features at different scales for enhanced global information fusion. To avoid losing small target details when upsampling the global segmentation mask to high-resolution input size, we develop progressive location refinement modules to incrementally refine small defect localization from the bottom up. Extensive experiments on our proposed small defect segmentation dataset and a public PCB dataset demonstrate that our method outperforms existing state-of-the-art methods in both performance and efficiency. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Incomplete Modalities Restoration via Hierarchical Adaptation for Robust Multimodal SegmentationabstractMultimodal semantic segmentation has significantly advanced the field of semantic segmentation by integrating data from multiple sources. However, this task often encounters missing modality scenarios due to challenges such as sensor failures or data transmission errors, which can result in substantial performance degradation. Existing approaches to addressing missing modalities predominantly involve training separate models tailored to specific missing scenarios, typically requiring considerable computational resources. In this paper, we propose a Hierarchical Adaptation framework to Restore Missing Modalities for Multimodal segmentation (HARM3), which enables frozen pretrained multimodal models to be directly applied to missing-modality semantic segmentation tasks with minimal parameter updates. Central to HARM3 is a text-instructed missing modality prompt module, which learns multimodal semantic knowledge by utilizing available modalities and textual instructions to generate prompts for the missing modalities. By incorporating a small set of trainable parameters, this module effectively facilitates knowledge transfer between high-resource domains and low-resource domains where missing modalities are more prevalent. Besides, to further enhance the model's robustness and adaptability, we introduce adaptive perturbation training and an affine modality adapter. Extensive experimental results demonstrate the effectiveness and robustness of HARM3 across a variety of missing modality scenarios. Weisheng Dong, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2025 | Exploring Hierarchical Spatial Layout Cues for 3D Point Cloud Based Scene Graph Predictionabstract3D scene graph prediction is important for intelligent agents to gather information and perceive semantics of their environments. However, constructing an effective graph is nontrivial given the complexity of natural scenes. Existing solutions for graph representation of 3D scenes still distinguish each detailed discrepancy among all the relationships as flat thinking, ignoring the mechanism used by humans to perform this task. Inspired by the role of the prefrontal cortex in hierarchical reasoning, we analyze this problem from a novel perspective: exploring hierarchical spatial layout cues in 3D space and navigating that hierarchy to make the 3D scene graph more accurate in a vertical division to horizontal propagation strategy. To this end, we first encode the contextual object features for fine-gained object category classification. Next, we build a bottom-up hierarchical graph to predict remarkably diverse support relationships in a single concept regardless of numerous irrelevant relationships. Finally, equipped with the spatially-true and semantically-meaningful support relationships, we focus on the local region layout to propagate the semantic features to predict the additional non-support relationships under the guidance of the given referred hierarchical graph nodes. Experiments on the challenging 3DSSG benchmark show that our algorithm outperforms existing state-of-the-art, and can also alleviate the impact of the long-tailed distribution of training data. Our code is available athttps://github.com/HHrEtvP/HSLC-3DSG/. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Multim. | 1 |
| 2025 | Category-Level Multi-Object 9D State Tracking Using Object-Centric Multi-Scale Transformer in Point Cloud StreamabstractCategory-level object pose estimation and tracking has achieved impressive progress in computer vision, augmented reality, and robotics. Existing methods either estimate the object states from a single observation or only track the 6-DoF pose of a single object. In this paper, we focus on category-level multi-object 9-Dimensional (9D) state tracking from the point cloud stream. We propose a novel 9D state estimation network to estimate the 6-DoF pose and 3D size of each instance in the scene. It uses our devised multi-scale global attention and object-level local attention modules to obtain representative latent features to estimate the 9D state of each object in the current observation. We then integrate our network estimation into a Kalman filter to combine previous states with the current estimates and achieve multi-object 9D state tracking. Experiment results on two public datasets show that our method achieves state-of-the-art performance on both category-level multi-object state estimation and pose tracking tasks. Furthermore, we directly apply the pre-trained model of our method to our air-ground robot system with multiple moving objects. Experiments on our collected real-world dataset show our method's strong generalization ability and real-time pose tracking performance. Yaonan Wang 0001, Mingtao Feng, Huimin Lu 0002, Xieyuanli Chen |
IEEE Trans. Multim. | 3 |
| 2025 | Pixel-Level Noise Mining for Weakly Supervised Salient Object DetectionabstractTraining a deep model for visual saliency detection requires the collection and labor-intensive annotation of overwhelmingly large data. We propose to learn saliency detection in a weakly supervised manner from single noisy label, which is easy to obtain from unsupervised handcrafted feature-based methods. However, deep networks tend to overfit such noises leading to a dramatic drop in accuracy. Given our goal, we address a natural question: can we identify outliers during network prediction and rectify the label noises? To this end, we propose a pixel-level noise mining framework for robust salient object detection (SOD) by exploiting its own knowledge, and without the need for external models. Specifically, during the early training stage, we progressively identify the outliers from a novel perspective during saliency detection, before the network overfits to the noisy labels, and generate a selection matrix in each iteration. Next, we adaptively rectify the label noises under the guidance of the selection matrix for better supervision in the later training stage. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our method showing its ability to learn saliency detection comparable to state-of-the-art fully supervised methods. Furthermore, our approach outperforms existing weakly supervised methods utilizing single noisy label and surpasses the half of existing weakly supervised methods employing multiple noisy labels. Our approach, which trains with multiple noisy labels, outperforms all other methods employing multiple noisy labels across four major datasets. Furthermore, we also evaluate the generalization ability of our method on the multiclass semantic segmentation (SS) task. Our code is available at https://github.com/kendongdong/NoiseMining. Kendong Liu, Mingtao Feng, Wei Zhao 0019, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Streamabstract3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However, the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper, we aim to handle a real-time multi-task for 6-DoF pose tracking of unknown objects, leveraging 3D-language pre-training scheme from a series of 3D point cloud video streams, while simultaneously performing 3D shape reconstruction in current observation. To this end, we present a generic Language-to-4D modeling paradigm termed L4D-Track, that tackles zero-shot 6-DoF Tracking and shape reconstruction by learning pairwise implicit 3D representation and multi-level multi-modal alignment. Our method constitutes two core parts. 1) Pairwise Implicit 3D Space Representation, that establishes spatial-temporal to language coherence descriptions across continuous 3D point cloud video. 2) Language-to-4D Association and Contrastive Alignment, enables multi-modality semantic connections between 3D point cloud video and language. Our method trained exclusively on public NOCS-REAL275 dataset, achieves promising results on both two publicly benchmarks. This not only shows powerful generalization performance, but also proves its remarkable capability in zero-shot inference. The project is released at L4D- Track. Yaonan Wang 0001, Mingtao Feng, Yulan Guo, Ajmal Mian, Zheng Shou 0001 |
CVPR | 3 |
| 2024 | External Knowledge Enhanced 3D Scene Generation from Sketch
Mingtao Feng, Yaonan Wang 0001, He Xie, Weisheng Dong, Bo Miao, Ajmal Mian |
ECCV (6) | 2 |
| 2024 | Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent FeatureabstractGeneralizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techniques, self-supervised learning methods, and the generative adversarial network framework to train policy neural networks with enhanced generalizability. However, current methods, upon extracting domain-general latent features, further utilize these features to train the reinforcement learning policy, resulting in a decline in the performance of the learned policy guiding the agent to accomplish tasks. To tackle these challenges, a framework of self-expert imitation with purifying latent features was devised, empowering the policy to achieve robust and stable zero-shot generalization performance in visually similar domains previously unseen, without diminishing the performance of guiding the agent to accomplish tasks. The extraction method of domain-general latent features is proposed to enhance their quality based on the variational autoencoder. Extensive experiments have shown that our policy, compared with state-of-the-art counterparts, does not diminish the performance of the policy guiding the agent to accomplish tasks after generalization. Lin Chen 0034, Jianan Huang 0002, Zhen Zhou 0003, Yaonan Wang 0001, Yang Mo, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Danwei Wang |
IROS | 8 |
| 2024 | Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement LearningabstractNavigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidance in multi-robot environments, comprising both static and dynamic obstacles, resulting in inadequate safety and efficiency in directing robot movement. In this study, we introduce a novel policy neural network called Spatial-Temporal RetNet (STR), designed to encode reciprocal collision avoidance states between robots in spatial and temporal dimensions. The goal is to improve the safety and efficacy of the policy neural network in directing robots to complete assigned tasks. The spatial state encoder module is built upon a parallel RetNet structure, which strengthens the neural network's capacity in extracting reciprocal collision avoidance states between robots in spatial dimensions. This module addresses the limitations of position encoding in transformer-based multi-robot navigation policy neural networks. We design a temporal state encoder utilizing a recurrent RetNet structure. This innovation bolsters the multi-robot navigation policy neural network's capability to capture features in the temporal dimension of multi-robot movements. It addresses the limitations of transformer-based multi-robot navigation policy neural networks, particularly in recurrently inferring information across time dimensions. Simulation experiments were conducted to showcase the superior safety and effectiveness of our proposed method compared to previous state-of-the-art approaches in guiding robots to accomplish tasks. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Yuanzhe Wang, Yang Mo, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang |
IROS | 4 |
| 2024 | Referring Human Pose and Mask Estimation In the WildabstractWe introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Bo Miao, Mingtao Feng, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
NeurIPS | 2 |
| 2024 | Toward Safe Distributed Multi-Robot Navigation Coupled With Variational Bayesian ModelabstractDesigning a safe and effective collision avoidance policy for multiple robots is essential in decentralized scenarios, where each robot is responsible for generating its own paths, to ensure their safe operation. Recently, the utilization of reinforcement learning to develop decentralized policies that enable multiple robots to move cooperatively and accomplish tasks has yielded positive outcomes. However, the presence of exploration unsafe actions during the reinforcement learning training process results in inadequate safety. We seek to enhance the safety of distributed multi-robot navigation policies and propose a new imitation learning framework based on the variational Bayesian model, which enables robots to learn safe actions by anticipating the subsequent state they are expected to reach. In addition, a new policy neural network structure for multi-robot navigation is proposed by introducing the transformer structure, which encodes the significance of nearby robots in relation to their forthcoming conditions. Experiments demonstrated that our policy can more safely guide robots to navigate in multi-robot environments under conditions of limited information, outperforming the state-of-the-art RL-RVO method in terms of success rate.Note to Practitioners—The motivation of this paper is to address the problem of collision avoidance in a multi-robot environment under limited information, which can also be applied to autonomous driving, crowd simulation, and other related fields. Positive outcomes have been observed in the utilization of reinforcement learning to create decentralized policies that enable multiple robots to move cooperatively and complete tasks. However, inadequate safety remains a challenging task due to the possibility of exploring hazardous actions during training. This article aims to enhance the safety of distributed policies guiding robots to accomplish navigation tasks in dynamic multi-robot environments. To begin with, we introduce a novel framework for imitation learning that is based on the variational Bayesian model. This framework facilitates the learning of safe actions by the policy to improve its performance and guide the robot in navigating and avoiding obstacles more securely. A loss function is proposed that enables the anticipation of the future state expected to be reached by the robot. By incorporating the transformer structure, a new neural network structure is designed for multi-robot navigation that encodes the significance of nearby robots concerning their upcoming conditions. This network structure employs a BiGRUs to facilitate the assimilation of observations from multiple agents by the policy. Compared to existing works such as GA3C-CADRL, SARL, and RL-RVO, our proposed method achieves a higher success rate. In our future research, we will investigate methods to enhance the policy’s performance in guiding robots to complete tasks by focusing on improving travel time and average speed, while also strictly ensuring safe navigation. Furthermore, we plan to extend this approach by addressing navigation challenges in more densely populated multi-robot environments. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | 3D Object Detection From Point Cloud via Voting Step Diffusionabstract3D object detection is a fundamental task in scene understanding. Numerous research efforts have been dedicated to better incorporate Hough voting into the 3D object detection pipeline. However, due to the noisy, cluttered, and partial nature of real 3D scans, existing voting-based methods often receive votes from the partial surfaces of individual objects together with severe noises, leading to sub-optimal detection performance. In this work, we focus on the distributional properties of point clouds and formulate the voting process as generating new points in the high-density region of the distribution of object centers. To achieve this, we propose a new method to move random 3D points toward the high-density region of the distribution by estimating the score function of the distribution with a noise conditioned score network. Specifically, we first generate a set of object center proposals to coarsely identify the high-density region of the object center distribution. To estimate the score function, we perturb the generated object center proposals by adding normalized Gaussian noise, and then jointly estimate the score function of all perturbed distributions. Finally, we generate new votes by moving random 3D points to the high-density region of the object center distribution according to the estimated score function. Extensive experiments on two large scale indoor 3D scene datasets, SUN RGB-D and ScanNet V2, demonstrate the superiority of our proposed method. The code will be released athttps://github.com/HHrEtvP/DiffVote. Haoran Hou, Mingtao Feng, Weisheng Dong, Qing Zhu 0003, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Unsupervised Homography Estimation With Pixel-Level SVDDabstractHomography estimation is a common image alignment method. Unsupervised learning, which uses unlabeled training and exhibits excellent performance, has attracted much attention in this field. When there are multiple planes in the scene, using features over the entire image for matching will lead to compromised results. However, existing methods for learning focused principal plane masks through deep neural networks lack explicit guidance. In this paper, we propose a novel unsupervised method to explicitly model anomaly descriptor removal and mask generation. Specifically, reliable feature descriptors are selected from a novel perspective, and regard the features that are not responsible for alignment as outliers. The pixel-level support vector data description (PL-SVDD) module is designed. This module learns the feature representation of image pixels and fits a hypersphere to exclude the feature redundancy information that is not responsible for alignment from the hypersphere, thereby optimizing the feature descriptor. Based on the optimized image features, a correlation learning (CL) module is designed. This module displays a generated mask through mathematical modeling to select reliable areas for homography estimation. Specifically, the feature descriptor of one unaligned images is modeled as a multivariate Gaussian distribution by Gaussian density estimation (GDE). Then, The Mahalanobis distance is combined with the multivariate Gaussian distribution of the model and the feature descriptor of another image to generate the mask. Experiments show that our method achieves good performance compared with previous methods. Zhen Zhou 0003, Qing Zhu 0003, Mingtao Feng, Yaonan Wang 0001, Jianqiao Luo, Zhiqiang Miao, Lin Chen 0034, Yang Mo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | A Depth Adaptive Feature Extraction and Dense Prediction Network for 6-D Pose Estimation in Robotic GraspingabstractEstimating the 6-D pose of an object is a vital and challenging task for robot vision systems in industrial robotic grasping. With the wide use of 3-D cameras, the additional acquired depth image provides geometric information of the scene to increase the pose estimation performance but leads to a challenge, fully leveraging the two-modal data, the color image and the depth image. Previous works usually adopt two individual strategies to handle the data, which suffer from limited accuracy and efficiency since the two complementary data are not fully explored. Thus, we propose a depth adaptive feature extraction and dense prediction network that decouples the scale-dependent and the scale-invariant information from the depth image. The former guides the network to perceive the 3-D structure of the scene, and the latter, together with color image, provides the scene textures for feature extraction. The proposed network not only fuses multimodal textures but also retains their 3-D structure. In addition, a dense prediction strategy is adopted to regress the object pose; this approach can mitigate the instability caused by outliers. We conduct various evaluations on a real-world industrial dataset to illustrate the advantages of the proposed approach; and a practical robotic grasping platform is presented to demonstrate its application performance. Xuebing Liu, Xiaofang Yuan, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | A Systematic Point Cloud Edge Detection Framework for Automatic Aircraft Skin MillingabstractThe edge detection technique is an essential step for aircraft skin milling in aviation manufacturing. Most of the current detection methods focus on traditionally defined edge extraction tasks but disregard the crucial systematic requirement of edge milling. In this article, we proposed a novel edge detection framework for automatic edge milling of aircraft skins. First, an edge probability detector is proposed by the spatial tangent continuity to provide the essential reference. Second, we propose a hierarchical branch searching method to hierarchically strip the desired milling edges from the raw point cloud, which consists of the following three graded progressive steps: branch backbone generation, branch extension, and branch pruning. We demonstrate the performance of the proposed method on both synthetic models and aircraft skin workpieces. The proposed method outperforms the other baselines and shows accurate edges for the edge milling task. Yaonan Wang 0001, He Xie, Mingtao Feng, Haotian Wu 0002, Chao Ding 0006, Ajmal Mian |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | PoseDiffusion: A Coarse-to-Fine Framework for Unseen Object 6-DoF Pose EstimationabstractAccurately estimating the six-degrees of freedom (DoF) pose of unseen objects is crucial for successful robotic manipulation in industrial automation. Some existing methods for this task rely on prior knowledge of individual objects, i.e., the model must be trained on the exact object instance or object category. Others perform unseen object pose estimation but are limited in their feature learning and pose refinement ability. To address these problems, we propose an unseen object pose estimation method that follows a coarse-to-fine framework and leverages the powerful learning ability of diffusion models. We introduce a diffusion model for generating object poses, and conduct a comparison between the generated poses and the original pose to determine the optimal one. We design a novel pose estimation module to provide coarse poses for the PoseDiffusion. This module comprises two feature extraction modules that extract global and masked features. In addition, we propose a strategy to estimate the pose by comparing the similarity between rendered and query poses. The renderings of an unseen object from various viewpoints are generated from its computer-aided design (CAD) model. Our method requires a CAD model of the unseen object only during inference, a scenario well suited to industrial applications. Experimental evaluation on benchmark datasets demonstrates that the proposed framework outperforms existing approaches, achieving state-of-the-art performance in six-DoF object pose estimation. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Chengzhong Wu, Xuebing Liu, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | 3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point CloudabstractIn-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide variety of challenging relationships, existing methods perform quite poorly with limited training samples. In this work, we find that the inherently hierarchical structures of physical space in 3D scenes aid in the automatic association of semantic and spatial arrangements, specifying clear patterns and leading to less ambiguous predictions. Thus, they well meet the challenges due to the rich variations within scene categories. To achieve this, we explicitly unify these structural cues of 3D physical spaces into deep neural networks to facilitate scene graph prediction. Specifically, we exploit an external knowledge base as a baseline to accumulate both contextualized visual content and textual facts to form a 3D spatial multimodal knowledge graph. Moreover, we propose a knowledge-enabled scene graph prediction module benefiting from the 3D spatial knowledge to effectively regularize semantic space of relationships. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art competitors. Our code is available at https://github.com/HHrEtvP/SMKA. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Ajmal Mian |
CVPR | 1 |
| 2023 | Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationabstractDiffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently ambiguous and lack details. In this paper, we propose a sketch and text guided probabilistic diffusion model for colored point cloud generation that conditions the denoising process jointly with a hand drawn sketch of the object and its textual description. We incrementally diffuse the point coordinates and color values in a joint diffusion process to reach a Gaussian distribution. Colored point cloud generation thus amounts to learning the reverse diffusion process, conditioned by the sketch and text, to iteratively recover the desired shape and color. Specifically, to learn effective sketch-text embedding, our model adaptively aggregates the joint embedding of text prompt and the sketch based on a capsule attention network. Our model uses staged diffusion to generate the shape and then assign colors to different parts conditioned on the appearance prompt while preserving precise shapes from the first stage. This gives our model the flexibility to extend to multiple tasks, such as appearance re-editing and part segmentation. Experimental results demonstrate that our model outperforms recent state-of-the-art in point cloud generation. Yaonan Wang 0001, Mingtao Feng, He Xie, Ajmal Mian |
ICCV | 3 |
| 2023 | Position and structure-aware graph learning
Guoqiang Ye, Juan Song, Mingtao Feng, Guangming Zhu 0001, Peiyi Shen, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 3 |
| 2023 | Exploring Spatio-Temporal Graph Convolution for Video-Based Human-Object Interaction RecognitionabstractVideo-based human-object interaction recognition is a challenging task since the state of objects as well as their correlations change constantly in the video. Existing methods mainly use 3DCNN or use separate components (e.g., GCN + RNN) to model the spatial correlation or the temporal correlation respectively, but ignore modeling spatio-temporal correlations simultaneously and long-term temporal dynamics of objects. In this paper, we propose a novel model, named Spatio-Temporal Interaction Graph Parsing Networks (STIGPN), for human-object interaction recognition in videos. STIGPN captures both spatial and temporal correlations simultaneously and thus can capture intra-frame and inter-frame dependencies efficiently and effectively. To model long-term temporal dynamics of objects, we introduce spatio-temporal feature enhancement, which can improve the detection of the salient human-object interaction pairs. We explore three types of spatio-temporal graph convolutions to simultaneously capture the spatio-temporal correlations and assess their effectiveness as the basic building block of STIGPN. Extensive experiments on CAD-120, Something-Else and Charades datasets show that our proposed solution leads to competitive results compared with the state-of-the-art methods. Code for STIGPN is available at:https://github.com/NingWang2049/STIGPN2 Ning Wang 0047, Guangming Zhu 0001, Hongsheng Li 0003, Mingtao Feng, Lan Ni, Peiyi Shen, Lin Mei 0001, Liang Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Transformer-Based Imitative Reinforcement Learning for Multirobot Path PlanningabstractMultirobot path planning leads multiple robots from start positions to designated goal positions by generating efficient and collision-free paths. Multirobot systems realize coordination solutions and decentralized path planning, which is essential for large-scale systems. The state-of-the-art decentralized methods utilize imitation learning and reinforcement learning methods to teach fully decentralized policies, dramatically improving their performance. However, these methods cannot enable robots to perform tasks efficiently in relatively dense environments without communication between robots. We introduce the transformer structure into policy neural networks for the first time, dramatically enhancing the ability of policy neural networks to extract features that facilitate collaboration between robots. It mainly focuses on improving the performance of policies in relatively dense multirobot environments under conditions where robots do not communicate with each other. Furthermore, a novel imitation reinforcement learning framework is proposed by combining contrastive learning and double deep Q-network to solve the problem of difficulty training policy neural networks after introducing the transformer structure. We present results in the simulation environment and compare the resulting policy against advanced multirobot path-planning methods in terms of success rate. Simulation results show that our policy achieves state-of-the-art performance when there is no communication between robots. Finally, we experimented with a real-world case using a total of three robots in our robotic laboratory. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Yang Mo, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Learning from Pixel-Level Noisy Label : A New Perspective for Light Field Saliency DetectionabstractSaliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obtained from unsupervised hand crafted featured-based saliency methods. Given this goal, a natural question is: can we efficiently incorporate the relationships among light field cues while identifying clean labels in a unified framework? We address this question by formulating the learning as a joint optimization of intra light field features fusion stream and inter scenes correlation stream to generate the predictions. Specially, we first introduce a pixel forgetting guided fusion module to mutually enhance the light field features and exploit pixel consistency across iterations to identify noisy pixels. Next, we introduce a cross scene noise penalty loss for better reflecting latent structures of training data and enabling the learning to be invariant to noise. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our framework showing that it learns saliency prediction comparable to state-of-the-art fully supervised light field saliency methods. Our code is available at h t tps://github.com/ OLobbCode/NoiseLF. Mingtao Feng, Kendong Liu, Liang Zhang 0010, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
CVPR | 1 |
| 2022 | ICK-Track: A Category-Level 6-DoF Pose Tracker Using Inter-Frame Consistent Keypoints for Aerial ManipulationabstractRobots that are supposed to interact with or manipulate objects in the world must be able to track the poses of objects in their sensor data. Thus, Detecting and tracking the 6-DoF poses of targeted objects is important for aerial manipulation and is still in the early stage due to the high dynamics and limited onboard capacity of such systems. In this paper, we propose ICK-Track, a novel method for onboard category-level object 6-DoF pose tracking that can be applied to aerial manipulation without using any pre-defined object CAD models. It first utilizes a semi-supervised video segmentation to detect objects in the eye-in-hand RGB-D camera stream to segment the 3D points of objects. Then, canonical keypoints are extracted using iterative farthest point sampling. We propose a novel inter-frame consistent keypoints generation network to generate the corresponding keypoint pairs, which are used together with ICP to estimate the pose changes of objects for tracking. Experimental results show that our method is more robust to viewpoint changes and runs faster than the state-of-the-art methods on category-level pose tracking. We further test our proposed method on a real aerial manipulator. A demo video showing the use of our method on a real aerial manipulator and the implementation of our method are available at: https://github.com/S-JingTao/ICK-Track. Yaonan Wang 0001, Mingtao Feng, Danwei Wang, Jiawen Zhao, Cyrill Stachniss, Xieyuanli Chen |
IROS | 3 |
| 2021 | Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloudabstract3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a freeform language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of point clouds. There are three main challenges in 3D object grounding: to find the main focus in the complex and diverse description; to understand the point cloud scene; and to locate the target object. In this paper, we address all three challenges. Firstly, we propose a language scene graph module to capture the rich structure and long-distance phrase correlations. Secondly, we introduce a multi-level 3D proposal relation graph module to extract the object-object and object-scene co-occurrence relationships, and strengthen the visual features of the initial proposals. Lastly, we develop a description guided 3D visual graph module to encode global contexts of phrases and proposals by a nodes matching strategy. Extensive experiments on challenging benchmark datasets (ScanRefer [3] and Nr3D [42]) show that our algorithm outperforms existing state-of-the-art. Our code is available at https://github.com/PNXD/FFL-3DOG. Mingtao Feng, Liang Zhang 0010, Guangming Zhu 0001, Hui Zhang 0023, Yaonan Wang 0001, Ajmal Mian |
ICCV | 1 |
| 2021 | Relation Graph Network for 3D Object Detection in Point CloudsabstractConvolutional Neural Networks (CNNs) have emerged as a powerful tool for object detection in 2D images. However, their power has not been fully realised for detecting 3D objects directly in point clouds without conversion to regular grids. Moreover, existing state-of-the-art 3D object detection methods aim to recognize objects individually without exploiting their relationships during learning or inference. In this article, we first propose a strategy that associates the predictions of direction vectors with pseudo geometric centers, leading to a win-win solution for 3D bounding box candidates regression. Secondly, we propose point attention pooling to extract uniform appearance features for each 3D object proposal, benefiting from the learned direction features, semantic features and spatial coordinates of the object points. Finally, the appearance features are used together with the position features to build 3D object-object relationship graphs for all proposals to model their co-existence. We explore the effect of relation graphs on proposals' appearance feature enhancement under supervised and unsupervised settings. The proposed relation graph network comprises a 3D object proposal generation module and a 3D relation module, making it an end-to-end trainable network for detecting 3D objects in point clouds. Experiments on challenging benchmark point cloud datasets (SunRGB-D, ScanNet and KITTI) show that our algorithm performs better than existing state-of-the-art. Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Liang Zhang 0010, Ajmal Mian |
IEEE Trans. Image Process. | 1 |
| 2020 | Point attention network for semantic segmentation of 3D point clouds
Mingtao Feng, Liang Zhang 0010, Xuefei Lin, Syed Zulqarnain Gilani, Ajmal Mian |
Pattern Recognit. | 1 |
| 2020 | Small Object Augmentation of Urban Scenes for Real-Time Semantic SegmentationabstractSemantic segmentation is a key step in scene understanding for autonomous driving. Although deep learning has significantly improved the segmentation accuracy, current highquality models such as PSPNet and DeepLabV3 are inefficient given their complex architectures and reliance on multi-scale inputs. Thus, it is difficult to apply them to real-time or practical applications. On the other hand, existing real-time methods cannot yet produce satisfactory results on small objects such as traffic lights, which are imperative to safe autonomous driving. In this paper, we improve the performance of real-time semantic segmentation from two perspectives, methodology and data. Specifically, we propose a real-time segmentation model coined Narrow Deep Network (NDNet) and build a synthetic dataset by inserting additional small objects into the training images. The proposed method achieves 65.7% mean intersection over union (mIoU) on the Cityscapes test set with only 8.4G floatingpoint operations (FLOPs) on 1024×2048 inputs. Furthermore, by re-training the existing PSPNet and DeepLabV3 models on our synthetic dataset, we obtained an average 2% mIoU improvement on small objects. Zhengeng Yang, Hongshan Yu, Mingtao Feng, Wei Sun 0028, Xuefei Lin, Mingui Sun, Zhi-Hong Mao, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2018 | 3D Face Reconstruction from Light Field Images: A Model-Free Approach
Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Ajmal Mian |
ECCV (10) | 1 |
| 2018 | Benchmark Data Set and Method for Depth Estimation From Light Field ImagesabstractConvolutional Neural Networks (CNN) have performed extremely well for many image analysis tasks. However, supervised training of deep CNN architectures requires huge amounts of labelled data which is unavailable for light field images. In this paper, we leverage on synthetic light field images and propose a two stream CNN network that learns to estimate the disparities of multiple correlated neighbourhood pixels from their Epipolar Plane Images (EPI). Since the EPIs are unrelated except at their intersection, a two stream network is proposed to learn convolution weights individually for the EPIs and then combine the outputs of the two streams for disparity estimation. The CNN estimated disparity map is then refined using the central RGB light field image as a prior in a variational technique. We also propose a new real world dataset comprising light field images of 19 objects captured with the Lytro Illum camera in outdoor scenes and their corresponding 3D pointclouds, as ground truth, captured with the 3dMD scanner. This dataset will be made public to allow more precise 3D pointcloud level comparison of algorithms in the future which is currently not possible. Experiments on the synthetic and real world datasets show that our algorithm outperforms existing state-of-the-art for depth estimation from light field images. Mingtao Feng, Yaonan Wang 0001, Jian Liu 0014, Liang Zhang 0010, Hasan Firdaus M. Zaki, Ajmal Mian |
IEEE Trans. Image Process. | 1 |