Mengyang Pu

dblp:228/1348 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Co-MixPL: An optimized semi-supervised learning method for tunnel water leakage detection
Xujie Long, Jing Teng, Shaobo Zhao, Mengyang Pu, Ruifeng Shi, Jonathan Li 0001, Guoqing Jing
Adv. Eng. Informatics5
2026 Intention-behavior consistency-based automated failure attribution for LLM-driven multi-agent systems
Hua Wu 0002, Wanhao Zheng, Xiaojing Bai, Mengyang Pu, Li Sun 0008, Yuhan Xu
Expert Syst. Appl.5
2026 Motion and Spatiotemporal Aggregation Network for Occlusion Edge Detection From Videos
abstract
Detecting occlusion edges from videos is a critical yet under-explored task, with recent works focusing on single-image occlusion edge detection while ignoring dynamic patterns in videos. In videos, occlusion edges typically occur when a moving object occludes either the background or another object, resulting in two types of occlusion edges:object-background (OB) edgesthat are characterized by appearance contrast and motion, and object-object (OO) edgesthat face ambiguity. Inspired by these observations, we propose a novel Motion and Spatio-Temporal Aggregated Network (MaSTAN) and treat the two edge types differently for more effectively detecting occlusion edges in videos. Specifically, we first extract spatial semantics and motion patterns from the videos and propose a novel Temporal Feature Propagation module (TFP) for temporal cue aggregation. Next, we put forward a Dual-branch Gated-attention Decoder (DG-Decoder) to generate edge-specific features for predicting the final occlusion edge maps. Extensive experiments on the OVIS-OE benchmark, the first large benchmark dedicated to video occlusion edge detection, demonstrate that MaSTAN achieves state-of-the-art performance, significantly advancing the capability of occlusion edge detection in video. The source code and benchmark will be made publicly available.
Mengyang Pu, Xiaohui Hou, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.2
2025 RINDNet++: Edge Detection for Discontinuity in Reflectance, Illumination, Normal, and Depth
Mengyang Pu, Qingji Guan, Haibin Ling
Int. J. Comput. Vis.1
2024 MuGE: Multiple Granularity Edge Detection
abstract
Edge segmentation is well-known to be subjective due to personalized annotation styles and preferred granular-ity. However, most existing deterministic edge detection methods produce only a single edge map for one input image. We argue that generating multiple edge maps is more reasonable than generating a single one considering the subjectivity and ambiguity of the edges. Thus motivated, in this paper we propose multiple granularity edge detection, called MuGE, which can produce a wide range of edge maps, from approximate object contours to fine texture edges. Specifically, we first propose to design an edge granularity network to estimate the edge granularity from an individual edge annotation. Subsequently, to guide the generation of diversified edge maps, we integrate such edge granularity into the multi-scale feature maps in the spatial domain. Meanwhile, we decompose the feature maps into low-frequency and high-frequency parts, where the encoded edge granularity is further fused into the high-frequency part to achieve more precise control over the details of the produced edge maps. Compared to previous methods, MuGE is able to not only generate multiple edge maps at different controllable granularities but also achieve a com-petitive performance on the BSDS500 and Multicue benchmark datasets.
Caixia Zhou, Mengyang Pu, Qingji Guan, Ruoxi Deng, Haibin Ling
CVPR3
2024 Learning Temporal Cues for Fine-Grained Action Recognition
abstract
Fine-grained action recognition aims to classify complex coarse actions into finer ones. It suffers due to varying temporal scales and subtle differences between categories, which brings significant challenges to the temporal perception capabilities of current models. In this paper, we proposed a novel Temporal Cues Transformer (TCT) to guide the recognition of fine-grained action by exploiting detailed temporal cues in video duration and action sequence. The proposed TCT consists of a duration-aware encoder and a hierarchical sequence aggregation decoder. In the encoder, we extract duration-aware representations from the video and its duration. In the decoder, we reinforce the learning of action sequence by first searching the fine-grained elements of action with a hierarchical element query module, then aggregating the elements with a sequence aggregation module to predict the action category. Extensive experiments on widely used Diving48 and FineGym datasets demonstrate the superiority of the proposed method compared to the state-of-the-art methods.
Mengyang Pu, Chao Deng 0002, Junlan Feng
ICIP5
2023 The Treasure Beneath Multiple Annotations: An Uncertainty-Aware Edge Detector
abstract
Deep learning-based edge detectors heavily rely on pixel-wise labels which are often provided by multiple annotators. Existing methods fuse multiple annotations using a simple voting process, ignoring the inherent ambiguity of edges and labeling bias of annotators. In this paper, we propose a novel uncertainty-aware edge detector (UAED), which employs uncertainty to investigate the subjectivity and ambiguity of diverse annotations. Specifically, we first convert the deterministic label space into a learnable Gaussian distribution, whose variance measures the degree of ambiguity among different annotations. Then we regard the learned variance as the estimated uncertainty of the predicted edge maps, and pixels with higher uncertainty are likely to be hard samples for edge detection. Therefore we design an adaptive weighting loss to emphasize the learning from those pixels with high uncertainty, which helps the network to gradually concentrate on the important pixels. UAED can be combined with various encoder-decoder backbones, and the extensive experiments demonstrate that UAED achieves superior performance consistently across multiple edge detection benchmarks. The source code is available at https://github.com/ZhouCX117/UAED.
Caixia Zhou, Mengyang Pu, Qingji Guan, Haibin Ling
CVPR3
2022 EDTER: Edge Detection with Transformer
abstract
Convolutional neural networks have made significant progresses in edge detection by progressively exploring the context and semantic features. However, local details are gradually suppressed with the enlarging of receptive fields. Recently, vision transformer has shown excellent capability in capturing long-range dependencies. Inspired by this, we propose a novel transformer-based edge detector, Edge Detection TransformER (EDTER), to extract clear and crisp object boundaries and meaningful edges by exploiting the full image context information and detailed local cues simultaneously. EDTER works in two stages. In Stage I, a global transformer encoder is used to capture long-range global context on coarse-grained image patches. Then in Stage II, a local transformer encoder works on fine-grained patches to excavate the short-range local cues. Each transformer encoder is followed by an elaborately designed Bi-directional Multi-Level Aggregation decoder to achieve high-resolution features. Finally, the global context and local cues are combined by a Feature Fusion Module and fed into a decision head for edge prediction. Extensive experiments on BSDS500, NYUDv2, and Multicue demonstrate the superiority of EDTER in comparison with state-of-the-arts. The source code is available at https://github.com/MengyangPu/EDTER.
Mengyang Pu, Qingji Guan, Haibin Ling
CVPR1
2022 Learning From Pixel-Level Label Noise: A New Perspective for Semi-Supervised Semantic Segmentation
abstract
This paper addresses semi-supervised semantic segmentation by exploiting a small set of images with pixel-level annotations (strong supervisions) and a large set of images with only image-level annotations (weak supervisions). Most existing approaches aim to generate accurate pixel-level labels from weak supervisions. However, we observe that those generated labels still inevitably contain noisy labels. Motivated by this observation, we present a novel perspective and formulate this task as a problem of learning with pixel-level label noise. Existing noisy label methods, nevertheless, mainly aim at image-level tasks, which can not capture the relationship between neighboring labels in one image. Therefore, we propose a graph-based label noise detection and correction framework to deal with pixel-level noisy labels. In particular, for the generated pixel-level noisy labels from weak supervisions by Class Activation Map (CAM), we train a clean segmentation model with strong supervisions to detect the clean labels from these noisy labels according to the cross-entropy loss. Then, we adopt a superpixel-based graph to represent the relations of spatial adjacency and semantic similarity between pixels in one image. Finally we correct the noisy labels using a Graph Attention Network (GAT) supervised by detected clean labels. We comprehensively conduct experiments on PASCAL VOC 2012, PASCAL-Context, MS-COCO and Cityscapes datasets. The experimental results show that our proposed semi-supervised method achieves the state-of-the-art performances and even outperforms the fully-supervised models on PASCAL VOC 2012 and MS-COCO datasets in some cases.
Rumeng Yi, Qingji Guan, Mengyang Pu, Runsheng Zhang
IEEE Trans. Image Process.4
2021 RINDNet: Edge Detection for Discontinuity in Reflectance, Illumination, Normal and Depth
abstract
As a fundamental building block in computer vision, edges can be categorised into four types according to the discontinuity in surface-Reflectance, Illumination, surface-Normal or Depth. While great progress has been made in detecting generic or individual types of edges, it remains under-explored to comprehensively study all four edge types together. In this paper, we propose a novel neural network solution, RINDNet, to jointly detect all four types of edges. Taking into consideration the distinct attributes of each type of edges and the relationship between them, RINDNet learns effective representations for each of them and works in three stages. In stage I, RINDNet uses a common backbone to extract features shared by all edges. Then in stage II it branches to prepare discriminative features for each edge type by the corresponding decoder. In stage III, an independent decision head for each type aggregates the features from previous stages to predict the initial results. Additionally, an attention module learns attention maps for all types to capture the underlying relations between them, and these maps are combined with initial results to generate the final edge detection results. For training and evaluation, we construct the first public benchmark, BSDS-RIND, with all four types of edges carefully annotated. In our experiments, RINDNet yields promising results in comparison with state-of-the-art methods. Additional analysis is presented in supplementary material.
Mengyang Pu, Qingji Guan, Haibin Ling
ICCV1
2021 Shadow Removal by a Lightness-Guided Network With Training on Unpaired Data
abstract
Shadow removal can significantly improve the image visual quality and has many applications in computer vision. Deep learning methods based on CNNs have become the most effective approach for shadow removal by training on either paired data, where both the shadow and underlying shadow-free versions of an image are known, or unpaired data, where shadow and shadow-free training images are totally different with no correspondence. In practice, CNN training on unpaired data is more preferred given the easiness of training data collection. In this paper, we present a new Lightness-Guided Shadow Removal Network (LG-ShadowNet) for shadow removal by training on unpaired data. In this method, we first train a CNN module to compensate for the lightness and then train a second CNN module with the guidance of lightness information from the first CNN module for final shadow removal. We also introduce a loss function to further utilise the colour prior of existing data. Extensive experiments on widely used ISTD, adjusted ISTD and USR datasets demonstrate that the proposed method outperforms the state-of-the-art methods with training on unpaired data.
Hui Yin 0002, Yang Mi, Mengyang Pu, Song Wang 0002
IEEE Trans. Image Process.4
2020 Detecting dense text in natural images
abstract
Most existing text detection methods are mainly motivated by deep learning‐based object detection approaches, which may result in serious overlapping between detected text lines, especially in dense text scenarios. It is because text boxes are not commonly overlapped, as different from general objects in natural scenes. Moreover, text detection requires higher localisation accuracy than object detection. To tackle these problems, the authors propose a novel dense text detection network (DTDN) to localise tighter text lines without overlapping. Their main novelties are: (i) propose an intersection‐over‐union overlap loss, which considers correlations between one anchor and GT boxes and measures how many text areas one anchor contains, (ii) propose a novel anchor sample selection strategy, named CMax‐OMin, to select tighter positive samples for training. CMax‐OMin strategy not only considers whether an anchor has the largest overlap with its corresponding GT box (CMax), but also ensures the overlapping between one anchor and other GT boxes as little as possible (OMin). Besides, they train a bounding‐box regressor as post‐processing to further improve text localisation performance. Experiments on scene text benchmark datasets and their proposed dense text dataset demonstrate that the proposed DTDN achieves competitive performance, especially for dense text scenarios.
Dianzhuan Jiang, Shengsheng Zhang, Qi Zou 0001, Xingyuan Zhang, Mengyang Pu
IET Comput. Vis.6
2020 Object Discovery From a Single Unlabeled Image by Mining Frequent Itemsets With Multi-Scale Features
abstract
The goal of our work is to discover dominant objects in a very general setting where only a single unlabeled image is given. This is far more challenge than typical colocalization or weakly-supervised localization tasks. To tackle this problem, we propose a simple but effective pattern mining-based method, called Object Location Mining (OLM), which exploits the advantages of data mining and feature representation of pretrained convolutional neural networks (CNNs). Specifically, we first convert the feature maps from a pre-trained CNN model into a set of transactions, and then discovers frequent patterns from transaction database through pattern mining techniques. We observe that those discovered patterns, i.e., co-occurrence highlighted regions, typically hold appearance and spatial consistency. Motivated by this observation, we can easily discover and localize possible objects by merging relevant meaningful patterns. Extensive experiments on a variety of benchmarks demonstrate that OLM achieves competitive localization performance compared with the state-of-the-art methods. We also evaluate our approach compared with unsupervised saliency detection methods and achieves competitive results on seven benchmark datasets. Moreover, we conduct experiments on finegrained classification to show that our proposed method can locate the entire object and parts accurately, which can benefit to improving the classification results significantly.
Runsheng Zhang, Mengyang Pu, Qingji Guan, Qi Zou 0001, Haibin Ling
IEEE Trans. Image Process.3
2018 GraphNet: Learning Image Pseudo Annotations for Weakly-Supervised Semantic Segmentation
abstract
Weakly-supervised semantic image segmentation suffers from lacking accurate pixel-level annotations. In this paper, we propose a novel graph convolutional network-based method, called GraphNet, to learn pixel-wise labels from weak annotations. Firstly, we construct a graph on the superpixels of a training image by combining the low-level spatial relation and high-level semantic content. Meanwhile, scribble or bounding box annotations are embedded into the graph, respectively. Then, GraphNet takes the graph as input and learns to predict high-confidence pseudo image masks by a convolutional network operating directly on graphs. At last, a segmentation network is trained supervised by these pseudo image masks. We comprehensively conduct experiments on the PASCAL VOC 2012 and PASCAL-CONTEXT segmentation benchmarks. Experimental results demonstrate that GraphNet is effective to predict the pixel labels with scribble or bounding box annotations. The proposed framework yields state-of-the-art results in the community.
Mengyang Pu, Qingji Guan, Qi Zou 0001
ACM Multimedia1