Mingmin Zhen

dblp:166/2746 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-8180-1023ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 52% Segmentation and scene understanding · 30% Video understanding and tracking · 9%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.232020
Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts · CVPR 2020
Multi-view based neural network for semantic segmentation on 3D scenes · Sci. China Inf. Sci. 2019
Learning Fully Dense Neural Networks for Image Semantic Segmentation · AAAI 2019
Computer vision › 3D vision
feature matching
0.922022
ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer · ECCV (32) 2022
Learning and Matching Multi-View Descriptors for Registration of Point Clouds · ECCV (15) 2018
Computer vision › 3D vision › feature matching › dense feature matching
detector-free matching
0.612022
ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer · ECCV (32) 2022
Computer vision › 3D vision
3d face reconstruction
0.412020
Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency · ECCV (15) 2020
Computer vision › Segmentation and scene understanding
boundary detection
0.412020
Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts · CVPR 2020
Computer vision › 3D vision
camera pose estimation
0.412020
KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering · CVPR 2020
Computer vision › 3D vision › visual localization
camera relocalization
0.412020
KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering · CVPR 2020
Computer vision › Segmentation and scene understanding
edge detection
0.412020
JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds · ECCV (20) 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian filtering
kalman filtering
0.412020
KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering · CVPR 2020
Computer vision › 3D vision
multi-view geometry
0.412020
Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency · ECCV (15) 2020
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.412020
JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds · ECCV (20) 2020
Computer vision › 3D vision › visual localization
scene coordinate regression
0.412020
KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering · CVPR 2020
Computer vision › Segmentation and scene understanding › boundary detection
semantic boundary detection
0.412020
Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts · CVPR 2020
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation
0.412020
Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation · ECCV (27) 2020
Computer vision › Video understanding and tracking
video object segmentation
0.412020
Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation · ECCV (27) 2020
Geometric modeling and processing
3d reconstruction
0.412020
Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction · ECCV (15) 2020
Geometric modeling and processing
bundle adjustment
0.412020
Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction · ECCV (15) 2020
Computer vision › Segmentation and scene understanding › 3d segmentation
3d scene segmentation
0.412019
Multi-view based neural network for semantic segmentation on 3D scenes · Sci. China Inf. Sci. 2019
Machine learning › Deep learning architectures and training
encoder-decoder architecture
0.412019
Learning Fully Dense Neural Networks for Image Semantic Segmentation · AAAI 2019
Computer vision › 3D vision › 3d shape representation › 3d shape representation learning
mesh representation learning
0.412019
Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface · CVPR 2019
Geometric modeling and processing › mesh processing › mesh signal processing
mesh convolution
0.412019
Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface · CVPR 2019
Computer vision › 3D vision
point cloud registration
0.312018
Learning and Matching Multi-View Descriptors for Registration of Point Clouds · ECCV (15) 2018
Computer vision › 3D vision
3d scene understanding
0.112019
Multi-view based neural network for semantic segmentation on 3D scenes · Sci. China Inf. Sci. 2019

Methods — techniques the papers use, named apart from their topics

transformer · 0.6attention · 0.6stochastic optimization · 0.4spatial gradient fusion · 0.4self-supervised learning · 0.4multi-task learning · 0.4kalman filtering · 0.4iterative pyramid context · 0.4discriminative feature learning · 0.4conditional random field · 0.4bayesian learning · 0.4texture map parameterization · 0.4cross-atlas convolution · 0.4
YearPublicationVenuePosition
2022 ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer
Zixin Luo, Lei Zhou 0011, Yurun Tian, Mingmin Zhen, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (32)5
2020 Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts
abstract
In this paper, we present a joint multi-task learning framework for semantic segmentation and boundary detection. The critical component in the framework is the iterative pyramid context module (PCM), which couples two tasks and stores the shared latent semantics to interact between the two tasks. For semantic boundary detection, we propose the novel spatial gradient fusion to suppress non-semantic edges. As semantic boundary detection is the dual task of semantic segmentation, we introduce a loss function with boundary consistency constraint to improve the boundary pixel accuracy for semantic segmentation. Our extensive experiments demonstrate superior performance over state-of-the-art works, not only in semantic segmentation but also in semantic boundary detection. In particular, a mean IoU score of 81.8% on Cityscapes test set is achieved without using coarse data or any external data for semantic segmentation. For semantic boundary detection, we improve over previous state-of-the-art works by 9.9% in terms of AP and 6.8% in terms of MF(ODS).
Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Shiwei Li 0001, Tianwei Shen, Jiaxiang Shang, Tian Fang, Long Quan
CVPR1
2020 KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering
abstract
Temporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the state-of-the-art one-shot approaches in terms of accuracy. In this work, we improve the temporal relocalization method by using a network architecture that incorporates Kalman filtering (KFNet) for online camera relocalization. In particular, KFNet extends the scene coordinate regression problem to the time domain in order to recursively establish 2D and 3D correspondences for the pose determination. The network architecture design and the loss formulation are based on Kalman filtering in the context of Bayesian learning. Extensive experiments on multiple relocalization benchmarks demonstrate the high accuracy of KFNet at the top of both one-shot and temporal relocalization approaches.
Lei Zhou 0011, Zixin Luo, Tianwei Shen, Mingmin Zhen, Yao Yao 0008, Tian Fang, Long Quan
CVPR5
2020 JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds
Zeyu Hu, Mingmin Zhen, Xuyang Bai, Hongbo Fu 0001, Chiew-Lan Tai
ECCV (20)2
2020 Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency
Jiaxiang Shang, Tianwei Shen, Shiwei Li 0001, Lei Zhou 0011, Mingmin Zhen, Tian Fang, Long Quan
ECCV (15)5
2020 Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation
Mingmin Zhen, Shiwei Li 0001, Lei Zhou 0011, Jiaxiang Shang, Haoan Feng, Tian Fang, Long Quan
ECCV (27)1
2020 Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction
Lei Zhou 0011, Zixin Luo, Mingmin Zhen, Tianwei Shen, Shiwei Li 0001, Zhuofei Huang, Tian Fang, Long Quan
ECCV (15)3
2019 Learning Fully Dense Neural Networks for Image Semantic Segmentation
abstract
Semantic segmentation is pixel-wise classification which retains critical spatial information. The “feature map reuse” has been commonly adopted in CNN based approaches to take advantage of feature maps in the early layers for the later spatial reconstruction. Along this direction, we go a step further by proposing a fully dense neural network with an encoderdecoder structure that we abbreviate as FDNet. For each stage in the decoder module, feature maps of all the previous blocks are adaptively aggregated to feedforward as input. On the one hand, it reconstructs the spatial boundaries accurately. On the other hand, it learns more efficiently with the more efficient gradient backpropagation. In addition, we propose the boundary-aware loss function to focus more attention on the pixels near the boundary, which boosts the “hard examples” labeling. We have demonstrated the best performance of the FDNet on the two benchmark datasets: PASCAL VOC 2012, NYUDv2 over previous works when not considering training on other datasets.
Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Tian Fang, Long Quan
AAAI1
2019 Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface
abstract
We present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the parameterization of texture map is not pre-determined, and depends on the surface topologies, we therefore introduce a novel cross-atlas convolution to recover the original mesh geodesic neighborhood, so as to achieve the invariance property to arbitrary parameterization. The proposed module is integrated into classification and segmentation architectures, which takes the input texture map of a mesh, and infers the output predictions. Our method not only shows competitive performances on classification and segmentation public benchmarks, but also paves the way for the broad mesh surfaces learning.
Shiwei Li 0001, Zixin Luo, Mingmin Zhen, Yao Yao 0008, Tianwei Shen, Tian Fang, Long Quan
CVPR3
2019 Multi-view based neural network for semantic segmentation on 3D scenes
Yonghua Lu, Mingmin Zhen, Tian Fang
Sci. China Inf. Sci.2
2018 Learning and Matching Multi-View Descriptors for Registration of Point Clouds
Lei Zhou 0011, Siyu Zhu 0001, Zixin Luo, Tianwei Shen, Mingmin Zhen, Tian Fang, Long Quan
ECCV (15)6
2016 Regional Subspace Projection Coding for Image Retrieval
abstract
For image retrieval task, hamming embedding, being proved to be one of the state-of-the-art methods, has been prevalently utilised. The basic idea is to project local features into orthogonal space randomly, in which the binary signature is generated based on a single partition of feature space. However, the binary signature generation process is coarse and heuristic. On the one hand, the same projection is carried out for all visual word space without consideration of difference among subspaces. On the other hand, the projection matrix is generated randomly regardless of the distribution of feature data. Therefore, the performance of hamming embedding is limited and far from the optimal. In this paper, we firstly analyse the limitation of hamming em- bedding and compare different orthogonal projection methods. Then we propose a regional subspace projection coding method that is based on the distribution of local features assigned to each visual word. Finally, our experiments on two benchmark datasets demonstrate that our proposed method outperforms current state-of-the-art methods.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
ICMR1
2015 Improved cluster center adaption for image classification
abstract
The feature coding algorithm, “Vector of Locally Aggregated Descriptors (VLAD)”, can be used effectively for large scale object instance retrieval. Despite its effectiveness and excellent performance, the existence of ambiguous cluster centers can reduce the performance. Though an idea to this problem has been proposed, it is not practical in fact. In this paper, we analyze possible situations that cause effect on the results and propose a novel approach to improve the VLAD method. The proposed method mainly focuses on the similarity measure between each two images. For each two images, we adapt the original cluster center to VLAD vectors. As we illustrate, our method has promising results with small vocabulary size on both datasets of 15 Scenes and VOC2007.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
ICIP1
2015 Improving VLAD with regional PCA whitening
abstract
In recent yeas, VLAD has been used to represent an image effectively and efficiently by just a few bytes in large-scale image retrieval. In spite of its remarkable performance, a series of modification methods have been presented. In addition, the redundancy between the features corresponding to the same cluster center could be improved. In this paper, a regional PCA Whitening method is proposed to decorrelate the features and reduce the dimensionality for each cluster with the consideration of mapping the descriptor into high dimensionality explicitly. Our method can also be embedded into original VLAD pipeline with global PCA very well. The experimental results on both Holidays and UKbench dataset show that our approach improves VLAD significantly.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
VCIP1