Mengyin Liu

dblp:286/0386 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Reconstructing Sparse-View Indoor Scenes in View Space With Global Monocular Prior Alignment
abstract
Although 3D Gaussian Splatting (3DGS) has greatly advanced the development of novel view synthesis (NVS) and surface reconstruction tasks, it still faces serious challenges in indoor scenes with only sparse views due to the lack of initialization point cloud. Existing GS-like methods are geometrically initialized using point clouds from Structure from Motion (SfM), and their performance is heavily limited by the point cloud quality. However, in indoor scenes containing large areas of textureless or weakly textured regions, the SfM algorithm struggles to generate point clouds in these areas, especially when only a limited number of input views are available. This situation can cause the GS optimization process to overfit the photometric error, resulting in severe geometric degradation in regions lacking proper initialization. To address the above problem, we propose a novel sparse-view 3DGS method using global scale-aligned monocular depth to initialize Gaussian primitives and optimize them in view space, named VGA-GS. First, we design a global scale-aligned monocular depth initialization strategy, which provides geometric priors in textureless or weakly textured regions while avoiding the local geometric distortions in vanilla alignments. Second, we propose a novel Gaussian primitive representation based on view space and a refinement strategy based on pixel errors, which can constrain each primitive to always be within the field of view where it is observed to be fully optimized. Finally, we extract the surface prior from monocular depth to regularize the rendering depth and normal, and further propagate the constraints of the training view to the neighboring pseudo-views via image warping. Our method achieves state-of-the-art reconstruction performance in both textureless regions and local details for indoor scenes with sparse views. Moreover, it is highly scalable, supporting integration with various 3DGS rasterizers and depth estimation methods. We release the code at https://github.com/XT5un/VGA-GS.
Xiaotian Sun 0005, Qingshan Xu 0001, Mengyin Liu, Cheng Wang 0003
IEEE Trans. Circuits Syst. Video Technol.3
2025 OSS-OCL: Occlusion Scenario Simulation and Occluded-edge Concentrated Learning for pedestrian detection
Keqi Lu, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
Pattern Recognit. Lett.3
2025 HA-FGOVD: Highlighting Fine-Grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) models are considered to be Large Multi-modal Models (LMM), due to their extensive training data and a large number of parameters. Mainstream OVD models prioritize object coarse-grained category rather than focus on their fine-grained attributes, e.g., colors or materials, thus failed to identify objects specified with certain attributes. Despite being pretrained on large-scale image-text pairs with rich attribute information, their latent feature space does not highlight these fine-grained attributes. In this paper, we introduce HA-FGOVD, a universal and explicit method that enhances the attribute-level detection capabilities of frozen OVD models by highlighting fine-grained attributes in explicit linear space. Our approach uses a LLM to extract attribute words in input text as a zero-shot task. Then, token attention masks are adjusted to guide text encoders in extracting both global and attribute-specific features, which are explicitly composited as two vectors in linear space to form a new attribute-highlighted feature for detection tasks. The composition weight scalars can be learned or transferred across different OVD models, showcasing the universality of our method. Experimental results show that HA-FGOVD achieves state-of-the-art performance on the FG-OVD benchmark and demonstrates promising generalization on the OVDEval benchmark, suggesting that our method addresses significant limitations in fine-grained attribute detection and has potential for broader fine-grained detection applications.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
IEEE Trans. Multim.2
2024 Unsupervised Multi-view Pedestrian Detection
abstract
With the prosperity of the intelligent surveillance, multiple cameras have been applied to localize pedestrians more accurately. However, previous methods rely on laborious annotations of pedestrians in every frame and camera view. Therefore, we propose in this paper an Unsupervised Multi-view Pedestrian Detection approach (UMPD) to learn an annotation-free detector via vision-language models and 2D-3D cross-modal mapping: 1) Firstly, Semantic-aware Iterative Segmentation (SIS) is proposed to extract unsupervised representations of multi-view images, which are converted into 2D masks as pseudo labels, via our proposed iterative PCA and zero-shot semantic classes from vision-language models; 2) Secondly, we propose Geometry-aware Volume-based Detector (GVD) to end-to-end encode multi-view 2D images into a 3D volume to predict voxel-wise density and color via 2D-to-3D geometric projection, trained by 3D-to-2D rendering losses with SIS pseudo labels; 3) Thirdly, for better detection results, i.e., the 3D density projected on Birds-Eye-View, we propose Vertical-aware BEV Regularization (VBR) to constrain pedestrians to be vertical like the natural poses. Extensive experiments on popular multi-view pedestrian detection benchmarks Wildtrack, Terrace, and MultiviewX, show that our proposed UMPD, as the first fully-unsupervised method to our best knowledge, performs competitively to the previous state-of-the-art supervised methods. Code is available at https://github.com/lmy98129/UMPD.
Mengyin Liu, Chao Zhu 0003, Shiqi Ren, Xu-Cheng Yin
ACM Multimedia1
2023 VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision
abstract
Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects often lead to wrong detections, and small scale or heavily occluded pedestrians are easily missed due to their unusual appearances. To address these challenges, only object regions are inadequate, thus how to fully utilize more explicit and semantic contexts becomes a key problem. Meanwhile, previous context-aware pedestrian detectors either only learn latent contexts with visual clues, or need laborious annotations to obtain explicit and semantic contexts. Therefore, we propose in this paper a novel approach via Vision-Language semantic self-supervision for context-aware Pedestrian Detection (VLPD) to model explicitly semantic contexts without any extra annotations. Firstly, we propose a self-supervised Vision-Language Semantic (VLS) segmentation method, which learns both fully-supervised pedestrian detection and contextual segmentation via self-generated explicit labels of semantic classes by vision-language models. Furthermore, a self-supervised Prototypical Semantic Contrastive (PSC) learning method is proposed to better discriminate pedestrians and other classes, based on more explicit and semantic contexts obtained from VLS. Extensive experiments on popular benchmarks show that our proposed VLPD achieves superior performances over the previous state-of-the-arts, particularly under challenging circumstances like small scale and heavy occlusion. Code is available at https://github.com/lmy98129/VLPD.
Mengyin Liu, Jie Jiang 0015, Chao Zhu 0003, Xu-Cheng Yin
CVPR1
2023 Towards Discriminative Semantic Relationship for Fine-grained Crowd Counting
abstract
As an extended task of crowd counting, fine-grained crowd counting aims to estimate the number of people in each semantic category instead of the whole in an image, and faces challenges including 1) inter-category crowd appearance similarity, 2) intra-category crowd appearance variations, and 3) frequent scene changes. In this paper, we propose a new fine-grained crowd counting approach named DSR to tackle these challenges by modeling Discriminative Semantic Relationship, which consists of two key components: Word Vector Module (WVM) and Adaptive Kernel Module (AKM). The WVM introduces more explicit semantic relationship information to better distinguish people of different semantic groups with similar appearance. The AKM dynamically adjusts kernel weights according to the features from different crowd appearance and scenes. The proposed DSR achieves superior results over state-of-the-art on the standard dataset. Our approach can serve as a new solid baseline and facilitate future research for the task of fine-grained crowd counting.
Shiqi Ren, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
ICME3
2023 Feature Implicit Enhancement via Super-Resolution for Small Object Detection
Zhehao Xu, Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
PRCV (12)2
2022 CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts Modeling
abstract
Image-text retrieval is an essential task of information retrieval, in which the models with the Vision-and-Language Pretraining(VLP) are able to achieve ideal accuracy compared with the ones without VLP. Among different VLP approaches, the single-stream models achieve the overall best retrieval accuracy, but slower inference speed. Recently, researchers have introduced the two-stage retrieval setting commonly used in the information retrieval field to the single-stream VLP model for a better accuracy/efficiency trade-off. However, the retrieval accuracy and efficiency are still unsatisfactory mainly due to the limitations of the patch-based visual unimodal encoder in these VLP models. The unimodal encoders are trained on pure visual data, so the visual features extracted by them are difficult to align with the textual features and it is also difficult for the multi-modal encoder to understand visual information. Under these circumstances, we propose an accurate and efficient two-stage image-text retrieval model via Contrastive Alignment and visual Contexts modeling(CAliC). In the first stage of the proposed model, the visual unimodal encoder is pretrained with cross-modal contrastive learning to extract easily aligned visual features, which improves the retrieval accuracy and the inference speed. In the second stage of the proposed model, we introduce a new visual contexts modeling task during pretraining to help the multi-modal encoder better understand the visual information and get more accurate predictions. Extensive experimental evaluation validates the effectiveness of our proposed approach, which achieves a higher retrieval accuracy while keeping a faster inference speed, and outperforms existing state-of-the-art retrieval methods on image-text retrieval tasks over Flickr30K and COCO benchmarks.
Chao Zhu 0003, Mengyin Liu, Weibo Gu, Hongfa Wang, Wei Liu 0005, Xu-Cheng Yin
ACM Multimedia3
2021 Adaptive Pattern-Parameter Matching for Robust Pedestrian Detection
abstract
Pedestrians with challenging patterns, e.g. small scale or heavy occlusion, appear frequently in practical applications like autonomous driving, which remains tremendous obstacle to higher robustness of detectors. Although plenty of previous works have been dedicated to these problems, properly matching patterns of pedestrian and parameters of detector, i.e., constructing a detector with proper parameter sizes for certain pedestrian patterns of different complexity, has been seldom investigated intensively. Pedestrian instances are usually handled equally with the same amount of parameters, which in our opinion is inadequate for those with more difficult patterns and leads to unsatisfactory performance. Thus, we propose in this paper a novel detection approach via adaptive pattern-parameter matching. The input pedestrian patterns, especially the complex ones, are first disentangled into simpler patterns for detection head by Pattern Disentangling Module (PDM) with various receptive fields. Then, Gating Feature Filtering Module (GFFM) dynamically decides the spatial positions where the patterns are still not simple enough and need further disentanglement by the next-level PDM. Cooperating with these two key components, our approach can adaptively select the best matched parameter size for the input patterns according to their complexity. Moreover, to further explore the relationship between parameter sizes and their performance on the corresponding patterns, two parameter selection policies are designed: 1) extending parameter size to maximum, aiming at more difficult patterns for different occlusion types; 2) specializing parameter size by group division, aiming at complex patterns for scale variations. Extensive experiments on two popular benchmarks, Caltech and CityPersons, show that our proposed method achieves superior performance compared with other state-of-the-art methods on subsets of different scales and occlusion types.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
AAAI1