Mingjie Wang 0002

dblp:122/8955-2 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-7346-1110ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mesoscopic Insights: Generative Content-Steering Video Codec for UAVs
abstract
ABSTRACT The proliferation of unmanned aerial vehicles (UAVs) in the low‐altitude economy has driven a surge in demand for high‐fidelity video transmission; however, existing video codec methods struggle to balance the conflicting requirements of resource‐constrained onboard platforms and high‐quality downstream visual perception. While conventional macroscopic codecs offer computational efficiency, they suffer from artefacts and bandwidth congestion in complex UAV environments. Conversely, microscopic neural video compression (NVC) models achieve superior rate‐distortion performance but incur prohibitive computational overhead. To reconcile these limitations, we propose the Generative Content‐steering Video Codec (GCVC), a novel framework designed from a mesoscopic perspective. GCVC synergises lightweight architectural design with generative semantic recovery by integrating streamlined encoding with a Stable Diffusion‐based (SD‐based) frame predictor. By employing a strategic frame‐skipping protocol to alleviate transmission bottlenecks and utilising the generative prior to synthesise missing scenes, our framework achieves state‐of‐the‐art reconstruction fidelity at significantly lower bitrates. Extensive experiments on public and self‐collected UAV datasets demonstrate that GCVC outperforms both traditional standards and cutting‐edge NVC approaches, establishing a new benchmark for efficient, perception‐aware video compression in intelligent UAV systems. The implementation codes will be publicly available at https://github.com/ZSTU‐CV‐Lab/GCVC .
Qincheng Wen, Yande Li, Mingjie Wang 0002
Expert Syst. J. Knowl. Eng.6
2026 GazeCLIP: Enhancing gaze estimation through text-guided multimodal learning
Jun Wang 0089, Hao Ruan, Liangjian Wen, Yong Dai 0001, Mingjie Wang 0002
Neurocomputing5
2026 VLCounting: Taming zero-shot counting via language-driven exemplar grounding
Mingjie Wang 0002, Yong Dai 0001, Eric Buys, Minglun Gong
Pattern Recognit.1
2025 Open-World 3D Scene Understanding with Cross-Modal Dual Consistency Learning
abstract
Large Vision-Language Pre-training models have achieved remarkable advancements in 2D zero-shot or few-shot visual tasks in the 2D domain. However, promoting their potential to benefit the 3D counterparts is full of challenging due to the notable hindrance of limited 3D-text pairs, which results in the open-world 3D scene understanding remaining an unexplored problem. In this paper, we pretrain a cross-modal dual consistency learning based 3D visual-language model to learn semantically-rich 3D point cloud representation. Specifically, we first introduce a Visual Feature Distribution Consistency strategy to bridge the gap between the point clouds and images. Then, a Visual-Semantic Enhancement Feature Distribution Consistency approach is developed to narrow the distance between enhanced visual and language information. Finally, under the supervision of profound knowledge from 2D Vsion-Language Models (VLMs), the learned 3D features achieve powerful generalization capability, facilitating open-world 3D scene understanding. Quantitative and qualitative evaluations on ScanNet and Matterport3D benchmarks demonstrate the effectiveness of our pre-trained method in open-world 3D semantic segmentation.
Xian-Feng Han, Yuhang Wang 0036, Mingjie Wang 0002
ICMR4
2025 UAVMamba: Elevating UAV's Crowd Counting Through a Synergistic Integration of Hybrid CNN and Mamba Paradigms
Longlong Zhu, Song Yuan, Mingjie Wang 0002
PRCV (15)5
2025 High-fidelity 3D Gaussian inpainting: Preserving multi-view consistency and photorealistic details
Jun Zhou 0023, Dinghao Li, Nannan Li 0002, Mingjie Wang 0002
Comput. Graph.4
2025 Fine-grained text and image guided point cloud completion with CLIP model
Jun Zhou 0023, Mingjie Wang 0002, Hongchen Tan, Nannan Li 0002, Xiuping Liu
Neurocomputing3
2025 Enhanced normal estimation of point clouds via fine-grained geometric information learning
Jun Zhou 0023, Mingjie Wang 0002, Nannan Li 0002, Weixiao Wang, Xiuping Liu
Mach. Vis. Appl.3
2025 FER-Former: Multimodal Transformer for Facial Expression Recognition
abstract
The ever-increasing demands for intuitive interactions in virtual reality have led to surging interests in facial expression recognition (FER). There are however several issues commonly seen in existing methods, including narrow receptive fields and homogenous supervisory signals. To address these issues, we propose in this paper a novel multimodal supervision-steering transformer for facial expression recognition in the wild, referred to as FER-former. Specifically, to address the limitation of narrow receptive fields, a hybrid feature extraction pipeline is designed by cascading both prevailing CNNs and transformers. To deal with the issue of homogenous supervisory signals, a heterogeneous domain-steering supervision module is proposed to incorporate text-space semantic correlations to enhance image features, based on the similarity between image and text features. Additionally, a FER-specific transformer encoder is introduced to characterize conventional one-hot label-focusing and CLIP-based text-oriented tokens in parallel for final classification. Based on the collaboration of multifarious token heads, global receptive fields with multimodal semantic cues are captured, delivering superb learning capability. Extensive experiments on popular benchmarks demonstrate the superiority of the proposed FER-former over the existing state-of-the-art methods.
Yande Li, Mingjie Wang 0002, Minglun Gong, Yonggang Lu, Li Liu 0001
IEEE Trans. Multim.2
2025 Draw What You Hear: High-Fidelity Image Generation and Manipulation via SoundAdapter
abstract
Currently, the text-to-image (T2I) generation has established itself as a cornerstone within the realm of AI-generated content (AIGC), due its remarkable success to the availability of extensive datasets comprising paired text-vision samples. Nevertheless, the absence of audio-visual pairs hinders the growth of audio-to-image (A2I). Although prior approaches have pioneered the A2I task, the tight entanglement between initial audio and image encoders imposes the challenge of gathering audio-visual samples, resulting in degraded performance and limited sound flexibility. Therefore, this article proposes a novel SoundAdapter to draw what you hear. Specifically, the SoundAdapter's structure is meticulously designed around transformer blocks, which are critical for capturing overarching patterns and dependencies within the data. In addition, it integrates a sophisticated multigranularity approach coupled with a hybrid supervisory signal, ensuring both fine-grained semantic alignment and seamless optimization across various levels of representation. Extensive tests demonstrate that the SoundAdapter excels in training, setting new benchmarks in zero-shot audio classification, as well as in creating and modifying images across a variety of datasets. The implementation code and several demos supporting this study are openly accessible at https://github.com/CV-MM-Lab/SoundAdapter, facilitating reproducibility and further research.
Mingjie Wang 0002, Song Yuan, Xian-Feng Han, Zili Yi
IEEE Trans. Neural Networks Learn. Syst.1
2024 GCNet: Probing self-similarity learning for Generalized Counting Network
Mingjie Wang 0002, Yande Li, Graham W. Taylor, Minglun Gong
Pattern Recognit.1
2024 Robust point cloud normal estimation via multi-level critical point aggregation
Jun Zhou 0023, Yaoshun Li, Mingjie Wang 0002, Nannan Li 0002, Zhiyang Li 0001, Weixiao Wang
Vis. Comput.3
2023 Dynamic Mixture of Counter Network for Location-Agnostic Crowd Counting
abstract
Crowd counting has attracted increasing attentions in recent years due to its challenges and wide societal applications. Despite persevering efforts made by the research community, most of existing methods require a large amount of location-level annotations. Collecting such type of fine-granularity supervisory signals is extremely time-consuming and labour-intensive, thereby hindering the well generalization of these location-adherent models. To shun this drawback, several pioneering studies open a promising research direction of location-agonistic crowd counting. Albeit the noticeable efforts, they somewhat ignore the merits of diverse learning paradigms and the issue of intractable density shift. To ameliorate these issues, in this paper, a novel Dynamic Mixture of Counter Network (DMCNet) is proposed for location-agnostic crowd counting. Specifically, our DMCNet inherits the hybrid advantages of CNNs (e.g. locality-oriented and pyramidal property) and MLP-based structure (e.g. global receptive fields and light weight). Particularly, the dynamic counter predictor and the mixture of counter heads are delicately designed to hammer at combating huge density shift and overfitting. Extensive experiments demonstrate that our DMCNet attains state-of-the-art performance against existing location-agnostic approaches and performs on par with many conventional location-adherent ones.
Mingjie Wang 0002, Hao Cai 0004, Yong Dai 0001, Minglun Gong
WACV1
2023 Improvement of normal estimation for point clouds via simplifying surface fitting
Jun Zhou 0023, Mingjie Wang 0002, Xiuping Liu, Zhiyang Li 0001
Comput. Aided Des.3
2023 CrowdMLP: Weakly-supervised crowd counting via multi-granularity MLP
Mingjie Wang 0002, Jun Zhou 0001, Hao Cai 0004, Minglun Gong
Pattern Recognit.1
2023 STNet: Scale Tree Network With Multi-Level Auxiliator for Crowd Counting
abstract
State-of-the-art approaches for crowd counting resort to deepneural networks to predict density maps. However, counting people in congested scenes remains a challenging task because the presence of drastic scale variation, density inconsistency, and complex background can seriously degrade their counting accuracy. To battle the ingrained issue of accuracy degradation, in this paper, we propose a novel and powerful network called Scale Tree Network (STNet) for accurate crowd counting. STNet consists of two key components: a Scale-Tree Diversity Enhancer and a Multi-level Auxiliator. Specifically, the Diversity Enhancer is designed to enrich scale diversity, which alleviates limitations of existing methods caused by insufficient level of scales. A novel tree structure is adopted to hierarchically parse coarse-to-fine crowd regions. Furthermore, a simple yet effective Multi-level Auxiliator is presented to aid in exploiting generalisable shared characteristics at multiple levels, allowing more accurate pixel-wise background cognition. The overall STNet is trained in an end-to-end manner, without the needs for manually tuning loss weights between the main and the auxiliary tasks. Extensive experiments on five challenging crowd datasets demonstrate the superiority of the proposed method.
Mingjie Wang 0002, Hao Cai 0004, Xian-Feng Han, Jun Zhou 0023, Minglun Gong
IEEE Trans. Multim.1
2022 Fast and Accurate Normal Estimation for Point Clouds Via Patch Stitching
Jun Zhou 0023, Mingjie Wang 0002, Xiuping Liu, Zhiyang Li 0001
Comput. Aided Des.3
2022 A robust framework for multi-view stereopsis
Wendong Mao, Mingjie Wang 0002, Hui Huang 0004, Minglun Gong
Vis. Comput.2
2021 Interlayer and intralayer scale aggregation for scale-invariant crowd counting
Mingjie Wang 0002, Hao Cai 0004, Jun Zhou 0023, Minglun Gong
Neurocomputing1
2021 Fine-grained talking face generation with video reinterpretation
Xin Huang 0030, Mingjie Wang 0002, Minglun Gong
Vis. Comput.2
2020 Stochastic Multi-Scale Aggregation Network for Crowd Counting
abstract
Crowd counting from unconstrained and congested scenes is an important task in computer vision. Its main difficulties stem from large scale/density variation and prone to over-fitting. This paper presents a novel end-to-end stochastic multi-scale aggregation network (SMANet) which carefully addresses these issues. Specifically, general features are first extracted by the front-end subnetwork and then fed into the back-end subnetwork which consists of stochastic multi-scale aggregation module, density map generator, and global prior encoder. The stochastic aggregation impels the multi-branch units to learn features at different scales effectively and reduces sensitivity to scale variations, whereas the global prior encoder is designed to encode global contextual information and guarantee density consistency of shared representations. Our proposed SMANet is the first work to fuse multi-scale features in a stochastic manner for crowd counting. Experimental results on four public datasets demonstrate that our SMANet consistently outperforms the state-of-the-arts.
Mingjie Wang 0002, Hao Cai 0004, Jun Zhou 0023, Minglun Gong
ICASSP1
2020 ADNet: Adaptively Dense Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) have demonstrated great success in vision tasks. However, most existing architectures still suffer from low feature reuse efficiency. In this paper, we present a layer attention based Adaptively Dense Network (ADNet) by adaptively determining the reuse status of hierarchical preceding features. Specifically, a dense residual aggregation strategy is developed to fuse multi-level internal representations in an effective manner. Furthermore, a novel layer attention mechanism is proposed to explicitly model the interrelationship among layers to automatically adjust the density of the network. It is worth noting that existing ResNets and DenseNets are both special cases of our ADNet. Extensive experiments demonstrate that the proposed architecture consistently and indubitably achieves competitive results in accuracy on benchmark datasets (CIFAR10, CIFAR100, and SVHN), while at the same time remarkably reduces computational costs and memory space. Visualization and analysis on layer-wise attention further provide better understanding on the density of feature reuse in Deep Networks.
Mingjie Wang 0002, Hao Cai 0004, Xin Huang 0030, Minglun Gong
WACV1
2020 No-reference image sharpness assessment based on discrepancy measures of structural degradation
Hao Cai 0004, Mingjie Wang 0002, Wendong Mao, Minglun Gong
J. Vis. Commun. Image Represent.2
2019 Semi-Dense Stereo Matching Using Dual CNNs
abstract
A robust solution for semi-dense stereo matching is presented. It utilizes two CNN models for computing stereo matching cost and performing confidence-based filtering, respectively. Compared to existing CNNs-based matching cost generation approaches, our method feeds additional global information into the network so that the learned model can better handle challenging cases, such as lighting changes and lack of textures. Through utilizing non-parametric transforms, our method is also more self-reliant than most existing semi-dense stereo approaches, which rely highly on the adjustment of parameters. The experimental results based on Middlebury Stereo dataset demonstrate that the proposed approach outperforms the state-of-the-art semi-dense stereo approaches.
Wendong Mao, Mingjie Wang 0002, Jun Zhou 0023, Minglun Gong
WACV2
2019 Multi-Scale Convolution Aggregation and Stochastic Feature Reuse for DenseNets
abstract
Recently, Convolution Neural Networks (CNNs) obtained huge success in numerous vision tasks. In particular, DenseNets have demonstrated that feature reuse via dense skip connections can effectively alleviate the difficulty of training very deep networks and that reusing features generated by the initial layers in all subsequent layers has strong impact on performance. To feed even richer information into the network, a novel adaptive Multi-scale Convolution Aggregation module is presented in this paper. Composed of layers for multi-scale convolutions, trainable cross-scale aggregation, maxout, and concatenation, this module is highly non-linear and can boost the accuracy of DenseNet while using much fewer parameters. In addition, due to high model complexity, the network with extremely dense feature reuse is prone to overfitting. To address this problem, a regularization method named Stochastic Feature Reuse is also presented. Through randomly dropping a set of feature maps to be reused for each mini-batch during the training phase, this regularization method reduces training costs and prevents co-adaptation. Experimental results on CIFAR-10, CIFAR-100 and SVHN benchmarks demonstrated the effectiveness of the proposed methods.
Mingjie Wang 0002, Jun Zhou 0023, Wendong Mao, Minglun Gong
WACV1