VLDB 2026 Research / reviewers in the wild / expert
Mengjie Xu
dblp:211/5571
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust piglet counting in crowded environments via small trajectory fusion and multi-line counting module
Mengjie Xu, Hailing Wu |
Vis. Comput. | 1 |
| 2025 | MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body ReconstructionabstractMultiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenges posed by self-occlusion under multiple views and the continuity of human body shape estimation. In this study, we introduce a method to reconstruct the 3D human body from multiple uncalibrated camera views. Initially, we utilize a pre-trained human body encoder to process each camera view individually, enabling the reconstruction of human body models and parameters for each view along with predicted camera positions. Rather than merely averaging the models across views, we develop a neural network trained to assign weights to individual views for all human body joints, based on the estimated distribution of joint distances from each camera. Additionally, we focus on the mesh surface of the human body for dynamic fusion, allowing for the seamless integration of facial expressions and body shape into a unified human body model. Our method has shown excellent performance in reconstructing the human body on two public datasets, advancing beyond previous work from the SMPL model to the SMPL-X model. This extension incorporates more complex hand poses and facial expressions, enhancing the detail and accuracy of the reconstructions. Crucially, it supports the flexible ad-hoc deployment of any number of cameras, offering significant potential for various applications. Yitao Zhu, Sheng Wang 0014, Mengjie Xu, Zixu Zhuang, Zhixin Wang, Kaidong Wang, Han Zhang 0002, Qian Wang 0001 |
AAAI | 3 |
| 2025 | MITracker: Multi-View Integration for Visual Object TrackingabstractMulti-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progress has been limited by the lack of comprehensive multi-view datasets and effective cross-view integration methods. To overcome these limitations, we compiled a Multi-View object Tracking (MVTrack) dataset of 234K high-quality annotated frames featuring 27 distinct objects across various scenes. In conjunction with this dataset, we introduce a novel MVOT method, Multi-View Integration Tracker (MITracker), to efficiently integrate multi-view object features and provide stable tracking outcomes. MI-Tracker can track any object in video frames of arbitrary length from arbitrary viewpoints. The key advancements of our method over traditional single-view approaches come from two aspects: (1) MITracker transforms 2D image features into a 3D feature volume and compresses it into a bird’s eye view (BEV) plane, facilitating inter-view information fusion; (2) we propose an attention mechanism that leverages geometric information from fused 3D feature volume to refine the tracking results at each view. MI-Tracker outperforms existing methods on the MVTrack and GMTD datasets, achieving state-of-the-art performance. The code and the new dataset will be available at mii-laboratory.github.io/MITracker. 1 Mengjie Xu, Yitao Zhu, Jiaming Li 0012, Zhenrong Shen 0001, Sheng Wang 0014, Haolin Huang, Han Zhang 0002, Qian Wang 0001 |
CVPR | 1 |
| 2025 | DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction ModuleabstractSpeech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that optimizes model performance by fine-tuning. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks. Haolin Huang, Yu Fang 0008, Mengjie Xu, Qian Wang 0001 |
ICASSP | 5 |
| 2025 | Multi-task Screening for Cervical Diseases via Feature Routing and Asymmetric Distillation
Haolin Huang, Jiangdong Cai, Mengjie Xu, Zhenrong Shen 0001, Manman Fei, Lichi Zhang, Qian Wang 0001 |
MICCAI (14) | 4 |
| 2025 | Med-LEGO: Editing and Adapting Toward Generalist Medical Image Diagnosis
Yitao Zhu, Jiaming Li 0012, Mengjie Xu, Zihao Zhao 0002, Honglin Xiong, Sheng Wang 0014, Qian Wang 0001 |
MICCAI (6) | 4 |
| 2024 | HLSFNet: Hybrid Long and Short-term collaboration based on Feature extraction for power generation forecasting in ChinaabstractAccurate prediction of power generation plays a crucial role in the planning of power resources. However, most current power generation prediction models are primarily focused on short-term, single-type scenarios, lacking for extracting both long and short-term information and features necessary for multi-type, medium and long-term forecasting. To address these limitations, this paper firstly clusters the regions, and then proposes Hybrid Long and Short-term based on Feature extraction (HLSFNet) for multi-type power generation forecasting in China. Our model combines SimpleRNN (Simple Recurrent Neural Network) with an appropriate number of layers of GRU (Gated Recurrent Unit) to capture short-term correlations and long-term dependencies. Furthermore, we incorporate a one-dimensional convolutional layer without pooling to effectively extract static spatial features. The addition of self-attention enhances the applicability of predictions across multiple variables. Experimental results show that HLSFNet outperforms multiple baseline models in terms of prediction accuracy for hydropower, thermal power, and wind power, as well as three types of regional power generation. Through ablation experiments, we have proven that the added modules (attention mechanism, convolutional layer, hybrid RNN series) have good effectiveness in improving prediction accuracy. Furthermore, we have carried out auxiliary analysis of features, and improved the analysis framework of power generation. Mengjie Xu, Chuanwang Sun |
CSCWD | 1 |
| 2023 | Beamforming design for max-min SINR in RIS-based hybrid relayingabstractAbstract Reconfigurable intelligent surface (RIS) technology is regarded as one of the important technologies for future 6G mobile communication networks. However, RIS can only passively reflect signals, rather than actively process received signals like the traditional relay. Therefore, a hybrid decode‐and‐forward (DF) relay and RIS‐assisted multiple‐input single‐output (MISO) multiuser downlink communication system is considered to combine the advantages of both. Then the beamforming at the base station (BS), the DF relay, and the phase vector of the RIS elements is jointly designed, aiming at achieving max‐min signal‐to‐interference‐plus‐noise ratio (SINR). Due to the non‐convexity caused by coupling between the beamforming vectors, the optimization problem is split into three semidefinite relaxation (SDR) sub‐problems, and then the alternate optimization method is adopted to solve the non‐convex problem. The numerical results verify the effectiveness and superiority of the proposed optimization algorithm in improving system performance compared with other benchmark schemes. Mengjie Xu, Minghe Mao, Tianhe Li |
IET Commun. | 1 |