EDBT 2026 Demo / reviewers in the wild / expert
Mengjie Hu 0002
dblp:170/4386-2
· DBLP profile ↗
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-7712-3322ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trajectory-Aware Attack: Explainable Adversarial Attack Against Multiple Object TrackersabstractMulti-Object Tracking (MOT) aims to build moving trajectories of objects within video sequences and serves as a critical component in autonomous driving systems. Recently, several studies have revealed the vulnerability of existing MOT methods by investigating adversarial attacks against MOT, raising significant safety concerns for real-world applications. These methods attack trackers by deliberately inserting false alarms, which mislead trajectories to drift from their correct paths. However, current MOT attack methods fail to propose efficient strategies for generating false alarms, as they either rely on computationally intensive optimization to determine the placement of false alarms, or crudely insert a large number of heuristically designed false alarms. In this paper, we propose an explainable and effective false alarm generation module, named Target Generating Module (TGM), that adaptively determines the location and size of false alarms by leveraging historical trajectory information. Based on this module, we design an attack method targeting mainstream MOT approaches, namedTrajectory-Aware Attack (TA Attack). TA Attack achieves effective disruption of MOT systems by combining detection erasure and false alarm generation, requiring only a few frames to successfully compromise trajectories. To exhibit the flexibility and effectiveness of our method, we conduct experiments using four multi-object trackers (ByteTrack, SORT, CenterTrack and FairMOT) which are enabled by two representative detectors (YOLOX and CenterNet). The results demonstrate our method achieves state of the art performance with 74.87% attack success rate on BDD100K, 81.7% attack success rate on MOT17 and 83.87% attack success rate on MOT20 while 4 frames being attacked averagely, revealing the vulnerability of association mechanism in MOT methods. Mengjie Hu 0002, Yufei Ding 0003, Chun Liu 0004, Qing Song 0006 |
IEEE Trans. Multim. | 1 |
| 2025 | Semantic Guided Matting NetabstractAbstract Human matting refers to extracting human parts from natural images with high quality, including human detail information such as hair, glasses, hats, etc. This technology plays an essential role in image synthesis and visual effects in the film industry. When the green screen is not available, the existing human matting methods need the help of additional inputs (such as trimap, background image, etc.), or the model with high computational cost and complex network structure, which brings great difficulties to the application of human matting in practice. To alleviate such problems, we use a segmentation network as the foundation and use multiple branches to achieve human segmentation, contour detail extraction, and information fusion. We also propose a foreground probability map module, which uses the feature maps in the segmentation network to pre-estimate the foreground probabilities of each pixel and obtain Semantic Guided Matting Net. Under the condition that only a single image is needed as the input, the human matting task can be realized by making full use of the semantic information in the image. We validate our method on the P3M-10k dataset. Compared with the benchmark, our method has made significant improvements in various evaluation indicators. Qing Song 0006, Wenfeng Sun, Donghan Yang, Mengjie Hu 0002, Chun Liu 0004 |
Comput. J. | 4 |
| 2025 | Beyond-Skeleton: Zero-shot Skeleton Action Recognition enhanced by supplementary RGB visual information
Yingchun Niu, Chun Liu 0004, Mengjie Hu 0002, Qing Song 0006 |
Expert Syst. Appl. | 5 |
| 2025 | E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP GuidanceabstractDiffusion-based image editing involves both preserving the source image content and generating new content or applying modifications. Although current editing approaches have made improvements under text guidance, they have two key drawbacks: overemphasis on retaining original image info, neglecting editability and text alignment, and inability to handle both structure-consistent and non-rigid editing tasks. In this paper, we propose a zero-shot image editing method, named Enhance Editability for text-based image Editing via Efficient CLIP guidance (E4C), which presents an innovative adaptive feature sharing mechanism to enable multi-task editing. Additionally, a novel random gateway mechanism is designed to efficiently introduce CLIP guidance into the multi-step sampling of diffusion, achieving high congruence between editing results and target text. Comprehensive quantitative and qualitative experiments demonstrate that our method effectively resolves the text alignment issues prevalent in existing methods while maintaining the fidelity to the source image, and performs well across a wide range of editing tasks. Tianrui Huang, Pu Cao, Lu Yang 0006, Chun Liu 0004, Mengjie Hu 0002, Zhiwei Liu 0004, Qing Song 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Object Tracking With Separation in Deep SpaceabstractIn deep space environment, some objects may split into several small fragments during movement, and these deep space objects often appear as points in satellite images. In this article, we conduct research on multi-object tracking (MOT) for these objects. First, we propose a simulation dataset, ScatterDataset, which simulates the movement and separation of objects in deep space background. By assigning two IDs to a trajectory, we describe the trajectory’s relationship before and after separation. Second, we present an end-to-end motion association model, ScatterNet, which encodes the position information of trajectories and detections into motion features. These features are processed through temporal aggregation by a Transformer encoder and spatial aggregation by a graph network; then, we get the association results by calculating the similarity between these features. Finally, we introduce a tracker, ScatterTracker, which is suitable for tracking in scenarios with object separation. Experiments with state-of-the-art tracking methods on ScatterDataset demonstrate that our approach has achieved significant performance improvements in deep space scenarios. The code is available at:https://github.com/wht-bupt/ScatterTrack. Mengjie Hu 0002, Binyu Li, Shixiang Cao, Tao Zhan 0002, Xiaotong Zhu, Chun Liu 0004, Qing Song 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | UV R-CNN: Stable and efficient dense human pose estimation
Wenhe Jia, Xuhan Zhu, Mengjie Hu 0002, Chun Liu 0004, Qing Song 0006 |
Multim. Tools Appl. | 4 |
| 2024 | CoT-MISR:Marrying convolution and transformer for multi-image super-resolution
Qing Song 0006, Mingming Xiu, Yang Nie, Mengjie Hu 0002, Chun Liu 0004 |
Multim. Tools Appl. | 4 |
| 2024 | Faster learning of temporal action proposal via sparse multilevel boundary generator
Qing Song 0006, Mengjie Hu 0002, Chun Liu 0004 |
Multim. Tools Appl. | 3 |
| 2023 | Fast and robust for texture-less feature registration via adaptive heterogeneous kernels
Yuandong Ma, Qing Song 0006, Hezheng Lin, Chun Liu 0004, Mengjie Hu 0002, Xiaotong Zhu |
Knowl. Based Syst. | 5 |
| 2023 | Rethinking the activation function in lightweight network
Lu Yang 0006, Qing Song 0006, Zimeng Fan 0003, Chun Liu 0004, Mengjie Hu 0002 |
Multim. Tools Appl. | 5 |
| 2023 | A continuation method for image registration based on dynamic adaptive kernel
Yuandong Ma, Hezheng Lin, Chun Liu 0004, Mengjie Hu 0002, Qing Song 0006 |
Neural Networks | 5 |
| 2023 | A Lightweight Neural Learning Algorithm for Real-Time Facial Feature Tracking System via Split-Attention and Heterogeneous Convolution
Yuandong Ma, Qing Song 0006, Mengjie Hu 0002, Xiaotong Zhu |
Neural Process. Lett. | 3 |
| 2023 | Correction: A Lightweight Neural Learning Algorithm for Real-Time Facial Feature Tracking System via Split-Attention and Heterogeneous Convolution
Yuandong Ma, Qing Song 0006, Mengjie Hu 0002, Xiaotong Zhu |
Neural Process. Lett. | 3 |
| 2023 | STDFormer: Spatial-Temporal Motion Transformer for Multiple Object TrackingabstractMainstream multi-object tracking methods exploit appearance information and/or motion information to achieve interframe association. However, dealing with similar appearance and occlusion is a challenge for appearance information, while motion information is limited by linear assumptions and is prone to failure in nonlinear motion patterns. In this work, we disregard appearance clues and propose a pure motion tracker to address the above issues. It dexterously utilizes Transformer to estimate complex motion and achieves high-performance tracking with low computing resources. Furthermore, contrastive learning is introduced to optimize feature representation for robust association. Specifically, we first exploit the long-range modeling capability of Transformer to mine intention information in temporal motion and decision information in spatial interaction and introduce prior detection to constrain the range of motion estimation. Then, we introduce contrastive learning as an auxiliary task to extract reliable motion features to compute affinity and introduce bidirectional matching to improve the affinity computation distribution. In addition, given that both tasks are dedicated to narrowing the embedding distance between the motion features of the tracked object and the detection features, we design a joint-motion-and-association framework to unify the above two tasks in one framework for optimization. The experimental results achieved with three benchmark datasets, MOT17, MOT20 and DanceTrack, verify the effectiveness of our proposed method. Compared with state-of-the-art methods, the proposed STDFormer sets a new state-of-the-art on DanceTrack and achieves competitive performance on MOT17 and MOT20. This demonstrates the advantage of our method in handling associations under similar appearance, occlusion or nonlinear motion. At the same time, the significant advantages of the proposed method over Transformer-based and contrastive learning-based methods suggest a new direction for the application of Transformer and contrastive learning in MOT. In addition, to verify the generalization of STDFormer in unmanned aerial vehicle (UAV) videos, we also evaluate STDFormer on VisDrone2019. The results show that STDFormer achieves state-of-the-art performance on VisDrone2019, which proves that it can handle small-scale object associations in UAV videos well. The code is available at https://github.com/Xiaotong-Zhu/STDFormer. Mengjie Hu 0002, Xiaotong Zhu, Shixiang Cao, Chun Liu 0004, Qing Song 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Double parallel branches FCOS for human detection in a crowd
Qing Song 0006, Lu Yang 0006, Xueshi Xin, Chun Liu 0004, Mengjie Hu 0002 |
Multim. Tools Appl. | 6 |
| 2021 | CPM R-CNN: Calibrating Point-guided Misalignment in Object DetectionabstractIn object detection, offset-guided and point-guided regression dominate anchor-based and anchor-free method separately. Recently, point-guided approach is introduced to anchor-based method. However, we observe points predicted by this way are misaligned with matched region of proposals and score of localization, causing a notable gap in performance. In this paper, we propose CPM R-CNN which contains three efficient modules to optimize anchor- based point-guided method. According to sufficient evaluations on the COCO dataset, CPM R-CNN is demonstrated efficient to improve the localization accuracy by calibrating mentioned misalignment. Compared with Faster R-CNN and Grid R-CNN based on ResNet-101 with FPN, our approach can substantially improve detection mAP by 3.3% and 1.5% respectively without whistles and bells. Moreover, our best model achieves improvement by a large margin to 49.9% on COCO test-dev. Code is available at https://github.com/zhubinQAQ/CPM-R-CNN. Qing Song 0006, Lu Yang 0006, Zhihui Wang 0011, Chun Liu 0004, Mengjie Hu 0002 |
WACV | 6 |
| 2021 | Hier R-CNN: Instance-Level Human Parts Detection and A New BenchmarkabstractDetecting human parts at instance-level is an essential prerequisite for the analysis of human keypoints, actions, and attributes. Nonetheless, there is a lack of a large-scale, rich-annotated dataset for human parts detection. We fill in the gap by proposing COCO Human Parts. The proposed dataset is based on the COCO 2017, which is the first instance-level human parts dataset, and contains images of complex scenes and high diversity. For reflecting the diversity of human body in natural scenes, we annotate human parts with (a) location in terms of a bounding-box, (b) various type including face, head, hand, and foot, (c) subordinate relationship between person and human parts, (d) fine-grained classification into right-hand/left-hand and left-foot/right-foot. A lot of higher-level applications and studies can be founded upon COCO Human Parts, such as gesture recognition, face/hand keypoint detection, visual actions, human-object interactions, and virtual reality. There are a total of 268,030 person instances from the 66,808 images, and 2.83 parts per person instance. We provide a statistical analysis of the accuracy of our annotations. In addition, we propose a strong baseline for detecting human parts at instance-level over this dataset in an end-to-end manner, call Hier(archy) R-CNN. It is a simple but effective extension of Mask R-CNN, which can detect human parts of each person instance and predict the subordinate relationship between them. Codes and dataset are publicly available (https://github.com/soeaver/Hier-R-CNN). Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004 |
IEEE Trans. Image Process. | 4 |
| 2020 | Renovating Parsing R-CNN for Accurate Multiple Human Parsing
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004, Xueshi Xin, Wenhe Jia, Songcen Xu |
ECCV (12) | 4 |
| 2019 | Attention Inspiring Receptive-Fields Network for Learning Invariant RepresentationsabstractIn this paper, we describe a simple and highly efficient module for image classification, which we term the "Attention Inspiring Receptive-fields" (Air) module. We effectively convert the spatial attention mechanism into a plug-in module. In addition, we reveal the relationship between the spatial attention mechanism and the receptive fields, indicating that the proper use of the spatial attention mechanism can effectively increase the receptive fields of the module, which is able to enhance translation invariance and scale invariance of the network. By integrating the Air module into advanced convolutional neural networks (such as ResNet and ResNeXt), we can construct AirNet architectures for learning invariant representations and gain significant improvements on challenging data sets. We present extensive experiments on CIFAR and ImageNet data sets to verify the effectiveness and feature invariance of the Air module and explore more concise and efficient designs of the proposed module. On ImageNet classification, our AirNet-50 and AirNet-101 (ResNet-50/101 with Air module) achieve 1.69% and 1.50% top-1 accuracy improvement with a small amount of extra computation and parameters compared with the original ResNet. We make models and code public available https://github.com/soeaver/AirNet-PyTorch. We further demonstrate that AirNet has a good ability for transfer learning and measure the performance on Microsoft Common Objects in Context object detection, instance segmentation, and pose estimation. Lu Yang 0006, Qing Song 0006, Yingqi Wu, Mengjie Hu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Cross Connected Network for Efficient Image Recognition
Lu Yang 0006, Qing Song 0006, Zuoxin Li, Yingqi Wu, Mengjie Hu 0002 |
ACCV (1) | 6 |