VLDB 2026 Research / reviewers in the wild / expert
Mengshi Qi
dblp:191/2586
· DBLP profile ↗
40ranked-venue papers
18as first author
30since 2021 · last 2026
0000-0002-6955-6635ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 15 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 7 first-author · 16 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Batch Normalization with Test-Time Adaptation for Robust Object Detection in Self-DrivingabstractIn open real-world autonomous driving scenarios, challenges such as sensor failure and extreme weather hinder the generalization of current autonomous driving perception models to these unseen domain, due to the domain shifts between the test and training data. As the parameter scale of autonomous driving perception models grows, traditional test-time adaptation (TTA) methods become unstable and often degrade model performance in most scenarios. To address these challenges, this paper proposes two new robust methods to improve the Batch Normalization with TTA for object detection in autonomous driving: (1) We introduce a new LearnableBN layer based on Geometric Confidence Maximization and Entropy Minimization. Specifically, we modify the traditional BN layer by incorporating auxiliary learnable parameters, which enables the BN layer to dynamically update the statistics according to the different input data. (2) We propose a novel semantic-consistency based dual-stage adaptation strategy, which encourages the model to iteratively search for the optimal solution and eliminates unstable samples during the adaptation process. Extensive experiments on the NuScenes-C dataset shows that our method achieves a maximum improvement of about 10\% using BEVFormer as the baseline across six corruption types and three levels of severity. Dacheng Liao, Mengshi Qi, Liang Liu 0001, Huadong Ma |
AAAI | 2 |
| 2026 | Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningabstractIn this paper, we propose a new Robust Disentangled Counterfactual Learning (RDCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects' physical commonsense based on both video and audio input, with the main challenge being how to imitate the reasoning ability of humans, even in the scenario of missing modalities. Most of the current methods fail to take full advantage of different characteristics in multi-modal data, and the lack of causal reasoning ability in models impedes the progress of implicit physical knowledge inference. To address these issues, our proposed RDCL method decouples videos into static (time-invariant) and dynamic (time-varying) factors in the latent space using the disentangled sequential encoder, which adopts a variational autoencoder (VAE) to maximize the mutual information with a contrastive loss function. Furthermore, we introduce a counterfactual learning module to augment the model's reasoning ability by modeling physical knowledge relationships among different objects under counterfactual intervention. To alleviate the incomplete modality data issue, we introduce a robust multimodal learning method to recover the missing data by decomposing the shared features and model-specific features. Our proposed method is a plug-and-play module that can be incorporated into any baseline, including VLMs. In experiments, we show that our proposed method improves the reasoning accuracy and robustness of baseline methods and achieves the state-of-the-art performance. Mengshi Qi, Changsheng Lv, Huadong Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | DC-SAM: In-Context Segment Anything in Images and Videos via Dual ConsistencyabstractGiven a single labeled examples, in-context segmentation aims to segment corresponding objects. This setting, known as one-shot segmentation in few-shot learning, explores the segmentation model's generalization ability and has been applied to various vision tasks, including scene understanding and image/video editing. While recent Segment Anything Models (SAMs) have achieved state-of-the-art results in interactive segmentation, these approaches are not directly applicable to in-context segmentation. In this work, we propose the Dual Consistency SAM (DC-SAM) method based on prompt-tuning to adapt SAM and SAM2 for in-context segmentation of both images and videos. Our key insights are to enhance the features of the SAM's prompt encoder in segmentation by providing high-quality visual prompts. When generating a mask prior from support images, we fuse the SAM features to better align the prompt encoder rather than relying solely on a pre-trained backbone. Then, we design a cycle-consistent cross-attention on fused features and initial visual prompts. This design leverages coarse masks from the SAM mask decoder to ensure consistency between features and visual prompts. Next, a dual-branch design is provided by using the discriminative positive and negative prompts in the prompt encoder. Furthermore, we design a simple mask-tube training strategy to adopt our proposed dual consistency method into the mask tube. Although the proposed DC-SAM is primarily designed for images, it can be seamlessly extended to the video domain with the support of SAM2. Given the absence of in-context segmentation in the video domain, we manually curate and construct the first benchmark from existing video segmentation datasets, namedIn-Context Video Object Segmentation (IC-VOS), to better assess the in-context capability of the model. Extensive experiments demonstrate that our method achieves 55.5 (+1.4) mIoU on COCO-20$^{i}$, 73.0 (+1.1) mIoU on PASCAL-5$^{i}$, and a$\mathcal {J\&F}$score of 71.52 on the proposed IC-VOS benchmark. Mengshi Qi, Pengfei Zhu 0001, Xiangtai Li, Xiaoyang Bi, Lu Qi 0001, Huadong Ma, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Chain-of-Evidence Multimodal Reasoning for Few-Shot Temporal Action Localization
Mengshi Qi, Hongwei Ji, Wulian Yun, Xianlin Zhang, Huadong Ma |
IEEE Trans. Image Process. | 1 |
| 2026 | Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts ReasoningabstractEvaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly concerned with what and where the action is, which is unable to meet the requirements. Meanwhile, most of the existing datasets lack the labels indicating the degree of action standardization, and the action quality assessment datasets lack explainability and detailed feedback. Therefore, we define a new Human Action Form Assessment (AFA) task, and introduce a new diverse dataset CoT-AFA, which contains a large scale of fitness and martial arts videos with multi-level annotations for comprehensive video analysis. We enrich the CoT-AFA dataset with a novel Chain-of-Thought explanation paradigm. Instead of offering isolated feedback, our explanations provide a complete reasoning process-from identifying an action step to analyzing its outcome and proposing a concrete solution. Furthermore, we propose a framework named Explainable Fitness Assessor, which can not only judge an action but also explain why and provide a solution. This framework employs two parallel processing streams and a dynamic gating mechanism to fuse visual and semantic information, thereby boosting its analytical capabilities. The experimental results demonstrate that our method has achieved improvements in explanation generation (e.g., + 16.0% in CIDEr),action classification (+ 2.7% in accuracy) and quality assessment (+ 2.1% in accuracy), revealing great potential of CoT-AFA for future studies. Our dataset and source code are available at https://github.com/MICLAB-BUPT/EFA. Mengshi Qi, Yeteng Wu, Wulian Yun, Xianlin Zhang, Huadong Ma |
IEEE Trans. Image Process. | 1 |
| 2026 | PITN: Physics-Informed Temporal Networks for Cuffless Blood Pressure EstimationabstractEstimating blood pressure (BP) plays a crucial role in healthcare. Traditionally, BP has been measured using a cuff, which is unsuitable for continuous BP monitoring. Recent advancements in cuffless wearable devices, such as graphene electronic tattoos and PPG sensors provide viable alternatives. However, existing BP estimation methods often overlook the importance of temporal modeling, despite the multi-periodicity and temporal dependencies inherent in cuffless BP signals. Moreover, continuous BP monitoring requires personalized modeling, which is further challenged by data scarcity. To address these two challenges, we introduce a novel Physics-Informed Temporal Network (PITN) with adversarial contrastive learning to enable precise BP estimation with very limited data for three different modalities (i.e, bioimpedance, PPG, millimeter wave). Specifically, we first introduce a novel Physics-Informed Temporal Network for investigating BP dynamics' multi-periodicity for cardiovascular cycle modeling and temporal variation. We then apply adversarial training to generate extra physiological time series data, improving PITN's robustness in the face of sparse subject-specific training data. Furthermore, we utilize contrastive learning to capture the discriminative variations of cardiovascular physiologic phenomena. This approach aggregates physiological signals with similar blood pressure values in latent space while separating clusters of samples with dissimilar blood pressure values. Experiments on three widely-adopted datasets with different modailties demonstrate the superiority and effectiveness of the proposed methods over previous state-of-the-art approaches. The code is available athttps://github.com/Zest86/ACL-PITN. Mengshi Qi, Yingxia Shao, Anfu Zhou, Huadong Ma |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Towards Efficient Object Re-Identification with a Novel Cloud-Edge Collaborative FrameworkabstractObject re-identification (ReID) is committed to searching for objects of the same identity across cameras, and its real-world deployment is gradually increasing. Current ReID methods assume that the deployed system follows the centralized processing paradigm, i.e., all computations are conducted in the cloud server and edge devices are only used to capture images. As the number of videos experiences a rapid escalation, this paradigm has become impractical due to the finite computational resources in the cloud server. Therefore, the ReID system should be converted to fit in the cloud-edge collaborative processing paradigm, which is crucial to boost its scalability and practicality. However, current works lack relevant research on this important specific issue, making it difficult to adapt them into a cloud-edge framework effectively. In this paper, we propose a cloud-edge collaborative inference framework for ReID systems, aiming to expedite the return of the desired image captured by the camera to the cloud server by learning the spatial-temporal correlations among objects. In the system, a Distribution-aware Correlation Modeling network (DaCM) is particularly proposed to embed the spatial-temporal correlations of the camera network implicitly into a graph structure, and it can be applied 1) in the cloud to regulate the size of the upload window and 2) on the edge device to adjust the sequence of images, respectively. Notably, the proposed DaCM can be seamlessly combined with traditional ReID methods, enabling their application within our proposed edge-cloud collaborative framework. Extensive experiments demonstrate that our method obviously reduces transmission overhead and significantly improves performance. Chuanming Wang, Yuxin Yang 0008, Mengshi Qi, Huadong Ma |
AAAI | 3 |
| 2025 | VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of ThingsabstractVideo Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To address the challenge, we build VIoTGPT, the framework based on LLMs to correctly interact with humans, query knowledge videos, and invoke vision models to analyze multimedia data collaboratively. To support VIoTGPT and related future works, we meticulously crafted the VIoT-Tool dataset, including the training dataset and the benchmark involving 11 representative vision models across three categories based on semi-automatic annotations. To guide LLM to act as the intelligent agent towards intelligent VIoT, we resort to ReAct instruction tuning method based on VIoT-Tool to learn the tool capability. Quantitative and qualitative experiments and analyses demonstrate the effectiveness of VIoTGPT. We believe VIoTGPT contributes to improving human-centered experiences in VIoT applications. Yaoyao Zhong, Mengshi Qi, Yuhan Qiu, Huadong Ma |
AAAI | 2 |
| 2025 | Global-Local Tree Search in VLMs for 3D Indoor Scene GenerationabstractLarge Vision-Language Models (VLMs), such as GPT-4, have achieved remarkable success across various fields. However, there are few studies on 3D indoor scene generation with VLMs. This paper considers this task as a planning problem subject to spatial and layout common sense constraints. To solve the problem with a VLM, we propose a new global-local tree search algorithm. Globally, the method places each object sequentially and explores multiple placements during each placement process, where the problem space is represented as a tree. To reduce the depth of the tree, we decompose the scene structure hierarchically, i.e. room level, region level, floor object level, and supported object level. The algorithm independently generates the floor objects in different regions and supported objects placed on different floor objects. Locally, we also decompose the sub-task, the placement of each object, into multiple steps. The algorithm searches the tree of problem space. To leverage the VLM model to produce positions of objects, we discretize the top-down view space as a dense grid and fill each cell with diverse emojis to make to cells distinct. We prompt the VLM with the emoji grid and the VLM produces a reasonable location for the object by describing the position with the name of emojis. The quantitative and qualitative experimental results illustrate our approach generates more plausible 3D scenes than state-of-the-art approaches. Our source code is available at https://github.com/dw-dengwei/TreeSearchGen. Wei Deng 0004, Mengshi Qi, Huadong Ma |
CVPR | 2 |
| 2025 | T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous DrivingabstractUnderstanding the traffic scenes and then generating highdefinition (HD) maps present significant challenges in autonomous driving. In this paper, we defined a novel Traffic Topology Scene Graph (T2SG), a unified scene graph explicitly modeling the lane, controlled and guided by different road signals (e.g., right turn), and topology relationships among them, which is always ignored by previous high-definition (HD) mapping methods. For the generation of T2SG, we propose TopoFormer, a novel one- stage Topology Scene Graph TransFormer with two newly-designed layers. Specifically, TopoFormer incorporates a Lane Aggregation Layer (LAL) that leverages the geometric distance among the centerline of lanes to guide the aggregation of global information. Furthermore, we proposed a Counterfactual Intervention Layer (CIL) to model the reasonable road structure (e.g., intersection, straight) among lanes under counterfactual intervention. Then the generated T2SG can provide a more accurate and explainable description of the topological structure in traffic scenes. Experimental results demonstrate that TopoFormer outperforms existing methods on the T2SG generation task, and the generated T2SG significantly enhances traffic topology reasoning in downstream tasks, achieving a state-of-the-art performance of 46.3 OLS on the OpenLane-V2 benchmark. Our source code is available at https://github.com/MICLAB-BUPT/T2SG. Changsheng Lv, Mengshi Qi, Liang Liu 0001, Huadong Ma |
CVPR | 2 |
| 2025 | SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented GenerationabstractIn this work, we study how vision-language models (VLMs) can be utilized to enhance the safety for the autonomous driving system, including perception, situational understanding, and path planning. However, existing research has largely overlooked the evaluation of these models in traffic safety-critical driving scenarios. To bridge this gap, we create the benchmark (SafeDrive228K) and propose a new baseline based on VLM with knowledge graph-based retrieval-augmented generation (SafeDriveRAG) for visual question answering (VQA). Specifically, we introduce SafeDrive228K, the first large-scale multimodal question-answering benchmark comprising 228K examples across 18 sub-tasks. This benchmark encompasses a diverse range of traffic safety queries, from traffic accidents and corner cases to common safety knowledge, enabling a thorough assessment of the comprehension and reasoning abilities of the models. Furthermore, we propose a plug-and-play multimodal knowledge graph-based retrieval-augmented generation approach that employs a novel multi-scale subgraph retrieval algorithm for efficient information retrieval. By incorporating traffic safety guidelines collected from the Internet, this framework further enhances the model's capacity to handle safety-critical situations. Finally, we conduct comprehensive evaluations on five mainstream VLMs to assess their reliability in safety-sensitive driving tasks. Experimental results demonstrate that integrating RAG significantly improves performance, achieving a +4.73% gain in Traffic Accidents tasks, +8.79% in Corner Cases tasks and +14.57% in Traffic Safety Commonsense across five mainstream VLMs, underscoring the potential of our proposed benchmark and methodology for advancing research in traffic safety. Our source code and data are available at https://github.com/Lumos0507/SafeDriveRAG. Mengshi Qi, Zhaohong Liu, Liang Liu 0001, Huadong Ma |
ACM Multimedia | 2 |
| 2025 | Synergistic Tensor and Pipeline ParallelismabstractIn the machine learning system, the hybrid model parallelism combining tensor parallelism (TP) and pipeline parallelism (PP) has become the dominant solution for distributed training of Large Language Models~(LLMs) and Multimodal LLMs (MLLMs). However, TP introduces significant collective communication overheads, while PP suffers from synchronization inefficiencies such as pipeline bubbles. Existing works primarily address these challenges from isolated perspectives, focusing either on overlapping TP communication or on flexible PP scheduling to mitigate pipeline bubbles. In this paper, we propose a new synergistic tensor and pipeline parallelism schedule that simultaneously reduces both types of bubbles. Our proposed schedule decouples the forward and backward passes in PP into fine-grained computation units, which are then braided to form a composite computation sequence. This compositional structure enables near-complete elimination of TP-related bubbles. Building upon this structure, we further design the PP schedule to minimize PP bubbles. Experimental results demonstrate that our approach improves training throughput by up to 12\% for LLMs and 16\% for MLLMs compared to existing scheduling methods. Our source code is avaiable at https://github.com/MICLAB-BUPT/STP. Mengshi Qi, Jiaxuan Peng 0001, Jie Zhang 0050, Juan Zhu, Huadong Ma |
NeurIPS | 1 |
| 2025 | Diffusion-driven Incomplete Multimodal Learning for Air Quality PredictionabstractPredicting air quality using multimodal data is crucial to comprehensively capture the diverse factors influencing atmospheric conditions. Therefore, this study introduces a multimodal learning framework that integrates outdoor images with traditional ground-based observations to improve the accuracy and reliability of air quality predictions. However, aligning and fusing these heterogeneous data sources poses a formidable challenge, further exacerbated by pervasive data incompleteness issues in practice. In this article, we propose a novel incomplete multimodal learning approach (iMMAir) to recovery missing data for robust air quality prediction. Specifically, we first design a shallow feature extractor to capture modal-specific features within the latent representation space. Then we develop a conditional diffusion-driven recovery module to mitigate the distribution gap between the recovered and true data. This module further incorporates two conditional constraints of temporal correlation and semantic consistency for effective modal completion. Finally, we reconstruct incomplete modalities and fuse available data using a multimodal transformer network to predict the air quality. To alleviate the modality imbalance problem, we employ an adaptive gradient modulation strategy to adjust the optimization of each modality. Experimental results demonstrate that iMMAir significantly reduces prediction errors, outperforming baseline models by an average of 5.6% and 2.5% in air quality regression and classification tasks. Our source code and data are available at https://github.com/pestasu/IMMAir . Jinxiao Fan, Mengshi Qi, Liang Liu 0001, Huadong Ma |
ACM Trans. Internet Things | 2 |
| 2025 | Action Quality Assessment via Hierarchical Pose-Guided Multi-Stage Contrastive RegressionabstractAction Quality Assessment (AQA), which aims at the automatic and fair evaluation of athletic performance, has gained increasing attention in recent years. However, athletes are often in rapid movement and the corresponding visual appearance variances are subtle, making it challenging to capture fine-grained pose differences and leading to poor estimation performance. Furthermore, most common AQA tasks, such as diving in sports, are usually divided into multiple sub-actions, each of which contains different durations. However, existing methods focus on segmenting the video into fixed frames, which disrupts the temporal continuity of sub-actions resulting in unavoidable prediction errors. To address these challenges, we propose a novel action quality assessment method through hierarchically pose-guided multi-stage contrastive regression. Firstly, we introduce a multi-scale dynamic visual-skeleton encoder to capture fine-grained spatio-temporal visual and skeletal features. Compared to mask or auxiliary visual features, skeletal features provide a more accurate representation during athletic movements. Then, a procedure segmentation network is introduced to separate different sub-actions and obtain segmented features. Afterwards, the segmented visual and skeletal features are both fed into a multi-modal fusion module as physics structural priors, to guide the model in learning refined activity similarities and variances. Finally, a multi-stage contrastive learning regression approach is employed to learn discriminative representations and output prediction results. In addition, we introduce a newly-annotated FineDiving-Pose Dataset to improve the current low-quality human pose labels. In experiments, the results on FineDiving and MTL-AQA datasets demonstrate the effectiveness and superiority of our proposed approach. Our source code and dataset are available at https://github.com/Lumos0507/HP-MCoRe. Mengshi Qi, Jiaxuan Peng 0001, Huadong Ma |
IEEE Trans. Image Process. | 1 |
| 2024 | SGFormer: Semantic Graph Transformer for Point Cloud-Based 3D Scene Graph GenerationabstractIn this paper, we propose a novel model called SGFormer, Semantic Graph TransFormer for point cloud-based 3D scene graph generation. The task aims to parse a point cloud-based scene into a semantic structural graph, with the core challenge of modeling the complex global structure. Existing methods based on graph convolutional networks (GCNs) suffer from the over-smoothing dilemma and can only propagate information from limited neighboring nodes. In contrast, SGFormer uses Transformer layers as the base building block to allow global information passing, with two types of newly-designed layers tailored for the 3D scene graph generation task. Specifically, we introduce the graph embedding layer to best utilize the global information in graph edges while maintaining comparable computation costs. Furthermore, we propose the semantic injection layer to leverage linguistic knowledge from large-scale language model (i.e., ChatGPT), to enhance objects' visual features. We benchmark our SGFormer on the established 3DSSG dataset and achieve a 40.94% absolute improvement in relationship prediction's R@50 and an 88.36% boost on the subset with complex scenes over the state-of-the-art. Our analyses further show SGFormer's superiority in the long-tail and zero-shot scenarios. Our source code is available at https://github.com/Andy20178/SGFormer. Changsheng Lv, Mengshi Qi, Zhengyuan Yang, Huadong Ma |
AAAI | 2 |
| 2024 | Weakly-Supervised Temporal Action Localization by Inferring Salient Snippet-FeatureabstractWeakly-supervised temporal action localization aims to locate action regions and identify action categories in untrimmed videos simultaneously by taking only video-level labels as the supervision. Pseudo label generation is a promising strategy to solve the challenging problem, but the current methods ignore the natural temporal structure of the video that can provide rich information to assist such a generation process. In this paper, we propose a novel weakly-supervised temporal action localization method by inferring salient snippet-feature. First, we design a saliency inference module that exploits the variation relationship between temporal neighbor snippets to discover salient snippet-features, which can reflect the significant dynamic change in the video. Secondly, we introduce a boundary refinement module that enhances salient snippet-features through the information interaction unit. Then, a discrimination enhancement module is introduced to enhance the discriminative nature of snippet-features. Finally, we adopt the refined snippet-features to produce high-fidelity pseudo labels, which could be used to supervise the training of the action localization network. Extensive experiments on two publicly available datasets, i.e., THUMOS14 and ActivityNet v1.3, demonstrate our proposed method achieves significant improvements compared to the state-of-the-art methods. Our source code is available at https://github.com/wuli55555/ISSF. Wulian Yun, Mengshi Qi, Chuanming Wang, Huadong Ma |
AAAI | 2 |
| 2024 | Semi-supervised Teacher-Reference-Student Architecture for Action Quality Assessment
Wulian Yun, Mengshi Qi, Fei Peng 0003, Huadong Ma |
ECCV (74) | 2 |
| 2024 | Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation
Mengshi Qi, Huadong Ma |
ECCV (29) | 2 |
| 2024 | Multi-Stage Contrastive Regression for Action Quality AssessmentabstractIn recent years, there has been growing interest in the video-based action quality assessment (AQA). Most existing methods typically solve AQA problem by considering the entire video yet overlooking the inherent stage-level characteristics of actions. To address this issue, we design a novel Multi-stage Contrastive Regression (MCoRe) framework for the AQA task. This approach allows us to efficiently extract spatial-temporal information, while simultaneously reducing computational costs by segmenting the input video into multiple stages or procedures. Inspired by the graph contrastive learning, we propose a new stage-wise contrastive learning loss function to enhance performance. As a result, MCoRe demonstrates the state-of-the-art result so far on the widely-adopted fine-grained AQA dataset. Our source code is available at https://github.com/Angel-1999/MCoRe. Mengshi Qi, Huadong Ma |
ICASSP | 2 |
| 2024 | RoboFormer: A Robust Multi-Modal Transformer for 3D Object Detection in Autonomous Driving
Yuang Liu, Dacheng Liao, Mengshi Qi, Liang Liu 0001, Huadong Ma |
MMAsia | 3 |
| 2024 | Following in the Footsteps: Predicting Human Trajectories Using Motion Pattern Memory
Yuxin Yang 0008, Pengfei Zhu 0001, Mengshi Qi, Huadong Ma |
MMAsia | 3 |
| 2024 | RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth CompletionabstractRaw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent materials frequently elude detection by depth sensors; surfaces may introduce measurement inaccuracies due to their polished textures, extended distances, and oblique incidence angles from the sensor. The presence of incomplete depth maps imposes significant challenges for subsequent vision applications, prompting the development of numerous depth completion techniques to mitigate this problem. Numerous methods excel at reconstructing dense depth maps from sparse samples, but they often falter when faced with extensive contiguous regions of missing depth values, a prevalent and critical challenge in indoor environments. To overcome these challenges, we design a novel two-branch end-to-end fusion network named RDFC-GAN, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure, by adhering to the Manhattan world assumption and utilizing normal maps from RGB-D information as guidance, to regress the local dense depth values from the raw depth map. The other branch applies an RGB-depth fusion CycleGAN, adept at translating RGB imagery into detailed, textured depth maps while ensuring high fidelity through cycle consistency. We fuse the two branches via adaptive fusion modules named W-AdaIN and train the model with the help of pseudo depth maps. Comprehensive evaluations on NYU-Depth V2 and SUN RGB-D datasets show that our method significantly enhances depth completion performance particularly in realistic indoor settings. Haowen Wang 0001, Zhengping Che, Mingyuan Wang 0003, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Mutual Distillation Learning for Person Re-IdentificationabstractWith the rapid advancements in deep learning technologies, person re-identification (ReID) has witnessed remarkable performance improvements. However, the majority of prior works have traditionally focused on solving the problem via extracting features solely from a single perspective, such as uniform partitioning, attention mechanisms, or semantic masks. While these approaches have demonstrated efficacy within specific contexts, they fall short in diverse situations. In this paper, we propose a novel approach, Mutual Distillation Learning For Person Re-identification (termed as MDPR), which addresses the challenging problem from multiple perspectives within a single unified model, leveraging the power of mutual distillation to enhance the feature representations collectively. Specifically, our approach encompasses two branches: a hard content branch to extract local features via a uniform horizontal partitioning strategy and a soft content branch to dynamically distinguish between foreground and background and facilitate the extraction of multi-granularity features via a carefully designed attention mechanism. To facilitate knowledge exchange between these two branches, a mutual distillation and fusion process is employed, promoting the capability of the outputs of each branch. Extensive experiments are conducted on widely used person ReID datasets to validate the effectiveness and superiority of our approach. Notably, our method achieves an impressive 88.7%/94.4% in mAP/Rank-1 on the DukeMTMC-reID dataset, surpassing the current state-of-the-art results. Our source code is available athttps://github.com/KuilongCui/MDPR. Huiyuan Fu, Kuilong Cui, Chuanming Wang, Mengshi Qi, Huadong Ma |
IEEE Trans. Multim. | 4 |
| 2023 | Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge EmbeddingabstractPredicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets also limits the potential for model training. To address these challenges, we are the first to introduce an unsupervised way to predict self-driving attention by uncertainty modeling and driving knowledge integration. Our approach’s Uncertainty Mining Branch (UMB) discovers commonalities and differences from multiple generated pseudo-labels achieved from models pre-trained on natural scenes by actively measuring the uncertainty. Meanwhile, our Knowledge Embedding Block (KEB) bridges the domain gap by incorporating driving knowledge to adaptively refine the generated pseudo-labels. Quantitative and qualitative results with equivalent or even more impressive performance compared to fully-supervised state-of-the-art approaches across all three public datasets demonstrate the effectiveness of the proposed method and the potential of this direction. The code is available at https://github.com/zaplm/DriverAttention. Pengfei Zhu 0001, Mengshi Qi, Weijian Li 0001, Huadong Ma |
ICCV | 2 |
| 2023 | Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningabstractIn this paper, we propose a Disentangled Counterfactual Learning (DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects’ physics commonsense based on both video and audio input, with the main challenge is how to imitate the reasoning ability of humans. Most of the current methods fail to take full advantage of different characteristics in multi-modal data, and lacking causal reasoning ability in models impedes the progress of implicit physical knowledge inferring. To address these issues, our proposed DCL method decouples videos into static (time-invariant) and dynamic (time-varying) factors in the latent space by the disentangled sequential encoder, which adopts a variational autoencoder (VAE) to maximize the mutual information with a contrastive loss function. Furthermore, we introduce a counterfactual learning module to augment the model’s reasoning ability by modeling physical knowledge relationships among different objects under counterfactual intervention. Our proposed method is a plug-and-play module that can be incorporated into any baseline. In experiments, we show that our proposed method improves baseline methods and achieves state-of-the-art performance. Our source code is available at https://github.com/Andy20178/DCL. Changsheng Lv, Yapeng Tian, Mengshi Qi, Huadong Ma |
NeurIPS | 4 |
| 2023 | GaitReload: A Reloading Framework for Defending Against On-Manifold Adversarial Gait SequencesabstractRecent on-manifold adversarial attacks can mislead gait recognition by generating adversarial walking postures (AWP) with image generation techniques. However, existing defense methods only eliminate adversarial perturbations on each frame isolatedly but ignore the temporal correlation of gait sequence, which leads to vulnerability of robust gait recognition. In this paper, we propose GaitReload, a post-processing adversarial defense method to defend against AWP for the gait recognition model with sequenced inputs. First, GaitReload utilizes sequenced entity recognition (SER) module to detect the adversarial frames by the temporal constraints of gait sequence. Then, we apply bayesian uncertainty filtering-based (BUF-based) gait interpolation to reform adversarial gait examples. After that, we reload the reformed gait sequence and rectify the recognition results with the guidance of reloading strategy. Specifically, SER has a bi-directional frame difference attention and a temporal feature aggregation to boost the detection performance. For training SER, we apply hidden posture selective attack (HPSA) to generate training samples. The extensive experimental results on CASIA-A, CASIA-B, and OU-ISIR demonstrate that GaitReload can defend against adversarial gait by large margins in both RGB and silhouette modes. Peilun Du, Xiaolong Zheng 0002, Mengshi Qi, Huadong Ma |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | RGB-Depth Fusion GAN for Indoor Depth CompletionabstractThe raw depth image captured by the indoor depth sen-sor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transparent objects and limited distance range. The incomplete depth map burdens many downstream vision tasks, and a rising number of depth completion methods have been proposed to alleviate this issue. While most existing meth-ods can generate accurate dense depth maps from sparse and uniformly sampled depth maps, they are not suitable for complementing the large contiguous regions of missing depth values, which is common and critical. In this paper, we design a novel two-branch end-to-end fusion network, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure to regress the local dense depth values from the raw depth map, with the help of local guidance information extracted from the RGB image. In the other branch, we propose an RGB-depth fusion GAN to transfer the RGB image to the fine-grained textured depth map. We adopt adaptive fusion modules named W-AdaIN to propagate the features across the two branches, and we append a confidence fusion head to fuse the two out-puts of the branches for the final depth map. Extensive ex-periments on NYU-Depth V2 and SUN RGB-D demonstrate that our proposed method clearly improves the depth completion performance, especially in a more realistic setting of indoor environments with the help of the pseudo depth map. Haowen Wang 0001, Mingyuan Wang 0003, Zhengping Che, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
CVPR | 6 |
| 2022 | Towards Adversarial Robust Representation Through Adversarial Contrastive DecouplingabstractAdversarial training can boost the robustness of the model by aligning discriminative features between natural and generated adversarial samples. However, the generated adversarial samples tend to have more features derived from changed patterns in other categories along with the training process, which prevents better feature alignment between natural and adversarial samples. Unfortunately, existing adversarial training methods ignore such dynamicity of generated adversarial samples. In this paper, we propose Adversarial Contrastive Decoupling (ACD) to filter the features derived from changed patterns. Specificity, we decouple the changed patterns from adversarial samples and then extract robust representations from remaining features. First, we introduce a decoupling module with a dynamic labeling strategy to explore the dynamicity of generated adversarial samples. Then, we propose a siamese network with contrastive learning mechanism to align remaining robust representations between adversarial and natural samples. Extensive experimental results demonstrate the superior performance of ACD over baselines. Peilun Du, Xiaolong Zheng 0002, Mengshi Qi, Liang Liu 0001, Huadong Ma |
ICME | 3 |
| 2021 | Latent Memory-augmented Graph Transformer for Visual StorytellingabstractVisual storytelling aims to automatically generate a human-like short story given an image stream. Most existing works utilize either scene-level or object-level representations, neglecting the interaction among objects in each image and the sequential dependency between consecutive images. In this paper, we present a novel Latent Memory-augmented Graph Transformer~(LMGT ), a Transformer based framework for visual story generation. LMGT directly inherits the merits from the Transformer, which is further enhanced with two carefully designed components, i.e., a graph encoding module and a latent memory unit. Specifically, the graph encoding module exploits the semantic relationships among image regions and attentively aggregates critical visual features based on the parsed scene graphs. Furthermore, to better preserve inter-sentence coherence and topic consistency, we introduce an augmented latent memory unit that learns and records highly summarized latent information as the story line from the image stream and the sentence history. Experimental results on three widely-used datasets demonstrate the superior performance of LMGT over the state-of-the-art methods. Mengshi Qi, Jie Qin 0004, Di Huang 0001, Yi Yang 0001, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2021 | Semantics-Aware Spatial-Temporal Binaries for Cross-Modal Video RetrievalabstractWith the current exponential growth of video-based social networks, video retrieval using natural language is receiving ever-increasing attention. Most existing approaches tackle this task by extracting individual frame-level spatial features to represent the whole video, while ignoring visual pattern consistencies and intrinsic temporal relationships across different frames. Furthermore, the semantic correspondence between natural language queries and person-centric actions in videos has not been fully explored. To address these problems, we propose a novel binary representation learning framework, named Semantics-aware Spatial-temporal Binaries ( [Formula: see text]Bin), which simultaneously considers spatial-temporal context and semantic relationships for cross-modal video retrieval. By exploiting the semantic relationships between two modalities, [Formula: see text]Bin can efficiently and effectively generate binary codes for both videos and texts. In addition, we adopt an iterative optimization scheme to learn deep encoding functions with attribute-guided stochastic training. We evaluate our model on three video datasets and the experimental results demonstrate that [Formula: see text]Bin outperforms the state-of-the-art methods in terms of various cross-modal video retrieval tasks. Mengshi Qi, Jie Qin 0004, Yi Yang 0001, Yunhong Wang 0001, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Imitative Non-Autoregressive Modeling for Trajectory Forecasting and ImputationabstractTrajectory forecasting and imputation are pivotal steps towards understanding the movement of human and objects, which are quite challenging since the future trajectories and missing values in a temporal sequence are full of uncertainties, and the spatial-temporally contextual correlation is hard to model. Yet, the relevance between sequence prediction and imputation is disregarded by existing approaches. To this end, we propose a novel imitative non-autoregressive modeling method to simultaneously handle the trajectory prediction task and the missing value imputation task. Specifically, our framework adopts an imitation learning paradigm, which contains a recurrent conditional variational autoencoder (RC-VAE) as a demonstrator, and a non-autoregressive transformation model (NART) as a learner. By jointly optimizing the two models, RC-VAE can predict the future trajectory and capture the temporal relationship in the sequence to supervise the NART learner. As a result, NART learns from the demonstrator and imputes the missing value in a non autoregressive strategy. We conduct extensive experiments on three popular datasets, and the results show that our model achieves state-of-the-art performance across all the datasets. Mengshi Qi, Jie Qin 0004, Yu Wu 0011, Yi Yang 0001 |
CVPR | 1 |
| 2020 | Few-Shot Ensemble Learning for Video Classification with SlowFast Memory NetworksabstractIn the era of big data, few-shot learning has recently received much attention in multimedia analysis and computer vision due to its appealing ability of learning from scarce labeled data. However, it has been largely underdeveloped in the video domain, which is even more challenging due to the huge spatial-temporal variability of video data. In this paper, we address few-shot video classification by learning an ensemble of SlowFast networks augmented with memory units. Specifically, we introduce a family of few-shot learners based on SlowFast networks which are used to extract informative features at multiple rates, and we incorporate a memory unit into each network to enable encoding and retrieving crucial information instantly. Furthermore, we propose a choice controller network to leverage the diversity of few-shot learners by learning to adaptively assign a confidence score to each SlowFast memory network, leading to a strong classifier for enhanced prediction. Experimental results on two widely-adopted video datasets demonstrate the effectiveness of the proposed method, as well as its superior performance over the state-of-the-art approaches. Mengshi Qi, Jie Qin 0004, Xiantong Zhen, Di Huang 0001, Yi Yang 0001, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2020 | Sports Video Captioning via Attentive Motion Representation and Group Relationship ModelingabstractSports video captioning refers to the task of automatically generating a textual description for sports events (football, basketball, or volleyball games). Although a great deal of previous work has shown promising performance in producing a coarse and a general description of a video but lack of professional sports knowledge, it is still quite challenging to caption a sports video with multiple fine-grained player's actions and complex group relationship between players. In this paper, we present a novel hierarchical recurrent neural network-based framework with an attention mechanism for sports video captioning, in which a motion representation module is proposed to capture individual pose attribute and dynamical trajectory cluster information with extra professional sports knowledge, and a group relationship module is employed to design a scene graph for modeling players' interaction by a gated graph convolutional network. Moreover, we introduce a new dataset called sports video captioning dataset-volleyball for evaluation. The proposed model is evaluated on three widely adopted public datasets and our collected new dataset, on which the effectiveness of our method is well demonstrated. Mengshi Qi, Yunhong Wang 0001, Annan Li, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | stagNet: An Attentive Semantic RNN for Group Activity and Individual Action RecognitionabstractIn real life, group activity recognition plays a significant and fundamental role in a variety of applications, e.g. sports video analysis, abnormal behavior detection, and intelligent surveillance. In a complex dynamic scene, a crucial yet challenging issue is how to better model the spatio-temporal contextual information and inter-person relationship. In this paper, we present a novel attentive semantic recurrent neural network (RNN), namely, stagNet, for understanding group activities and individual actions in videos, by combining the spatio-temporal attention mechanism and semantic graph modeling. Specifically, a structured semantic graph is explicitly modeled to express the spatial contextual content of the whole scene, which is further incorporated with the temporal factor through structural-RNN. By virtue of the “factor sharing” and “message passing” mechanisms, our stagNet is capable of extracting discriminative and informative spatio-temporal representations and capturing inter-person relationships. Moreover, we adopt a spatio-temporal attention model to focus on key persons/frames for improved recognition performance. Besides, a body-region attention and a global-part feature pooling strategy are devised for individual action recognition. In experiments, four widely-used public datasets are adopted for performance evaluation, and the extensive results demonstrate the superiority and effectiveness of our method. Mengshi Qi, Yunhong Wang 0001, Jie Qin 0004, Annan Li, Jiebo Luo 0001, Luc Van Gool |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | STC-GAN: Spatio-Temporally Coupled Generative Adversarial Networks for Predictive Scene ParsingabstractPredictive scene parsing is a task of assigning pixellevel semantic labels to a future frame of a video. It has many applications in vision-based artificial intelligent systems, e.g., autonomous driving and robot navigation. Although previous work has shown its promising performance in semantic segmentation of images and videos, it is still quite challenging to anticipate future scene parsing with limited annotated training data. In this paper, we propose a novel model called STC-GAN, Spatio-Temporally Coupled Generative Adversarial Networks for predictive scene parsing, which employ both convolutional neural networks and convolutional long short-term memory (LSTM) in the encoderdecoder architecture. By virtue of STC-GAN, both spatial layout and semantic context can be captured by the spatial encoder effectively, while motion dynamics are extracted by the temporal encoder accurately. Furthermore, a coupled architecture is presented for establishing joint adversarial training where the weights are shared and features are transformed in an adaptive fashion between the future frame generation model and predictive scene parsing model. Consequently, the proposed STCGAN is able to learn valuable features from unlabeled video data. We evaluate our proposed STC-GAN on two public datasets, i.e., Cityscapes and CamVid. Experimental results demonstrate that our method outperforms the state-of-the-art. Mengshi Qi, Yunhong Wang 0001, Annan Li, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Attentive Relational Networks for Mapping Images to Scene GraphsabstractScene graph generation refers to the task of automatically mapping an image into a semantic structural graph, which requires correctly labeling each extracted object and their interaction relationships. Despite the recent success in object detection using deep learning techniques, inferring complex contextual relationships and structured graph representations from visual data remains a challenging topic. In this study, we propose a novel Attentive Relational Network that consists of two key modules with an object detection backbone to approach this problem. The first module is a semantic transformation module utilized to capture semantic embedded relation features, by translating visual features and linguistic features into a common semantic space. The other module is a graph self-attention module introduced to embed a joint graph representation through assigning various importance weights to neighboring nodes. Finally, accurate scene graphs are produced by the relation inference module to recognize all entities and corresponding relations. We evaluate our proposed method on the widely-adopted Visual Genome Dataset, and the results demonstrate the effectiveness and superiority of our model. Mengshi Qi, Weijian Li 0001, Zhengyuan Yang, Yunhong Wang 0001, Jiebo Luo 0001 |
CVPR | 1 |
| 2019 | KE-GAN: Knowledge Embedded Generative Adversarial Networks for Semi-Supervised Scene ParsingabstractIn recent years, scene parsing has captured increasing attention in computer vision. Previous works have demonstrated promising performance in this task. However, they mainly utilize holistic features, whilst neglecting the rich semantic knowledge and inter-object relationships in the scene. In addition, these methods usually require a large number of pixel-level annotations, which is too expensive in practice. In this paper, we propose a novel Knowledge Embedded Generative Adversarial Networks, dubbed as KE-GAN, to tackle the challenging problem in a semi-supervised fashion. KE-GAN captures semantic consistencies of different categories by devising a Knowledge Graph from the large-scale text corpus. In addition to readily-available unlabeled data, we generate synthetic images to unveil rich structural information underlying the images. Moreover, a pyramid architecture is incorporated into the discriminator to acquire multi-scale contextual information for better parsing results. Extensive experimental results on four standard benchmarks demonstrate that KE-GAN is capable of improving semantic consistencies and learning better representations for scene parsing, resulting in the state-of-the-art performance. Mengshi Qi, Yunhong Wang 0001, Jie Qin 0004, Annan Li |
CVPR | 1 |
| 2018 | stagNet: An Attentive Semantic RNN for Group Activity Recognition
Mengshi Qi, Jie Qin 0004, Annan Li, Yunhong Wang 0001, Jiebo Luo 0001, Luc Van Gool |
ECCV (10) | 1 |
| 2017 | Online Cross-Modal Scene Retrieval by Binary Representation and Semantic GraphabstractIn recent years, cross-modal scene retrieval has attracted more attention. However, most existing approaches neglect the semantic relationship between objects in a scene together with the embedded spatial layouts. Moreover, these methods mostly apply the batch learning strategy, which is not suitable for processing streaming data. To address the aforementioned problems, we propose a new framework for online cross-modal scene retrieval based on binary representations and semantic graph. Specially, we adopt the cross-modal hashing based on the quantization loss of different modalities. By introducing the semantic graph, we are able to extract wealthy semantics and measure their correlation across different modalities. Further more, we propose a two-step optimization procedure based on stochastic gradient descent for online update. Experimental results on four datasets show the superiority of our approach over the state-of-the-art. Mengshi Qi, Yunhong Wang 0001, Annan Li |
ACM Multimedia | 1 |
| 2016 | DEEP-CSSR: Scene classification using category-specific salient region with deep featuresabstractResearches in neuroscience and biological vision have shown that the bio-inspired methods have excellent recognition performance, such as the salient detection, artificial neural network and the ganglion cell inspired image feature. In this paper, we introduce a novel framework towards scene classification using category-specific salient region(CSSR) with deep CNN features, called Deep-CSSR. Firstly, by using the salient region detection algorithm, we extract a set of image patches which contain the salient regions. Also we apply DERF, a novel bio-inspired image descriptor, to represent the salient patches and clustering all of them to remove the outliers. Then we learn the CSSR filters and construct the CSSR representation. Further more, we do scene image classification using CSSR representation concatenate with the deep CNN features extracted from the whole images. By using this new pipeline, we obtain better results than recent methods over MIT Indoor 67 and Sun397 databases. Mengshi Qi, Yunhong Wang 0001 |
ICIP | 1 |