Mingang Chen

dblp:22/1557 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-0718-8187ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
abstract
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability.
Jingyu Gong, Kunkun Tong, Zhuoran Chen, Chuanhan Yuan, Mingang Chen, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
AAAI5
2026 FocusPatch AD: Few-Shot Multi-Class Anomaly Detection With Unified Keywords Patch Prompts
abstract
Industrial few-shot anomaly detection (FSAD) requires identifying various abnormal states by leveraging as few normal samples as possible (abnormal samples are unavailable during training). However, current methods often require training a separate model for each category, leading to increased computation and storage overhead. Thus, designing a unified anomaly detection model that supports multiple categories remains a challenging task, as such a model must recognize anomalous patterns across diverse objects and domains. To tackle these challenges, this paper introduces FocusPatch AD, a unified anomaly detection framework based on vision-language models, achieving anomaly detection under few-shot multi-class settings. FocusPatch AD links anomaly state keywords to highly relevant discrete local regions within the image, guiding the model to focus on cross-category anomalies while filtering out background interference. This approach mitigates the false detection issues caused by global semantic alignment in vision-language models. We evaluate the proposed method on the MVTec, VisA, and Real-IAD datasets, comparing them against several prevailing anomaly detection methods. In both image-level and pixel-level anomaly detection tasks, FocusPatch AD achieves significant gains in classification and localization performance, demonstrating excellent generalization and adaptability.
Xicheng Ding, Xiaofan Li 0008, Mingang Chen, Jingyu Gong, Yuan Xie 0006
IEEE Trans. Image Process.3
2025 MCM: A Multi-Agent Collaborative Multimodal Framework For Traditional Chinese Medicine Diagnosis
abstract
The advancement of information technology and the rise of generative AI have paved the way for the development of Large Language Models (LLMs) tailored for TCM diagnostics. However, existing LLMs in the field of TCM face challenges in interpretability, limited modality in interaction, and robustness. To address these limitations, we propose MCM, a Multi-Agent Collaborative Multimodal Framework for TCM Diagnosis. This framework enables robust and interpretable multimodal diagnosis through multi-agent collaboration, offering novel methodologies for applying LLMs in the TCM domain. Experimental results demonstrate that the model within the MCM framework improved performance after fine-tuning, with additional capability gains under the MCM framework’s support, effectively addressing the challenges faced by LLMs in TCM, including interpretability, limited data modality, and lack of robustness. The code is open-sourced at: https://github.com/JerryMazeyu/MCM.
Chendan Liang, Zeyu Ma 0003, Wanying Wang, Minjie Ding, Mingang Chen
ICIP6
2025 Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
abstract
Large vision-language models (LVLMs) are vulnerable to harmful input compared to their language-only backbones. We investigated this vulnerability by exploring LVLMs internal dynamics, framing their inherent safety understanding in terms of three key capabilities. Specifically, we define these capabilities as safety perception, semantic understanding, and alignment for linguistic expression, and experimentally pinpointed their primary locations within the model architecture. The results indicate that safety perception often emerges before comprehensive semantic understanding, leading to the reduction in safety. Motivated by these findings, we propose Self-Aware Safety Augmentation (SASA), a technique that projects informative semantic representations from intermediate layers onto earlier safety-oriented layers. This approach leverages the model's inherent semantic understanding to enhance safety recognition without fine-tuning. Then, we employ linear probing to articulate the model's internal semantic comprehension to detect the risk before the generation process. Extensive experiments on various datasets and tasks demonstrate that SASA significantly improves the safety of LVLMs, with minimal impact on the utility.
Wanying Wang, Zeyu Ma 0003, Xin Tan 0002, Mingang Chen
ACM Multimedia5
2024 An Evaluation System for Large Language Models based on Open-Ended Questions
abstract
We designed a large language model evaluation system based on open-ended questions. The system accomplished multidimensional evaluation of LLMs using open-ended questions, and it presented evaluation results with evaluation reports. Currently, the evaluation of large-scale language models often exists with two prominent limitations: (1) The evaluation methods are often single-minded, resulting in less credible results. (2) Most evaluations are based on datasets with closed-ended questions, treating generative large language models as discriminative models, which fails to adequately reflect the high output flexibility characteristic of these models. For these two limitations, we proposed an evaluation system for LLMs based on open-ended questions. Our experiments on the adapted open-source datasets demonstrated the effectiveness of this system. The code of the system was released on https://github.com/JerryMazeyu/GreatLibrarian.
Zeyu Ma 0003, Mingang Chen
CSCloud3
2024 Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection
abstract
Industrial anomaly detection (I AD) has garnered signif-icant attention and experienced rapid development. However, the recent development of I AD approach has encountered certain difficulties due to dataset limitations. On the one hand, most of the state-of-the-art methods have achieved saturation (over 99% in AUROC) on mainstream datasets such as MVTec, and the differences of methods cannot be well distinguished, leading to a significant gap between public datasets and actual application scenarios. On the other hand, the research on various new practical anomaly detection settings is limited by the scale of the dataset, posing a risk of overfitting in evaluation results. Therefore, we propose a large-scale, Real-world, and multi-view Industrial Anomaly Detection dataset, named Real- I AD, which contains 150K high-resolution images of 30 different objects, an order of magnitude larger than existing datasets. It has a larger range of defect area and ratio proportions, making it more challenging than previous datasets. To make the dataset closer to real application scenarios, we adopted a multi-view shooting method and proposed sample-level evaluation metrics. In addition, beyond the general unsupervised anomaly detection setting, we propose a new setting for Fully Unsupervised Indus-trial Anomaly Detection (FUIAD) based on the observation that the yield rate in industrial production is usually greater than 60%, which has more practical application value. Finally, we report the results of popular I AD methods on the Real- I AD dataset, providing a highly challenging benchmark to promote the development of the I AD field.
Chengjie Wang 0001, Wenbing Zhu, Bin-Bin Gao, Zhenye Gan, Jiangning Zhang, Shuguang Qian, Mingang Chen, Lizhuang Ma
CVPR8
2024 UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
Qingdong He, Jinlong Peng, Zhengkai Jiang 0001, Xiaozhong Ji, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Mingang Chen, Yunsheng Wu
IJCAI9
2024 Glass Makes Blurs: Learning the Visual Blurriness for Glass Surface Detection
abstract
Glass surface detection is challenging as glass normally borrows similar visual appearances from the arbitrary objects/scenes behind it. Although some methods have been proposed to address this problem, they may fail if the reference objects are nonexistent or the additional annotations are missing. This article aims to address the glass surface detection problem by utilizing the intrinsic glass properties without reference objects and additional annotations. We observe glass makes blurs naturally. Based on the investigation of this intrinsic visual blurriness cue, we propose a novel visual blurriness aggregation module to model visual blurriness as a learnable residual in order to extract and aggregate multiscale valuable visual blurriness features used for guiding the backbone features to detect glass precisely. Besides, we note the ratio of the blurred area assists in utilizing the visual blurriness cue caused by glass and propose a visual blurriness driven refinement module to refine glass maps with this ratio to better leverage the visual blurriness information. Extensive experiments show that the proposed method achieves state-of-the-art performance on popular glass surface datasets.
Fulin Qi, Xin Tan 0002, Zhizhong Zhang 0001, Mingang Chen, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Ind. Informatics4
2023 Learning to Measure the Point Cloud Reconstruction Loss in a Representation Space
abstract
For point cloud reconstruction-related tasks, the reconstruction losses to evaluate the shape differences between reconstructed results and the ground truths are typically used to train the task networks. Most existing works measure the training loss with point-to-point distance, which may introduce extra defects as predefined matching rules may deviate from the real shape differences. Although some learning-based works have been proposed to overcome the weaknesses of manually-defined rules, they still measure the shape differences in 3D Euclidean space, which may limit their ability to capture defects in reconstructed shapes. In this work, we propose a learning-based Contrastive Adver-sarial Loss (CALoss) to measure the point cloud reconstruction loss dynamically in a non-linear representation space by combining the contrastive constraint with the adversarial strategy. Specifically, we use the contrastive constraint to help CALoss learn a representation space with shape similarity, while we introduce the adversarial strategy to help CALoss mine differences between reconstructed results and ground truths. According to experiments on reconstruction-related tasks, CALoss can help task networks improve re-construction performances and learn more representative representations.
Tianxin Huang, Zhonggan Ding, Jiangning Zhang, Ying Tai, Zhenyu Zhang 0005, Mingang Chen, Chengjie Wang 0001, Yong Liu 0007
CVPR6
2023 Multi-domain mixup for scenario-universal face anti-spoofing
Shitao Lu, Shice Liu, Keyue Zhang, Mingang Chen, Xin Tan 0002, Lizhuang Ma
Comput. Graph.4
2022 Optimal Transport for Label-Efficient Visible-Infrared Person Re-Identification
Jiangming Wang, Zhizhong Zhang 0001, Mingang Chen, Cong Wang 0039, Bin Sheng 0001, Yanyun Qu, Yuan Xie 0006
ECCV (24)3
2022 ScatterNet: Point Cloud Learning via Scatters
abstract
Design of point cloud shape descriptors is a challenging problem in practical applications due to the sparsity and the inscrutable distribution of the point clouds. In this paper, we propose ScatterNet, a novel 3D local feature learning approach for exploring and aggregating hypothetical scatters of the point clouds. Scatters of relational points are first organized in point cloud via guided explorations, and then propagated back to extend the capacity in representing the point-wise characteristics. We provide an practical implementation of the ScatterNet, which involves an unique scatter exploration operator and a scatter convolution operator. Our method achieves the state-of-the-art performance on several point cloud analysis tasks like classification, part segmentation and normal estimation. The source code of ScatterNet is available in supplementary materials.
Nianjuan Jiang, Jiangbo Lu, Mingang Chen, Ran Yi 0002, Lizhuang Ma
ACM Multimedia4
2022 Joint Learning Content and Degradation Aware Feature for Blind Super-Resolution
abstract
To achieve promising results on blind image super-resolution (SR),some attempts leveraged the low resolution (LR) images to predict the kernel and improve the SR performance. However, these Supervised Kernel Prediction (SKP) methods are impractical due to the unavailable real-world blur kernels. Although some Unsupervised Degradation Prediction (UDP) methods are proposed to bypass this problem, the inconsistency between degradation embedding and SR feature is still challenging. By exploring the correlations between degradation embedding and SR feature, we observe that jointly learning the content and degradation aware feature is optimal. Based on this observation, a Content and Degradation aware SR Network dubbed CDSR is proposed. Specifically, CDSR contains three newly-established modules: (1) a Lightweight Patch-based Encoder (LPE) is applied to jointly extract content and degradation features; (2) a Domain Query Attention based module (DQA) is employed to adaptively reduce the inconsistency; (3) a Codebook-based Space Compress module (CSC) that can suppress the redundant information. Extensive experiments on several benchmarks demonstrate that the proposed CDSR outperforms the existing UDP models and achieves competitive performance on PSNR and SSIM even compared with the state-of-the-art SKP methods.
Chuming Lin, Donghao Luo 0001, Yong Liu 0032, Ying Tai, Chengjie Wang 0001, Mingang Chen
ACM Multimedia7
2022 Automatic Collaborative Testing of Applications Integrating Text Features and Priority Experience Replay
abstract
With the popularity of deep reinforcement learning(DRL), people have great interest in using deep reinforcement learning for application automated testing. However, most automated testing methods based on reinforcement learning ignore text information, use random sampling in experience replay and ignore the characteristics of Android automated testing. To solve above problem, this paper proposes ITPRTesting(Integrated Text feature information and Priority experience in Testing). It extracts the text information in the interface and uses the BERT algorithm to generate sentence vectors. It fuses the interactive control feature diagram(ICFD), which is mentioned in the previous work, and text information as the state required by reinforcement learning. And in reinforcement learning, the priority experience replay is combined, also the traditional priority experience replay is improved. This paper has carried out experiments on 10 open source applications. The experimental results show that ITPRTesting is superior to other methods in statement coverage and branch coverage.
Lizhi Cai, Mingang Chen, Jilong Wang 0009
QRS3
2021 Situational Awareness Platform Based on Multi-source Vulnerability Fusion
abstract
With the increasingly severe cyberspace security situation, situational awareness has become a focus in the field of Cyberspace Security. The effect of situational awareness decision-making depends on the breadth and depth of data, the rationality of data processing, and the clarity and intuition of the presentation method. This paper designs a situational awareness platform, introduces vulnerability identification and fusion, situational awareness and data visualization. In particular, the data fusion method of online and offline multi-source vulnerability is proposed to enrich the data sources, and the soft splicing technology of display arrays is used to improve the visual compatibility under high resolution. In this paper, the data processing method is continuously optimized through experiments, which improves the overall processing efficiency and prediction accuracy of the platform.
Mingang Chen, Jiayu Gong
ICIS2
2021 An Efficient Method to Measure Robustness of ReLU-Based Classifiers via Search Space Pruning
abstract
Deep Neural Networks (DNNs) have achieved high accuracy on image classification. However, a small disturbance to an input may fool the networks to misclassify the label, which can cause a series of security and social problems. Thus, the robustness of DNNs must be ensured, particularly to those safety-critical systems. In this paper, we focus on the problem of measuring the robustness of ReLU-based DNNs, which can be equivalently formulated to solve a Mixed Integer Linear Programming problem (MILP). The complexity of solving MILP is directly related to the number of integer variables. We propose an efficient method for robustness measurement and verification by pruning the search space of MILP problems. Particularly, we design a greedy algorithm based on linear programming (LP) to determine the reasonable boundary. Then the search space is pruned by setting the boundary to integer variables in MILP. The comparison experiments on five classifiers trained on MNIST and CIFAR-10 datasets show our method outperforms other related tools in terms of efficiency and accuracy.
Xinping Wang, Liangyu Chen 0001, Tong Wang 0042, Mingang Chen, Min Zhang 0002
IJCNN4
2021 English Cloze Test Based on BERT
Minjie Ding, Mingang Chen, Lizhi Cai
KSEM2
2019 Classification of Skin Lesions Based on Data Collaboration Under Imbalance Dataset
Weijia Ji, Lizhi Cai, Mingang Chen, Naiqi Wang
CollaborateCom3
2012 Efficient video cutout based on adaptive multilevel banded method
Mingang Chen, BoCong Sui, Yan Gao 0004, Lizhuang Ma
Sci. China Inf. Sci.1
2012 An improved approach to the efficient construction of and search operations in motion graphs
Yan Gao 0004, Mingang Chen, Yang Shen 0011
Sci. China Inf. Sci.3
2012 A gradient-domain-based edge-preserving sharpen filter
Rynson W. H. Lau, Yan Gui, Mingang Chen, Lizhuang Ma
Vis. Comput.4