Mingjie Sun

dblp:54/3913 · DBLP profile ↗
← Back
54ranked-venue papers
15as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 11 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 22 since 2021Software engineering, systems software and programming languages · 3Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Align³GR: Unified Multi-Level Alignment for LLM-based Generative Recommendation
abstract
Large Language Models (LLMs) demonstrate significant advantages in leveraging structured world knowledge and multi-step reasoning capabilities. However, fundamental challenges arise when transforming LLMs into real-world recommendation systems due to semantic and behavioral misalignment. To bridge this gap, we propose Align³GR, a novel framework that unifies token-level, behavior modeling-level, and preference-level alignment. Our approach introduces: Dual tokenization fusing user-item semantic and collaborative signals. Enhanced behavior modeling with bidirectional semantic alignment. Progressive DPO strategy combining self-play (SP-DPO) and real-world feedback (RF-DPO) for dynamic preference adaptation. Experiments show Align³GR outperforms the SOTA baseline by +17.8% in Recall@10 and +20.2% in NDCG@10 on the public dataset, with significant gains in online A/B tests and full-scale deployment on an industrial large-scale recommendation platform.
Wencai Ye, Mingjie Sun, Wenjin Wu, Peng Jiang 0002
AAAI2
2026 Rethinking hard training sample generation for medical image segmentation
Zhibin Wan, Mingjie Sun, Cao Min, Guohong Fu
Pattern Recognit.3
2026 MvP-Diff: Multivariate yet precise diffusion for anomaly images synthesis and segmentation
Siyue Yao, Eng Gee Lim, Siyue Yu, Jimin Xiao, Mingjie Sun
Pattern Recognit.6
2025 ConSense: Continually Sensing Human Activity with WiFi via Growing and Picking
abstract
WiFi-based human activity recognition (HAR) holds significant application potential across various fields. To handle dynamic environments where new activities are continuously introduced, WiFi-based HAR systems must adapt by learning new concepts without forgetting previously learned ones. Furthermore, retaining knowledge from old activities by storing historical exemplar is impractical for WiFi-based HAR due to privacy concerns and limited storage capacity of edge devices. In this work, we propose ConSense, a lightweight and fast-adapted exemplar-free class incremental learning framework for WiFi-based HAR. The framework leverages the transformer architecture and involves dynamic model expansion and selective retraining to preserve previously learned knowledge while integrating new information. Specifically, during incremental sessions, small-scale trainable parameters that are trained specifically on the data of each task are added in the multi-head self-attention layer. In addition, a selective retraining strategy that dynamically adjusts the weights in multilayer perceptron based on the performance stability of neurons across tasks is used. Rather than training the entire model, the proposed strategies of dynamic model expansion and selective retraining reduce the overall computational load while balancing stability on previous tasks and plasticity on new tasks. Evaluation results on three public WiFi datasets demonstrate that ConSense not only outperforms several competitive approaches but also requires fewer parameters, highlighting its practical utility in class-incremental scenarios for HAR.
Tao Deng 0003, Siwei Feng, Mingjie Sun, Juncheng Jia
AAAI4
2025 DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System
abstract
Semantic IDs are discrete identifiers generated by quantizing the Multi-modal Large Language Models embeddings, enabling efficient multi-modal content integration in recommendation systems. However, their lack of collaborative signals results in a misalignment with downstream discriminative and generative recommendation objectives. Recent studies have introduced various alignment mechanisms to address this problem, but their two-stage framework design still leads to two main limitations: (1) inevitable information loss during alignment, and (2) inflexibility in applying adaptive alignment strategies, consequently constraining the mutual information maximization during the alignment process.
Wencai Ye, Mingjie Sun, Shaoyun Shi, Wenjin Wu, Peng Jiang 0002
CIKM2
2025 Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
abstract
Deep learning has revolutionized medical image segmentation, yet its full potential remains constrained by the paucity of annotated datasets. While diffusion models have emerged as a promising approach for generating synthetic image-mask pairs to augment these datasets, they paradoxically suffer from the same data scarcity challenges they aim to mitigate. Traditional mask-only models frequently yield low-fidelity images due to their inability to adequately capture morphological intricacies, which can critically compromise the robustness and reliability of segmentation models. To alleviate this limitation, we introduce Siamese-Diffusion, a novel dual-component model comprising Mask-Diffusion and Image-Diffusion. During training, a Noise Consistency Loss is introduced between these components to enhance the morphological fidelity of Mask-Diffusion in the parameter space. During sampling, only Mask-Diffusion is used, ensuring diversity and scalability. Comprehensive experiments demonstrate the superiority of our method. Siamese-Diffusion boosts SANet’s mDice and mIoU by 3.6% and 4.4% on the Polyps, while UNet improves by 1.52% and 1.64% on the ISIC2018.
Kunpeng Qiu, Zhiying Zhou, Mingjie Sun, Yongxin Guo 0002
CVPR4
2025 Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic Segmentation
Siyue Yu, Bingfeng Zhang, Mingjie Sun, Yi Dong 0002, Jimin Xiao
ICCV4
2025 SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Zhentao Tan, Ben Xue, Jian Jia, Wencai Ye, Shaoyun Shi, Mingjie Sun, Wenjin Wu, Quan Chen 0006, Peng Jiang 0002
ICCV7
2025 Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
abstract
Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show for the first time that these concentrated massive values consistently emerge in specific regions of attention queries (Q) and keys (K) while not having such patterns in values (V) in various modern transformer-based LLMs. Through extensive experiments, we further demonstrate that these massive values play a critical role in interpreting contextual knowledge (i.e., knowledge obtained from the current context window) rather than in retrieving parametric knowledge stored within the model’s parameters. Our further investigation of quantization strategies reveals that ignoring these massive values leads to a pronounced drop in performance on tasks requiring rich contextual understanding, aligning with our analysis. Finally, we trace the emergence of concentrated massive values and find that such concentration is caused by Rotary Positional Encoding (RoPE) and it appears since very first layers. These findings shed new light on how Q and K operate in LLMs and offer practical insights for model design and optimization. The code is available at https://github.com/MingyuJ666/Rope_with_LLM.
Mingyu Jin, Kai Mei, Wujiang Xu, Mingjie Sun, Ruixiang Tang, Mengnan Du, Zirui Liu 0001, Yongfeng Zhang 0003
ICML4
2025 Idiosyncrasies in Large Language Models
abstract
In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) – unique patterns in their outputs that can be used to distinguish the models. To do so, we consider a simple classification task: given a particular text output, the objective is to predict the source LLM that generates the text. We evaluate this synthetic task across various groups of LLMs and find that simply fine-tuning text embedding models on LLM-generated texts yields excellent classification accuracy. Notably, we achieve 97.1% accuracy on held-out validation data in the five-way classification problem involving ChatGPT, Claude, Grok, Gemini, and DeepSeek. Our further investigation reveals that these idiosyncrasies are rooted in word-level distributions. These patterns persist even when the texts are rewritten, translated, or summarized by an external LLM, suggesting that they are also encoded in the semantic content. Additionally, we leverage LLM as judges to generate detailed, open-ended descriptions of each model’s idiosyncrasies. Finally, we discuss the broader implications of our findings, including training on synthetic data, inferring model similarity, and robust evaluation of LLMs.
Mingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter, Zhuang Liu 0003
ICML1
2025 Free Lunch of Image-mask Alignment for Anomaly Image Generation and Segmentation
abstract
This paper aims at generating anomalous images and their segmentation labels to address the lack of real-world anomaly samples and privacy issues. Departing from conventional approaches that use masks solely to guide the generation of anomaly images, we propose a dual-branch training strategy for the generative model. This strategy enables the simultaneous production of anomaly images and masks, with an alignment regularization loss that ensures the coherence between the generated images and their masks. During inference, only the image-generation branch is activated to produce synthetic samples for training the downstream segmentation model. Furthermore, we propose to integrate the well-trained generative model into the training of segmentation models, utilizing a generative feedback loss to refine the segmentation model's performance. Experiments show our method's IoU metrics exceed previous methods by 5.03%, 5.68% and 16.63% on Real-IAD (industrial), polyp (medical), and Floor Dirty (indoor) datasets. The code is publicly accessible at https://github.com/huan-yin/anomaly-alignment.
Xiangyue Li, Xiaoyang Wang 0007, Zhibin Wan, Yupei Wu, Mingjie Sun
IJCAI7
2025 DriftRemover: Hybrid Energy Optimizations for Anomaly Images Synthesis and Segmentation
abstract
This paper tackles the challenge of anomaly image synthesis and segmentation to generate various anomaly images and their segmentation labels to mitigate the issue of data scarcity. Existing approaches employ the precise mask to guide the generation, relying on additional mask generators, leading to increased computational costs and limited anomaly diversity. Although a few works use coarse masks as the guidance to expand diversity, they lack effective generation of labels for synthetic images, thereby reducing their practicality. Therefore, our proposed method simultaneously generates anomaly images and their corresponding masks by utilizing coarse masks and anomaly categories. The framework utilizes attention maps from synthesis process as mask labels and employs two optimization modules to tackle drift challenges, which are mismatches between synthetic results and real situations. Our evaluation demonstrates that our method improves pixel-level AP by 1.3% and F1-MAX by 1.8% in anomaly detection tasks on the MVTec dataset. Additionally, its successful application in practical scenarios highlights its effectiveness, improving IoU by 37.2% and F-measure by 25.1% with the Floor Dirt dataset. The code is available at https://github.com/JJessicaYao/DriftRemover.
Siyue Yao, Mingjie Sun, Siyue Yu, Jimin Xiao, Eng Gee Lim
IJCAI3
2025 Training-Free Clothing Region of Interest Self-correction for Virtual Try-On
Shengjie Lu, Zhibin Wan, Jiejie Liu, Mingjie Sun
PRICAI5
2025 Topological Fusion Model for Molecular Property Prediction
abstract
Abstract The prominence of 3D molecular property prediction arises from its ability to provide insights into the drug discovery and design, material science and chemical synthesis. Transformer-based models have been widely adopted to autonomously learn long-range atom-to-atom interactions on a global scale, resulting in significant success. However, these models may struggle to capture intricate substructure details (e.g., covalent bond and functional group). In this work, topological simplices defined on nodes, links, triangles are extracted from the atoms’ 3D positional information to provide comprehensive representations of the local substructure information, such as atoms, covalent bonds and functional groups. We then propose a topological fusion network, which enhances each atom’s features not only through global atom-to-atom interactions but also by incorporating the fine-grained topological substructure information. In comparison to existing popular methods, our proposed method outperforms the state-of-the-art (SOTA) method by 1.2%, 3.0%, 2.4%, 2.7% on BBBP, BACE, ClinTox, MUV datasets for classification task and 0.048, 0.022, 3.8 on FreeSolv, Lipo and QM7 datasets for regression task, respectively. The code will be released soon.
Xia Rong, Junwei Wu 0001, Zhang Shufei, Mingjie Sun, Jiejie Liu
Appl. Intell.5
2025 Class agnostic and specific consistency learning for weakly-supervised point cloud semantic segmentation
Junwei Wu 0001, Mingjie Sun, Chenru Jiang, Wuwei Ma
Pattern Recognit.2
2025 Auxiliary captioning: Bridging image-text matching and image captioning
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001
Signal Process. Image Commun.3
2025 Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-Scarce Scenarios
abstract
The scarcity of data in various scenarios, such as medical, industry and autonomous driving, leads to model overfitting and dataset imbalance, thus hindering effective detection and segmentation performance. Existing studies employ the generative models to synthesize more training samples to mitigate data scarcity. However, these synthetic samples are repetitive or simplistic and fail to provide "crucial information" that targets the downstream model's weaknesses. Additionally, these methods typically require separate training for different objects, leading to computational inefficiencies. To address these issues, we propose Crucial-Diff, a domain-agnostic framework designed to synthesize crucial samples. Our method integrates two key modules. The Scene Agnostic Feature Extractor (SAFE) utilizes a unified feature extractor to capture target information. The Weakness Aware Sample Miner (WASM) generates hard-to-detect samples using feedback from the detection results of downstream model, which is then fused with the output of SAFE module. Together, our Crucial-Diff framework generates diverse, high-quality training data, achieving a pixel-level AP of 83.63% and an F1-MAX of 78.12% on MVTec. On polyp dataset, Crucial-Diff reaches an mIoU of 81.64% and an mDice of 87.69%. Code is publicly available at https://github.com/JJessicaYao/Crucial-diff.
Siyue Yao, Mingjie Sun, Eng Gee Lim, Ran Yi 0002, Baojiang Zhong, Moncef Gabbouj
IEEE Trans. Image Process.2
2025 Trust Online Over-the-Air Computation for Wireless Federated Learning
abstract
Using the wireless waveform superposition property, over-the-air computation (OAC) enables federated learning (FL) to achieve fast model aggregation. However, this computing paradigm is vulnerable to poisoning attacks due to the openness of a wireless channel over time, where malicious mobile devices can introduce cumulative errors for the global FL model in a time-varying wireless environment for each communication round. This article presents a trust online OAC (TO-OAC) scheme to minimize impacts on the global model introduced by malicious devices adjusting to dynamic attack and wireless channel fluctuations over time. TO-OAC achieves this by utilizing trustworthy security quantification of OAC for each FL training round. To optimize the cumulative training loss at the aggregation node with the long-term power and trust constraints of mobile devices, we propose a joint trust, power, and channel-aware algorithm to flexibly update local and global models in response to the dynamic changes in the wireless and secure environment. We analyze the performance limits for the aggregation of trust models, considering metrics for computation and communication through time. We then propose another trust online regularization over-the-air computation (TOR-OAC) as an improved version of the TO-OAC scheme to decrease convergence time while ensuring long-term trust and power limitation. Experimental results performed on real-life datasets show that the two proposed schemes (TO-OAC and TOR-OAC) outperform prior works, especially in noisy, time-varying wireless channels and malicious attacks.
Mingjie Sun, Jie Zheng 0005, Hongyang Du 0001, Haijun Zhang 0001, Dusit Niyato, Jiawen Kang 0001, Jiacheng Wang 0001, Jie Ren 0007, Zheng Wang 0001
IEEE Trans. Mob. Comput.1
2025 NetEventCause: Event-Driven Root Cause Analysis for Large Network System Without Topology
abstract
Root cause analysis (RCA) is a crucial technique in network systems for uncovering the abnormal nodes that lead to the network alarm flood. Within private cloud network systems, the calling chains and topologies among entities, such as hosts, routes, and services, are always incomplete due to nonstandardized management. Existing topology-free RCA techniques, which rely on the casual discovery, are inapplicable when the scale of the network system is extremely large or the number of triggered alarms is sparse. This article proposes NetEventCause (NEC), an event-driven, unsupervised, and nonintrusive RCA algorithm for large network systems, where the network topology is unknown. NEC learns from historical alarm events to model the occurrences of various alarm types using a multivariate neural temporal point process (TPP). Based on the conditional intensity predicted by the learned TPP, NEC can identify the root alarms from a cascade of alarm events and locate the causal alarms of derivative alarms using the attribution method. The experimental section evaluates the NEC using both a synthetic event dataset and a large real-world dataset. The real-world dataset is exported from the Huawei Shennong Intelligent Maintenance and Operation Center (IMOC), a platform deployed at one of China's largest airports and manages over 200000 entities. Results obtained from the two datasets demonstrate that NEC outperforms most state of the art (SOTA) TPP models in modeling alarm events and surpasses general RCA methods in terms of identifying root alarms and recovering transmission chains of anomalies.
Zhaolin Yuan, Wenjia Wei, Mingjie Sun, Duxin Chen
IEEE Trans. Neural Networks Learn. Syst.5
2025 SFBM: Shared Feature Bias Mitigating for Long-Tailed Image Recognition
abstract
Long-tailed distribution exists in real-world scenario and compromises the performance of recognition models. In this article, we point out that a neural network classifier has a shared feature bias, which tends to regard the shared features among different classes as head-class discriminative features, leading to misclassifications on tail-class samples under long-tailed scenarios. To solve this issue, we propose a shared feature bias mitigating (SFBM) framework. Specifically, we create two parallel classifiers trained concurrently with the baseline classifier, using our special training loss. The parallel classifier weight sums are then used for estimating the shared feature components in baseline classifier weights. Finally, we rectify the baseline classifier by removing the estimated shared feature components from it while supplementing the parallel classifier weights class by class to the rectified classifier weights, mitigating shared feature bias. Our proposed SFBM demonstrates broad compatibility with nearly all recognition methods while maintaining high computational efficiency, as it introduces no additional computation during inference. Extensive experiments on CIFAR10/100-LT, ImageNet-LT, and iNaturalist 2018 demonstrate that simply incorporating SFBM during the training phase consistently boosts the performance of various state-of-the-art methods by significant margins. The complete source code will be made publicly available at https://github.com/bzbz-bot/SFBM.
Xinqiao Zhao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001, Jimin Xiao
IEEE Trans. Neural Networks Learn. Syst.2
2024 PromptCD: Coupled and Decoupled Prompt Learning for Vision-Language Models
abstract
Large-scale pre-trained vision-language models (VLMs), like CLIP, have presented striking generalizability for adapting to image classification in a few shot setting. Most existing methods explore a set of learnable tokens, such as prompt learning, on data-efficient utilization for task adaptation. However, they focus on either the coupled-modality property by prompt projection or decoupled-modality characteristic by prompt consistency, which ignores effective interaction between prompts. To model the deep yet sufficient cross-modal interaction and enhance the generalization between both seen and unseen tasks, in this paper, we propose a novel coupled and decoupled prompt learning framework, dubbed PromptCD, for vision-language models. Specifically, we introduce a bi-directional coupled-modality mechanism to intensify the interaction between both vision and language branches. Additionally, we propose mixture consistency to further improve the generalization and discrimination of the models on unseen tasks. The integration of such a mechanism and consistency facilitates the proposed framework adaptation for various downstream tasks. We conduct extensive experiments on 11 image classification datasets under a range of evaluation protocols, including base-to-novel and domain generalization, and cross-dataset recognition. Experimental results demonstrate that our proposed PromptCD overall outperforms state-of-the-art methods.
Junjie Wu 0005, Mingjie Sun, Chen Gong 0004, Guohong Fu
ECAI2
2024 Adversarial Erasing Transformer for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation has attracted a lot of attention recently. Previous methods can be divided into two types, which are single-stage training and multi-stage training. In this paper, we focus on multi-stage training for image-level weakly supervised semantic segmentation. Many recent methods have tried to use transformer architecture as the backbone for CAM generation since it can capture global relationships to refine CAM accurately. However, we observe that such a backbone still fails to generate complete and smooth CAM. We argue that this is because the attention mechanism in the transformer can only pay attention to the most discriminative relationships. It is difficult to capture semantic-level long-range pair-wise relationships under image-level supervision. Thus, we propose an adversarial erasing transformer network called AETN, where an erasing attention mechanism is designed to establish more extensive pair-wise relationships. To cope with erasing, more target features will be forced to activate. Thus, better feature representation can be obtained for more accurate CAM generation. Besides, to further help our network learn better feature representation, we propose a self-consistent learning mechanism based on different augmentations. In this way, our AETN outperforms recent methods. Our AETN achieves 73.0 mIoU on the PASCAL VOC 2012 val set and 73.9 mIoU on the PASCAL VOC 2012 test set. Code is available a https://github.com/siyueyu/AETN.
Bingfeng Zhang, Siyue Yu, Xuru Gao, Mingjie Sun, Eng Gee Lim, Jimin Xiao
ECAI4
2024 A Simple and Effective Pruning Approach for Large Language Models
abstract
As their size increases, Large Languages Models (LLMs) are natural candidates for network pruning methods: approaches that drop a subset of network weights while striving to preserve performance. Existing methods, however, require either retraining, which is rarely affordable for billion-scale LLMs, or solving a weight reconstruction problem reliant on second-order information, which may also be computationally expensive. In this paper, we introduce a novel, straightforward yet effective pruning method, termed Wanda (Pruning by Weights and activations), designed to induce sparsity in pretrained LLMs. Motivated by the recent observation of emergent large magnitude features in LLMs, our approach prunes weights with the smallest magnitudes multiplied by the corresponding input activations, on a per-output basis. Notably, Wanda requires no retraining or weight update, and the pruned LLM can be used as is. We conduct a thorough evaluation of our method Wanda on LLaMA and LLaMA-2 across various language benchmarks. Wanda significantly outperforms the established baseline of magnitude pruning and performs competitively against recent method involving intensive weight update.
Mingjie Sun, Zhuang Liu 0003, Anna Bair, J. Zico Kolter
ICLR1
2024 Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
abstract
Recently, Miller et al. (2021) and Baek et al. (2022) empirically demonstrated strong linear correlations between in-distribution (ID) versus out-of-distribution (OOD) accuracy and agreement. These trends, coined accuracy-on-the-line (ACL) and agreement-on-the-line (AGL), enable OOD model selection and performance estimation without labeled data. However, these phenomena also break for certain shifts, such as CIFAR10-C Gaussian Noise, posing a critical bottleneck. In this paper, we make a key finding that recent test-time adaptation (TTA) methods not only improve OOD performance, but it drastically strengthen the ACL and AGL trends in models, even in shifts where models showed very weak correlations before. To analyze this, we revisit the theoretical conditions from Miller et al. (2021) that outline the types of distribution shifts needed for perfect ACL in linear models. Surprisingly, these conditions are satisfied after applying TTA to deep models in the penultimate feature embedding space. In particular, TTA causes the data distribution to collapse complex shifts into those can be expressed by a singular "scaling" variable in the feature space. Our results show that by combining TTA with AGL-based estimation methods, we can estimate the OOD performance of models with high precision for a broader set of distribution shifts. This lends us a simple system for selecting the best hyperparameters and adaptation strategy without any OOD labeled data. Code is available at https://github.com/EungyeupKim/TTALine.
Eungyeup Kim, Mingjie Sun, Christina Baek, Aditi Raghunathan, J. Zico Kolter
NeurIPS2
2024 Context-based local-global fusion network for 3D point cloud classification and segmentation
Junwei Wu 0001, Mingjie Sun, Chenru Jiang, Jiejie Liu, Jeremy S. Smith
Expert Syst. Appl.2
2024 Prototype Guided Pseudo Labeling and Perturbation-based Active Learning for domain adaptive semantic segmentation
Junkun Peng, Mingjie Sun, Eng Gee Lim, Qiufeng Wang 0001, Jimin Xiao
Pattern Recognit.2
2024 Unified Multi-Modality Video Object Segmentation Using Reinforcement Learning
abstract
The main task we aim to tackle is the multi-modality video object segmentation (VOS), which can be divided into two sub-tasks: mask-referred and language-referred VOS, where the first-frame mask-level or language-level label is utilized to provide the target information, respectively. Due to the huge gap between different modalities, existing works never come up with a unified framework for these two sub-tasks. In this work, such a unified framework is designed, where the visual and linguistic inputs are first spilt into a number of image patches and words, and then mapped into same-size tokens, which are equally processed by a self-attention based segmentation model. Furthermore, to highlight the significant information and discard the non-target or ambiguous one, unified multi-modality filter networks are further designed, and reinforcement learning is adopted to optimize such networks. Experiments show that new state-of-the-art performances are achieved by the proposed method: 52.8% ofJ&Fon Ref-YoutubeVOS dataset and 83.2% ofJSon YoutubeVOS dataset, respectively. The code will be released.
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Cairong Zhao, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Single Image Backdoor Inversion via Robust Smoothed Classifiers
abstract
Backdoor inversion, the process of finding a backdoor “trigger” inserted into a machine learning model, has become the pillar of many backdoor detection and defense methods. Previous works on backdoor inversion often recover the backdoor through an optimization process to flip a support set of clean images into the target class. However, it is rarely studied and understood how large this support set should be to recover a successful backdoor. In this work, we show that one can reliably recover the backdoor trigger with as few as a single image. Specifically, we propose the SmoothInv method, which first constructs a robust smoothed version of the backdoored classifier and then performs guided image synthesis towards the target class to reveal the backdoor pattern. SmoothInv requires neither an explicit modeling of the backdoor via a mask variable, nor any complex regularization schemes, which has become the standard practice in backdoor inversion methods. We perform both quantitaive and qualitative study on backdoored classifiers from previous published backdoor attacks. We demonstrate that compared to existing methods, SmoothInv is able to recover successful backdoors from single images, while maintaining high fidelity to the original backdoor. We also show how we identify the target backdoored class from the backdoored classifier. Last, we propose and analyze two countermeasures to our approach and show that SmoothInv remains robust in the face of an adaptive attacker. Our code is available at https://github.com/locuslab/smoothinv.
Mingjie Sun, J. Zico Kolter
CVPR1
2023 (Certified!!) Adversarial Robustness for Free!
Nicholas Carlini, Florian Tramèr, Krishnamurthy Dvijotham, Leslie Rice, Mingjie Sun, J. Zico Kolter
ICLR5
2023 Dance with You: The Diversity Controllable Dancer Generation via Diffusion Models
abstract
Recently, digital humans for interpersonal interaction in virtual environments have gained significant attention. In this paper, we introduce a novel multi-dancer synthesis task called partner dancer generation, which involves synthesizing virtual human dancers capable of performing dance with users. The task aims to control the pose diversity between the lead dancer and the partner dancer. The core of this task is to ensure the controllable diversity of the generated partner dancer while maintaining temporal coordination with the lead dancer. This scenario varies from earlier research in generating dance motions driven by music, as our emphasis is on automatically designing partner dancer postures according to pre-defined diversity, the pose of lead dancer, as well as the accompanying tunes. To achieve this objective, we propose a three-stage framework called Dance-with-You (DanY). Initially, we employ a 3D Pose Collection stage to collect a wide range of basic dance poses as references for motion generation. Then, we introduce a hyper-parameter that coordinates the similarity between dancers by masking poses to prevent the generation of sequences that are over-diverse or consistent. To avoid the rigidity of movements, we design a Dance Pre-generated stage to pre-generate these masked poses instead of filling them with zeros. After that, a Dance Motion Transfer stage is adopted with leader sequences and music, in which a multi-conditional sampling formula is rewritten to transfer the pre-generated poses into a sequence with a partner style. In practice, to address the lack of multi-person datasets, we introduce AIST-M, a new dataset for partner dancer generation, which is publicly availiable at https://github.com/JJessicaYao/AIST-M-Dataset. Comprehensive evaluations on our AIST-M dataset demonstrate that the proposed DanY can synthesize satisfactory partner dancer results with controllable diversity.
Siyue Yao, Mingjie Sun, Bingliang Li, Fengyu Yang 0005, Junle Wang, Ruimao Zhang
ACM Multimedia2
2023 Fully and Weakly Supervised Referring Expression Segmentation With End-to-End Learning
abstract
Referring Expression Segmentation (RES), which is aimed at localizing and segmenting the target according to the given language expression, has drawn increasing attention. Existing methods jointly consider the localization and segmentation steps, which rely on the fused visual and linguistic features for both steps. We argue that the conflict between the purpose of identifying an object and generating a mask limits the RES performance. To solve this problem, we propose a parallel position-kernel-segmentation pipeline to better isolate and then interact the localization and segmentation steps. In our pipeline, linguistic information will not directly contaminate the visual feature for segmentation. Specifically, the localization step localizes the target object in the image based on the referring expression, and then the visual kernel obtained from the localization step guides the segmentation step. This pipeline also enables us to train RES in a weakly-supervised way, where the pixel-level segmentation labels are replaced by click annotations on center and corner points. The position head is fully-supervised and trained with the click annotations as supervision, and the segmentation head is trained with weakly-supervised segmentation losses. To validate our framework on a weakly-supervised setting, we annotated three RES benchmark datasets (RefCOCO, RefCOCO+ and RefCOCOg) with click annotations. Our method is simple but surprisingly effective, outperforming all previous state-of-the-art RES methods on fully- and weakly-supervised settings by a large margin. The code and dataset will be released onhttps://github.com/detectiveli/PKS.git.
Hui Li 0085, Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Plausible Proxy Mining With Credibility for Unsupervised Person Re-Identification
abstract
One effective way to address unsupervised person re-identification is to use a clustering-based contrastive learning approach. Existing state-of-the-art methods adopt clustering algorithms (e.g., DBSCAN) and camera ID information to divide all person images into several camera-aware proxies. Then, for each person image, the extracted feature representation is pulled closer to the centroids of its pseudo-positive proxies (the proxies that share the same pseudo-identity label with this image) and pushed away from the centroids of other pseudo-negative proxies (the proxies that share the different pseudo-identity label with this image). However, the quality of the proxy centroid is significantly affected by the proxy impurity issue and thus deteriorates the learned feature representations. On the premise that we cannot introduce superior supervision signals by thoroughly solving the proxy impurity issue, for a person image, identifying its plausible proxies: the pseudo-negative proxies which potentially include its wrongly-clustered instances (the instances with the same ground-truth identity with this image), and further fixing the resulted incorrect supervision signals become an urgent and challenging problem. This paper proposes a simple yet effective approach to address this problem. With a given image, our method can effectively locate its plausible proxies. Then we introduce credibility to measure how much we should treat the centroid of each mined plausible proxy as a positive supervision signal rather than entirely negative. Extensive experiments on three widely-used person re-ID datasets validate the effectiveness of our proposed approach. Codes will be available at:https://github.com/Dingyuan-Zheng/PPCL.
Dingyuan Zheng, Jimin Xiao, Mingjie Sun, Huihui Bai 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.3
2023 Cycle-Free Weakly Referring Expression Grounding With Self-Paced Learning
abstract
In this paper, we are tackling the weakly referring expression grounding task to localize the target object in an image according to a given query sentence, where the mapping between the query sentence and image regions is blind during the training period. Previous methods all follow a cyclic forward-backward pipeline to handle this task, where the query sentence is firstly converted to the result region through the forward module, and then the result region is converted back to a sentence through the backward module, with the difference between the reconstructed sentence and original query used as the loss to optimize the entire network. These existing methods, however, suffer from the deviation issue when the result region, generated through the forward module, totally deviates from the target area, but the backward module still reconstructs a similar sentence. The aforementioned loss function cannot penalize this kind of deviation because of the consistent prediction of the sentence. To overcome this limitation, we propose a cycle-free pipeline, where a region describer network is designed to predict the textual description for each candidate region, and a result region is selected according to the similarity between the predicted description and the query sentence. Furthermore, a self-paced learning mechanism is designed to avoid the drift issue during the warm-up period of the optimization process. The proposed method achieves a higher average accuracy on RefCOCO and RefCOCO+ datasets, compared with all previous state-of-the-art methods.
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001
IEEE Trans. Multim.1
2023 Starting Point Selection and Multiple-Standard Matching for Video Object Segmentation With Language Annotation
abstract
In this study, we investigate language-level video object segmentation, where first-frame language annotation is used to describe the target object. Because a language label is typically compatible with all frames in a video, the proposed method can choose the most suitable starting frame to mitigate initialization failure. Apart from extracting the visual feature from a static video frame, a motion-language score based on optical flow is also proposed to describe moving objects more accurately. Scores of multiple standards are then aggregated using an attention-based mechanism to predict the final result. The proposed method is evaluated on four widely-used video object segmentation datasets, including the DAVIS 2017, DAVIS 2016, SegTrack V2 and YouTubeObject datasets, and a novel accuracy measured as mean region similarity is obtained on both the DAVIS 2017 (67.2%) and DAVIS 2016 (83.5%) datasets. The code will be published.
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001
IEEE Trans. Multim.1
2022 An Unsupervised Sentence Embedding Method by Maximizing the Mutual Information of Augmented Text Representations
Tianye Sheng, Lisong Wang, Zongfeng He, Mingjie Sun, Guohua Jiang
ICANN (2)4
2022 Chinese Named Entity Recognition Using the Improved Transformer Encoder and the Lexicon Adapter
Mingjie Sun, Lisong Wang, Tianye Sheng, Zongfeng He, Yuhua Huang
ICANN (2)1
2022 Spatial-Temporal Attention Network for Crime Prediction with Adaptive Graph Learning
Mingjie Sun, Peng Yuan Zhou, Yong Liao 0003, Haiyong Xie 0001
ICANN (2)1
2022 Multi-relation Word Pair Tag Space for Joint Entity and Relation Extraction
Mingjie Sun, Lisong Wang, Tianye Sheng, Zongfeng He, Yuhua Huang
ICONIP (2)1
2022 MLSAN: Mixed-Lattice Self-Attention Network for Chinese Named Entity Recognition
abstract
Named entity recognition (NER) is an essential subtask in natural language processing field. Recent studies have demonstrated that character-word lattice models are efficient for Chinese NER, which can leverage useful word boundary information to enhance the representation of characters. However, previous models only consider the integration of local matched word features and neglect the semantic interactions with long-range matched words. Moreover, prior methods solely achieve superficial fusion in the character-word feature space with simple methods, such as feature concatenation, but fail to implement fine-grained semantic fusion. In this paper, we propose a mixed-lattice self-attention network (MLSAN) to integrate richer word boundary information, which can explicitly capture the fine-grained correlations across characters and long-range matched words and achieve the integration of global word features. In addition, we design an end-to-end method for incorporating lexical information at the bottom layer of BERT. Compared with existing methods, our model achieves a deeper lexical knowledge fusion, which makes MLSAN perform well on cases of a small number of samples. Experimental results on four Chinese NER datasets show that our model obtains competitive performance.
Zongfeng He, Lisong Wang, Tianye Sheng, Mingjie Sun, Liang Liu 0006
ICPR4
2022 Test Time Adaptation via Conjugate Pseudo-labels
abstract
Test-time adaptation (TTA) refers to adapting neural networks to distribution shifts, specifically with just access to unlabeled test samples from the new domain at test-time. Prior TTA methods optimize over unsupervised objectives such as the entropy of model predictions in TENT (Wang et al., 2021), but it is unclear what exactly makes a good TTA loss. In this paper, we start by presenting a surprising phenomenon: if we attempt to $\textit{meta-learn}$ the ``best'' possible TTA loss over a wide class of functions, then we recover a function that is $\textit{remarkably}$ similar to (a temperature-scaled version of) the softmax-entropy employed by TENT. This only holds, however, if the classifier we are adapting is trained via cross-entropy loss; if the classifier is trained via squared loss, a different ``best'' TTA loss emerges.To explain this phenomenon, we analyze test-time adaptation through the lens of the training losses's $\textit{convex conjugate}$. We show that under natural conditions, this (unsupervised) conjugate function can be viewed as a good local approximation to the original supervised loss and indeed, it recovers the ``best'' losses found by meta-learning. This leads to a generic recipe than be used to find a good TTA loss for $\textit{any}$ given supervised training loss function of a general class. Empirically, our approach dominates other TTA alternatives over a wide range of domain adaptation benchmarks. Our approach is particularly of interest when applied to classifiers trained with $\textit{novel}$ loss functions, e.g., the recently-proposed PolyLoss (Leng et al., 2022) function, where it differs substantially from (and outperforms) an entropy-based loss. Further, we show that our conjugate based approach can also be interpreted as a kind of self-training using a very specific soft label, which we refer to as the $\textit{conjugate pseudo-label}$. Overall, therefore, our method provides a broad framework for better understanding and improving test-time adaptation. Code is available at https://github.com/locuslab/tta_conjugate.
Sachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico Kolter
NeurIPS2
2022 Transformer-Based Language-Person Search With Multiple Region Slicing
abstract
Language-person search is an essential technique for applications like criminal searching, where it is more feasible for a witness to provide language descriptions of a suspect than providing a photo. Most existing works treat the language-person pair as a black-box, neither considering the inner structure in a person picture, nor the correlations between image regions and referring words. In this work, we propose a transformer-based language-person search framework with matching conducted between words and image regions, where a person picture is vertically separated into multiple regions using two different ways, including the overlapped slicing and the key-point-based slicing. The co-attention between linguistic referring words and visual features are evaluated via transformer blocks. Besides the obtained outstanding searching performance, the proposed method enables to provide interpretability by visualizing the co-attention between image parts in the person picture and the corresponding referring words. Without bells and whistles, we achieve the state-of-the-art performance on the CUHK-PEDES dataset with Rank-1 score of 57.67% and the PA100K dataset with mAP of 22.88%, with simple yet elegant design. Code is available onhttps://github.com/detectiveli/T-MRS.
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
abstract
In this paper, we are tackling the proposal-free referring expression grounding task, aiming at localizing the target object according to a query sentence, without relying on off-the-shelf object proposals. Existing proposal-free methods employ a query-image matching branch to select the highest-score point in the image feature map as the target box center, with its width and height predicted by another branch. Such methods, however, fail to utilize the contextual relation between the target and reference objects, and lack interpretability on its reasoning procedure. To solve these problems, we propose an iterative shrinking mechanism to localize the target, where the shrinking direction is decided by a reinforcement learning agent, with all contents within the current image patch comprehensively considered. Besides, the sequential shrinking processes enable to demonstrate the reasoning about how to iteratively find the target. Experiments show that the proposed method boosts the accuracy by 4.32% against the previous state-of-the- art (SOTA) method on the RefCOCOg dataset, where query sentences are long and complex with many targets referred by other reference objects.
Mingjie Sun, Jimin Xiao, Eng Gee Lim
CVPR1
2021 Can Shape Structure Features Improve Model Robustness under Diverse Adversarial Settings?
abstract
Recent studies show that convolutional neural networks (CNNs) are vulnerable under various settings, including adversarial attacks, common corruptions, and backdoor attacks. Motivated by the findings that human visual sys-tem pays more attention to global structure (e.g., shapes) for recognition while CNNs are biased towards local texture features in images, in this work we aim to analyze whether "edge features" could improve the recognition robustness in these scenarios, and if so, to what extent? To answer these questions and systematically evaluate the global structure features, we focus on shape features and pro-pose two edge-enabled pipelines EdgeNetRob and Edge-GANRob, forcing the CNNs to rely more on edge features. Specifically, EdgeNetRob and EdgeGANRob first explicitly extract shape structure features from a given image via an edge detection algorithm. Then EdgeNetRob trains down-stream learning tasks directly on the extracted edge features, while EdgeGANRob reconstructs a new image by refilling the texture information with a trained generative adversarial network (GANs). To reduce the sensitivity of edge detection algorithms to perturbations, we additionally propose a robust edge detection approach Robust Canny based on vanilla Canny. Based on our evaluation, we find that EdgeNetRob can help boost model robustness under different attack scenarios at the cost of the clean model accuracy. EdgeGANRob, on the other hand, is able to improve the clean model accuracy compared to EdgeNetRob while preserving the robustness. This shows that given such edge features, how to leverage them matters for robustness, and it also depends on data properties. Our systematic studies on edge structure features under different settings will shed light on future robust feature exploration and optimization.
Mingjie Sun, Zichao Li 0009, Chaowei Xiao, Haonan Qiu, Bhavya Kailkhura, Mingyan Liu, Bo Li 0026
ICCV1
2021 Discriminative Triad Matching and Reconstruction for Weakly Referring Expression Grounding
abstract
In this paper, we are tackling the weakly-supervised referring expression grounding task, for the localization of a referent object in an image according to a query sentence, where the mapping between image regions and queries are not available during the training stage. In traditional methods, an object region that best matches the referring expression is picked out, and then the query sentence is reconstructed from the selected region, where the reconstruction difference serves as the loss for back-propagation. The existing methods, however, conduct both the matching and the reconstruction approximately as they ignore the fact that the matching correctness is unknown. To overcome this limitation, a discriminative triad is designed here as the basis to the solution, through which a query can be converted into one or multiple discriminative triads in a very scalable way. Based on the discriminative triad, we further propose the triad-level matching and reconstruction modules which are lightweight yet effective for the weakly-supervised training, making it three times lighter and faster than the previous state-of-the-art methods. One important merit of our work is its superior performance despite the simple and neat design. Specifically, the proposed method achieves a new state-of-the-art accuracy when evaluated on RefCOCO (39.21 percent), RefCOCO+ (39.18 percent) and RefCOCOg (43.24 percent) datasets, that is 4.17, 4.08 and 7.8 percent higher than the previous one, respectively. The code is available at https://github.com/insomnia94/DTWREG.
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu 0001, John Yannis Goulermas
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Progressive sample mining and representation learning for one-shot person re-identification
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001
Pattern Recognit.3
2020 Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation Approach
abstract
Weakly supervised semantic segmentation is a challenging task as it only takes image-level information as supervision for training but produces pixel-level predictions for testing. To address such a challenging task, most recent state-of-the-art approaches propose to adopt two-step solutions, i.e. 1) learn to generate pseudo pixel-level masks, and 2) engage FCNs to train the semantic segmentation networks with the pseudo masks. However, the two-step solutions usually employ many bells and whistles in producing high-quality pseudo masks, making this kind of methods complicated and inelegant. In this work, we harness the image-level labels to produce reliable pixel-level annotations and design a fully end-to-end network to learn to predict segmentation maps. Concretely, we firstly leverage an image classification branch to generate class activation maps for the annotated categories, which are further pruned into confident yet tiny object/background regions. Such reliable regions are then directly served as ground-truth labels for the parallel segmentation branch, where a newly designed dense energy loss function is adopted for optimization. Despite its apparent simplicity, our one-step solution achieves competitive mIoU scores (val: 62.6, test: 62.9) on Pascal VOC compared with those two-step state-of-the-arts. By extending our one-step method to two-step, we get a new state-of-the-art performance on the Pascal VOC (val: 66.3, test: 66.5).
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, Kaizhu Huang
AAAI4
2020 Fast Template Matching and Update for Video Object Tracking and Segmentation
abstract
In this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the selection of the matching method to predict the result as well as to decide whether to update the target template using the newly predicted result. The existing methods, however, make these selections in a rough and inflexible way, compromising their performance. To overcome this limitation, we propose a novel approach which utilizes reinforcement learning to make these two decisions at the same time. Specifically, the reinforcement learning agent learns to decide whether to update the target template according to the quality of the predicted result. The choice of the matching method will be determined at the same time, based on the action history of the reinforcement learning agent. Experiments show that our method is almost 10 times faster than the previous state-of-the-art method with even higher accuracy (region similarity of 69.1% on DAVIS 2017 dataset).
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang, Yao Zhao 0001
CVPR1
2020 Feature Representation Matters: End-to-End Learning for Reference-Based Image Super-Resolution
Yanchun Xie, Jimin Xiao, Mingjie Sun, Kaizhu Huang
ECCV (4)3
2020 Denoised Smoothing: A Provable Defense for Pretrained Classifiers
abstract
We present a method for provably defending any pretrained image classifier against $\ell_p$ adversarial attacks. This method, for instance, allows public vision API providers and users to seamlessly convert pretrained non-robust classification services into provably robust ones. By prepending a custom-trained denoiser to any off-the-shelf image classifier and using randomized smoothing, we effectively create a new classifier that is guaranteed to be $\ell_p$-robust to adversarial examples, without modifying the pretrained classifier. Our approach applies to both the white-box and the black-box settings of the pretrained classifier. We refer to this defense as denoised smoothing, and we demonstrate its effectiveness through extensive experimentation on ImageNet and CIFAR-10. Finally, we use our approach to provably defend the Azure, Google, AWS, and ClarifAI image classification APIs. Our code replicating all the experiments in the paper can be found at: https://github.com/microsoft/denoised-smoothing.
Hadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor, J. Zico Kolter
NeurIPS2
2020 Adaptive ROI generation for video object segmentation using reinforcement learning
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yanchun Xie, Jiashi Feng
Pattern Recognit.1
2019 Rethinking the Value of Network Pruning
Zhuang Liu 0003, Mingjie Sun, Tinghui Zhou, Gao Huang 0001, Trevor Darrell
ICLR (Poster)2
2008 Extending PSC for Monitoring the Timed Properties in Composite Services
abstract
Due to the dynamically evolving attribute, validation of composite services must be extended from design time to run-time. Dynamical verification techniques, such as runtime monitoring, have been first class activities to be performed during the execution of composite services. For a kind of composite services, nonfunctional properties, such as timed properties, are as important as functional properties and need to be monitored at run-time. However, using traditional logic and formalism, these timed properties are not easily represented for general software engineers. In order to deal with this problem, we first extend a novel notation (Property Sequence Chart) with time constructs. Then, we give its semantics in terms of timed Buchi automata and measure its expressiveness based on recently proposed real-time specification patterns. Finally, we propose a novel framework to monitor two kinds of timed properties in composite services: the accomplished time of basic service operations and some additional timed assumptions of the composition process. Our framework provides a completely graphical front-end which can friendly help general software engineers to monitor the timed properties in composite services.
Pengcheng Zhang 0001, Bixin Li, Zhiyong Su, Mingjie Sun
APSEC4
2008 A PSC-Based Approach to Monitor the Timed Properties in Web Service Compositions
abstract
Runtime monitoring is significantly essential for web service compositions. For a kind of composite services, nonfunctional properties, such as timed properties, are as important as functional properties and need to be monitored in runtime. In this paper, we extend properly sequence chart into timed properly sequence chart and propose a new approach to monitor two kinds of timed properties in web service compositions: the accomplished time of basic service operations and some additional timed assumptions of the composition process. Our approach is more intuitive than traditional monitoring approaches.
Pengcheng Zhang 0001, Bixin Li, Mingjie Sun, Xufang Gong
COMPSAC3
2008 Data-Enriched Modeling and Verification of WS-CDL Based on UML Models
abstract
The Web Services Choreography Description Language (WS-CDL) is a specification developed by the W3C that can be viewed as a blueprint for the development of end-point services. Considering that it is the W3C candidate recommendation for web service choreography, it is worth providing a systematic approach for its modeling, analysis and verification. The Unified Modeling Language (UML) is the de facto industry standard for modeling. Applying UML to model WS-CDL is obviously a promising solution to bring together academics and practitioners in through a unique standard language. This paper proposes to use different UML diagrams to model WS-CDL. Given the UML specification of WS-CDL, we then provide a systematic way of formally analyzing and verifying WS-CDL.
Pengcheng Zhang 0001, Bixin Li, Henry Muccini, Yu Zhou 0010, Mingjie Sun
ICWS5