Songmin Dai

dblp:230/3847 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
11since 2021 · last 2024
0000-0002-2048-0748ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Generating and Reweighting Dense Contrastive Patterns for Unsupervised Anomaly Detection
abstract
Recent unsupervised anomaly detection methods often rely on feature extractors pretrained with auxiliary datasets or on well-crafted anomaly-simulated samples. However, this might limit their adaptability to an increasing set of anomaly detection tasks due to the priors in the selection of auxiliary datasets or the strategy of anomaly simulation. To tackle this challenge, we first introduce a prior-less anomaly generation paradigm and subsequently develop an innovative unsupervised anomaly detection framework named GRAD, grounded in this paradigm. GRAD comprises three essential components: (1) a diffusion model (PatchDiff) to generate contrastive patterns by preserving the local structures while disregarding the global structures present in normal images, (2) a self-supervised reweighting mechanism to handle the challenge of long-tailed and unlabeled contrastive patterns generated by PatchDiff, and (3) a lightweight patch-level detector to efficiently distinguish the normal patterns and reweighted contrastive patterns. The generation results of PatchDiff effectively expose various types of anomaly patterns, e.g. structural and logical anomaly patterns. In addition, extensive experiments on both MVTec AD and MVTec LOCO datasets also support the aforementioned observation and demonstrate that GRAD achieves competitive anomaly detection accuracy and superior inference speed.
Songmin Dai, Yifan Wu 0011, Xiaoqiang Li 0002, Xiangyang Xue 0001
AAAI1
2023 GradPU: Positive-Unlabeled Learning via Gradient Penalty and Positive Upweighting
abstract
Positive-unlabeled learning is an essential problem in many real-world applications with only labeled positive and unlabeled data, especially when the negative samples are difficult to identify. Most existing positive-unlabeled learning methods will inevitably overfit the positive class to some extent due to the existence of unidentified positive samples. This paper first analyzes the overfitting problem and proposes to bound the generalization errors via Wasserstein distances. Based on that, we develop a simple yet effective positive-unlabeled learning method, GradPU, which consists of two key ingredients: A gradient-based regularizer that penalizes the gradient norms in the interpolated data region, which improves the generalization of positive class; An unnormalized upweighting mechanism that assigns larger weights to those positive samples that are hard, not-well-fitted and less frequently labeled. It enforces the training error of each positive sample to be small and increases the robustness to the labeling bias. We evaluate our proposed GradPU on three datasets: MNIST, FashionMNIST, and CIFAR10. The results demonstrate that GradPU achieves state-of-the-art performance on both unbiased and biased positive labeling scenarios.
Songmin Dai, Xichen Ye
AAAI1
2023 OEST: Outlier Exposure by Simple Transformations for Out-of-Distribution Detection
abstract
Although the previous works for out-of-distribution(OOD) detection have achieved great improvements, they are still highly dependent on the specific selection of the outliers from external datasets or that transformed by certain data augmentations, and hence cannot be applied in a wide range of domains. To solve this problem, in this paper, we propose a simple, yet effective method called Outlier Exposure by Simple Transformations (OEST), which aims at exposing the outliers by the composition of several simple transformations of data augmentations via energy score. In addition, our training scheme can make full use of nearly all considered data augmentations in previous works, even though some of them are generally regarded as useless. And we also find that, for simple data augmentation, our training scheme is less time-consuming and better in performance than relative works. Furthermore, our experiments validate that our method outperforms the state-of-the-art methods.
Yifan Wu 0011, Songmin Dai, Dengye Pan, Xiaoqiang Li 0002
ICIP2
2023 Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https://github.com/Wu0409/HSC_WSSS.
Yuanchen Wu, Xiaoqiang Li 0002, Songmin Dai, Jide Li, Tong Liu 0001, Shaorong Xie
IJCAI3
2023 Active Negative Loss Functions for Learning with Noisy Labels
abstract
Robust loss functions are essential for training deep neural networks in the presence of noisy labels. Some robust loss functions use Mean Absolute Error (MAE) as its necessary component. For example, the recently proposed Active Passive Loss (APL) uses MAE as its passive loss function. However, MAE treats every sample equally, slows down the convergence and can make training difficult. In this work, we propose a new class of theoretically robust passive loss functions different from MAE, namely *Normalized Negative Loss Functions* (NNLFs), which focus more on memorized clean samples. By replacing the MAE in APL with our proposed NNLFs, we improve APL and propose a new framework called *Active Negative Loss* (ANL). Experimental results on benchmark and real-world datasets demonstrate that the new set of loss functions created by our ANL framework can outperform state-of-the-art methods. The code is available at https://github.com/Virusdoll/Active-Negative-Loss.
Xichen Ye, Xiaoqiang Li 0002, Songmin Dai, Tong Liu 0001, Weiqin Tong
NeurIPS3
2023 Semi-supervised medical imaging segmentation with soft pseudo-label fusion
Xiaoqiang Li 0002, Yuanchen Wu, Songmin Dai
Appl. Intell.3
2023 Multiscale features integration based multiple-in-single-out network for object detection
Kequan Yang, Jide Li, Songmin Dai, Xiaoqiang Li 0002
Image Vis. Comput.3
2023 Multi-Sourced Knowledge Integration for Robust Self-Supervised Facial Landmark Tracking
abstract
Expensive annotation costs significantly hinder the development of facial landmark tracking owing to the frame-by-frame labeling of dense landmarks. The most promising approach to address this problem is to develop a self-supervised tracker for large-scale unlabeled videos. However, existing self-supervised trackers trained using single-sourced knowledge are unstable under unconstrained environments. Herein, we propose multi-sourced knowledge integration (MSKI), a robust self-supervised tracking method. It integrates knowledge from multiple sources to provide supervisory signals, thereby improving the stability of the self-supervised tracker. Specifically, the proposed MSKI comprises two complementary modules: a temporal knowledge reasoning (TempRes) module and an interactive knowledge distillation (KnowDist) module. The TempRes module enforces the tracker to achieve cycle-consistent tracking, allowing the tracker to learn temporal correspondence based on the cycle-consistency of time. To exploit facial geometry knowledge against various occlusions, our tracker imposes a multi-level shape constraint over the structure of facial landmarks by leveraging adversarial shape learning, thereby enabling the tracking of occluded faces. Moreover, the tracker interacts with an initialization detector to further develop complementary knowledge via KnowDist. The KnowDist module distills the spatial and temporal knowledge provided by the detector and tracker to generate plausible labels automatically. Finally, these generated labels are utilized to fine-tune the detector, such that it provides high-quality initial landmarks for the cycle-consistent tracking of the tracker on unlabeled videos. The experimental results show that the proposed MSKI can stabilize the tracking trajectory and improve the robustness against various occlusions.
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong
IEEE Trans. Multim.4
2022 Reasoning structural relation for occlusion-robust facial landmark localization
abstract
In facial landmark localization tasks, various occlusions heavily degrade the localization accuracy due to the partial observability of facial features . This paper proposes a structural relation network (SRN) for occlusion-robust landmark localization. Unlike most existing methods that simply exploit the shape constraint, the proposed SRN aims to capture the structural relations among different facial components. These relations can be considered a more powerful shape constraint against occlusion. To achieve this, a hierarchical structural relation module (HSRM) is designed to hierarchically reason the structural relations that represent both long- and short-distance spatial dependencies . Compared with existing network architectures ,the HSRM can efficiently model the spatial relations by leveraging its geometry-aware network architecture, which reduces the semantic ambiguity caused by occlusion. Moreover, the SRN augments the training data by synthesizing occluded faces. To further extend our SRN for occluded video data, we formulate the occluded face synthesis as a Markov decision process (MDP). Specifically, it plans the movement of the dynamic occlusion based on an accumulated reward associated with the performance degradation of the pre-trained SRN. This procedure augments hard samples for robust facial landmark tracking. Extensive experimental results indicate that the proposed method achieves outstanding performance on occluded and masked faces. Code is available at https://github.com/zhuccly/SRN
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong
Pattern Recognit.4
2021 Improving Robustness of Facial Landmark Detection by Defending against Adversarial Attacks
abstract
Many recent developments in facial landmark detection have been driven by stacking model parameters or augmenting annotations. However, three subsequent challenges remain, including 1) an increase in computational overhead, 2) the risk of overfitting caused by increasing model parameters, and 3) the burden of labor-intensive annotation by humans. We argue that exploring the weaknesses of the detector so as to remedy them is a promising method of robust facial landmark detection. To achieve this, we propose a sample-adaptive adversarial training (SAAT) approach to interactively optimize an attacker and a detector, which improves facial landmark detection as a defense against sample-adaptive black-box attacks. By leveraging adversarial attacks, the proposed SAAT exploits adversarial perturbations beyond the handcrafted transformations to improve the detector. Specifically, an attacker generates adversarial perturbations to reflect the weakness of the detector. Then, the detector must improve its robustness to adversarial perturbations to defend against adversarial attacks. Moreover, a sample-adaptive weight is designed to balance the risks and benefits of augmenting adversarial examples to train the detector. We also introduce a masked face alignment dataset, Masked-300W, to evaluate our method. Experiments show that our SAAT performed comparably to existing state-of-the-art methods. The dataset and model are publicly available at https://github.com/zhuccly/SAAT.
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai
ICCV4
2021 Point cloud super-resolution based on geometric constraints
abstract
Abstract Among all digital representations we have for real physical objects, three‐dimensional ( 3D) is arguably the most expressive encoding. But due to the limitations of 3D scanning equipment, point cloud often becomes sparse or partially missing. A point cloud super‐resolution (PCSR) method based on geometric constraints is proposed to solve the sparse problem of point clouds: it allows dense point clouds to be generated by sparse point clouds. The method is based on the conditional generative adversarial network including redesigned generator and discriminator for point cloud data specially. Moreover, the method can maintain the shape of the dense point cloud by adding geometric constraints. The contributions of our work are as follows: (1) a PCSR method based on geometric constraints is proposed; (2) add a module for obtaining point cloud neighbourhood information in the generator, called K‐nn operation module; and (3) feature aggregation is performed using the weighted pooling to process the neighbourhood information obtained by the K‐nn operation module. Extensive experimental results demonstrate the effectiveness of the proposed method.
Xiaoqiang Li 0002, Jitao Liu, Songmin Dai
IET Comput. Vis.3
2020 Semi-Supervised Semantic Segmentation Constrained by Consistency Regularization
abstract
In this paper, we propose a self-training based method for semi-supervised semantic segmentation. Our method utilizes k perturbed images of each unlabeled image to generate a new mask through the proposed vote operation which in turn is used as a supervision signal to train the model. The k predicted masks of perturbed images can provide nontrivial knowledge that is not captured by a single prediction and the proposed vote operation enables the model to output a low entropy prediction. Consistency regularization is applied between generated mask and k masks in terms of MSE loss in our network. Extensive experiments are conducted on different datasets to evaluate the effectiveness of our method. We show that the proposed method surpasses previous state-of the-art semi-supervised methods on ISIC 2017 dataset and achieves competitive performance on PASCAL VOC 2012 dataset.
Songmin Dai, Pin Wu, Weiqin Tong
ICME3
2020 Scale specified single shot multibox detector
abstract
Detecting objects at vastly different scales is a fundamental challenge in computer vision. To solve this, some approaches (e.g. TridentNet) investigate the effect of receptive fields, whereas other approaches (e.g. SNIP, SNIPER) are based on the image pyramid strategy. In this study, a novel single‐shot based detector, called scale specified single‐shot multibox detector (4SD) is proposed. It aims to predict objects of a specific scale range separately by using feature maps of different sizes. First, a parallel multi‐branch architecture with feature maps of different sizes is generated by scale specific inference module. Then, the authors propose a scale specific training scheme to specialise each branch by sampling object instances of proper scales for training. Results are shown on both PASCAL VOC and COCO detection. The proposed method can achieve a mean average precision of 83.1% on PASCAL VOC 2007, and 36.9% on MS‐COCO at a speed of 28 frames per second, which is superior to most single‐stage detectors.
Xiaoqiang Li 0002, Chuanwei Liu, Songmin Dai, Huichen Lian, Guangtai Ding
IET Comput. Vis.3
2020 Robust landmark-free head pose estimation by learning to crop and background augmentation
abstract
It is well known that the performance of head pose estimation is greatly affected by the bounding box margin of the face and its background. Traditionally, researchers will manually choose a suitable bounding box margin to strike a balance between ensuring sufficient information and minimising background noise. However, head pose estimation is still worse when the background is complex in reality or when the box margin changes slightly. To make estimation results more robust, the authors propose two methods to improve it: (i) a convolutional cropping module that can learn to crop the input image to an attentional area for head pose regression. (ii) Background augmentation that can make the network more robust to the background noise. Rather than using the face landmarking to calculate head pose angles, they use another convolutional neural network to regress the head pose angles, which is independent of the landmark detection results. They evaluate the method on BIWI and AFLW2000 dataset and experimental results show that their approach outperforms many other methods. Besides, they evaluate the method on Pointing′04 dataset using head pose accuracy. Furthermore, the approach is more robust and has a lower variance in realistic scenarios.
Aoru Xue, Songmin Dai, Xiaoqiang Li 0002
IET Image Process.3
2019 Learning Segmentation Masks with the Independence Prior
abstract
An instance with a bad mask might make a composite image that uses it look fake. This encourages us to learn segmentation by generating realistic composite images. To achieve this, we propose a novel framework that exploits a new proposed prior called the independence prior based on Generative Adversarial Networks (GANs). The generator produces an image with multiple category-specific instance providers, a layout module and a composition module. Firstly, each provider independently outputs a category-specific instance image with a soft mask. Then the provided instances’ poses are corrected by the layout module. Lastly, the composition module combines these instances into a final image. Training with adversarial loss and penalty for mask area, each provider learns a mask that is as small as possible but enough to cover a complete category-specific instance. Weakly supervised semantic segmentation methods widely use grouping cues modeling the association between image parts, which are either artificially designed or learned with costly segmentation labels or only modeled on local pairs. Unlike them, our method automatically models the dependence between any parts and learns instance segmentation. We apply our framework in two cases: (1) Foreground segmentation on category-specific images with box-level annotation. (2) Unsupervised learning of instance appearances and masks with only one image of homogeneous object cluster (HOC). We get appealing results in both tasks, which shows the independence prior is useful for instance segmentation and it is possible to unsupervisedly learn instance masks with only one image.
Songmin Dai, Pin Wu, Weiqin Tong
AAAI1
2019 PCCN: POINT Cloud Colorization Network
abstract
This paper investigates into the colorization problem which generates color value to point clouds. It is difficult to apply image colorization methods to three dimension directly. Currently, Generative Adversarial Network (GAN) is the most reliable solution to the colorization problem. However, it cannot directly process point cloud data with geometric data structures. In this paper, we propose a scheme for point cloud colorization which can effectively solve the permutation invariance of points as well as produce realistic color effect. In order to achieve a better result, a well-designed network that aggregates the point features as global feature is built to guide color generation. After computing the global point cloud feature vector, we integrates both the local and global information by concatenating the global feature with each of the point features. Extensive experimental results demonstrate the effectiveness of the proposed algorithm.
Jitao Liu, Songmin Dai
ICIP2