VLDB 2026 Research / reviewers in the wild / expert
Yubao Sun
dblp:45/8108
· DBLP profile ↗
36ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-0462-3729ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse-view CT image reconstruction using conditional embedding fusion diffusion model
Chenchun Zhou, Yubao Sun, Jia Liu 0034, Qingshan Liu 0001 |
Neurocomputing | 2 |
| 2026 | Person image generation via regional style rectification
Guiyu Xia, Zhedong Jin, Yubao Sun |
Pattern Recognit. | 4 |
| 2026 | Vision-Language-Driven Prompt Learning for Weakly Supervised Semantic SegmentationabstractThe primary challenges in image-level weakly supervised semantic segmentation (WSSS) lie in addressing the under-activation issue of target pixels and mitigating the co-occurrence phenomenon in class activation maps. In recent years, Vision-Language Models (VLM) have demonstrated exceptional performance across various vision tasks, primarily attributed to their cross-modal semantic alignment capabilities achieved through contrastive learning mechanisms. Leveraging VLM’s capability to capture fine-grained visual-textual correspondences, this paper proposes a novel Vision-Language Driven Prompt Learning (VLD-PL) framework that addresses two fundamental challenges in WSSS by establishing explicit semantic correspondences between textual descriptors and visual components, ultimately enabling efficient semantic segmentation. The VLD-PL framework consists of two core components Auxiliary Class Matching (ACM) and Background Class Filtering (BCF). The ACM module dynamically identifies semantically relevant auxiliary classes through feature alignment between image and textual embeddings, effectively enlarging target activation while mitigating co-occurrence interference by expanding semantic coverage. Simultaneously, the BCF constructs image-specific background prompts and adaptively refines background feature representations, achieving precise suppression of irrelevant background regions. These dual mechanisms synergistically address both target localization accuracy and background noise suppression, achieving state-of-the-art performance on both the PASCAL VOC 2012 and MS COCO 2014 benchmarks. Junxia Li, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Adversarial Pruning Networks for Compact 3D Gaussian Splattingabstract3D Gaussian Splatting holds significant potential for high-quality visual scene rendering. However, the large number of Gaussian primitives it requires poses challenges in memory consumption and practical deploy. Existing methods often rely on empirical criteria to prune Gaussians, which inevitably compromises visual quality. To address this, we propose Adversarial Pruning Networks (APNet), a framework that employs adversarial learning to balances the reduction of redundant Gaussians with the preservation of visual fidelity. APNet comprises a Gaussian Learning and Pruning Network (GLPN) and a Discriminative Network. GLPN incorporates the geometric information into the learning of Gaussians and prunes these Gaussians through a data-driven mask. Meanwhile, the Discriminative Network is trained to distinguish between synthesized and real images, acting as an adversary. Through adversarial pruning, APNet significantly reduces the number of Gaussians while rendering high-quality images. Extensive experiments on the Mip-NeRF360, Tanks & Temples, and Deep Blending datasets demonstrate that APNet achieves up to a 90% reduction in the original 3DGS while maintaining high rendering quality. Hui Shuai, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | 3D Scenes Motion Planning and Generation with Motion Diffusion Probabilistic ModelabstractGenerating natural and realistic human motion sequences under the constraints of 3D scenes is a highly challenging task, requiring not only the precise modeling of dynamic variations in human joints but also the rigorous consideration of intricate interactions between the human body and the surrounding environment. While recent advances in deep generative models show great potential in tackling these challenges, existing methods often result in unnatural human motions and human–environment penetration during generation. In order to cope with these issues, we propose a novel approach that divides human motion generation into two stages. The first stage employs a bidirectional long short-term memory network incorporated with full-connected layers to generate motion trajectory under the input conditions including the starting and ending positions and orientations of the human model and scene feature point clouds extracted from the surrounding environment. In the second stage, we design a conditional diffusion model, guided by the trajectory generated in the first stage and the embedding of 3D scene information, to generate human motion sequences within 3D scenes. We evaluate our framework through extensive experiments on the PROX datasets, which validates its effectiveness. The results show that our method significantly outperforms existing ones in enhancing human motion naturalness and reasonableness, and reducing human penetration. Yubao Sun, Guiyu Xia, Qingshan Liu 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | ROA-BEV: 2D Region-Oriented Attention for BEV-based 3D Object DetectionabstractVision-based Bird’s-Eye-View (BEV) 3D object detection has recently become popular in autonomous driving. However, objects with a high similarity to the background from a camera perspective cannot be detected well by existing methods. In this paper, we propose a BEV-based 3D Object Detection Network with 2D Region-Oriented Attention (ROA-BEV), which enables the backbone to focus more on feature learning of the regions where objects exist. Moreover, our method further enhances the information feature learning ability of ROA through multi-scale structures. Each block of ROA utilizes a large kernel to ensure that the receptive field is large enough to catch information about large objects. Experiments on nuScenes show that ROA-BEV improves the performance based on BEVDepth. The source codes of this work will be available at https://github.com/DFLyan/ROA-BEV. Yubao Sun, Laiyan Ding, Rui Huang 0001 |
IROS | 2 |
| 2025 | Text-driven human image generation with texture and pose control
Zhedong Jin, Guiyu Xia, Paike Yang, Mengxiang Wang, Yubao Sun, Qingshan Liu 0001 |
Neurocomputing | 5 |
| 2025 | Enhancing semantics consistency via hybrid attention fusion in multimodal sentiment analysis of short videos
Xuanchi Gong, Ziyang Xue, Tengjun Liu, Yubao Sun |
Multim. Syst. | 4 |
| 2025 | Geometric transformation supervised disentanglement of pose and expression for talking face generation
Mengxiang Wang, Guiyu Xia, Zhedong Jin, Paike Yang, Yubao Sun |
Multim. Syst. | 5 |
| 2025 | Boosting Adversarial Transferability via Relative Feature Importance-Aware AttacksabstractModern deep neural networks are known highly vulnerable to adversarial examples. As a pioneering work, the fast gradient sign method (FGSM) is proved more transferable in black-box attacks than its multi-small-step extension, i.e., iterative-FGSM, particularly being restricted by a limited number of iterations. This paper revisits their early, representative successor MI-FGSM as a baseline, i.e., iterative-FGSM with momentum, and introduces an innovative boosting idea different from either FGSM-inspired algorithms or other mainstream methods. For one thing, during gradient backpropogation of MI-FGSM, the proposed approach merely requires amending the chain rule with respect to adversarial images using the counterpart original images. For another, a credible analysis has revealed that such a naively boosted MI-FGSM essentially performs a special kind of intermediate-layer attacks. In specific, the notable finding in the paper is a new principle of adversarial transferability guided by the relative feature importance, emphasizing the significance of semantically non-critical information for the first time in the literature, although originally thought to be weak in large. Experimental results on various leading victim models, both undefended and defended, demonstrate that the new approach incorporating robust gradients has indeed attained stronger adversarial transferability than state-of-the-art works. The code is available at:https://github.com/ljwooo/RFIA-main. Wenze Shao, Yubao Sun, Li-Qian Wang, Qi Ge, Liang Xiao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Source Information-Assisted UV-Space Transformation Network for Person Image GenerationabstractPerson image generation is widely used in many fields, but it still faces some challenges. Most of current person image generation methods suffer from an intractable problem of handling the spatial deformation caused by the pose change in the generation process, while convolution-based generative model is not good at handling the region-unaligned task. Therefore, we propose a novel UV-space transformation network to implement the primary generation of person image in the UV-space. This framework can effectively avoid the spatial deformation problems in the generation process and instead transfer them to the preceding pose estimation stage. Within the framework, we propose the self-reconstruction-assisted UV texture transformation blocks which aim to exploit the self-reconstruction of source texture map to guide and assist the generation of target UV texture map. In addition, after obtaining the target person image from the generated UV texture map, we use the correlations between the source and generated images to further improve the details of the generated person images. Superior experiment results compared with other state-of-the-art methods demonstrate the effectiveness of the proposed method. Guiyu Xia, Zhedong Jin, Dongdong Fang, Yubao Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Motion Compression Using Structurally Connected Neural NetworkabstractMotion compression technologies can significantly reduce the redundant information of motion data and increase the efficiency of storage and transmission. Current methods mainly utilize some ready-made universal algorithms, such as signal processing and dimensionality reduction, to model the statistical characteristics of motion data, while the individual structure of motion data is ignored. In this paper, we propose to use a deep neural network with specially designed architecture to represent motion data considering the similarity between the articulated structure of a human skeleton and the architecture of neural networks. The network parameters are then taken as the compressed data. We design a structurally connected network which just looks like a human skeleton. Within the network, only the neurons corresponding to the joints connected to each other in a human skeleton are connected. It effectively exploits the correlations between connected joints to cut down the unnecessary connections between the neurons, which leads to the significant improvement of compression efficiency. Additionally, we extract the two inherent DOFs instead of the original three DOFs of each joint by representing its movement on a sphere according to the rigidity of the articulated human skeleton. This actually achieves the theoretically lossless pre-compression with the ratio of 3:2. Extensive experiment results demonstrate the superior performances of the proposed model at the high compression ratios over other state-of-the-art methods. Guiyu Xia, Wenkai Ye, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Highway Visibility Level Prediction Using Geometric and Visual Features Driven Dual-Branch Fusion NetworkabstractAutomatic prediction of visibility from surveillance images can provide timely warnings for transportation management departments and drivers, which is of great significance for improving the driving safety of highways in foggy weather. The currently mainstream prediction models are based on deep networks, which mainly learn visual features as clues to predict visibility levels. However, the geometric features of highways, such as the lane lines that can be observed from surveillance images are also important clues to reflect the visibility. Therefore, we propose a dual-branch fusion network driven by both geometric and visual features to achieve robust and effective visibility prediction. Specifically, we first exploit dual branches to learn geometric features of highway and deep visual features from the foggy surveillance images, respectively. We then design a fused classification module to fuse the dual-branch features to predict the visibility level. In order to simultaneously purify features during the fusion process, it utilizes a road attention block to highlight the deep visual features corresponding to the highway road area, and a lane length estimation block to extract the feature of the length of observable lane lines. Therefore, the dual-branch features can be adaptively fused to boost prediction performance. Meanwhile, we construct a real-scene foggy image dataset, which are all gathered from the surveillance video of real highways in China. We validate the effectiveness of the proposed network on this real-scene dataset and the synthetic dataset FRIDA. The experimental results show that our method can predict visibility levels more accurately than multiple existing methods. Yubao Sun, Jihui Tang, Qingshan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | A Deep Learning Framework for Start-End Frame Pair-Driven Motion SynthesisabstractA start-end frame pair and a motion pattern-based motion synthesis scheme can provide more control to the synthesis process and produce content-various motion sequences. However, the data preparation for the motion training is intractable, and concatenating feature spaces of the start-end frame pair and the motion pattern lacks theoretical rationality in previous works. In this article, we propose a deep learning framework that completes automatic data preparation and learns the nonlinear mapping from start-end frame pairs to motion patterns. The proposed model consists of three modules: action detection, motion extraction, and motion synthesis networks. The action detection network extends the deep subspace learning framework to a supervised version, i.e., uses the local self-expression (LSE) of the motion data to supervise feature learning and complement the classification error. A long short-term memory (LSTM)-based network is used to efficiently extract the motion patterns to address the speed deficiency reflected in the previous optimization-based method. A motion synthesis network consists of a group of LSTM-based blocks, where each of them is to learn the nonlinear relation between the start-end frame pairs and the motion patterns of a certain joint. The superior performances in action detection accuracy, motion pattern extraction efficiency, and motion synthesis quality show the effectiveness of each module in the proposed framework. Guiyu Xia, Qingshan Liu 0001, Yubao Sun |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Polarity-aware attention network for image sentiment analysis
Qiming Yan, Yubao Sun, Shaojing Fan, Liling Zhao |
Multim. Syst. | 2 |
| 2023 | 3D Information Guided Motion Transfer via Sequential Image Based Human Model Refinement and Face-Attention GANabstractImage and video based human motions can be regarded as the deformation processes of person appearances, so motion transfer is usually treated as a pose guided image generation task and implemented in the 2D image plane. However, the 2D plane image generation lacks guidance of the original 3D motion information, which results in blur and shape distortions of the generated motion images. Therefore, we propose to simulate the generation process of real motion images by projecting the 3D human models, which are reconstructed from the training motion images and driven with target poses, into the 2D plane. We then take the 2D projections as the pose representations and input them into the generation model as they naturally inherit the 3D information from the original motions. Considering the unreliability on the invisible surface of the single image based human model reconstruction, we propose a sequential image based human model refinement module which exploits the complementary information between adjacent motion frames to refine the 3D human model. Furthermore, we propose a face-attention GAN model to conduct the final motion transfer, in which we use the Gaussian distribution to match the elliptical face region and design a face enhancement loss function since the faces in the generated motion images influence the performances very much. The generated motion images with reliable depth information, accurate shapes and clear faces demonstrate the effectiveness of the proposed method. Guiyu Xia, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Transformer-based Residual Network for Hyperspectral Snapshot Compressive ReconstructionabstractThe core problem of hyperspectral snapshot compressive imaging is to achieve high-quality reconstruction from a single snapshot measurement. Existing reconstruction methods are mainly based on convolutional networks. Inspired by the exciting success of Transformer in high-level vision tasks, we introduce Transformer into hyperspectral snapshot compressive imaging, and propose the Transformer-based residual network to learn reconstruction mapping. Specifically, the proposed network cascades multiple basic modules, each of which consists of two local-enhanced window (Lewin) Transformer blocks and two residual blocks. The advantage of this basic module is ability to exploit both local feature maps and long-range dependencies for resonstruction. The Transformer blocks treat the feature map at each pixel location as a token and capture the spatial-spectral context of each pixel within a local window through self-attention, thus effectively improving the spectral fidelity of reconstructed hyperspectral images at modest computational cost. We conduct extensive experiments on both simulation and real data. The experimental results show that the proposed network is concise and effective, achieving the best reconstruction results compared with several state-of-the-art methods. Junru Huang, Yubao Sun, Jiaxuan Wen, Qingshan Liu 0001 |
ICPR | 2 |
| 2022 | Visual saliency prediction using multi-scale attention gated network
Yubao Sun, Kai Hu 0006, Shaojing Fan |
Multim. Syst. | 1 |
| 2022 | Video Snapshot Compressive Imaging Using Residual Ensemble NetworkabstractVideo snapshot compressive imaging (SCI) system enables high-frame-rate imaging by projecting multiple frames into a 2D snapshot measurement during a single exposure, and the original video frames can be reconstructed by solving an optimization problem. However, existing methods usually cannot achieve a good balance between reconstruction time and reconstruction quality, which has become a major obstacle for practical application of video SCI. In order to cope with this issue, we propose a residual ensemble network to learn the explicit inverse mapping from the 2D snapshot measurement to the original video. Specifically, the proposed network aims to exploit the spatiotemporal correlations between video frames for improving reconstruction quality. The spatiotemporal correlations of video frames demonstrate multiple types, including intra-frame spatial correlation, inter-frame forward and backward temporal correlation. With the purpose of fully capturing these differentiated correlations, we design four sub-networks, namely, a pseudo-3D U-shape sub-network, two residual sub-networks, and a serial forward and backward recurrent sub-network, and further assemble these four sub-networks into an ensemble network through alternate residual links. This ensemble network can effectively fuse the predictions of each sub-network and maintain spatiotemporal consistency between video frames. We further design a compound loss function to guide the network learning, and the new video can be fast reconstructed by simply feeding its 2D snapshot measurement into the learned network. The experimental results demonstrate that our network can significantly improve the reconstruction quality while maintaining low computational cost. Yubao Sun, Xunhao Chen, Mohan Kankanhalli, Qingshan Liu 0001, Junxia Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Complementarity-Aware Attention Network for Salient Object DetectionabstractIn this article, we tackle the saliency detection task from an interesting perspective: we focus both on salient regions (or foreground) detection and nonsalient regions (or background) detection instead of only the foreground and propose a novel complementarity-aware attention network. It is a unified framework with two branches, namely, positive attention module (PAM) and negative attention module (NAM), for the foreground and background detection, respectively. More specifically, the PAM exploits a position self-attention mechanism to enhance the discriminant ability of feature representation, which can detect most of the salient object regions. Meanwhile, the NAM is designed to detect the background regions, aiming to pop out the missing object parts and details in the prediction map produced by the PAM. By fusing these two attention modules together, NAM can provide complementary cues to assist PAM for precise object detection. Furthermore, in order to capture more multiscale contextual information, we introduce a bidirectional structure with multisupervision to the proposed complementarity-aware attention module for performance improvement. Experiments on five benchmark datasets show that the proposed framework achieves comparable results compared with the state-of-the-art saliency detection methods. Junxia Li, Qingshan Liu 0001, Yubao Sun |
IEEE Trans. Cybern. | 5 |
| 2022 | Unsupervised Spatial-Spectral Network Learning for Hyperspectral Compressive Snapshot ReconstructionabstractHyperspectral compressive imaging takes advantage of compressive sensing theory to achieve coded aperture snapshot measurement without temporal scanning, and the entire 3-D spatial–spectral data is captured by a 2-D projection during a single integration period. Its core issue is how to reconstruct the underlying hyperspectral image (HSI) using compressive sensing reconstruction algorithms. Due to the diversity in the spectral response characteristics and wavelength range of different spectral imaging devices, previous works are often inadequate to capture complex spectral variations or lack the adaptive capacity to new hyperspectral imagers. In order to address these issues, we propose an unsupervised spatial–spectral network to reconstruct HSIs only from the compressive snapshot measurement. The proposed network acts as a conditional generative model conditioned on the snapshot measurement, and it exploits the spatial–spectral attention module to capture the joint spatial–spectral correlation of HSIs. The network parameters are optimized to make sure that the network output can closely match the given snapshot measurement according to the imaging model, thus the proposed network can adapt to different imaging settings, which can inherently enhance the applicability of the network. Extensive experiments upon multiple datasets demonstrate that our network can achieve better reconstruction results than the state-of-the-art methods. Yubao Sun, Qingshan Liu 0001, Mohan Kankanhalli |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Local Self-Expression Subspace Learning Network for Motion Capture DataabstractDeep subspace learning is an important branch of self-supervised learning and has been a hot research topic in recent years, but current methods do not fully consider the individualities of temporal data and related tasks. In this paper, by transforming the individualities of motion capture data and segmentation task as the supervision, we propose the local self-expression subspace learning network. Specifically, considering the temporality of motion data, we use the temporal convolution module to extract temporal features. To implement the local validity of self-expression in temporal tasks, we design the local self-expression layer which only maintains the representation relations with temporally adjacent motion frames. To simulate the interpolatability of motion data in the feature space, we impose a group sparseness constraint on the local self-expression layer to impel the representations only using selected keyframes. Besides, based on the subspace assumption, we propose the subspace projection loss, which is induced from distances of each frame projected to the fitted subspaces, to penalize the potential clustering errors. The superior performances of the proposed model on the segmentation task of synthetic data and three tasks of real motion capture data demonstrate the feature learning ability of our model. Guiyu Xia, Huaijiang Sun, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | 3D Tensor Auto-encoder with Application to Video CompressionabstractAuto-encoder has been widely used to compress high-dimensional data such as the images and videos. However, the traditional auto-encoder network needs to store a large number of parameters. Namely, when the input data is of dimension n , the number of parameters in an auto-encoder is in general O ( n ). In this article, we introduce a network structure called 3D Tensor Auto-Encoder (3DTAE). Unlike the traditional auto-encoder, in which a video is represented as a vector, our 3DTAE considers videos as 3D tensors to directly pass tensor objects through the network. The weights of each layer are represented by three small matrices, and thus the number of parameters in 3DTAE is just O ( n 1/3). The compact nature of 3DTAE fits well the needs of video compression. Given an ensemble of high-dimensional videos, we represent them as 3DTAE networks plus some small core tensors, and we further quantize the network parameters and the core tensors to get the final compressed data. Experimental results verify the efficiency of 3DTAE. Yang Li 0039, Guangcan Liu, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Learning Memory Augmented Cascading Network for Compressed Sensing of Images
Yubao Sun, Qingshan Liu 0001, Rui Huang 0001 |
ECCV (22) | 2 |
| 2020 | Learning image compressed sensing with sub-pixel convolutional generative adversarial network
Yubao Sun, Qingshan Liu 0001, Guangcan Liu |
Pattern Recognit. | 1 |
| 2020 | Dual-Path Attention Network for Compressed Sensing Image ReconstructionabstractAlthough deep neural network methods achieved much success in compressed sensing image reconstruction in recent years, they still have some issues, especially in preserving texture details. In this paper, we propose a new dual-path attention network for compressed sensing image reconstruction, which is composed of a structure path, a texture path and a texture attention module. Motivated by the classical paradigm of image structure-texture decomposition, the structure path aims to reconstruct the dominant structure component of the original image, and the texture path targets at recovering the remaining texture details. To better bridge the information between two paths, the texture attention module is designed to deliver the useful structure information to the texture path and predict the texture region, thereby facilitating the recovery of texture details. Two paths are optimized with a unified loss function. In the testing phase, given the measurement vector of a new image, it can be well reconstructed by carrying out the well trained dual-path attention network and integrating the outputs of the structure path and the texture path. Experimental results on the SET5, SET11 and BSD68 testing datasets demonstrate that the proposed method achieves comparable or better results compared with some state-of-the-art deep learning based methods and conventional iterative optimization based methods in terms of reconstruction quality and robustness to noise. Yubao Sun, Qingshan Liu 0001, Bo Liu 0005, Guodong Guo |
IEEE Trans. Image Process. | 1 |
| 2020 | Learning Non-Locally Regularized Compressed Sensing Network With Half-Quadratic SplittingabstractDeep learning-based Compressed Sensing (CS) reconstruction attracts much attention in recent years, due to its significant superiority of reconstruction quality. Its success is mainly attributed to the employment of a large dataset for pre-training the network to learn a reconstruction mapping. In this paper, we propose a non-locally regularized compressed sensing network for reconstructing image sequences, which can achieve high reconstruction quality without pre-training. Specifically, the proposed method attempts to learn a deep network prior for the reconstruction of an individual instance under the constraint that the network output can well match the given CS measurement. The non-local prior is designed to guide the network to capture the long-range dependencies by exploiting the self-similarities among images, and it can also make the network noise-aware. In order to deal with the compound of non-local prior and deep network prior, we construct a half-quadratic splitting based optimization method for network learning, in which the two priors are decoupled into two simple sub-problems by introducing an auxiliary variable and a quadratic fidelity constraint. Extensive experimental results demonstrate that our method is competitive to the popular methods, including sparsity prior based methods and deep learning based methods, even better than them in the cases of low measurement rates. Yubao Sun, Qingshan Liu 0001, Xiao-Tong Yuan, Guodong Guo |
IEEE Trans. Multim. | 1 |
| 2019 | Moving object detection via segmentation and saliency constrained RPCA
Yang Li 0039, Guangcan Liu, Qingshan Liu 0001, Yubao Sun, Shengyong Chen |
Neurocomputing | 4 |
| 2018 | Fast subspace segmentation via Random Sample Probing
Yang Li 0039, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
Neurocomputing | 2 |
| 2018 | Multi-component group sparse RPCA model for motion object detection under complex dynamic background
Yubao Sun, Renlong Hang, Qingshan Liu 0001, Guangcan Liu |
Neurocomputing | 2 |
| 2017 | Elastic Net Hypergraph Learning for Image Clustering and Semi-Supervised ClassificationabstractGraph model is emerging as a very effective tool for learning the complex structures and relationships hidden in data. In general, the critical purpose of graph-oriented learning algorithms is to construct an informative graph for image clustering and classification tasks. In addition to the classical K-nearest-neighbor and r-neighborhood methods for graph construction, l1-graph and its variants are emerging methods for finding the neighboring samples of a center datum, where the corresponding ingoing edge weights are simultaneously derived by the sparse reconstruction coefficients of the remaining samples. However, the pairwise links of l1-graph are not capable of capturing the high-order relationships between the center datum and its prominent data in sparse reconstruction. Meanwhile, from the perspective of variable selection, the l1norm sparse constraint, regarded as a LASSO model, tends to select only one datum from a group of data that are highly correlated and ignore the others. To simultaneously cope with these drawbacks, we propose a new elastic net hypergraph learning model, which consists of two steps. In the first step, the robust matrix elastic net model is constructed to find the canonically related samples in a somewhat greedy way, achieving the grouping effect by adding the l2penalty to the l1constraint. In the second step, hypergraph is used to represent the high order relationships between each datum and its prominent samples by regarding them as a hyperedge. Subsequently, hypergraph Laplacian matrix is constructed for further analysis. New hypergraph learning algorithms, including unsupervised clustering and multi-class semi-supervised classification, are then derived. Extensive experiments on face and handwriting databases demonstrate the effectiveness of the proposed method. Qingshan Liu 0001, Yubao Sun, Cantian Wang, Tongliang Liu, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2016 | Matrix-Based Discriminant Subspace Ensemble for Hyperspectral Image Spatial-Spectral Feature FusionabstractSpatial-spectral feature fusion is well acknowledged as an effective method for hyperspectral (HS) image classification. Many previous studies have been devoted to this subject. However, these methods often regard the spatial-spectral high-dimensional data as 1-D vector and then extract informative features for classification. In this paper, we propose a new HS image classification method. Specifically, matrix-based spatial-spectral feature representation is designed for each pixel to capture the local spatial contextual and the spectral information of all the bands, which can well preserve the spatial-spectral correlation. Then, matrix-based discriminant analysis is adopted to learn the discriminative feature subspace for classification. To further improve the performance of discriminative subspace, a random sampling technique is used to produce a subspace ensemble for final HS image classification. Experiments are conducted on three HS remote sensing data sets acquired by different sensors, and experimental results demonstrate the efficiency of the proposed method. Renlong Hang, Qingshan Liu 0001, Huihui Song 0002, Yubao Sun |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Low rank driven robust facial landmark regression
Jiankang Deng, Yubao Sun, Qingshan Liu 0001, Hanqing Lu |
Neurocomputing | 2 |
| 2014 | Learning Discriminative Dictionary for Group Sparse RepresentationabstractIn recent years, sparse representation has been widely used in object recognition applications. How to learn the dictionary is a key issue to sparse representation. A popular method is to use l1 norm as the sparsity measurement of representation coefficients for dictionary learning. However, the l1 norm treats each atom in the dictionary independently, so the learned dictionary cannot well capture the multisubspaces structural information of the data. In addition, the learned subdictionary for each class usually shares some common atoms, which weakens the discriminative ability of the reconstruction error of each subdictionary. This paper presents a new dictionary learning model to improve sparse representation for image classification, which targets at learning a class-specific subdictionary for each class and a common subdictionary shared by all classes. The model is composed of a discriminative fidelity, a weighted group sparse constraint, and a subdictionary incoherence term. The discriminative fidelity encourages each class-specific subdictionary to sparsely represent the samples in the corresponding class. The weighted group sparse constraint term aims at capturing the structural information of the data. The subdictionary incoherence term is to make all subdictionaries independent as much as possible. Because the common subdictionary represents features shared by all classes, we only use the reconstruction error of each class-specific subdictionary for classification. Extensive experiments are conducted on several public image databases, and the experimental results demonstrate the power of the proposed method, compared with the state-of-the-arts. Yubao Sun, Qingshan Liu 0001, Jinhui Tang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2012 | Perceptual image quality assessment based on structural similarity and visual masking
Xuan Fei, Liang Xiao 0001, Yubao Sun, Zhihui Wei |
Signal Process. Image Commun. | 3 |
| 2009 | Compressed sensing image reconstruction based on morphological component analysisabstractCompressed sensing (CS) is a new area of signal processing for simultaneous signal sampling and compression. Most of existing methods for CS image reconstruction are suitable for piecewise smooth image, but do not behave well on texture-rich natural image. In this paper, a new optimization problem for CS image reconstruction is proposed, in which different regularization terms are introduced for different morphological components of image. Furthermore, an alternating iterative algorithm is presented to solve the relevant optimization problem. Experimental results show that the proposed method can be applied to reconstruct texture-rich images besides piecewise smooth ones, and outperforms the existing methods on preserving detail feature. Xingxiu Li, Zhihui Wei, Liang Xiao 0001, Yubao Sun, Jian Yang 0003 |
ICIP | 4 |