VLDB 2026 Research / reviewers in the wild / expert
Xin Sun 0003
dblp:20/3535-3
· DBLP profile ↗
88ranked-venue papers
23as first author
39since 2021 · last 2026
0000-0002-2125-2595ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 9 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge Self-Adversarial Augmentation Enhances Graph Contrastive Learning Against Neighborhood InconsistencyabstractRecent studies have shown that unsupervised graph contrastive learning (GCL) is vulnerable to adversarial attacks. Automatic adversarial augmentation techniques are proposed to improve both the effectiveness and robustness of GCL. Existing methods typically regard unsupervised contrastive loss as the adversarial goal, essentially aiming to maximize inter-view instance-wise discrepancies between adversarial and original views. However, such attacks overlook intra-view neighborhood inconsistency, which hinders the robustness of GCL models against local neighborhood noises, resulting in performance degradation on low-homophily graphs. To tackle this issue, we propose a novel adversarial contrastive paradigm, named Edge self-aDversarial Augmentation for Graph Contrastive Learning (EDA-GCL). We theoretically establish that the adversarial objective of the intra-view neighborhood is equivalent to maximizing the discrepancy between bidirectional edge features. Hence, we build our adversarial framework based on edge self-adversarial learning. It generates pairwise adversarial augmentations from the original view by learning distinct neighborhood connectivity structures. The learned pairwise adversarial views are utilized for GCL model training in the minimization stage. Notably, this edge-level adversarial approach reduces the computational complexity to the level of the edge number. Experiments on various graph tasks and complex noise scenarios demonstrate the superiority and robustness of our EDA-GCL. Chunchun Chen, Chenrun Wang, Yiwei Fu, Xin Sun 0003, Rui Fan 0001, Wei Ye 0001 |
AAAI | 7 |
| 2026 | SSDMamba: A spectral-spatial dual-branch mamba for hyperspectral image classification
Zhaopeng Deng, Gengshen Wu, Xin Sun 0003 |
Neurocomputing | 5 |
| 2026 | Adaptive gradient-oriented sampling and dynamic-blended shadow generation for realistic face relighting
Hanzhao Pan, Gengshen Wu, Xin Sun 0003 |
Neurocomputing | 3 |
| 2026 | MEF-DETR: Multi-scale edge-enhanced and feature-fused transformer for robust aerial small object detection
Zejv Wu, Zhichao Qi, Junyu Dong, Xin Sun 0003 |
Neurocomputing | 5 |
| 2026 | Bootstrap Deep Spectral Clustering With Optimal TransportabstractSpectral clustering is a leading clustering method. Two of its major shortcomings are the disjoint optimization process and the limited representation capacity. To address these issues, we propose a deep spectral clustering model (named BootSC), which jointly learns all stages of spectral clustering—affinity matrix construction, spectral embedding, and$k$-means clustering—using a single network in an end-to-end manner. BootSC leverages effective and efficient optimal-transport-derived supervision to bootstrap the affinity matrix and the cluster assignment matrix. Moreover, a semantically-consistent orthogonal re-parameterization technique is introduced to orthogonalize spectral embeddings, significantly enhancing the discrimination capability. Experimental results indicate that BootSC achieves state-of-the-art clustering performance. For example, it accomplishes a notable 16% NMI improvement over the runner-up method on the challenging ImageNet-Dogs dataset. Our code is available athttps://github.com/spdj2271/BootSC. Wengang Guo, Wei Ye 0001, Chunchun Chen, Xin Sun 0003, Christian Böhm 0001, Claudia Plant, Susanto Rahardja |
IEEE Trans. Multim. | 4 |
| 2026 | D3BSR: Blind Super-Resolution via Diffusion-Based Disentangled Degradation RepresentationabstractExisting Blind Super-Resolution (BSR) methods are mostly trained on artificial synthetic degradation data pairs or rely on specific degradation priors, which lead to poor performance due to the trained degradation mismatch between other unknown complex degradations in real-world scenarios. To tackle this problem, we propose a novel Diffusion-based Disentangled Degradation representation method for BSR, dubbed D3BSR, which disentangles arbitrary unknown degradation into structure and texture degradations to enhance perception and fidelity quality individually. Specifically, the structure degradation is optimized by degradation distribution transition with a self-supervised collaborative learning strategy to recursively minimize the perception error. The texture degradation is restored through posterior sampling controlled by a fidelity coefficient to leverage rich texture priors encapsulated in a pre-trained diffusion model for preserving fidelity. The degraded image is super-resolved using an analytical solution with the pseudo inverse of the structural and texture degradation, which achieves a controllable trade-off between perception and fidelity and does not rely on any degradation priors or extra-supervised training. Extensive experiments on the nine heavily degraded synthetic and real-world natural and face datasets demonstrate that our D3BSR outperforms SOTA methods on the diverse metrics in reconstruction faithfulness and perceptual quality. Wei Yu 0004, Qinglin Liu, Quanling Meng, Chenyang Wang 0002, Xin Sun 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | OTPNet: ODE-inspired Tuning-free Proximal Network for Remote Sensing Image FusionabstractRemote sensing image fusion aims to reconstruct a high spatial and spectral resolution image by integrating the spatial and spectral information from multiple remote sensing sensor data. Despite the remarkable progress of deep learning-based fusion methods, most existing methods rely on manual network architecture design and hyperparameter tuning, lacking sufficient interpretability and adaptability. To address this limitation, we propose a novel neural Ordinary Differential Equation (ODE)-inspired tuning-free proximal splitting algorithm, which splits remote sensing image fusion as two optimization problems regularized by deep priors to model the fusion of spatial and spectral. Firstly, based on the physical properties of spatial and spectral information, the two problems are optimized by two proximal splitting operators to iteratively integrate spatial-spectral complementary information, eliminating or suppressing redundant information to reduce fusion errors. Secondly, considering the efficiency of neural ODE in reducing optimization error, we utilize a high-order numerical scheme to customize the proximal operator theoretically without additional handcrafted design and parameter tuning. Finally, by incorporating the numerical scheme as a solver into the proximal optimization algorithm, we derive an ODE-inspired Tuning-free Proximal Network, dubbed OTPNet, which achieves efficient and robust fusion reconstruction. Extensive experiments on nine datasets across three different remote sensing image fusion tasks show that our OTPNet outperforms existing state-of-the-art approaches, which validates the effectiveness of our method. Wei Yu 0002, Zonglin Li 0004, Qinglin Liu, Xin Sun 0003 |
AAAI | 4 |
| 2025 | ProsodyTalker: 3D Visual Speech Animation via Prosody DecompositionabstractMost existing 3D visual speech animation methods synthesize lip movements synchronized with speech, which however neglect head poses and therefore degrade the animation realism. The animation of head poses presents two primary challenges: (1) the intricate mapping between speech and head poses remains poorly understood and (2) the absence of 4D face datasets featuring realistic head poses. Inspired by prosody decomposition in speech processing, we discern that head movements correlate with the fundamental frequency (F0) of speech prosody, while lip movements align with the language content. These observations motivate us to propose a novel framework, dubbed ProsodyTalker, that concurrently synthesizes lip and head movements, grounded in the principles of prosody decomposition. The core idea is first to adopt information perturbation to explicitly decompose the speech prosody into pose-related F0 and lip-related language content. Then, an autoregressive content-oriented fusion decoder is employed to enhance lip synchronization in the synthesized facial sequences. To synthesize head poses, we design a transformer-based variational autoencoder to learn a latent distribution of facial sequences and propose an F0-conditioned latent diffusion model to establish a probabilistic mapping from F0 to pose-related latent codes. Furthermore, we contribute a large-scale 4D face dataset containing bunches of variations in identities, head poses and facial motions. Extensive experiments show that our method achieves more realistic animation than state-of-the-art methods. Zonglin Li 0004, Xiaoqian Lv, Qinglin Liu, Quanling Meng, Xin Sun 0003, Shengping Zhang |
AAAI | 5 |
| 2025 | Path-Adaptive Matting for Efficient Inference Under Various Computational Cost ConstraintsabstractIn this paper, we explore a novel image matting task aimed at achieving efficient inference under various computational cost constraints, specifically FLOP limitations, using a single matting network. Existing matting methods which have not explored scalable architectures or path-learning strategies, fail to tackle this challenge. To overcome these limitations, we introduce Path-Adaptive Matting (PAM), a framework that dynamically adjusts network paths based on image contexts and computational cost constraints. We formulate the training of the computational cost-constrained matting network as a bilevel optimization problem, jointly optimizing the matting network and the path estimator. Building on this formalization, we design a path-adaptive matting architecture by incorporating path selection layers and learnable connect layers to estimate optimal paths and perform efficient inference within a unified network. Furthermore, we propose a performance-aware path-learning strategy to generate path labels online by evaluating a few paths sampled from the prior distribution of optimal paths and network estimations, enabling robust and efficient online path learning. Experiments on five image matting datasets demonstrate that the proposed PAM framework achieves competitive performance across a range of computational cost constraints. Qinglin Liu, Zonglin Li 0004, Xiaoqian Lv, Xin Sun 0003, Ru Li 0002, Shengping Zhang |
AAAI | 4 |
| 2025 | Multi-view Consistent 3D Panoptic Scene Understandingabstract3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machine-generated 2D panoptic segmentation masks as training labels. However, these generated masks often exhibit multi-view inconsistencies, leading to ambiguities during the optimization process. To address this, we present Multi-view Consistent 3D Panoptic Scene Understanding (MVC-PSU), featuring two key components: 1) Probabilistic Semantic Aligner, which associates semantic information of corresponding pixels across multiple views by probabilistic alignment to ensure that predicted panoptic segmentation masks are consistent across different views. 2) Geometric Consistency Enforcer, which uses multi-view projection and monocular depth consistency to ensure that the geometry of the reconstructed scene is accurate and consistent across different views. Experimental results demonstrate that the proposed MVC-PSU surpasses state-of-the-art methods on the ScanNet, Replica, and HyperSim datasets. Xianzhu Liu, Xin Sun 0003, Haozhe Xie, Zonglin Li 0004, Ru Li 0002, Shengping Zhang |
AAAI | 2 |
| 2025 | An Efficient CNN-Transformer Architecture for Facial Expression Recognition in Cloud-Edge SystemsabstractFacial Expression Recognition (FER) is a critical technology in the field of human-computer interaction, with broad prospects in cloud service applications such as intelligent monitoring, remote education, and healthcare. However, in practical cloud-edge collaborative deployment scenarios, models need to balance high accuracy with computational efficiency. Traditional Convolutional Neural Networks (CNNs), while adept at extracting local features, struggle to effectively model the long-range spatial dependencies necessary for distinguishing expressions. Conversely, although Vision Transformer models excel at global modeling, their high computational overhead presents challenges for deployment on resource-constrained edge devices. Therefore, this paper proposes a hybrid CNN-Transformer architecture that utilizes a pre-trained ResNet-18, optimized for low-resolution input images, as the backbone for local feature extraction. These features are then fed into a compact Transformer encoder to model global relationships. To accelerate training convergence in the cloud and mitigate overfitting, we introduce several advanced training strategies, including Mixup and CoarseDropout. Experimental results on the public FER2013plus dataset demonstrate that the proposed model achieves a test accuracy of 84.31%, significantly outperforming traditional methods. Hanzhao Pan, Gengshen Wu, Xin Sun 0003 |
CloudCom | 3 |
| 2025 | A comb concatenation diffusion model for hyperspectral image super-resolution
Yinghao Xu 0003, Hao Wang 0192, Xin Sun 0003, Qianlong Xie, Peng Ren 0001, Fei Zhou 0007, Susanto Rahardja |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Boosting accuracy of student models via Masked Adaptive Self-Distillation
Shuwen Tian, Zhaopeng Deng, Xin Sun 0003, Junyu Dong |
Neurocomputing | 5 |
| 2025 | MSFFT-Net: A multi-scale feature fusion transformer network for underwater image enhancement
Zeju Wu, Kaiming Chen, Panxin Ji, Xin Sun 0003 |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Fast tone mapping operator for high dynamic range image using prior information
Xueyu Han, Xin Sun 0003, Susanto Rahardja |
Signal Process. Image Commun. | 2 |
| 2025 | MambaHSISR: Mamba Hyperspectral Image Super-ResolutionabstractOne of the main challenges facing hyperspectral image super-resolution is the complex high dimensional data processing. Mamba leverages its ability to model long-range dependencies of linear complexity to capture the global spatial and spectral information of high-dimensional data while maintaining linear complexity. However, its visual state space equation mainly focuses on the band dimension mapping of the image, while ignoring the modeling of the spatial dimension. To overcome this limitation, we develop a Mamba hyperspectral image super-resolution framework, which comprises three essential components. The first component, i.e., spatial Mamba sub-network, models the spatial dimensions of hyperspectral data. It captures long-range dependencies in the pixel space, thereby integrating global spatial information into the framework. The second component, i.e., spectral Mamba sub-network, serves to capture long-range spectral dependencies. The third component, i.e., reconstruction, generates hyperspectral images with rich spatial and spectral details through pixel interpolation. Our Mamba framework fully develops the potential of the Mamba model in hyperspectral image super-resolution, significantly enhancing the restoration quality and accuracy of hyperspectral images. Extensive experiments on the Houston and QUST-1 datasets show that our framework outperforms state-of-the-art methods in both quantitative metrics and visual quality across diverse scenarios. We release our source code at https://gitee.com/xu_yinghao/MambaHSISR for public evaluations. Yinghao Xu 0003, Hao Wang 0192, Fei Zhou 0007, Chunbo Luo, Xin Sun 0003, Susanto Rahardja, Peng Ren 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Distributional Shortest-Path Graph KernelsabstractTraditional shortest-path graph kernels generate for each graph a histogram-like feature map, whose elements represent the number of occurrences of non-isomorphic shortest paths in this graph. The histogram-like feature map does not contain the distributions of the shortest paths within and across graphs, causing inaccurate graph similarities. To this end, we propose a novel graph kernel called the Distributional Shortest-Path (DSP) graph kernel to embrace both types of distribution information. Since the distribution of substructures (e.g., the shortest paths) follows a power law like that of words in natural language, we utilize neural language models to learn each node's distributional shortest-path feature map, encompassing the distributions and dependencies of the shortest paths in each graph. Moreover, we design the Partition Kernel (PK) to capture the dataset-wide distribution information of the shortest paths. PK projects similar (i.e., belonging to the same partition) distributional shortest-path node feature maps to the same point in the Reproducing Kernel Hilbert Space. Finally, Kernel Mean Embedding (KME) is applied to compute graph feature maps and efficiently construct the DSP graph kernel. Empirical experiments demonstrate that DSP outperforms state-of-the-art graph kernels on most benchmark datasets. Wei Ye 0001, Wengang Guo, Shuhao Tang, Xin Sun 0003, Xiaofeng Cao 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | ADOD: Adaptive Density Outlier DetectionabstractOutlier detection plays a dual role in data analysis: cleansing data to optimize the performance of downstream tasks and identifying potentially rare valuable events or patterns. Proximity-based methods, which are independent of data distribution assumptions, are plagued by parameter selection and performance challenges when handling data with varying densities. This study proposed a novel unsupervised algorithm named Adaptive Density Outlier Detection (ADOD) to address these challenges. The core innovation of ADOD involves two main aspects: adaptive neighborhood boundaries and density consistency scoring. First, instead of relying on a predefined fixed radius, ADOD employs perplexity to calculate the local scale of each data point. It then dynamically adjusts the neighborhood boundaries according to this scale to adapt to data with varying densities. Second, ADOD estimates local density using a mutual neighbor graph and combines the density differences between data points and their neighbors to compute outlier scores, effectively distinguishing outliers that significantly deviate from their surroundings. This study evaluated ADOD on one synthetic and 32 real datasets, and compared it with 14 classical and state-of-the-art algorithms from different categories. Extensive experimental results demonstrated the superior performance of ADOD, achieving the highest average accuracy across ROC, P@N, and AP metrics. This study promotes the development of outlier detection techniques and expands their potential for real-time applications. Li Qian 0001, Xin Sun 0003, Wengang Guo, Christian Böhm 0001 |
ICDM | 3 |
| 2024 | Shape-Guided Clothing Warping for Virtual Try-OnabstractImage-based virtual try-on aims to seamlessly fit in-shop clothing to a person image while maintaining pose consistency. Existing methods commonly employ the thin plate spline (TPS) transformation or appearance flow to deform in-shop clothing for aligning with the person's body. Despite their promising performance, these methods often lack precise control over fine details, leading to inconsistencies in shape between clothing and the person's body as well as distortions in exposed limb regions. To tackle these challenges, we propose a novel shape-guided clothing warping method for virtual try-on, dubbed SCW-VTON, which incorporates global shape constraints and additional limb textures to enhance the realism and consistency of the warped clothing and try-on results. To integrate global shape constraints for clothing warping, we devise a dual-path clothing warping module comprising a shape path and a flow path. The former path captures the clothing shape aligned with the person's body, while the latter path leverages the mapping between the pre- and post-deformation of the clothing shape to guide the estimation of appearance flow. Furthermore, to alleviate distortions in limb regions of try-on results, we integrate detailed limb guidance by developing a limb reconstruction network based on masked image modeling. Through the utilization of SCW-VTON, we are able to generate try-on results with enhanced clothing shape consistency and precise control over details. Extensive experiments demonstrate the superiority of our approach over state-of-the-art methods both qualitatively and quantitatively. Shunyuan Zheng, Zonglin Li 0004, Chenyang Wang 0002, Xin Sun 0003, Quanling Meng |
ACM Multimedia | 5 |
| 2024 | Attention-Based Multi-Kernelized and Boundary-Aware Network for image semantic segmentation
Xuanchen Zhou, Gengshen Wu, Xin Sun 0003, Pengpeng Hu, Yi Liu 0038 |
Neurocomputing | 3 |
| 2024 | Toward Open-World Text-Driven Face Generation and Manipulation via StyleGAN3abstractMost existing text-driven face image generation and manipulation methods are based on StyleGAN2, which is inherently limited to aligned faces and therefore makes these methods fail to preserve the highly variable face placement. Additionally, these methods directly leverage a pairwise loss to learn the correspondence between the image and text, which can not handle complex text descriptions, e.g., the text with multiple captions describes multiple facial attributes. To address these issues, we explore the feasibility of applying the more advanced StyleGAN3 to generate and manipulate the face images in an Open-World setup, e.g., the target face image is not required to be aligned and the text description contains multiple captions. To this end, we first design an improved iterative refinement strategy that adaptively predicts the generator weight offsets rather than residuals for the inverted latent code via a hypernetwork, which efficiently finds a desired generator with no image-specific optimization. We further analyze the disentanglement of different StyleGAN3 latent spaces and demonstrate that the${\mathcal {S}}$space learns a more semantically-disentangled representation. To enable complex edits mentioned by the multi-caption text, we propose a cross-modal feature filtration module with a probability adaptation strategy to capture the image-text correspondences. Finally, we incorporate a channel-wise attention mechanism to obtain a global latent manipulation direction, which learns to assign importance weights to different channels. Extensive experiments demonstrate the superior performance of our proposed method compared against the state-of-the-art methods. Zonglin Li 0004, Peiqiang Liu, Qinglin Liu, Xin Sun 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Attention-ConvNet Network for Ocean-Front Prediction via Remote Sensing SST ImagesabstractOcean front is one typical geophysical phenomenon acting as oases in the ocean for fishes and marine mammals. Accurate ocean-front prediction is critical for fishery and navigation safety. However, the formation and evolution of ocean fronts are inherently nonlinear and are influenced by various factors such as ocean currents, wind fields, and temperature changes, making ocean-front prediction a considerable challenge. This study proposes a temporal-sensitive network named Attention-ConvNet to address this challenge. Ocean fronts exhibit significant multiscale characteristics, requiring analysis and prediction across various temporal and spatial scales. The proposed network designs a hierarchical attention mechanism (HAM) that efficiently prioritizes relevant spatial and temporal information to meet the specific requirement. What is more, the proposed network uses a complex hierarchical branching convolutional network (HBCNet) architecture, which allows our network to leverage the complementary strengths of spatial and temporal information, effectively capturing the dynamic and complex variations in ocean fronts. In general, the network prioritizes and focuses on the most relevant information of front dynamics, which ensures its ability to effectively predict the ocean front. External experiments demonstrate that our network significantly outperforms conventional methods, confirming its capability for precise ocean-front prediction. The codes will be publicly available athttps://github.com/yuhudeyue/Ocean-Front-Prediction-Model. Yuting Yang 0001, Xin Sun 0003, Junyu Dong, Kin-Man Lam 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | High-order paired-ASPP for deep semantic segmentation networks
Xin Sun 0003, Yu Zhang 0165, Changrui Chen, Sihang Xie, Junyu Dong |
Inf. Sci. | 1 |
| 2023 | Dual autoencoder based zero shot learning in special domain
Eric Rigall, Xin Sun 0003, Kin-Man Lam 0001, Junyu Dong |
Pattern Anal. Appl. | 3 |
| 2023 | SurroundNet: Towards effective low-light image enhancement
Fei Zhou 0007, Xin Sun 0003, Junyu Dong, Xiao Xiang Zhu 0001 |
Pattern Recognit. | 2 |
| 2023 | Network Embedding via Deep Prediction ModelabstractNetwork-structured data becomes ubiquitous in daily life and is growing at a rapid pace. It presents great challenges to feature engineering due to the high non-linearity and sparsity of the data. The local and global structure of the real-world networks can be reflected by dynamical transfer behaviors among nodes. This paper proposes a network embedding framework to capture the transfer behaviors on structured networks via deep prediction models. We first design a degree-weight biased random walk model to capture the transfer behaviors on the network. Then a deep network embedding method is introduced to preserve the transfer possibilities among the nodes. A network structure embedding layer is added into conventional deep prediction models, including Long Short-Term Memory Network and Recurrent Neural Network, to utilize the sequence prediction ability. To keep the local network neighborhood, we further perform a Laplacian supervised space optimization on the embedding feature representations. Experimental studies are conducted on various datasets including social networks, citation networks, biomedical network, collaboration network and language network. The results show that the learned representations can be effectively used as features in a variety of tasks, such as clustering, visualization, classification, reconstruction and link recovery, and achieve promising performance compared with state-of-the-arts. Xin Sun 0003, Zenghui Song, Junyu Dong, Claudia Plant, Christian Böhm 0001 |
IEEE Trans. Big Data | 1 |
| 2023 | Adaptive Morphology Filter: A Lightweight Module for Deep Hyperspectral Image ClassificationabstractDeep neural network models significantly outperform classical algorithms in the hyperspectral image (HSI) classification task. These deep models improve generalization but incur significant computational demands. This article endeavors to alleviate the computational distress in a depthwise manner through the use of morphological operations. We propose the adaptive morphology filter (AMF) to effectively extract spatial features like the conventional depthwise convolution layer. Furthermore, we reparameterize AMF into its equivalent form, i.e., a traditional binary morphology filter, which drastically reduces the number of parameters in the inference phase. Finally, we stack multiple AMFs to achieve a large receptive field and construct a lightweight AMNet for classifying HSIs. It is noteworthy that we prove the deep stack of depthwise AMFs to be equivalent to structural element decomposition. We test our model on five benchmark datasets. Experiments show that our approach outperforms state-of-the-art methods with fewer parameters (${\approx }10 k$). The codes will be publicly available athttps://github.com/zhu-xlab/Adaptive-Morphology-Filter. Fei Zhou 0007, Xin Sun 0003, Chengze Sun, Junyu Dong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Average Activation Network for Weakly Supervised Semantic SegmentationabstractMost current image-level weakly supervised semantic segmentation (WSSS) methods are based on class activation map (CAM). However, the main limitation of weakly supervised semantic segmentation is that the CAMs generated by the WSSS network always focus on the most discriminative parts of the object, limiting the CAM to capture the holistic object. So we propose an average activation network to generate CAMs of the holistic object with weakly supervision. We restrain the highest activation regions of the CAM by continuously splitting the training image on the maximum activation point of the CAM during the training process. In this way, our network pays more attention to the whole of the object. In addition, we propose a CAM similarity loss function to narrow the gap between fully-supervised semantic segmentation (FSSS) and WSSS. We conducted experiments on the PASCAL VOC 2012 dataset to validate the effectiveness of our method. Zhenkun Fan, Xin Sun 0003, Junyu Dong |
ICPR | 2 |
| 2022 | Self-attention neural architecture search for semantic image segmentation
Zhenkun Fan, Guosheng Hu, Xin Sun 0003, Gaige Wang, Junyu Dong, Chi Su |
Knowl. Based Syst. | 3 |
| 2022 | Multi-instance semantic similarity transferring for knowledge distillation
Xin Sun 0003, Junyu Dong, Hui Yu 0001, Gaige Wang |
Knowl. Based Syst. | 2 |
| 2022 | Gaussian Dynamic Convolution for Efficient Single-Image SegmentationabstractInteractive single-image segmentation is ubiquitous in the scientific and commercial imaging software. Lightweight neural network is one practical and effective way to accomplish the single-image segmentation task. This work focuses on the single-image segmentation problem only with some seeds such as scribbles. Inspired by the dynamic receptive field in the human being’s visual system, we propose the Gaussian dynamic convolution (GDC) to fast and efficiently aggregate the contextual information for neural networks. The core idea is randomly selecting the spatial sampling area according to the Gaussian distribution offsets. Our GDC can be easily used as a module to build lightweight or complex segmentation networks. We adopt the proposed GDC to address the typical single-image segmentation tasks. Furthermore, we also build a Gaussian dynamic pyramid Pooling to show its potential and generality in common semantic segmentation. Experiments demonstrate that the GDC outperforms other existing convolutions on three benchmark segmentation datasets including Pascal-Context, Pascal-VOC 2012, and Cityscapes. Additional experiments are also conducted to illustrate that the GDC can produce richer and more vivid features compared with other convolutions. In general, our GDC is conducive to the convolutional neural networks to form an overall impression of the image. Xin Sun 0003, Changrui Chen, Junyu Dong, Huiyu Zhou 0001, Sheng Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Highlight Every Step: Knowledge Distillation via Collaborative TeachingabstractHigh storage and computational costs obstruct deep neural networks to be deployed on resource-constrained devices. Knowledge distillation (KD) aims to train a compact student network by transferring knowledge from a larger pretrained teacher model. However, most existing methods on KD ignore the valuable information among the training process associated with training results. In this article, we provide a new collaborative teaching KD (CTKD) strategy which employs two special teachers. Specifically, one teacher trained from scratch (i.e., scratch teacher) assists the student step by step using its temporary outputs. It forces the student to approach the optimal path toward the final logits with high accuracy. The other pretrained teacher (i.e., expert teacher) guides the student to focus on a critical region that is more useful for the task. The combination of the knowledge from two special teachers can significantly improve the performance of the student network in KD. The results of experiments on CIFAR-10, CIFAR-100, SVHN, Tiny ImageNet, and ImageNet datasets verify that the proposed KD method is efficient and achieves state-of-the-art performance. Xin Sun 0003, Junyu Dong, Changrui Chen, Zihe Dong |
IEEE Trans. Cybern. | 2 |
| 2022 | Parallel Complement Network for Real-Time Semantic Segmentation of Road ScenesabstractReal-time semantic segmentation is in intense demand for the application of autonomous driving. Most of the semantic segmentation models tend to use large feature maps and complex structures to enhance the representation power for high accuracy. However, these inefficient designs increase the amount of computational costs, which hinders the model to be applied on autonomous driving. In this paper, we propose a lightweight real-time segmentation model, named Parallel Complement Network (PCNet), to address the challenging task with fewer parameters. A Parallel Complement layer is introduced to generate complementary features with a large receptive field. It provides the ability to overcome the problem of similar feature encoding among different classes, and further produces discriminative representations. With the inverted residual structure, we design a Parallel Complement block to construct the proposed PCNet. Extensive experiments are carried out on challenging road scene datasets, i.e., CityScapes and CamVid, to make comparison against several state-of-the-art real-time segmentation models. The results show that our model has promising performance. Specifically, PCNet* achieves 72.9% Mean IoU on CityScapes using only 1.5M parameters and reaches 79.1 FPS with$1024\times 2048$resolution images on GTX 2080Ti. Moreover, our proposed system achieves the best accuracy when being trained from scratch. Qingxuan Lv, Xin Sun 0003, Changrui Chen, Junyu Dong, Huiyu Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Hierarchical Triplet Attention Pooling for Graph ClassificationabstractRecent years, graph neural network have been introduced to handle the structural data in non-Euclidean space. Graph neural network learns the representation of the whole network through the representation of nodes. The pooling layer plays an important role in gradually reducing the network to a sufficient small coarse graph. This work will propose an end-to-end pooling method, which builds pseudo-edge weights to form a subset of important features. It updates the coarse-grained node embedding by focusing on the common information at different locations, to improve the attention to the valuable information. Moreover, we propose a transfer operator that integrates into the convolutional layer, which not only integrates the intrinsic features of the node but also learns rich community properties. Experimental results show that our proposed pooling method can be combined with multiple convolutional layers to achieve optimal results in graph classification tasks. In addition, our proposed model method is generally superior to most baseline graph classification methods. Liande Bi, Xin Sun 0003, Fei Zhou 0007, Junyu Dong |
ICTAI | 2 |
| 2021 | Fusing attributed and topological global-relations for network embedding
Xin Sun 0003, Junyu Dong, Claudia Plant, Christian Böhm 0001 |
Inf. Sci. | 1 |
| 2021 | Is It Easy to Recognize Baby's Age and Gender?
Yang Liu 0119, Ruili He, Xiaoqian Lv, Wei Wang 0107, Xin Sun 0003, Shengping Zhang |
J. Comput. Sci. Technol. | 5 |
| 2021 | Knowledge distillation via instance-level sequence learning
Xin Sun 0003, Junyu Dong, Zihe Dong |
Knowl. Based Syst. | 2 |
| 2021 | GPNet: Gated pyramid network for semantic segmentation
Yu Zhang 0165, Xin Sun 0003, Junyu Dong, Changrui Chen, Qingxuan Lv |
Pattern Recognit. | 2 |
| 2021 | A Deep Framework for Eddy Detection and Tracking From Satellite Sea Surface Height DataabstractOcean eddies, as a ubiquitous phenomenon of the global ocean, are extremely important for ocean energy and material exchanges. Therefore, efficient eddy detection and tracking are crucial for advancing our understanding of ocean dynamics. This work presents a framework for automatic ocean eddy detection and tracking by leveraging state-of-the-art machine learning algorithms. First, we propose a new convolutional neural network model for multieddies detection. This model is capable of extracting accurate boundary information of eddies and fitting the gap between semantic context and sea surface height (SSH). Second, a tracking algorithm is designed to track eddies lasting a number of days and provide visualization of the dynamical processes governing eddies' movements. Finally, we have made our data set publicly available, which is named SCSE-Eddy and can be used as a benchmark to evaluate the performances of artificial intelligence (AI)-based eddy detection methods. The data set covers daily remotely sensed SSH data located in the South China Sea and its eastern sea areas over a period of 15 years. The experimental results show that our methods achieve promising performances compared to existing approaches, especially for the eddies with indistinct geographical border. We believe that this work opens a new avenue for oceanographers to better discover and understand the physical properties of ocean eddies. Xin Sun 0003, Junyu Dong, Redouane Lguensat, Yuting Yang 0001, Xirong Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Learning Deep Relations to Promote Saliency DetectionabstractThough saliency detectors has made stunning progress recently. The performances of the state-of-the-art saliency detectors are not acceptable in some confusing areas, e.g., object boundary. We argue that the feature spatial independence should be one of the root cause. This paper explores the ubiquitous relations on the deep features to promote the existing saliency detectors efficiently. We establish the relation by maximizing the mutual information of the deep features of the same category via deep neural networks to break this independence. We introduce a threshold-constrained training pair construction strategy to ensure that we can accurately estimate the relations between different image parts in a self-supervised way. The relation can be utilized to further excavate the salient areas and inhibit confusing backgrounds. The experiments demonstrate that our method can significantly boost the performance of the state-of-the-art saliency detectors on various benchmark datasets. Besides, our model is label-free and extremely efficient. The inference speed is 140 FPS on a single GTX1080 GPU. Changrui Chen, Xin Sun 0003, Yang Hua 0001, Junyu Dong, Hongwei Xv |
AAAI | 2 |
| 2020 | RSAN: Residual Subtraction and Attention Network for Single Image Super-ResolutionabstractThe single-image super-resolution (SISR) aims to recover a potential high-resolution image from its low-resolution version. Recently, deep learning-based methods have played a significant role in super-resolution field due to its effectiveness and efficiency. However, most of the SISR methods neglect the importance among the feature map channels. Moreover, they can not eliminate the redundant noises, making the output image be blurred. In this paper, we propose the residual subtraction and attention network (RSAN) for powerful feature expression and channels importance learning. More specifically, RSAN firstly implements one redundance removal module to learn noise information in the feature map and subtract noise through residual learning. Then it introduces the channel attention module to amplify high-frequency information and suppress the weight of effectless channels. Experimental results on extensive public benchmarks demonstrate our RSAN achieves significant improvement over the previous SISR methods in terms of both quantitative metrics and visual quality. Shuo Wei, Xin Sun 0003, Junyu Dong |
ICPR | 2 |
| 2020 | Underwater Enhancement Model via Reverse Dark Channel Prior
Xin Sun 0003, Yu Zhang 0165, Junyu Dong |
PRCV (1) | 3 |
| 2020 | Exploring ubiquitous relations for boosting classification and localization
Xin Sun 0003, Changrui Chen, Junyu Dong, Guosheng Hu |
Knowl. Based Syst. | 1 |
| 2020 | A self-adjusting quantum key renewal management scheme in classical network symmetric cryptography
Jiawei Han 0006, Yanheng Liu 0001, Xin Sun 0003, Aiping Chen |
J. Supercomput. | 3 |
| 2020 | Correction to: A self-adjusting quantum key renewal management scheme in classical network symmetric cryptography
Jiawei Han 0006, Yanheng Liu 0001, Xin Sun 0003, Aiping Chen |
J. Supercomput. | 3 |
| 2019 | Network Structure and Transfer Behaviors Embedding via Deep Prediction ModelabstractNetwork-structured data is becoming increasingly popular in many applications. However, these data present great challenges to feature engineering due to its high non-linearity and sparsity. The issue on how to transfer the link-connected nodes of the huge network into feature representations is critical. As basic properties of the real-world networks, the local and global structure can be reflected by dynamical transfer behaviors from node to node. In this work, we propose a deep embedding framework to preserve the transfer possibilities among the network nodes. We first suggest a degree-weight biased random walk model to capture the transfer behaviors of the network. Then a deep embedding framework is introduced to preserve the transfer possibilities among the nodes. A network structure embedding layer is added into the conventional Long Short-Term Memory Network to utilize its sequence prediction ability. To keep the local network neighborhood, we further perform a Laplacian supervised space optimization on the embedding feature representations. Experimental studies are conducted on various real-world datasets including social networks and citation networks. The results show that the learned representations can be effectively used as features in a variety of tasks, such as clustering, visualization and classification, and achieve promising performance compared with state-of-the-art models. Xin Sun 0003, Zenghui Song, Junyu Dong, Claudia Plant, Christian Böhm 0001 |
AAAI | 1 |
| 2019 | Deep pixel-to-pixel network for underwater image enhancement and restorationabstractTurbid underwater environment poses great difficulties for the applications of vision technologies. One of the biggest challenges is the complicated noise distribution of the underwater images due to the serious scattering and absorption. To alleviate this problem, this work proposes a deep pixel‐to‐pixel networks model for underwater image enhancement by designing an encoding–decoding framework. It employs the convolution layers as encoding to filter the noise, while uses deconvolution layers as decoding to recover the missing details and refine the image pixel by pixel. Moreover, skip connection is introduced in the networks model in order to avoid low‐level features losing while accelerating the training process. The model achieves the image enhancement in a self‐adaptive data‐driven way rather than considering the physical environment. Several comparison experiments are carried out with different datasets. Results show that it outperforms the state‐of‐the‐art image restoration methods in underwater image defogging, denoising and colour enhancement. Xin Sun 0003, Lipeng Liu, Junyu Dong, Estanislau Lima, Ruiying Yin |
IET Image Process. | 1 |
| 2019 | A procedural texture generation framework based on semantic descriptions
Junyu Dong, Jun Liu 0055, Ying Gao 0005, Lin Qi 0004, Xin Sun 0003 |
Knowl. Based Syst. | 6 |
| 2019 | Inpainting of Remote Sensing SST Images With Deep Convolutional Generative Adversarial NetworkabstractCloud occlusion is a common problem in the satellite remote sensing (RS) field and poses great challenges for image processing and object detection. Most existing methods for cloud occlusion recovery extract the surrounding information from the single corrupted image rather than the historical RS image records. Moreover, the existing algorithms can only handle small and regular-shaped obnubilation regions. This letter introduces a deep convolutional generative adversarial network to recover the RS sea surface temperature images with cloud occlusion from the big historical image records. We propose a new loss function for the inpainting network, which adds a supervision term to solve our specific problem. Given a trained generative model, we search for the closest encoding of the corrupted image in the low-dimensional space using our inpainting loss function. This encoding is then passed through the generative model to infer the missing content. We conduct experiments on the RS image data set from the national oceanic and atmospheric administration. Compared with traditional and machine learning methods, both qualitative and quantitative results show that our method has advantages over existing methods. Junyu Dong, Ruiying Yin, Xin Sun 0003, Yuting Yang 0001, Xukun Qin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | A Multiscale Deep Framework for Ocean Fronts Detection and Fine-Grained LocationabstractOcean front plays an important role in marine fishery production and biogeochemical cycling. This letter proposes a multiscale deep framework to meet the need for automatic ocean front detection and fine-grained location. The framework mainly focuses on bringing a well-trained deep learning model into front detection and location on the global satellite sea surface temperature image. First, a multiscale scanner is designed to divide the ocean into small areas of different scales. Then, we introduce the deep model to determine that a front has occurred, and translate the global image into binary ones of various grained. Here, an overlapping scanning way is suggested to locate the front in a small region. Finally, all the binary images are scale-weighted fused into one image, which presents the center and periphery with different brightness levels. Experimental illustrations on three typical areas of the ocean are featured with six scanning scales to show the effectiveness and practical use of the proposed framework. Moreover, the comparison experiments with the traditional method also show its advantages. Xin Sun 0003, Changgang Wang, Junyu Dong, Estanislau Lima, Yuting Yang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Perception-driven procedural texture generation from examples
Jun Liu 0055, Yanhai Gan, Junyu Dong, Lin Qi 0004, Xin Sun 0003, Muwei Jian, Hui Yu 0001 |
Neurocomputing | 5 |
| 2018 | Transferring deep knowledge for object recognition in Low-quality underwater videos
Xin Sun 0003, Junyu Shi, Lipeng Liu, Junyu Dong, Claudia Plant, Huiyu Zhou 0001 |
Neurocomputing | 1 |
| 2018 | Saliency detection based on background seeds by object proposals and extended random walk
Muwei Jian, Runxia Zhao, Xin Sun 0003, Hanjiang Luo, Wenyin Zhang, Huaxiang Zhang 0001, Junyu Dong, Yilong Yin, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | A CFCC-LSTM Model for Sea Surface Temperature PredictionabstractSea surface temperature (SST) prediction is not only theoretically important but also has a number of practical applications across a variety of ocean-related fields. Although a large amount of SST data obtained via remote sensor are available, previous work rarely attempted to predict future SST values from history data in spatiotemporal perspective. This letter regards SST prediction as a sequence prediction problem and builds an end-to-end trainable long short term memory (LSTM) neural network model. LSTM naturally has the ability to learn the temporal relationship of time series data. Besides temporal information, spatial information is also included in our LSTM model. The local correlation and global coherence of each pixel can be expressed and retained by patches with fixed dimensions. The proposed model essentially combines the temporal and spatial information to predict future SST values. Its structure includes one fully connected LSTM layer and one convolution layer. Experimental results on two data sets, i.e., one Advanced Very High Resolution Radiometer SST data set covering China Coastal waters and one National Oceanic and Atmospheric Administration High-Resolution SST data set covering the Bohai Sea, confirmed the effectiveness of the proposed model. Yuting Yang 0001, Junyu Dong, Xin Sun 0003, Estanislau Lima, Quanquan Mu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Saliency detection using quaternionic distance based weber local descriptor and level priors
Muwei Jian, Qiang Qi, Junyu Dong, Xin Sun 0003, Yujuan Sun, Kin-Man Lam 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Non-rigid Object Tracking via Deformable Patches Using Shape-Preserved KCF and Level SetsabstractPart-based trackers are effective in exploiting local details of the target object for robust tracking. In contrast to most existing part-based methods that divide all kinds of target objects into a number of fixed rectangular patches, in this paper, we propose a novel framework in which a set of deformable patches dynamically collaborate on tracking of non-rigid objects. In particular, we proposed a shape-preserved kernelized correlation filter (SP-KCF) which can accommodate target shape information for robust tracking. The SP-KCF is introduced into the level set framework for dynamic tracking of individual patches. In this manner, our proposed deformable patches are target-dependent, have the capability to assume complex topology, and are deformable to adapt to target variations. As these deformable patches properly capture individual target subregions, we exploit their photometric discrimination and shape variation to reveal the trackability of individual target subregions, which enables the proposed tracker to dynamically take advantage of those subregions with good trackability for target likelihood estimation. Finally the shape information of these deformable patches enables accurate object contours to be computed as the tracking output. Experimental results on the latest public sets of challenging sequences demonstrate the effectiveness of the proposed method. Xin Sun 0003, Ngai-Man Cheung, Hongxun Yao, Yiluan Guo |
ICCV | 1 |
| 2017 | Attributed Graph Clustering with Unimodal Normalized Cut
Wei Ye 0001, Linfei Zhou, Xin Sun 0003, Claudia Plant, Christian Böhm 0001 |
ECML/PKDD (1) | 3 |
| 2017 | Learning and Transferring Convolutional Neural Network Knowledge to Ocean Front RecognitionabstractIn this letter, we investigated how to apply a deep learning method, in particular convolutional neural networks (CNNs), to an ocean front recognition task. Exploring deep CNNs knowledge to ocean front recognition is a challenging task, because the training data is very scarce. This letter overcomes this challenge using a sequence of transfer learning steps via fine-tuning. The core idea is to extract deep knowledge of the CNN model from a large data set and then transfer the knowledge to our ocean front recognition task on limited remote sensing (RS) images. We conducted experiments on two different RS image data sets, with different visual properties, i.e., colorful and gray-level data, which were both downloaded from the National Oceanic and Atmospheric Administration (NOAA). The proposed method was compared with the conventional handcraft descriptor with bag-of-visual-words, original CNN model, and last-layer fine-tuned CNN model. Our method showed a significantly higher accuracy than other methods in both datasets. Estanislau Lima, Xin Sun 0003, Junyu Dong, Yuting Yang 0001, Lipeng Liu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Encoding Spectral and Spatial Context Information for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a popular yet challenging research topic in the remote sensing community. This letter attempts to encode both spectral and spatial information into deep features for HSI classification. We first propose a semisupervised method for training the stacked autoencoder to obtain discriminative deep features. A batch training scheme is introduced to constrain the label consistency on a neighborhood region. Second, a mean pooling procedure is suggested to further fuse the spectral and local spatial information for deep feature generation. The experimental results on two hyperspectral scenes show that the proposed method achieves promising classification performance. Xin Sun 0003, Fei Zhou 0007, Junyu Dong, Feng Gao 0005, Quanquan Mu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Prediction of Sea Surface Temperature Using Long Short-Term MemoryabstractThis letter adopts long short-term memory (LSTM) to predict sea surface temperature (SST), and makes short-term prediction, including one day and three days, and long-term prediction, including weekly mean and monthly mean. The SST prediction problem is formulated as a time series regression problem. The proposed network architecture is composed of two kinds of layers: an LSTM layer and a full-connected dense layer. The LSTM layer is utilized to model the time series relationship. The full-connected layer is utilized to map the output of the LSTM layer to a final prediction. The optimal setting of this architecture is explored by experiments and the accuracy of coastal seas of China is reported to confirm the effectiveness of the proposed method. The prediction accuracy is also tested on the SST anomaly data. In addition, the model's online updated characteristics are presented. Qin Zhang 0008, Junyu Dong, Guoqiang Zhong 0001, Xin Sun 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Modeling Information Diffusion over Social Networks for Temporal Dynamic PredictionabstractModeling the process of information diffusion is a challenging problem. Although numerous attempts have been made in order to solve this problem, very few studies are actually able to simulate and predict temporal dynamics of the diffusion process. In this paper, we propose a novel information diffusion model, namely GT model, which treats the nodes of a network as intelligent and rational agents and then calculates their corresponding payoffs, given different choices to make strategic decisions. By introducing time-related payoffs based on the diffusion data, the proposed GT model can be used to predict whether or not the user's behaviors will occur in a specific time interval. The user's payoff can be divided into two parts: social payoff from the user's social contacts and preference payoff from the user's idiosyncratic preference. We here exploit the global influence of the user and the social influence between any two users to accurately calculate the social payoff. In addition, we develop a new method of presenting social influence that can fully capture the temporal dynamics of social influence. Experimental results from two different datasets, Sina Weibo and Flickr demonstrate the rationality and effectiveness of the proposed prediction method with different evaluation metrics. Shengping Zhang, Xin Sun 0003, Huiyu Zhou 0001, Sheng Li 0003, Xuelong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Ocean Front Detection From Instant Remote Sensing SST ImagesabstractIdentifying fronts manually from satellite images is a tedious and subjective task. Accordingly, edge detection algorithms are introduced for automatic detection of fronts. However, traditional algorithms cannot be applied to cloud-contaminated images, because missing data caused by occasional cloud coverage interferes with front detection. To diminish this risk, this letter proposes a new algorithm for a quick and an accurate detection of fronts from an instant cloud-contaminated sea surface temperature (SST) image, instead of depending on the daily or weekly averaged SST images. This algorithm adopts a data-driven analog interpolation method, which estimates missing values from the historical data of the same region. After reducing the contour between the interpolated data and the original data, an instant front detection algorithm is proposed based on microcanonical multiscale formalism (MMF). The algorithm utilizes MMF to detect singularity exponents (SEs), and then enhances the features detected in a cloud-contaminated region. Finally, a threshold is set to extract fronts from SE. Experimental results on an AVHRR satellite SST image of 12:00 o'clock covering China Coastal waters confirmed the effectiveness of the proposed algorithm. Yuting Yang 0001, Junyu Dong, Xin Sun 0003, Redouane Lguensat, Muwei Jian |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Human fall detection in surveillance video based on PCANet
Shengke Wang, Long Chen 0019, Zixi Zhou, Xin Sun 0003, Junyu Dong |
Multim. Tools Appl. | 4 |
| 2015 | Histograms of locally aggregated oriented gradientsabstractMotivated by the Vector of Locally Aggregated Descriptors (VLAD), we propose a new Histograms of Locally Aggregated Oriented Gradients descriptor (called HLAOG). In the Histograms of Oriented Gradients descriptor (HOG), the zero-order information of the gradients is captured. By contrast, in the HLAOG descriptor we accumulate the differences between gradient orientations and their nearest bin centers, which characterizes the distribution of the gradient orientations in regard to the bin centers. The HLAOG descriptor is demonstrated to be complementary to HOG in the experiments. Then, for setting the weights of the votes on different bins in a better way, we choose Gaussian function as the weighting method and present another new Gaussian Weighted Histograms of Oriented Gradients descriptor (called GWHOG) based on HOG. Evaluations on two public object recognition datasets (Caltech-101 and VOC2007) show that the combination of HOG and HLAOG outperforms HOG and the combination of HLAOG and GWHOG gets the best result. Xiusheng Lu, Shengping Zhang, Hongxun Yao, Xin Sun 0003, Yanhao Zhang 0001 |
ICIP | 4 |
| 2015 | Non-Rigid Object Contour Tracking via a Novel Supervised Level Set ModelabstractWe present a novel approach to non-rigid objects contour tracking in this paper based on a supervised level set model (SLSM). In contrast to most existing trackers that use bounding box to specify the tracked target, the proposed method extracts the accurate contours of the target as tracking output, which achieves better description of the non-rigid objects while reduces background pollution to the target model. Moreover, conventional level set models only emphasize the regional intensity consistency and consider no priors. Differently, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the targets we want to track. Therefore, the SLSM can ensure a more accurate convergence to the exact targets in tracking applications. In particular, we firstly construct the appearance model for the target in an online boosting manner due to its strong discriminative power between the object and the background. Then, the learnt target model is incorporated to model the probabilities of the level set contour by a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, the accurate target region qualifies the samples fed to the boosting procedure as well as the target model prepared for the next time step. We firstly describe the proposed mechanism of two-phase SLSM for single target tracking, then give its generalized multi-phase version for dealing with multi-target tracking cases. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the proposed method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
IEEE Trans. Image Process. | 1 |
| 2014 | Structure-aware multi-object discovery for weakly supervised trackingabstractRecent progress on tracking has focused on designing robust statistical model or proposing effective appearance features to improve precision. This paper addresses another problem, namely the discovery and tracking of generic multi-object which have the similar appearance and motion pattern based on limited human annotations. We present a model-free tracking method that can automatically discover and track multi-object sharing the same spatial and motion structure, and update the structure during the tracking without prior acknowledge. The candidate objects are first selected by a SVM classifier trained on histogram-of-gradient (HOG) features. Then a segment algorithm is exploited to decide the suitable sizes of tracking boxes. The structure constrains are updated in a real-time manner according to the motion measure among the specified object and corresponding candidates. Experimental results reveal significant convenience and remarkable performance of our approach for the task of structure preserving multi-object discovery and tracking. Yuankai Qi, Hongxun Yao, Xiaoshuai Sun, Xin Sun 0003, Yanhao Zhang 0001, Qingming Huang |
ICIP | 4 |
| 2014 | Action recognition based on overcomplete independent components analysis
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Kuanquan Wang, Jun Zhang 0017, Xiusheng Lu, Yanhao Zhang 0001 |
Inf. Sci. | 3 |
| 2014 | A refined particle filter based on determined level set model for robust contour tracking
Xin Sun 0003, Hongxun Yao |
Mach. Vis. Appl. | 1 |
| 2013 | Real-time visual tracking using ℓ2 norm regularization based collaborative representationabstractRecently, sparse representation based visual tracking have been attracting increasing interests. Although reported desired performance, whether the sparse representation constrain is really useful is not clear. In addition, the high computation complexity also limits their usage in real-time applications. In this paper, we proposed a real-time visual tracking framework using ℓ2norm regularization based collaborative representation. Our framework represents any target candidate using a set of target templates and a set of background templates respectively, then combines their reconstruction errors to track the target accurately. By constraining ℓ2norm regularization on the representation coefficients, the coefficients can be solved analytically, which makes the proposed method run in real-time. The experimental results demonstrate that the proposed approach outperforms several state-of-the-art trackers. Xiusheng Lu, Hongxun Yao, Xin Sun 0003, Xuesong Jiang |
ICIP | 3 |
| 2013 | Non-rigid object tracking by adaptive data-driven kernelabstractWe derive an adaptive data-driven kernel in this paper to simultaneously address the kernel scale/orientation selection problem as well as the constant kernel shape in deformable object tracking applications. Level set technique is novelly introduced into the mean shift sample space to implement kernel evolution and update. Since the active contour model is designed to drive the kernel constantly to the direction that maximizes target likelihood, the kernel can adapt to target shape variation simultaneously with the mean shift iterations. Thus, it can give a better estimation bias to produce accurate shift of the mean and successfully avoid performance loss stemmed from pollution of the non-object regions hiding inside the kernel. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang, Mingui Sun |
ICIP | 1 |
| 2013 | An efficient diagnosis system for detection of Parkinson's disease using fuzzy k-nearest neighbor approach
Huiling Chen 0001, Xin-Gang Yu, Xin Sun 0003, Gang Wang 0013 |
Expert Syst. Appl. | 5 |
| 2013 | Robust visual tracking based on online learning sparse representation
Shengping Zhang, Hongxun Yao, Huiyu Zhou 0001, Xin Sun 0003, Shaohui Liu |
Neurocomputing | 4 |
| 2013 | Selection of interdependent genes via dynamic relevance analysis for cancer diagnosis
Xin Sun 0003, Yanheng Liu 0001, Da Wei, Mantao Xu, Huiling Chen 0001, Jiawei Han 0006 |
J. Biomed. Informatics | 1 |
| 2013 | Electrical characterization of integrated passive devices using thin film technology for 3D integrationabstractWith the development of 3D integration technology, microsystems with vertical interconnects are attracting attention from researchers and industry applications. Basic elements of integrated passive devices (IPDs), including inductors, capacitors, and resistors, could dramatically save the footprint of the system, optimize the form factor, and improve the performance of radio frequency (RF) systems. In this paper, IPDs using thin film built-up technology are introduced, and the design and characterization of coplanar waveguides (CPWs), inductors, and capacitors are presented. Xin Sun 0003, Yunhui Zhu, Zhen Hua Liu, Qing-hu Cui, Shenglin Ma, Min Miao, Yufeng Jin |
J. Zhejiang Univ. Sci. C | 1 |
| 2013 | Feature selection using dynamic weights for classification
Xin Sun 0003, Yanheng Liu 0001, Mantao Xu, Huiling Chen 0001, Jiawei Han 0006 |
Knowl. Based Syst. | 1 |
| 2013 | Sparse coding based visual tracking: Review and experimental comparison
Shengping Zhang, Hongxun Yao, Xin Sun 0003, Xiusheng Lu |
Pattern Recognit. | 3 |
| 2012 | Using cooperative game theory to optimize the feature selection problem
Xin Sun 0003, Yanheng Liu 0001, Jianqi Zhu, Xuejie Liu, Huiling Chen 0001 |
Neurocomputing | 1 |
| 2012 | Feature evaluation and selection with cooperative game theory
Xin Sun 0003, Yanheng Liu 0001, Jianqi Zhu, Huiling Chen 0001, Xuejie Liu |
Pattern Recognit. | 1 |
| 2012 | Robust Visual Tracking Using an Effective Appearance Model Based on Sparse CodingabstractIntelligent video surveillance is currently one of the most active research topics in computer vision, especially when facing the explosion of video data captured by a large number of surveillance cameras. As a key step of an intelligent surveillance system, robust visual tracking is very challenging for computer vision. However, it is a basic functionality of the human visual system (HVS). Psychophysical findings have shown that the receptive fields of simple cells in the visual cortex can be characterized as being spatially localized, oriented, and bandpass, and it forms a sparse, distributed representation of natural images. In this article, motivated by these findings, we propose an effective appearance model based on sparse coding and apply it in visual tracking. Specifically, we consider the responses of general basis functions extracted by independent component analysis on a large set of natural image patches as features and model the appearance of the tracked target as the probability distribution of these features. In order to make the tracker more robust to partial occlusion, camouflage environments, pose changes, and illumination changes, we further select features that are related to the target based on an entropy-gain criterion and ignore those that are not. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance between the distributions of the target model and a candidate using Newton-style iterations. The experimental results validate that the proposed method is more robust and effective than three state-of-the-art methods. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | A novel supervised level set method for non-rigid object trackingabstractWe present a novel approach to non-rigid object tracking based on a supervised level set model (SLSM). In contrast with conventional level set models, which emphasize the intensity consistency only and consider no priors, the curve evolution of the proposed SLSM is object-oriented and supervised by the specific knowledge of the target we want to track. Therefore, the SLSM can ensure a more accurate convergence to the target in tracking applications. In particular, we firstly construct the appearance model for the target in an on-line boosting manner due to its strong discriminative power between objects and background. Then the probability of the contour is modeled by considering both the region and edge cues in a Bayesian manner, leading the curve converge to the candidate region with maximum likelihood of being the target. Finally, accurate target region qualifies the samples fed the boosting procedure as well as the target model prepared for the next time step. Positive decrease rate is used to adjust the learning pace over time, enabling tracking to continue under partial and total occlusion. Experimental results on a number of challenging sequences validate the effectiveness of the technique. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
CVPR | 1 |
| 2011 | Contour tracking via on-line discriminative appearance modeling based level setsabstractA novel level set method based on on-line discriminative appearance modeling (DAMLSM) is presented for contour tracking. In contrast with traditional level set models which emphasize the intensity consistent segmentation and consider no priors, the proposed DAMLSM takes the context of tracking into account and use a discriminative patch based target model to guide the curve evolution. By modeling both the region and edge cues in a Bayesian manner, the proposed level set method can lead an accurate convergence to the candidate region with maximum likelihood of being the target. Finally, we update the target model to adapt to the appearance variation, enabling tracking to continue under occlusion. Experiments confirm the robustness and reliability of our method. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
ICIP | 1 |
| 2011 | Robust visual tracking via context objects computingabstractOcclusions are challenging issue for robust visual tracking. In this paper, motivated by the fact that a tracked object is usual- ly embedded into context that provides useful information for estimating the target, we propose a novel tracking algorithm named Tracking with Context Prediction (TCP). The context here includes the neighboring objects and specific parts of tar- get. The proposed method simultaneously track the target and context objects using the existing tracking methods. The positions of the context objects are used to predict the position of the target. Thus, the target can be stably tracked even when it is partially or fully occluded. By computing the probability of each prediction being target, our algorithm allows the drifting of context objects during tracking and do not require predictions from all context objects are correct. Experiments on challenging sequences show significant improvements especially in the case of occlusions and appearance changes. Zhongqian Sun, Hongxun Yao, Shengping Zhang, Xin Sun 0003 |
ICIP | 4 |
| 2011 | Design of a Robot Cloud CenterabstractService-oriented architecture and cloud computing have become the prevalent computing paradigm. In this paradigm, computing resources can be accessed like other utility services available in today's society. In the meantime, robotics applications are joining the trend. More and more robot applications are shifting from manufacture to non-manufacture and service industries. However, for the on-demand supply of the large-scale heterogeneous robots, It is still a problem have not yet been studied, including the fundamental management and efficiency issues in using of these resources. In this paper, we design a framework of "Robot Cloud Center" (RCC) following the general cloud computing paradigm to address the current limitations in capacity and versatility of robotic applications. In this framework, a robot can be provided as a service just like a public utility service so that everyone can access the powerful robotic services easily, efficiently, and cheaply. Based on a given scenario, a robot scheduling algorithm in RCC is proposed to take advantage of the heterogeneous robot resources to meet the end user's requirement with the minimum cost. Zhihui Du, Weiqiang Yang, Yinong Chen 0004, Xin Sun 0003, Xiaoying Wang 0002 |
ISADS | 4 |
| 2011 | The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature
Eric W. T. Ngai, Yong Hu 0002, Y. H. Wong, Xin Sun 0003 |
Decis. Support Syst. | 5 |
| 2010 | Real-Time Service-Oriented Cloud ComputingabstractCloud computing has received significant attention recently. This paper presents real-time issues related to cloud computing, such as multi-tenancy architecture, scheduling, paralleled computing and proposes a framework for real-time service-oriented cloud computing. Specially, we propose a novel real time architecture which solve the new challenges in Cloud Computing. Wei-Tek Tsai, Qihong Shao, Xin Sun 0003, Jay Elston |
SERVICES | 3 |
| 2010 | A refined particle filter method for contour trackingabstractTraditional particle filter which uses simple geometric shapes for representation cannot track objects with complex shape accurately. In this paper, we propose a refined particle filter method for contour tracking based on a binary level set model. In contrast with other previous work, the computational efficiency is greatly improved due to the simple form of the level set function. In addition, we perform curve evolution in the update step to make good use of the observation at current time. Finally, we consider some appearance information as well as the energy function to measure the weight for particles, which can identify the target more accurately. Experiment results on several challenging video sequences have verified the proposed algorithm is efficient and effective in many complicated scenes. Xin Sun 0003, Hongxun Yao, Shengping Zhang |
VCIP | 1 |
| 2010 | Robust object tracking based on sparse representationabstractIn this paper, we propose a novel and robust object tracking algorithm based on sparse representation. Object tracking is formulated as a object recognition problem rather than a traditional search problem. All target candidates are considered as training samples and the target template is represented as a linear combination of all training samples. The combination coefficients are obtained by solving for the minimum l1-norm solution. The final tracking result is the target candidate associated with the non-zero coefficient. Experimental results on two challenging test sequences show that the proposed method is more effective than the widely used mean shift tracker. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
VCIP | 3 |
| 2008 | Teaching Service-Oriented Computing and STEM Topics via Robotic GamesabstractThis paper proposes a new approach to teach the STEM (Science, Technology, Engineering, and Mathematics) knowledge informally via robotic games. In this approach, a robotic playground is built to provide a hands-on programming and playing experience with robots controlled by Service-Oriented Computing (SOC) software, which is based on a new approach that uses reusable services (components) with standard interfaces and platform-independent interoperability. Services in the repository are annotated with STEM knowledge to enforce the required contents. In this way, students can learn computing and STEM in an entertaining manner. Wei-Tek Tsai, Xin Sun 0003, Yinong Chen 0004, Qian Huang 0002, Gary Bitter, Mary White |
ISORC | 2 |