Xia Li 0006

dblp:97/30-6 · DBLP profile ↗
← Back
69ranked-venue papers
2as first author
42since 2021 · last 2026
0000-0002-8043-9966ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 22 since 2021Artificial intelligence and machine learning · 24 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Fast and Effective Video Inpainting via Implicit Motion-Guided Propagation and Sparse Attention
abstract
Video inpainting aims to reconstruct missing or corrupted regions in video frames, with applications in video editing, restoration, and special effects. Current deep video inpainting methods rely on optical flow to guide the propagation of effective features and spatiotemporal attention mechanisms to model relationships between frames. However, as an explicit motion representation, the optical flow extracted offline in preceding steps often suffers from instability and errors during estimation. These errors accumulate during subsequent content hallucination, resulting in artifacts and blurring. Meanwhile, although traditional spatiotemporal attention effectively captures frame relationships, its dense computational nature introduces redundant information, disrupting inpainting tasks and reducing efficiency. To address these issues, we propose an implicit motion-guided approach for efficient video inpainting. Instead of relying on optical flow, our method uses implicit motion in the latent feature space to guide the dual-domain propagation of images and features end-to-end, avoiding error accumulation from the independent optical flow estimation process. Additionally, we introduce a self-correcting module that enables feedback between image and feature propagation, reducing errors during propagation. Furthermore, we design an adaptive sparse video attention mechanism to focus on highly relevant regions, minimizing the impact of irrelevant information. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches both qualitatively and quantitatively, while also delivering superior efficiency.
Yuanman Li, Bin Li 0011, Yanshan Li, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.6
2025 DAFN: A Robust Dual-Modal Framework for Online Video Source Platform Identification
abstract
Identifying the source platform of an online video is a critical yet challenging task in digital forensics, complicated by proprietary transcoding and adversarial post-processing. The preceding forensic methods are limited to a single modality, analyzing either fragile container metadata or easily distorted spatiotemporal artifacts severely compromise their robustness. This paper pioneers a multi-modal framework that deeply integrates both information domains. We contend that naive fusion is insufficient as it fails to address cross-modal ambiguity-instances where modalities provide conflicting evidence. To resolve this, we propose DAFN, a novel Dual-modality Ambiguity-Aware Fingerprinting Network which extract features from both container structure and video content. At its core, DAFN introduces an adaptive fusion mechanism guided by a measure of cross-modal ambiguity. This mechanism, which incorporates a variational module to quantify the discrepancy between modalities, enables the model to intelligently arbitrate between evidence sources. We further contribute CNSNVD, a large-scale dataset with nine major platforms and six post-processing types. Extensive experiments show that DAFN significantly outperforms existing baselines, establishing a new state-of-the-art by effectively resolving modal ambiguity to achieve superior accuracy and resilience.
Yulong Zheng, Xia Li 0006, Yuanman Li
CloudCom3
2025 Flow-Aware Dynamic Fusion for Video Inpainting Detection
abstract
The malicious misuse of deep learning-based video inpainting techniques poses significant security risks, highlighting the critical importance of accurately detecting inpainted regions in digital video content. However, existing methods suffer from limited accuracy in complex motion scenarios, and their cross-modal feature fusion efficiency is low due to simplistic integration strategies. To address these challenges, we propose FAD-Net, an end-to-end dual-branch framework that performs video inpainting localization by jointly modeling optical flow and RGB modalities to complement each other's capabilities. The optical flow branch employs a flow consistency error constraint and flow anomaly awareness module to mitigate the impact of inaccurate optical flow estimation, while the parallel RGB branch utilizes an encoder for spatial texture feature extraction and cross-frame attention for long-term temporal modeling. A bidirectional dynamic fusion module then adaptively integrates complementary features using motion-aware weights. Experimental results demonstrate that FAD-Net outperforms existing methods in accurately localizing the inpainted regions, particularly exhibiting excellent performance in complex dynamic scenarios and with unknown inpainting types, enabling reliable forensic analysis of tampered videos.
Wenze Zheng, Xia Li 0006, Yuanman Li
CloudCom3
2025 Frequency-Enhanced Multi-Scale Progressive Detection for Document Image Manipulation
abstract
Images playa vital role in modern communication, cultural exchange, and information dissemination. Document images, as a digital form of textual content, are widely used in government, finance, and e-commerce scenarios. However, tampered with document images has become increasingly common and is often exploited for identity fraud and financial scams, posing serious threats to platform credibility and public interests. Compared to natural images, document forgeries typically involve small, visually inconspicuous regions against uniform backgrounds, making RGB-based detection methods less effective. To address these challenges, we propose MFPD-Net (Multi-scale Frequency Progressive Detector), which integrates a Transformer-based frequency feature enhancement module combined with an adaptive feature fusion strategy. Additionally, a Progressive multi-scale feedback decoder is designed to refine mask generation and improve localization accuracy. Extensive experiments on multiple document tampered detection tasks demonstrate that our method outperforms existing approaches in both accuracy and robustness, showing strong practicality and generalization capabilities.
Qiyue Zhong, Xia Li 0006, Yuanman Li
CloudCom2
2025 Semantic-Guided Residual Learning for the Quality Assessment of Enhanced Images
abstract
Image enhancement algorithms are essential for improving visual quality but often introduce new distortions, highlighting the need for reliable image quality assessment (IQA). However, existing IQA methods typically focus on semantic information or distortion-prone regions while ignoring their interactions, resulting in unsatisfactory performance. To address this issue, we propose to integrate semantic information with edge residual learning and design a semantic-guided residual learning IQA framework tailored for enhanced images across diverse scenarios. Specifically, the proposed framework utilizes a covariance-guided encoder to extract semantic information, which is then enhanced using a semantic refinement module. The refined semantic information is subsequently utilized to guide edge residual feature learning in the decoder. Extensive experiments on multiple tasks such as deraining, dehazing, and low-light enhancement demonstrate that our method outperforms state-of-the-art approaches.
Shishun Tian, Zhiwei Lan, Ting Su 0004, Xia Li 0006, Lu Zhang 0037
ICME5
2025 Dynamic vision-based machine vibration sensing and fault diagnosis with signal alignment and feature clustering
Ruiyi Guang, Xia Li 0006, Yaguo Lei, Bin Yang 0014, Naipeng Li
Eng. Appl. Artif. Intell.2
2025 Class-discriminative domain generalization for semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Rong You, Wenbin Zou, Xia Li 0006
Image Vis. Comput.7
2025 A memory-augmented multi-task collaborative framework for unsupervised traffic anomaly detection in driving videos
Rongqin Liang, Yuanman Li, Yingxin Yi, Jiantao Zhou 0001, Xia Li 0006
Pattern Recognit.5
2025 Dual-Stream Image Sharing Chain Detection via Dynamic Information Compensation
abstract
Image Sharing Chain Detection (ISCD) aims to reconstruct the complete trajectory of an image's dissemination across social platforms and is an important task in multimedia forensics. Current methods using DCT histograms are insufficient in uncovering platform compression traces and exhibit limitations in detecting weak trace platforms. In this letter, we propose an innovative dual-stream ISCD framework via dynamic information compensation. This framework integrates features from both the frequency domain and the residual domain to extract compression characteristics. Unlike existing methods, we employ binary stereo DCT in the frequency domain to focus on the spatiality of compression operations. Additionally, we design a dynamic information compensation mechanism to enhance platform traces by storing compensation fingerprints of the sharing chains. Furthermore, we develop a new dataset, F-4OSN-SC, encompassing 4 platforms to simulate more realistic social networking scenarios. Experimental results demonstrate that our model outperforms existing methods across multiple datasets.
Xinyi Su, Yuanman Li, Yulong Zheng, Xia Li 0006
IEEE Signal Process. Lett.4
2025 Image Copy-Move Forgery Detection via Deep PatchMatch and Pairwise Ranking Learning
abstract
Recent advances in deep learning algorithms have shown impressive progress in image copy-move forgery detection (CMFD). However, these algorithms lack generalizability in practical scenarios where the copied regions are not present in the training images, or the cloned regions are part of the background. Additionally, these algorithms utilize convolution operations to distinguish source and target regions, leading to unsatisfactory results when the target regions blend well with the background. To address these limitations, this study proposes a novel end-to-end CMFD framework that integrates the strengths of conventional and deep learning methods. Specifically, the study develops a deep cross-scale PatchMatch (PM) method that is customized for CMFD to locate copy-move regions. Unlike existing deep models, our approach utilizes features extracted from high-resolution scales to seek explicit and reliable point-to-point matching between source and target regions. Furthermore, we propose a novel pairwise rank learning framework to separate source and target regions. By leveraging the strong prior of point-to-point matches, the framework can identify subtle differences and effectively discriminate between source and target regions, even when the target regions blend well with the background. Our framework is fully differentiable and can be trained end-to-end. Comprehensive experimental results highlight the remarkable generalizability of our scheme across various copy-move scenarios, significantly outperforming existing methods.
Yuanman Li, Yingjie He 0003, Changsheng Chen 0001, Li Dong 0006, Bin Li 0011, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Image Process.7
2025 Class-Balanced Sampling and Discriminative Stylization for Domain Generalization Semantic Segmentation
abstract
Existing domain generalization semantic segmentation (DGSS) methods have achieved remarkable performance on unseen domains by generating stylized images to increase the diversity of training data. However, since the training data is usually class-imbalanced, uniform style randomization is unable to generate diverse minority classes. This means that models may overfit to the minority classes, resulting in suboptimal performance on the minority classes. In addition, the image-level style randomization may also corrupt the class-discriminative regions of objects, leading to a loss of the class-discriminative representation. To address these issues, a novel class-balanced sampling and discriminative stylization (CSDS) approach is proposed for DGSS. Specifically, first, a pixel-level class-balanced sampling (PCS) strategy is proposed to adaptively sample patches of the minority classes from the source domain images and paste the sampled patches on the input images. Unlike existing class sampling strategies that fix the minority classes, the PCS strategy dynamically determines the minority classes by estimating the class distribution after each sampling. Then, a class-discriminative style randomization (CSR) strategy is proposed to increase the style diversity of the sampled patches while preserving the class-discriminative regions. Finally, since the pasting positions of the sampled patches are uncertain, which may confuse the semantic relations between the classes, a semantic consistency constraint is proposed to ensure the learning of reliable semantic relations. Extensive experiments demonstrate that the proposed approach achieves superior performance compared to existing DGSS methods on multiple benchmarks. The source code has been released onhttps://github.com/seabearlmx/CSDS.
Muxin Liao, Shishun Tian, Binbin Wei, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.6
2025 Dual Residual-Guided Interactive Learning for the Quality Assessment of Enhanced Images
abstract
Image enhancement algorithms can facilitate computer vision tasks in real applications. However, various distortions may also be introduced by image enhancement algorithms. Therefore, the image quality assessment (IQA) plays a crucial role in accurately evaluating enhanced images to provide dependable feedback. Current enhanced IQA methods are mainly designed for single specific scenarios, resulting in limited performance in other scenarios. Besides, no-reference methods predict quality utilizing enhanced images alone, which ignores the existing degraded images that contain valuable information, are not reliable enough. In this work, we propose a degraded-reference image quality assessment method based on dual residual-guided interactive learning (DRGQA) for the enhanced images in multiple scenarios. Specifically, a global and local feature collaboration module (GLCM) is proposed to imitate the perception of observers to capture comprehensive quality-aware features by using convolutional neural networks (CNN) and Transformers in an interactive manner. Then, we investigate the structure damage and color shift distortions that commonly occur in the enhanced images and propose a dual residual-guided module (DRGM) to make the model concentrate on the distorted regions that are sensitive to human visual system (HVS). Furthermore, a distortion-aware feature enhancement module (DEM) is proposed to improve the representation abilities of features in deeper networks. Extensive experimental results demonstrate that our proposed DRGQA achieves superior performance with lower computational complexity compared to the state-of-the-art IQA methods.
Shishun Tian, Tiantian Zeng, Wenbin Zou, Xia Li 0006
IEEE Trans. Multim.5
2024 A Unified Environmental Network for Pedestrian Trajectory Prediction
abstract
Accurately predicting pedestrian movements in complex environments is challenging due to social interactions, scene constraints, and pedestrians' multimodal behaviors. Sequential models like long short-term memory fail to effectively integrate scene features to make predicted trajectories comply with scene constraints due to disparate feature modalities of scene and trajectory. Though existing convolution neural network (CNN) models can extract scene features, they are ineffective in mapping these features into scene constraints for pedestrians and struggle to model pedestrian interactions due to the loss of target pedestrian information. To address these issues, we propose a unified environmental network based on CNN for pedestrian trajectory prediction. We introduce a polar-based method to reflect the distance and direction relationship between any position in the environment and the target pedestrian. This enables us to simultaneously model scene constraints and pedestrian social interactions in the form of feature maps. Additionally, we capture essential local features in the feature map, characterizing potential multimodal movements of pedestrians at each time step to prevent redundant predicted trajectories. We verify the performance of our proposed model on four trajectory prediction datasets, encompassing both short-term and long-term predictions. The experimental results demonstrate the superiority of our approach over existing methods.
Yuanman Li, Wei Wang 0077, Jiantao Zhou 0001, Xia Li 0006
AAAI5
2024 PDA: Progressive Domain Adaptation for Semantic Segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Knowl. Based Syst.6
2024 Considering representation diversity and prediction consistency for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Knowl. Based Syst.6
2024 MI-RPN: Integrating multi-modalities and multi-scales information for region proposal
Shishun Tian, Wenbin Zou, Xia Li 0006
Multim. Tools Appl.4
2024 Transformer-Based Image Inpainting Detection via Label Decoupling and Constrained Adversarial Training
abstract
Image inpainting based on generative adversarial networks (GANs) has achieved great success in producing visually plausible images and plays an important role in many real tasks. However, the techniques of image inpainting might also be maliciously used, e.g., altering or removing interesting objects to report fake news. Despite the promising performance of recently developed inpainting detection algorithms, they are built on convolutional neural networks (CNNs) with limited receptive fields. Consequently, they fail to fully capture the disparity between the inpainted regions and untouched regions and thus are ineffective in obtaining fine-grained detection results. In this work, we develop a new image inpainting detection approach. First, we propose a locally enhanced transformer architecture tailored for image inpainting detection. Unlike previous CNN-based methods, our approach leverages both the short-range and long-range dependencies of pixels, enabling the learning of diverse statistical behaviors of inpainted and untouched regions. Second, to mitigate the distraction caused by near-edge pixels with a mixed nature during training, we propose decoupling the label into a body map and a soft-edge map, and then a cross-modality attention module is designed to propagate their information interactively. It demonstrates that our decoupling strategy outperforms the conventional edge supervision in enhancing detection accuracy. Finally, we devise a constrained adversarial training methodology in consideration of the confrontational generation procedure of deep image inpainting methods. It shows that our constrained adversarial training further enhances the detection performance by adaptively introducing interference noise in the inpainted regions. Extensive experiments validate the superiority of our scheme compared to existing CNN-based methods, showcasing its desirable detection generalizability for both deep inpainting and traditional inpainting algorithms.
Yuanman Li, Liangpei Hu, Li Dong 0006, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.7
2024 Text-Driven Traffic Anomaly Detection With Temporal High-Frequency Modeling in Driving Videos
abstract
Traffic anomaly detection (TAD) in driving videos is critical for ensuring the safety of autonomous driving and advanced driver assistance systems. Previous single-stage TAD methods primarily rely on frame prediction, making them vulnerable to interference from dynamic backgrounds induced by the rapid movement of the dashboard camera. While two-stage TAD methods appear to be a natural solution to mitigate such interference by pre-extracting background-independent features (such as bounding boxes and optical flow) using perceptual algorithms, they are susceptible to the performance of first-stage perceptual algorithms and may result in error propagation. In this paper, we introduce TTHF, a novel single-stage method aligning video clips with text prompts, offering a new perspective on traffic anomaly detection. Unlike previous approaches, the supervised signal of our method is derived from languages rather than orthogonal one-hot vectors, providing a more comprehensive representation. Further, concerning visual representation, we propose to model the high frequency of driving videos in the temporal domain. This modeling captures the dynamic changes of driving scenes, enhances the perception of driving behavior, and significantly improves the detection of traffic anomalies. In addition, to better perceive various types of traffic anomalies, we carefully design an attentive anomaly focusing mechanism that visually and linguistically guides the model to adaptively focus on the visual context of interest, thereby facilitating the detection of traffic anomalies. It is shown that our proposed TTHF achieves promising performance, outperforming state-of-the-art competitors by +5.4% AUC on the DoTA dataset and achieving high generalization on the DADA dataset.
Rongqin Liang, Yuanman Li, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.4
2024 Preserving Label-Related Domain-Specific Information for Cross-Domain Semantic Segmentation
abstract
Unsupervised domain adaptation semantic segmentation (UDASS) methods aim to learn domain-invariant information for alleviating the distribution shift problem between the source and target domains. However, ignoring the learning of domain-specific information that is label-related may limit the class discriminability on the target domain. We argue that a good representation for the UDASS task not only contains domain-invariant information but also preserves label-related domain-specific information. In this paper, a novel frequency spectrum domain adaptation approach via meta-learning (ML-FSDA) is proposed to achieve this goal for improving the class discriminability and generalization ability. ML-FSDA contains a frequency-spectrum meta-learning framework (FMF) and a class-aware domain-specific memory bank (CDMB). Specifically, first, inspired by the observation that the high-frequency component is consistent across different domains while the low-frequency component is much more domain-specific, the FMF aims to respectively learn label-related domain-specific and domain-invariant information from low-frequency and high-frequency images in a unified framework via the meta-learning strategy. Second, the CDMB is designed to preserve the label-related domain-specific information of each class in an external memory bank while the CDMB is updated in every iteration of the meta-training stage. Finally, the CDMB is utilized to embed the label-related domain-specific information into domain-invariant information at the class level during the meta-testing stage to enhance the class discriminability on the target domain. Extensive experiments demonstrate the effectiveness of ML-FSDA on two challenging cross-domain semantic segmentation benchmarks. Notably, for the GTA5 to Cityscapes task and the SYNTHIA to Cityscapes task, the proposed ML-FSDA achieves superior performance with 77.3% mIoU and 68.8% mIoU, respectively. The source code is released at https://github.com/seabearlmx/FSL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.6
2024 Calibration-Based Multi-Prototype Contrastive Learning for Domain Generalization Semantic Segmentation in Traffic Scenes
abstract
Prototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features for domain generalization semantic segmentation. These methods assume that the prototypes in different domains are invariant. However, the prototypes in different domains have discrepancies as well. First, the prototypes of the same class in different domains may be different. Second, the prototypes of different classes may be similar. To address these issues, a calibration-based multi-prototype contrastive learning (CMPCL) approach is proposed, which contains an uncertainty-guided multi-prototype contrastive learning (UMPCL) and a hard-weighted multi-prototype contrastive learning (HMPCL). Specifically, the UMPCL uses an uncertainty probability matrix, derived from element-wise discrepancies between the prototypes of the same class, to calibrate the weights of prototypes for alleviating the discrepancy between the prototypes of the same class in different domains. The HMPCL uses a hard-weighted matrix that is generated by the similarity between the prototypes of different classes, to calibrate the weights of the hard-aligned prototypes for alleviating the issue of similar prototypes between different classes, with hard-aligned prototypes referring to those exhibiting such similarity. Furthermore, since the learned class-wise domain-invariant features may overfit the prototype in the source domain, multi-prototype contrastive learning is used in the UMPCL and HMPCL to avoid this risk. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on multiple benchmarks of domain generalization semantic segmentation. The source code has been released onhttps://github.com/seabearlmx/CMPCL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.6
2024 "Where Does the Devil Lie?": Multimodal Multitask Collaborative Revision Network for Trusted Road Segmentation
abstract
Road segmentation is an essential component of navigation systems. Although recent advancements in road segmentation, the occurrence of failure segmentations remains inevitable. For safety-critical tasks, e.g., navigation, knowing when and where road segmentation fails is crucial. In this paper, we propose a novel trusted road segmentation architecture, namely Multimodal Multitask Collaborative Revision Network (M2CRN), to improve the trust of road segmentation. Our approach incorporates two strategies to predict and rectify segmentation errors. Firstly, a joint learning framework is devised to generate road segmentation results while estimating failure segmentation masks. Secondly, the road segmentation branch is equipped with an Uncertainty-Aware Revision Module (UARM), which eliminates the error in road segmentation. Additionally, we suppress the response of error regions in the road segmentation branch with an innovative design, called Adaptive Soft Error Suppression (ASES). To validate our methods, extensive experiments are conducted on three benchmark road segmentation datasets. The results demonstrate significant performance improvements with a real-time inference speed of 33.3 FPS, reaffirming the soundness of our revision model.
Guoguang Hua, Dalian Zheng, Shishun Tian, Wenbin Zou, Shenglan Liu 0001, Xia Li 0006
IEEE Trans. Multim.6
2024 STGlow: A Flow-Based Generative Framework With Dual-Graphormer for Pedestrian Trajectory Prediction
abstract
The pedestrian trajectory prediction task is an essential component of intelligent systems. Its applications include but are not limited to autonomous driving, robot navigation, and anomaly detection of monitoring systems. Due to the diversity of motion behaviors and the complex social interactions among pedestrians, accurately forecasting their future trajectory is challenging. Existing approaches commonly adopt generative adversarial networks (GANs) or conditional variational autoencoders (CVAEs) to generate diverse trajectories. However, GAN-based methods do not directly model data in a latent space, which may make them fail to have full support over the underlying data distribution. CVAE-based methods optimize a lower bound on the log-likelihood of observations, which may cause the learned distribution to deviate from the underlying distribution. The above limitations make existing approaches often generate highly biased or inaccurate trajectories. In this article, we propose a novel generative flow-based framework with a dual-graphormer for pedestrian trajectory prediction (STGlow). Different from previous approaches, our method can more precisely model the underlying data distribution by optimizing the exact log-likelihood of motion behaviors. Besides, our method has clear physical meanings for simulating the evolution of human motion behaviors. The forward process of the flow gradually degrades complex motion behavior into simple behavior, while its reverse process represents the evolution of simple behavior into complex motion behavior. Furthermore, we introduce a dual-graphormer combined with the graph structure to more adequately model the temporal dependencies and the mutual spatial interactions. Experimental results on several benchmarks demonstrate that our method achieves much better performance compared to previous state-of-the-art approaches.
Rongqin Liang, Yuanman Li, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Neural Networks Learn. Syst.4
2023 Image Sharing Chain Detection VIA Sequence-To-Sequence Model
abstract
Image sharing chain detection aims to recover the sharing history of an image downloaded from online social networks (OSNs), including the ever-shared OSNs and their orders, which is an important task in the multimedia forensics community. Most of the existing algorithms directly treat the sharing chain detection as a classification problem by simply assigning a unique label to each sharing chain. Such a strategy though seems straightforward, it ignores the inherent properties of the sharing chain which can be regarded as a time sequence that carries the sharing history of an online image. In this paper, we suggest a new sharing chain detection framework via Sequence-to-Sequence (Seq2Seq) model. Different from previous classification based approaches, our model detects the sharing chain of online image progressively via a decoder. This progressive manner can fully utilize the decoded chain, which is embedded into a series of learned representations. Experimental results show that our method can detect sharing chains involving up to three OSNs, and exhibits much better performance than conventional ones.
Jiaxiang You, Yuanman Li, Rongqin Liang, Yuxuan Tan, Jiantao Zhou 0001, Xia Li 0006
ICASSP6
2023 TRG-DQA: Texture Residual-Guided Dehazed Image Quality Assessment
abstract
Image dehazing algorithms have emerged to solve the visual impairment caused by haze. It is important to establish dehazed image quality assessment (DQA) methods that can accurately evaluate the dehazed image quality and the performance of dehazing algorithms. However, classical image quality assessment (IQA) and most hand-crafted feature based DQA methods may not be able to adequately measure complex distortions of dehazed images. To address this issue, this paper proposes a Texture Residual-Guided Dehazed image Quality Assessment (TRG-DQA) method. Specifically, we first introduce a global and local feature extraction module employing a combination of the Transformer and convolutional neural networks (CNN) for extracting the comprehensive features. Considering that texture residual maps represent haze density and artifact distortion information, we propose a residual-guided module to guide the model for efficient learning. Additionally, to mitigate the information loss issue that occurs in deeper networks, a distortion-aware feature enhancement module is proposed. Extensive experiments on six DQA databases demonstrate the proposed TRG-DQA achieves superior performance among all the state-of-the-art methods.
Tiantian Zeng, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Shishun Tian
ICIP4
2023 Image Copy-Move Forgery Detection via Deep Cross-Scale PatchMatch
abstract
The recently developed deep algorithms achieve promising progress in the field of image copy-move forgery detection (CMFD). However, they have limited generalizability in some practical scenarios, where the copy-move objects may not appear in the training images or cloned regions are from the background. To address the above issues, in this work, we propose a novel end-to-end CMFD framework by integrating merits from both conventional and deep methods. Specifically, we design a deep cross-scale patchmatch method tailored for CMFD to localize copy-move regions. In contrast to existing deep models, our scheme aims to seek explicit and reliable point-to-point matching between source and target regions using features extracted from high-resolution scales. Further, we develop a manipulation region location branch for source/target separation. The proposed CMFD framework is completely differentiable and can be trained in an end-to-end manner. Extensive experimental results demonstrate the high generalizability of our method to different copy-move contents, and the proposed scheme achieves significantly better performance than existing approaches.
Yingjie He 0003, Yuanman Li, Changsheng Chen 0001, Xia Li 0006
ICME4
2023 Multiple degraded image restoration via degradation history estimation
abstract
Image restoration is a fundamental task in low-level computer vision. Most existing algorithms assume that the input image has a single known degradation type. In reality, images usually contain multiple degradations, making the restoration challenging. Though recent works restore the multiple degraded images, they assume that the degradation history is known. Obviously, such an ideal assumption often does not hold in real applications. This work proposes a novel restoration framework for multiple degraded images via degradation history estimation. Specifically, we first develop a sequential model to estimate the degradation history, including both the degradation operation chain and the corresponding parameters. By resorting to designed self-attention and cross-attention mechanisms, our method can effectively model the correlation of the input image, degradation operation chain, and parameters. Then, we apply our estimation framework for the multiple degraded image restoration, without requiring the degradation history. Experiment results demonstrate much better performance than existing approaches.
Minhua Liu, Yuanman Li, Rongqin Liang, Jiaxiang You, Xia Li 0006
ICME5
2023 Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic Segmentation
abstract
Prototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features recently. These methods are based on the assumption that the prototypes, which are represented as the central value of the same class in a certain domain, are domain-invariant. Since the prototypes of different domains have discrepancies as well, the class-wise domain-invariant features learned from the source domain by PCL need to be aligned with the prototypes of other domains simultaneously. However, the prototypes of the same class in different domains may be different while the prototypes of different classes may be similar, which may affect the learning of class-wise domain-invariant features. Based on these observations, a calibration-based dual prototypical contrastive learning (CDPCL) approach is proposed to reduce the domain discrepancy between the learned class-wise features and the prototypes of different domains for domain generalization semantic segmentation. It contains an uncertainty-guided PCL (UPCL) and a hard-weighted PCL (HPCL). Since the domain discrepancies of the prototypes of different classes may be different, we propose an uncertainty probability matrix to represent the domain discrepancies of the prototypes of all the classes. The UPCL estimates the uncertainty probability matrix to calibrate the weights of the prototypes during the PCL. Moreover, considering that the prototypes of different classes may be similar in some circumstances, which means these prototypes are hard-aligned, the HPCL is proposed to generate a hard-weighted matrix to calibrate the weights of the hard-aligned prototypes during the PCL. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on domain generalization segmentation tasks. The source code will be released at https://github.com/seabearlmx/CDPCL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
ACM Multimedia6
2023 Multi-scale Target-Aware Framework for Constrained Splicing Detection and Localization
abstract
Constrained image splicing detection and localization (CISDL) is a fundamental task of multimedia forensics, which detects splicing operation between two suspected images and localizes the spliced region on both images. Recent works regard it as a deep matching problem and have made significant progress. However, existing frameworks typically perform feature extraction and correlation matching as separate processes, which may hinder the model's ability to learn discriminative features for matching and can be susceptible to interference from ambiguous background pixels. In this work, we propose a multi-scale target-aware framework to couple feature extraction and correlation matching in a unified pipeline. In contrast to previous methods, we design a target-aware attention mechanism that jointly learns features and performs correlation matching between the probe and donor images. Our approach can effectively promote the collaborative learning of related patches, and perform mutual promotion of feature learning and correlation matching. Additionally, in order to handle scale transformations, we introduce a multi-scale projection method, which can be readily integrated into our target-aware framework that enables the attention process to be conducted between tokens containing information of varying scales. Our experiments demonstrate that our model, which uses a unified pipeline, outperforms state-of-the-art methods on several benchmark datasets and is robust against scale transformations.
Yuxuan Tan, Yuanman Li, Limin Zeng, Jiaxiong Ye, Wei Wang 0077, Xia Li 0006
ACM Multimedia6
2023 Domain-invariant information aggregation for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Neurocomputing6
2023 Image Operation Chain Detection with Machine Translation Framework
abstract
The aim of operation chain detection for a given manipulated image is to reveal the operations involved and the order in which they were applied, which is significant for image processing and multimedia forensics. Currently,allexisting approaches simply treat image operation chain detection as a classification problem and consider only chains of at most two operations. Considering the complex interplay between operations and the exponentially increasing solution space, detecting longer operation chains is extremely challenging. To address this issue, in this work, we devise a new methodology for image operation chain detection. Different from existing approaches based on classification modeling, we strategically conduct operation chain detection within a machine translation framework. Specifically, the chain in our work is modeled as a sentence in a target language, with each possible operation represented by a word in that language. When executing chain detection, we propose first transforming the input image into a sentence in a latent source language from the learned deep features. Then, we propose translating the latent language into the target language within a machine translation framework and finally decoding all operations, arranged in order. Besides, a chain inversion strategy and a bi-directional modeling mechanism are developed to improve the detection performance. We further design a weighted cross-entropy loss to alleviate the problems presented by imbalance among chain lengths and chain categories. Our method can detect operation chains containing up to seven operations and obtains very promising results in various scenarios for the detection of both short and long chains.
Yuanman Li, Jiaxiang You, Jiantao Zhou 0001, Wei Wang 0077, Xin Liao 0001, Xia Li 0006
IEEE Trans. Multim.6
2023 Expensive Multiobjective Optimization Based on Information Transfer Surrogate
abstract
Objective value estimation based on computationally efficient surrogate models is widely used to reduce the computational cost in solving expensive multiobjective optimization problems (MOPs). However, due to the scarcity of training data and the lack of data sharing between training tasks in a surrogate-based system, the estimation effectiveness of the surrogate models might not be satisfactory. In this study, we present a novel surrogate methodology based on information transfer to deal with this problem. Particularly, in the proposed framework, the objectives of an MOP that may have little apparent similarity or correlation are linearly mapped to a number of related tasks. Afterward, the related tasks are used to train a multitask Gaussian process (MTGP). MTGP expands the training data leading to more confident learning of the parameters of the model. The predicted values of the objective functions can be obtained by a reverse mapping from the learned MTGP model. In this way, the computational burden of the expensive objective functions of an MOP can be substantially reduced while maintaining good estimation accuracy. MTGP facilitates mutual information transfer across tasks, avoids learning from scratch for new tasks, and captures the underlying structural information between tasks. The proposed surrogate approach is merged into MOEA/D to address MOPs. Experimental tests under various scenarios indicate that the resultant algorithm outperforms other state-of-the-art surrogate-based multiobjective optimization algorithms.
Jianping Luo, YongFei Dong, Zexuan Zhu 0001, Wenming Cao 0001, Xia Li 0006
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Synchronous Bi-directional Pedestrian Trajectory Prediction with Error Compensation
Ce Xie, Yuanman Li, Rongqin Liang, Li Dong 0006, Xia Li 0006
ACCV (6)5
2022 Exploring more concentrated and consistent activation regions for cross-domain semantic segmentation
Muxin Liao, Guoguang Hua, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006
Neurocomputing6
2022 RGB-D Gate-guided edge distillation for indoor semantic segmentation
Wenbin Zou, Yingqing Peng, Shishun Tian, Xia Li 0006
Multim. Tools Appl.5
2022 Robust Matrix Factorization via Minimum Weighted Error Entropy Criterion
abstract
Learning the intrinsic low-dimensional subspace from high-dimensional data is a key step for many social systems of artificial intelligence. In practical scenarios, the observed data are usually corrupted by many types of noise, which brings a great challenge for social systems to analyze data. As a commonly utilized subspace learning technique, robust low-rank matrix factorization (LRMF) focuses on recovering the underlying subspaces in a noisy environment. However, most of the existing approaches simply assume that the noise contaminating the data is independent identically distributed (i.i.d.), such as Gaussian and Laplacian noises. This assumption, though greatly simplifies the underlying learning problem, may not hold for more complex non-i.i.d. noise widely existed in social systems. In this work, we suggest a robust LRMF approach to deal with various types of noise in a unified manner. Different from traditional algorithms, noise in our framework is modeled using an independent and piecewise identically distributed (i.p.i.d.) source, which employs a collection of distributions, instead of a single one to characterize the statistical behavior of the underlying noise. Assisted by the generic noise model, we then design a robust LRMF algorithm under the information-theoretic learning (ITL) framework through a new minimization criterion. By adopting the half-quadratic optimization paradigm, we further deliver an optimization strategy for our proposed method. Experimental results on both synthetic and real data are provided to demonstrate the superiority of our proposed scheme.
Yuanman Li, Jiantao Zhou 0001, Junyang Chen 0001, Jinyu Tian 0001, Li Dong 0006, Xia Li 0006
IEEE Trans. Comput. Soc. Syst.6
2022 Novel Multitask Conditional Neural-Network Surrogate Models for Expensive Optimization
abstract
Multiple-related tasks can be learned simultaneously by sharing information among tasks to avoid tabula rasa learning and to improve performance in the no transfer case (i.e., when each task learns in isolation). This study investigates multitask learning with conditional neural process (CNP) networks and proposes two multitask learning network models on the basis of CNPs, namely, the one-to-many multitask CNP (OMc-MTCNP) and the many-to-many MTCNP (MMc-MTCNP). Compared with existing multitask models, the proposed models add an extensible correlation learning layer to learn the correlation among tasks. Moreover, the proposed multitask CNP (MTCNP) networks are regarded as surrogate models and applied to a Bayesian optimization framework to replace the Gaussian process (GP) to avoid the complex covariance calculation. The proposed Bayesian optimization framework simultaneously infers multiple tasks by utilizing the possible dependencies among them to share knowledge across tasks. The proposed surrogate models augment the observed dataset with a number of related tasks to estimate model parameters confidently. The experimental studies under several scenarios indicate that the proposed algorithms are competitive in performance compared with GP-, single-task-, and other multitask model-based Bayesian optimization methods.
Jianping Luo, Xia Li 0006, Qingfu Zhang 0001
IEEE Trans. Cybern.3
2022 Trajectory Forecasting Based on Prior-Aware Directed Graph Convolutional Neural Network
abstract
Predicting the motion trajectories of moving agents in complex traffic scenes, such as crossroads and roundabouts, plays an important role in cooperative intelligent transportation systems. Nevertheless, accurately forecasting the motion behavior in a dynamic scenario is challenging due to the complex cooperative interactions between moving agents. Graph Convolutional Neural Network has recently been employed to deal with the cooperative interactions between agents. Despite the promising performance of resulting trajectory prediction algorithms, many existing graph-based approaches model interactions with an undirected graph, where the strength of influence between agents is assumed to be symmetric. However, such an assumption often does not hold in reality. For example, in pedestrian or vehicle interaction modeling, the moving behavior of a pedestrian or vehicle is highly affected by the ones ahead, while the ones ahead usually pay less attention to the ones behind. To fully exploit the asymmetric attributes of the cooperative interactions in intelligent transportation systems, in this work, we present a directed graph convolutional neural network for multiple agents trajectory prediction. First, we propose three directed graph topologies, i.e., view graph, direction graph, and rate graph, by encoding different prior knowledge of a cooperative scenario, which endows the capability of our framework to effectively characterize the asymmetric influence between agents. Then, a fusion mechanism is devised to jointly exploit the asymmetric mutual relationships embedded in constructed graphs. Furthermore, a loss function based on Cauchy distribution is designed to generate multimodal trajectories. Experimental results on complex traffic scenes demonstrate the superior performance of our proposed model when compared with existing approaches.
Jie Du 0001, Yuanman Li, Xia Li 0006, Rongqin Liang, Zhongyun Hua, Jiantao Zhou 0001
IEEE Trans. Intell. Transp. Syst.4
2021 Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-Supervision
abstract
Predicting human motion behavior in a crowd is important for many applications, ranging from the natural navigation of autonomous vehicles to intelligent security systems of video surveillance. All the previous works model and predict the trajectory with a single resolution, which is relatively ineffective and difficult to simultaneously exploit the long-range information (e.g., the destination of the trajectory), and the short-range information (e.g., the walking direction and speed at a certain time) of the motion behavior. In this paper, we propose a temporal pyramid network for pedestrian trajectory prediction through a squeeze modulation and a dilation modulation. Our hierarchical framework builds a feature pyramid with increasingly richer temporal information from top to bottom, which can better capture the motion behavior at various tempos. Furthermore, we propose a coarse-to-fine fusion strategy with multi-supervision. By progressively merging the top coarse features of global context to the bottom fine features of rich local context, our method can fully exploit both the long-range and short-range information of the trajectory. Experimental results on two benchmarks demonstrate the superiority of our method. Our code and models will be available upon acceptance.
Rongqin Liang, Yuanman Li, Xia Li 0006, Yi Tang 0008, Jiantao Zhou 0001, Wenbin Zou
AAAI3
2021 A Transformer based Approach for Image Manipulation Chain Detection
abstract
Image manipulation chain detection aims to identify the existence of involved operations and also their orders, playing an important role in multimedia forensics and image analysis. However,all the existing algorithms model the manipulation chain detection as a classification problem, and can only detect chains containing up to two operations. Due to the exponentially increased solution space and the complex interactions among operations, how to reveal a long chain from a processed image remains a long-standing problem in the multimedia forensic community. To address this challenge, in this paper, we propose a new direction for manipulation chain detection. Different from previous works, we treat the manipulation chain detection as a machine translation problem rather than a classification one, where we model the chains as the sentences of a target language, and each word serves as one possible image operation. Specifically, we first transform the manipulated image into a deep feature space, and further model the traces left by the manipulation chain as a sentence of a latent source language. Then, we propose to detect the manipulation chain through learning the mapping from the source language to the target one under a machine translation framework. Our method can detect manipulation chains consisting of up to five operations, and we obtain promising results on both the short-chain detection and the long-chain detection.
Jiaxiang You, Yuanman Li, Jiantao Zhou 0001, Zhongyun Hua, Weiwei Sun 0009, Xia Li 0006
ACM Multimedia6
2021 Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges
Neurocomputing4
2021 STA3D: Spatiotemporally attentive 3D network for video saliency prediction
Wenbin Zou, Shengkai Zhuo, Yi Tang 0008, Shishun Tian, Xia Li 0006, Chen Xu 0004
Pattern Recognit. Lett.5
2021 Dual-Stream Multi-Path Recursive Residual Network for JPEG Image Compression Artifacts Reduction
abstract
JPEG is the most widely used lossy image compression standard. When using JPEG with high compression ratios, visual artifacts cannot be avoided. These artifacts not only degrade the user experience but also negatively affect many low-level image processing tasks. Recently, convolutional neural network (CNN)-based compression artifact removal approaches have achieved significant success, however, at the cost of high computational complexity due to an enormous number of parameters. To address this issue, we propose a dual-stream recursive residual network (STRRN) which consists of structure and texture streams for separately reducing the specific artifacts related to high-frequency or low-frequency image components. The outputs of these streams are combined and fed into an aggregation network to further enhance the restored images. By using parameter sharing, the proposed network reduces the total number of training parameters significantly. Moreover, experiments conducted on five commonly used datasets confirm that the proposed STRRN can efficiently reduce the compression artifacts, while using up to 4.6 times less training parameters and 5 times less running time compared to the state-of-the-art approaches.
Zhi Jin 0002, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2020 Video salient object detection via spatiotemporal attention neural networks
Yi Tang 0008, Wenbin Zou, Yang Hua 0001, Zhi Jin 0002, Xia Li 0006
Neurocomputing5
2020 A many-objective particle swarm optimizer based on indicator and direction vectors for many-objective optimization
Jianping Luo, Xiongwen Huang, Xia Li 0006, Zhenkun Wang 0001, Jiqiang Feng
Inf. Sci.4
2020 Deep and joint learning of longitudinal data for Alzheimer's disease prediction
Bai Ying Lei, Mengya Yang, Peng Yang 0011, Feng Zhou 0003, Wen Hou, Wenbin Zou, Xia Li 0006, Tianfu Wang 0001, Xiaohua Xiao, Shuqiang Wang
Pattern Recognit.7
2020 A Flexible Deep CNN Framework for Image Restoration
abstract
Image restoration is a long-standing problem in image processing and low-level computer vision. Recently, discriminative convolutional neural network (CNN)-based approaches have attracted considerable attention due to their superior performance. However, most of these frameworks are designed for one specific image restoration task; hence, they seldom show high performance on other image restoration tasks. To address this issue, we propose a flexible deep CNN framework that exploits the frequency characteristics of different types of artifacts. Hence, the same approach can be employed for a variety of image restoration tasks by adjusting the architecture. For reducing the artifacts with similar frequency characteristics, a quality enhancement network that adopts residual and recursive learning is proposed. Residual learning is utilized to speed up the training process and boost the performance; recursive learning is adopted to significantly reduce the number of training parameters as well as boost the performance. Moreover, lateral connections transmit the extracted features between different frequency streams via multiple paths. One aggregation network combines the outputs of these streams to further enhance the restored images. We demonstrate the capabilities of the proposed framework with three representative applications: image compression artifacts reduction (CAR), image denoising, and single image super-resolution (SISR). Extensive experiments confirm that the proposed framework outperforms the state-of-the-art approaches on benchmark datasets for these applications.
Zhi Jin 0002, Dmytro Bobkov, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Multim.5
2019 A novel particle swarm optimizer for many-objective optimization
abstract
A novel many-objective particle swarm optimization (PSO) algorithm called IDMOPSO is presented in this study to robustly and effectively address many-objective optimization problems (MaOPs). IDMOPSO is based on a performance indicator and direction vectors. A selection strategy based on the quality indicator Iε+ and Pareto dominance for personal best (pbest) particles is proposed to ensure the convergence and diversity of the algorithm and enhance the capability of local exploration. An external archive based on Iε+ and direction vectors is used to preserve the diversity of non-dominated solutions found in the search process. A multi-global optimal (gbest) particle selection method is developed to increase global search ability and ensure the particles' diversity. This method allows each particle to be assigned to a different gbest particle. This method differs from the traditional method, wherein only one gbest particle is allocated for the whole population of PSO. We aim to design a robust multi-objective evolutionary algorithm to deal with MaOPs. Extensive comparative experiments on DTLZ and DTLZ-1problems with varied numbers of objectives show that IDMOPSO is effective and flexible in addressing MaOPs. The influences and effectiveness of the proposed strategies are also analyzed in detail.
Jianping Luo, Xiongwen Huang, Xia Li 0006, Kai-Zhou Gao
CEC3
2019 An Efficient Quality Enhancement Solution for Stereo Images
Yingqing Peng, Zhi Jin 0002, Wenbin Zou, Yi Tang 0008, Xia Li 0006
ICIG (3)5
2019 Robust Plane Detection Using Depth Information From a Consumer Depth Camera
abstract
The emerging of depth-camera technology is paving the way for a variety of new applications and it is believed that plane detection is one of them. In fact, planes are common in man-made living structures, thus their accurate detection can benefit many visual-based applications. The use of depth information allows detecting planes characterized by complex pattern and texture, where the texture-based plane detection algorithms usually fail. In this paper, we propose a robust depth-driven plane detection (DPD) algorithm which consists of two parts: the growing-based plane detection and a two-stage refinement. The proposed approach starts from the seed patch with the highest planarity and uses the estimated equation of the growing plane and a dynamic threshold function to steer the growing process. Aided with this mechanism, each seed patch can grow to its maximum extent, and then the next seed patch starts to grow. This process is iteratively repeated so as to detect all the planes. Moreover, the refinement is proposed to tackle two common problems suffered by growing-based approaches, the over-growing problem, and the under-growing problem. Validated by extensive experiments, the proposed DPD algorithm is able to accurately detect planes and robust to various testing conditions. In terms of applications, it can be used as the pre-processing step for a variety of applications, such as, planar object recognition, super-resolution of the time-of-flight depth images with intrinsically low resolution.
Zhi Jin 0002, Tammam Tillo, Wenbin Zou, Yao Zhao 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.5
2019 Weakly Supervised Salient Object Detection With Spatiotemporal Cascade Neural Networks
abstract
Recently, deep learning techniques have substantially boosted the performance of salient object detection in still images. However, the salient object detection in videos by using traditional handcrafted features or deep learning features is not fully investigated, probably due to the lack of sufficient manually labeled video data for saliency modeling, especially for the data-driven deep learning. This paper proposes a novel weakly supervised approach to the salient object detection in a video, which can learn a robust saliency prediction model by using very limited manually labeled data and a large amount of weakly labeled data that could be easily generated in a supervised approach. Furthermore, we propose a spatiotemporal cascade neural network architecture for saliency modeling, in which two fully convolutional networks are cascaded to evaluate the visual saliency from both spatial and temporal cues to lead the optimal video saliency prediction. The proposed approach is extensively evaluated on the widely used challenging data sets, and the experiments demonstrate that our proposed approach substantially outperforms the state-of-the-art salient object detection models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Yuhuan Chen, Yang Hua 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.6
2018 Video Salient Object Detection via Multiple Time-scale Analysis
abstract
This paper focuses on salient object detection in video by multiple time-scale analysis, which exploits the temporally consistent information under three different scales. In the first time-scale, we define an effective measure called motion contrast from both low-level cues and the optical flow fields. In the second time-scale, we propose a novel approach to repair the inaccurate motion contrast due to the mistake of optical flow. In the third time-scale, considering the low-contrast objects that stop moving for a certain amount of time and cannot remain prominent, we present a robust motion detection method based on point-tracking and trajectories clustering. Finally, the outcomes from the three time-scales jointly formulate the saliency detection by Bayesian inference. The proposed model is evaluated on the widely-used DAVIS and FBMS benchmark. Experiments demonstrate that our proposed model substantially outperforms the state-of-the-art saliency detection models.
Yuhuan Chen, Limin Huang, Wenbin Zou, Xia Li 0006, Guoping Qiu
ICPR4
2018 Multi-Scale Spatiotemporal Conv-LSTM Network for Video Saliency Detection
abstract
Recently, deep neural networks have been crucial techniques for image salient detection. However, two difficulties prevent the development of deep learning in video saliency detection. The first one is that the traditional static network cannot conduct a robust motion estimation in videos. The other is that the data-driven deep learning is in lack of sufficient manually annotated pixel-wise ground truths for video saliency network training. In this paper, we propose a multi-scale spatiotemporal convolutional LSTM network (MSST-ConvLSTM) to incorporate spatial and temporal cues for video salient objects detection. Furthermore, as manually pixel-wised labeling is very time-consuming, we sign lots of coarse labels, which are mixed with fine labels to train a robust saliency prediction model. Experiments on the widely used challenging benchmark datasets (e.g., FBMS and DAVIS) demonstrate that the proposed approach has competitive performance of video saliency detection compared with the state-of-the-art saliency models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Xia Li 0006
ICMR4
2018 A new hybrid memetic multi-objective optimization algorithm for multi-objective optimization
Jianping Luo, Qiqi Liu, Xia Li 0006, Min-Rong Chen, Kai-Zhou Gao
Inf. Sci.4
2018 Objective reduction for many-objective optimization problems using objective subspace extraction
Naili Luo, Xia Li 0006, Qiuzhen Lin
Soft Comput.2
2018 SCOM: Spatiotemporal Constrained Optimization for Salient Object Detection
abstract
This paper presents a novel model for video salient object detection called spatiotemporal constrained optimization model (SCOM), which exploits spatial and temporal cues, as well as a local constraint, to achieve a global saliency optimization. For a robust motion estimation of salient objects, we propose a novel approach to modeling the motion cues from optical flow field, the saliency map of the prior video frame and the motion history of change detection, which is able to distinguish the moving salient objects from diverse changing background regions. Furthermore, an effective objectness measure is proposed with intuitive geometrical interpretation to extract some reliable object and background regions, which provided as the basis to define the foreground potential, background potential, and the constraint to support saliency propagation. These potentials and the constraint are formulated into the proposed SCOM framework to generate an optimal saliency map for each frame in a video. The proposed model is extensively evaluated on the widely used challenging benchmark data sets. Experiments demonstrate that our proposed SCOM substantially outperforms the state-of-the-art saliency models.
Yuhuan Chen, Wenbin Zou, Yi Tang 0008, Xia Li 0006, Chen Xu 0004, Nikos Komodakis
IEEE Trans. Image Process.4
2017 Personalized recommendation based on time perception and users' feedback
abstract
As the basis of human interactions, trust has been playing an influential role in addressing information sharing, experience communication, and public opinions. Trust-aware recommender systems are an effective solution to the information overload problem, especially in the online world where we are constantly faced with inordinately many choices. Based traditional rating prediction approach, this study focuses on constructing personalized recommendation by considering time perception and users' feedback. We present technical details about modeling trust evolution and perform experiments to show how the exploitation of trust evolution can help improve the performance of rating prediction and bring more robust solutions to the cold-start problem.
Zhaonan Chen, Na Wang 0001, Xia Li 0006, Piao Chen
CSCWD3
2017 Multi-modal metric learning for vehicle re-identification in traffic surveillance environment
abstract
Vehicle re-identification (Re-Id) aims to retrieve the same vehicle captured by disjoint cameras at different time instants from different locations, and is a challenging task mainly due to the high similarity among the captured vehicle images in surveillance environment. With the rapid development of Convolutional Neural Network (CNN), learning-based deep features have been adopted to combine with hand-crafted features to re-identify vehicles in traffic surveillance environment. However, the two kinds of features are in different feature space, and if they are fused directly together, their complementary correlation is not able to be fully explored. To address such an issue, this paper proposes a multi-modal metric learning architecture to fuse deep features and hand-crafted ones in an end-to-end optimization network, which achieves a more robust and discriminative feature representation for vehicle re-identification. The extensive experiments on a large-scale traffic surveillance vehicle dataset demonstrate that our proposed approach substantially outperforms the state-of-the-art methods on vehicle Re-Id.
Yi Tang 0008, Di Wu 0009, Zhi Jin 0002, Wenbin Zou, Xia Li 0006
ICIP5
2017 A CNN cascade for quality enhancement of compressed depth images
abstract
Transmitting depth images along with the corresponding textures enables a wide range of receiver-side 3D applications. Since each pixel on the depth images represents a corresponding 3D scene geometric information, when compressed during transmission the compression artifacts will lead to severe geometry distortions and visual perceptual degradation. To solve this problem, in this paper we proposed a convolutional neural network (CNN) cascade for suppressing the compression artifacts on depth images. According to the feature of depth images, we furthermore, adopt a weighted loss function for network training which can adaptively improve the learning efficiency and accuracy. Meanwhile, in order to overcome the limited training data problem, we audaciously trained our network on textures first and then finetune on the target depth images. To our best knowledge, few works have applied CNN on depth images targeting for compression artifacts reduction (CAR). Through extensive experiments, our proposed solution achieves higher quality for both reconstructed depth images and synthesized virtual views than the state-of-the-art methods.
Zhi Jin 0002, Lei Luo 0003, Yi Tang 0008, Wenbin Zou, Xia Li 0006
VCIP5
2017 Stalling Assessment for Wireless Online Video Streams via ISP Traffic Monitoring
abstract
Nowadays, ISPs witness an increasing network traffic generated by wireless online video services. Facing the intense mutual competition, it is important for an ISP to improve QoE of wireless video services. Since it is difficult for ISPs to collect user experience information at routers and switches, we propose an approach based on ISP traffic monitoring to access the stalling of a video stream, which is the key factor to video QoE. Concretely, we passively sniff out the following parameters to reconstruct video playing and stalling: playback rate of a slice, start-up delay and playback resuming threshold. Our difference from previous works is that we need not intrude into packets to extract HTTP meta information for these parameters. Moreover, we develop a video stalling assessment system based on Apache Storm, which can process the network traffic data in real time. Experiments on popular online video services (Tencent Video, iQiyi and Youku) in China as well as Youtube in US show not bad absolute difference of stalling times and total stalling duration between our traffic monitoring approach and real measurement.
Na Wang 0001, Xia Li 0006
WCNC5
2017 Global and local scatter based semi-supervised dimensionality reduction with active constraints selection in ensemble subspaces
Na Wang 0001, Xia Li 0006, Piao Chen
Pattern Anal. Appl.2
2017 Enhanced Autofocusing in Optical Scanning Holography Based on Hologram Decomposition
abstract
Optical scanning holography is a compact and powerful method for capturing hologram of a wide three-dimensional (3-D) view scene. After a hologram has been taken, it is often necessary to determine the locations of the focal plane on which the objects are residing, so that the 3-D scene can be numerically reconstructed for further analysis or processing. Recent research has shown that automatic detection of the depth (focal plane) of objects represented in a hologram can be conducted with entropy minimization method. Despite the success, the method could fail if the entropy information of objects in a hologram are interfering with each other. In this paper, we propose a method based on the hologram decomposition to overcome this problem. Briefly, the hologram is decomposed into subholograms and the focal plane distance is determined separately for each subobject. Simulation results reveal that our proposed method has good accuracy and reliability.
Shuming Jiao, Peter Wai-Ming Tsang, Ting-Chung Poon, Jung-Ping Liu, Wenbin Zou, Xia Li 0006
IEEE Trans. Ind. Informatics6
2015 A novel hybrid shuffled frog leaping algorithm for vehicle routing problem with time windows
Jianping Luo, Xia Li 0006, Min-Rong Chen
Inf. Sci.2
2015 A Feature Point Matching Based on Spatial Order Constraints Bilateral-Neighbor Vote
abstract
Feature point matching is a fundamental and challenging problem in many computer vision applications. In this paper, a robust feature point matching algorithm named spatial order constraints bilateral-neighbor vote (SOCBV) is proposed to remove outliers for a set of matches (including outliers) between two images. A directed k nearest neighbor (knn) graph of match sets is generated, and the problem of feature point matching is formulated as a binary discrimination problem. In the discrimination process, the class labeled matrix is built via the spatial order constraints defined on the edges that connect a point to its knn. Then, the posterior inlier class probability of each match is estimated with the knn density estimation and spatial order constraints. The vote of each match is determined by averaging all posterior class probabilities that originate from its associative inliers set and is used for removing outliers. The algorithm iteratively removes outliers from the directed graph and recomputes the votes until the stopping condition is satisfied. Compared with other popular algorithms, such as RANSAC, RSOC, GTM, SOC and WGTM, experiments under various testing data sets demonstrate strong robustness for the proposed algorithm.
Fanyang Meng, Xia Li 0006, Jihong Pei
IEEE Trans. Image Process.2
2014 A novel Artificial Bee Colony algorithm with integration of extremal optimization for numerical optimization problems
abstract
Artificial Bee Colony (ABC) algorithm is an optimization algorithm based on a particular intelligent behaviour of honeybee swarms. The standard ABC is weak at the local-search capability and precision. Extremal Optimization (EO) is a general-purpose heuristic method which has strong local-search capability and has been successfully applied to a wide variety of hard optimization problems. In order to strengthen the local-search capability of ABC, this work proposes a novel hybrid optimization method, called ABC-EO algorithm, through introducing EO to ABC. The simulation results show that the performance of the proposed method is as good as or superior to those of the state-of-the-art algorithms in complex numerical optimization problems.
Min-Rong Chen, Xia Li 0006, Jianping Luo
IEEE Congress on Evolutionary Computation4
2014 Hybrid shuffled frog leaping algorithm for energy-efficient dynamic consolidation of virtual machines in cloud data centers
Jianping Luo, Xia Li 0006, Min-Rong Chen
Expert Syst. Appl.2
2012 An improved shuffled frog-leaping algorithm with extremal optimisation for continuous optimisation
Xia Li 0006, Jianping Luo, Min-Rong Chen, Na Wang 0001
Inf. Sci.1
2010 Fast three-dimensional Otsu thresholding with shuffled frog-leaping algorithm
Na Wang 0001, Xia Li 0006, Xiaohong Chen 0001
Pattern Recognit. Lett.2
2008 Semi-supervised kernel-based fuzzy C-means with pairwise constraints
abstract
Clustering with constraints is an active area in machine learning and data mining. In this paper, a semi-supervised kernel-based fuzzy C-means algorithm called PCKFCM is proposed which incorporates both semi-supervised learning technique and the kernel method into traditional fuzzy clustering algorithm. The clustering is achieved by minimizing a carefully designed objective function. A kernel-based fuzzy term defined by the violation of constraints is included. The proposed PCKFCM is compared with other clustering techniques on benchmark and the experimental results convince that effective use of constraints improves the performance of kernel-based clustering. As for the effect of key parameter selection and the non-linear capability, it outperforms a similar semi-supervised fuzzy clustering approach Pairwise Constrained Competitive Agglomeration (PCCA).
Na Wang 0001, Xia Li 0006, Xuehui Luo
IJCNN2
2007 A Fast Training Algorithm for SVM Via Clustering Technique and Gabriel Graph
Xia Li 0006, Na Wang 0001, Shu-Yuan Li
ICIC (3)1