Xin Li 0005

dblp:09/1365-5 · DBLP profile ↗
← Back
178ranked-venue papers
38as first author
93since 2021 · last 2026
0000-0003-2067-2763ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 126 · 37 first-author · 53 since 2021Artificial intelligence and machine learning · 70 · 1 first-author · 55 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 4 since 2021Security and privacy · 5Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention Network
abstract
The task of image feature matching aims to establish correct correspondences between images from two different views. While approaches based on attention mechanisms have demonstrated remarkable advancements in image feature matching, they still encounter substantial limitations. Specifically, current graph attention network approaches face performance bottlenecks in complex scenarios, such as low-texture regions or occlusions. This limitation stems from the self-attention mechanism, which, when lacking effective guidance, can lead to divergent attention weights or incorrect focus on regions with low discriminability, resulting in matching failures in low-texture environments. Inspired by how humans focus on distinctive regions when performing cross-view matching, we enhance attention to singular points in images that are salient, unique and have high cross-view matching potential during information aggregation, thereby improving matching capability. To realize the aforementioned strategies, we develop a novel Singularity-enhanced Graph Attention Network (SGAT). SGAT leverages Co-potentiality and Multi-Scale Singularity as prior guidance, and designs a Singularity-aware Attention mechanism and a Co-potentiality Guided Attention mechanism , specifically enhancing the perception of singularity and matching potential during feature interaction. Experimental results on multiple datasets, including ScanNet1500, demonstrate that our method outperforms current state-of-the-art sparse matching methods. In particular, the improvement is most pronounced in complex scenarios such as low-texture environments, significantly enhancing the accuracy and robustness of image matching and its downstream tasks.
Kun Sun 0002, Chang Tang, Yuanyuan Liu 0004, Xin Li 0005
AAAI5
2026 Multimodal chain-of-thought reasoning with large language models to protect children from age-inappropriate apps
Chuanbo Hu, Bin Liu 0045, Minglei Yin, Yilu Zhou, Xin Li 0005
Inf. Manag.5
2026 COBRA: A Continual Learning Approach to Vision-Brain Understanding
Xuan-Bac Nguyen, Manuel Serna-Aguilera, Arabinda Kumar Choudhary, Pawan Sinha, Xin Li 0005, Khoa Luu
Int. J. Comput. Vis.5
2026 CoCoFR: Collaborative codebooks learning with soft matching strategy for blind face restoration
Teng Feng, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi
Neural Networks7
2026 RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045
Neural Networks7
2026 Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation
Yuejiang Li, Weisheng Dong, Peng Wu 0015, Lichao Mou, Xin Li 0005
Neural Networks7
2026 BRACTIVE: A Brain Activation Approach to Human Visual Brain Learning
abstract
The human brain is a highly efficient processing unit, and understanding how it works can inspire new algorithms and architectures in machine learning. In this work, we introduce a novel framework named Brain Activation Network (BRACTIVE), a transformer-based approach to studying the human visual brain. The primary objective of BRACTIVE is to align the visual features of subjects with their corresponding brain representations using functional Magnetic Resonance Imaging (fMRI) signals. It enables us to identify the brain's Regions of Interest (ROIs) in the subjects. Unlike previous brain research methods, which can only identify ROIs for one subject at a time and are limited by the number of subjects, BRACTIVE automatically extends this identification to multiple subjects and ROIs. Our experiments demonstrate that BRACTIVE effectively identifies person-specific regions of interest, such as face and body-selective areas, aligning with neuroscience findings and indicating potential applicability to various object categories. More importantly, we found that leveraging human visual brain activity to guide deep neural networks enhances performance across various benchmarks. It encourages the potential of BRACTIVE in both neuroscience and machine intelligence studies.
Xuan-Bac Nguyen, Hojin Jang, Xin Li 0005, Samee Ullah Khan, Pawan Sinha, Khoa Luu
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Visible-infrared joint image deraining for harsh rain conditions with cross-modal semantic consistency
Xin Li 0005, Chengpei Xu, Zhenyu Wang 0008, Weisheng Dong
Pattern Recognit.2
2026 Uncertainty-Driven Generative Prior Learning for Sparse Model-Guided Hyperspectral Image Fusion
abstract
As an alternative to acquiring high-resolution hyperspectral images (HR-HSI), Hyperspectral Image Fusion (HIF) aims to recover clean HR-HSIs by fusing degraded low spatial resolution hyperspectral images and high spatial resolution multispectral images. Among existing HIF approaches, model-guided HIF methods stand out by integrating physical degradation constraints with the learning capabilities of data-driven networks. However, most of them learn deep priors only from degraded-clean pairs without degradation-free knowledge, making them struggle with severe or unseen degradations. To address these issues, we propose a Vector-Quantized Prior-Guided Network (VPG-Net), an unfolding-based HIF framework enhanced by sparse representation and novel uncertainty-driven generative priors. Specifically, VPG-Net unfolds the Maximum A Posteriori (MAP) estimation with a sparse representation model into an uncertainty-aware VQ prior-guided network implementation. Within this framework, the sparse representation prior is integrated into the MAP formulation to improve noise resistance. As the core of our method, we leverage a high-quality vector-quantized (VQ) prior, which serves as a powerful degradation-free generative prior for the HIF process. We pre-train a discrete codebook and encoder on clean HR-HSIs to generate a VQ-prior representation (VQPR), which preserves complete spatial-spectral information. To effectively bridge the gap between degraded inputs and the learned degradation-free codebook, we further incorporate a novel uncertainty-driven probabilistic matching strategy that improves feature alignment and suppresses artifacts. The learned VQPR is then incorporated into the deep prior module as dynamic modulation parameters to enhance the fidelity and realism of the reconstructed results, particularly for severely degraded inputs. Extensive experiments on clean and degraded synthetic and real-world datasets demonstrate that our approach outperforms state-of-the-art HIF methods in both quantitative metrics and visual quality.
Teng Feng, Zhenxuan Fang, Weisheng Dong, Xin Li 0005
IEEE Trans. Image Process.9
2025 Training-Free Image Manipulation Localization Using Diffusion Models
abstract
Image manipulation localization (IML) is a critical technique in media forensics, focusing on identifying tampered regions within manipulated images. Most existing IML methods require extensive training on labeled datasets with both image-level and pixel-level annotations. These methods often struggle with new manipulation types and exhibit low generalizability. In this work, we propose a training-free IML approach using diffusion models. Our method adaptively selects an appropriate number of diffusion timesteps for each input image in the forward process and performs both conditional and unconditional reconstructions in the backward process without relying on external conditions. By comparing these reconstructions, we generate a localization map highlighting regions of manipulation based on inconsistencies. Extensive experiments were conducted using sixteen state-of-the-art (SoTA) methods across six IML datasets. The results demonstrate that our training-free method outperforms SoTA unsupervised and weakly-supervised techniques. Furthermore, our method competes effectively against fully-supervised methods on novel (unseen) manipulation types.
Zhenfei Zhang, Ming-Ching Chang, Xin Li 0005
AAAI3
2025 Parameterized Blur Kernel Prior Learning for Local Motion Deblurring
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Guangming Shi
CVPR6
2025 Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classes
abstract
Recent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impact of distributional shifts. In this paper, we observe that classification errors arising from distribution shifts tend to cluster near the true values, suggesting that misclassifications commonly occur in semantically similar, neighboring categories. Furthermore, robust advanced vision foundation models maintain larger inter-class distances while preserving semantic consistency, making them less vulnerable to such shifts. Building on these findings, we propose a new method called GFN (Gain From Neighbors), which uses gradient priors from neighboring classes to perturb input images and incorporates an inter-class distance-weighted loss to improve class separation. This approach encourages the model to learn more resilient features from data prone to errors, enhancing its robustness against shifts in diverse settings. In extensive experiments across various model architectures and benchmark datasets, GFN consistently demonstrated superior performance. For instance, compared to the current state-of-the-art TAPADL method, our approach achieved a higher corruption robustness of 41.4% on ImageNet-C (+2.3%), without requiring additional parameters and using only minimal data.
Mingtao Feng, Weisheng Dong, Xin Li 0005, Guangming Shi
CVPR6
2025 Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven Adaptation
abstract
Given image (or text), remote sensing image-text retrieval (RSITR) aims to retrieve corresponding text (or image) within diverse remote sensing data. However, due to the complex scenes and compact distribution of targets in remote sensing data, existing methods, particularly those leveraging large models like CLIP, often generate features with high intra-modal similarity and insufficient distinctive characteristics, thus resulting in suboptimal retrieval performance. To address these issues, we pro- pose a novel dictionary-based RSITR method that jointly models image and text feature estimation. Specifically, by incorporating a general dictionary and the corresponding sparse coefficients, our method more effectively captures the correlations between the learned features. Furthermore, we introduce adaptive weighted metric learning on a sample-by-sample basis to encourage the model to focus on more challenging negative samples, promoting fine-grained feature alignment. The extensive experiments on the RSICD and RSITMD datasets demonstrate the effectiveness of our method, demonstrating significant improvements in retrieval performance.
Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005
ICASSP5
2025 RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation
abstract
Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler.
Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045
IJCAI5
2025 PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing
abstract
Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. Pattern CIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12.40% to 62.03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here.
Zhechun Liang, Shiwen Xue, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi
IJCAI7
2025 Beyond Visual Quality: Fidelity-Oriented Diffusion Model for Real-world Image Super-Resolution
abstract
Although existing diffusion-based image super-resolution methods have achieved remarkable visual quality, they often struggle with fidelity issues, particularly in preserving consistency with the original input image. This issue arises because using low-quality images as conditional inputs introduces substantial errors in the diffusion backward denoising process, making the restored features deviate from target features and thus degrade image fidelity. To improve the accuracy of noise estimation, we propose a dual-memory module to reinforce the input low-quality conditional features, which consists of a pre-trained high-quality memory bank to enrich the structural information and a degradation memory to remove the degradation components. Furthermore, we develop an uncertainty-aware noise estimation framework, utilizing an extra branch in the denoising network to predict pixel-wise uncertainty values, thus dynamically adjust the optimization weights for high-uncertainty regions. This adaptive strategy effectively improves the accuracy of noise estimation in challenging reconstruction areas. Experimental results demonstrate that our method significantly enhances the fidelity while preserving high visual quality of diffusion-based super-resolution, improving the reliability of diffusion applications.
Zhenxuan Fang, Shuaibo Wang, Weisheng Dong, Xin Li 0005, Guangming Shi
ACM Multimedia6
2025 TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth Estimation
abstract
Recent diffusion-based methods have shown strong ability in the depth estimation task, but they largely overlook the rich textual priors embedded in pretrained diffusion models that can enhance both performance and robustness in diverse scenes. In this paper, we propose TPDepth, a diffusion-based, affine-invariant monocular depth estimator that incorporates textual semantics via a Text-Prompted ControlNet. While directly injecting text into the diffusion U-Net can cause the network to over-attend to local semantic cues and compromise global structural modeling, TPDepth processes textual features through a separate ControlNet branch, allowing semantic information to be incorporated without disrupting the spatial reasoning pipeline. Prompt-conditioned features are modulated by an Adaptive Control Scale Module(ACSM) and injected into decoder of the diffusion UNet with skip connections. The model is fine-tuned with a fixed timestep for deterministic prediction. TPDepth achieves state-of-the-art results on NYUv2, KITTI, and ScanNet, and demonstrates competitive performance on two additional zero-shot benchmarks using only 61K training images. Code and models can be found on our https://github.com/Lioely/TPDepth project page.
Yu Liu 0136, Kun Sun 0002, Chang Tang, Xin Li 0005
ACM Multimedia5
2025 Exploring Global Correlations via Polarity Memory for Multispectral Demosaicing
abstract
Multispectral image demosaicing aims to reconstruct full band multispectral images from a compressed spectral mosaic images. Although existing learning-based methods have made progress in multispectral image demosaicing, there still exist intrinsic performance bottlenecks due to the heavy undersampling according to mosaic pattern. To address this issue, we propose Polarity memory network with quant attention to establish global correlation, thus reconstructing high-quality multispectral images from compressed spectral mosaic images. Our proposed Polarity memory network adaptively encapsulates reconstruction-oriented representations, then amplifies relevant ones and reducing noise from irrelevant ones in a polarity-aware manner to better cater to the enhancement of different spectral information with linear computational complexity. Moreover, considering existing methods' inability to adequately compensate for long distance interactions in reconstruction, we introduce a quant attention paradigm that categorize tokens into semantic-aware groups using an efficient quant operation for attention computation. Experimental results show our method achieves state-of-the-art performance on various simulation datasets and better vision results on real-world datasets.
Mengzu Liu, Xin Li 0005, Weisheng Dong
ACM Multimedia6
2025 A Semantically Impactful Image Manipulation Dataset: Characterizing Image Manipulations Using Semantic Significance
abstract
We investigate how to characterize semantic significance (SS) in detecting image manipulations (IMD) for media forensics. We introduce the Characterization of Seman-tic Impact for IMD (CSI-IMD) dataset, which focuses on localizing and evaluating the semantic impact of image manipulations to counter advanced generative techniques. Our evaluation of 10 state-of-the-art IMD and localization methods on CSI-IMD reveals key insights. Unlike existing datasets, CSI-IMD provides detailed semantic annotations beyond traditional manipulation masks, aiding in the development of new defensive strategies. The dataset features manipulations from advanced generation methods, offering various levels of semantic significance. It is divided into two parts: a gold-standard set of 1,000 manu-ally annotated manipulations with high-quality control, and an extended set of 500,000 automated manipulations for large-scale training and analysis. We also propose a new SS-focused task to assess the impact of semantically targeted manipulations. Our experiments show that current IMD methods struggle with manipulations created using stable diffusion, with TruFor and Cat-Net performing the best among those tested. The CSI-IMD dataset will become available at https://github.com/csiimd/csiimd.
Ming-Ching Chang, Matthias Kirchner, Zhenfei Zhang, Xin Li 0005, Arslan Basharat, Anthony Hoogs
WACV5
2025 Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
abstract
Abstract Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data. However, current conversational models still lack knowledge about visual insects since they are often trained on the general knowledge of vision-language data. Meanwhile, understanding insects is a fundamental problem in precision agriculture, helping to promote sustainable development in agriculture. Therefore, this paper proposes a novel multimodal conversational model, Insect-LLaVA, to promote visual understanding in insect-domain knowledge. In particular, we first introduce a new large-scale Multimodal Insect Dataset with Visual Insect Instruction Data that enables the capability of learning the multimodal foundation models. Our proposed dataset enables conversational models to comprehend the visual and semantic features of the insects. Second, we propose a new Insect-LLaVA model, a new general Large Language and Vision Assistant in Visual Insect Understanding. Then, to enhance the capability of learning insect features, we develop an Insect Foundation Model by introducing a new micro-feature self-supervised learning with a Patch-wise Relevant Attention mechanism to capture the subtle differences among insect images. We also present Description Consistency loss to improve micro-feature learning via text descriptions. The experimental results evaluated on our new Visual Insect Question Answering benchmarks illustrate the effective performance of our proposed approach in visual insect understanding and achieve State-of-the-Art performance on standard benchmarks of insect-related tasks Project Page: https://uarkcviu.github.io/projects/insectfoundation .
Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling, Xin Li 0005, Khoa Luu
Int. J. Comput. Vis.5
2025 Brainformer: Mimic human visual brain functions to machine vision models via fMRI
Xuan-Bac Nguyen, Xin Li 0005, Pawan Sinha, Samee Ullah Khan, Khoa Luu
Neurocomputing2
2025 Growing-before-pruning: A progressive neural architecture search strategy via group sparsity and deterministic annealing
abstract
Network pruning is a widely studied technique of obtaining compact representations from over-parameterized deep convolutional neural networks . Existing pruning methods are based on finding an optimal combination of pruned filters in the fixed search space . However, the optimality of those methods is often questionable due to limited search space and pruning choices - e.g., the difficulty with removing the entire layer and the risk of unexpected performance degradation . Inspired by the exploration vs. exploitation trade-off in reinforcement learning, we propose to reconstruct the filter space without increasing the model capacity and prune them by exploiting group sparsity . Our approach challenges the conventional wisdom by advocating the strategy of Growing-before-Pruning (GbP), which allows us to explore more space before exploiting the power of architecture search. Meanwhile, to achieve more efficient pruning, we propose to measure the importance of filters by global group sparsity , which extends the existing Gaussian scale mixture model. Such global characterization of sparsity in the filter space leads to a novel deterministic annealing strategy for progressively pruning the filters. We have evaluated our method on several popular datasets and network architectures. Our extensive experiment results have shown that the proposed method advances the current state-of-the-art.
Xiaotong Lu, Weisheng Dong, Zhenxuan Fang, Jie Lin 0008, Xin Li 0005, Guangming Shi
Pattern Recognit.5
2025 Distilling Hierarchical Knowledge From Multimodal Fusion for Unimodal Image Segmentation
abstract
The application of multimodal image fusion has become increasingly widespread across various fields in the era of deep learning. Existing fusion methods integrate infrared and visible images to provide complementary content and enhance the robustness of complex real-world scenes for high-level visual tasks, such as semantic segmentation and object detection. In return, high-level visual tasks facilitate the fusion of infrared and visible by providing mid-level semantic information. However, such frameworks rely heavily on multimodal data and require strict registration of images from different modalities before fusion, seriously limiting their practical applications due to the common realistic situations of missing modalities or misregistration. To move beyond this limitation, we propose a novel hierarchical knowledge distillation (HKD) framework tailored for unimodal image segmentation with the guidance of multi-modality. This framework aims to retain as much diverse information from multimodal image fusion as possible, thereby enhancing downstream high-level visual tasks when only the unimodal images are available during the inference phase. Our proposed method is two-stage, and we construct a robust multimodal fusion and segmentation interaction network in the first stage as a powerful teacher model. In the second stage, we design a hierarchical distillation method to transfer the fused and segmented multi-layer knowledge from the multimodal teacher model to the unimodal student model. Extensive experimental results on two public datasets, i.e., MFNet and FMB, demonstrate that the proposed hierarchical knowledge distillation framework can effectively transfuse multimodal knowledge into the unimodal student model for image enhancement and segmentation under incomplete multimodal conditions, and achieves considerably competitive results compared to multimodal image fusion and segmentation models.
Weisheng Dong, Shuaibo Wang, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.6
2025 CDS-Net: Contextual Difference Sensitivity Network for Pixel-Wise Road Crack Detection
abstract
Road crack detection is a key computer vision task that identifies and locates cracks in road surface images, which usually have an irregular shape and contain only a few pixels in width. Generative and unsupervised methods are popular these years, but generative methods require a lot of training data and computational power while unsupervised methods are not so satisfactory in pixel-level segmentation. The process is challenged by the irregularity of crack shapes and complex road image backgrounds. To alleviate these problems, we propose a novel method in this paper, CDS-Net, that significantly improves road crack detection performance through multiple practical modules, including the Multi-Directional Hierarchical Attention (MDHA) module and the Difference Sensitivity Reconstruction Block (DSRB). Specifically, the MDHA module employs a multi-directional feature extraction strategy to capture detailed information of cracks, thereby enhancing the discriminative power of the features. The DSRB module, designed to address the inefficiency of traditional skip-connections, utilizes masked convolution and graph convolution attention to reconstruct and refine feature representations. Additionally, we propose an improved weighted cross-entropy loss function to address the inherent class imbalance problem in road crack detection. Extensive experiments on five public datasets demonstrate that CDS-Net achieves superior performance compared to other state-of-the-art methods, showcasing its effectiveness and robustness in road crack detection. It also has a stronger generalization ability compared with other methods. Code is available athttps://github.com/ttttqz/CDS-Net/tree/master.
Qinzhong Tan, Weisheng Dong, Xin Li 0005, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.5
2025 Scale-Aware Crowd Counting Network With Annotation Error Modeling
abstract
Traditional crowd-counting networks suffer from information loss when feature maps are reduced by pooling layers, leading to inaccuracies in counting crowds at a distance. Existing methods often assume correct annotations during training, disregarding the impact of noisy annotations, especially in crowded scenes. Furthermore, using a fixed Gaussian density model does not account for the varying pixel distribution of the camera distance. To overcome these challenges, we propose a Scale-Aware Crowd Counting Network (SACC-Net) that introduces a scale-aware loss function with error-compensation capabilities of noisy annotations. For the first time, we simultaneously model labeling errors (mean) and scale variations (variance) by spatially varying Gaussian distributions to produce fine-grained density maps for crowd counting. Furthermore, the proposed scale-aware Gaussian density model can be dynamically approximated with a low-rank approximation, leading to improved convergence efficiency with comparable accuracy. To create a smoother scale-aware feature space, this paper proposes a novel Synthetic Fusion Module (SFM) and an Intra-block Fusion Module (IFM) to generate fine-grained heat maps for better crowd counting. The lightweight version of our model, named SACC-LW, enhances the computational efficiency while retaining accuracy. The superiority and generalization properties of scale-aware loss function are extensively evaluated for different backbone architectures and performance metrics on six public datasets: UCF-QNRF, UCF CC 50, NWPU, ShanghaiTech A, ShanghaiTech B, and JHU. Experimental results also demonstrate that SACC-Net outperforms all state-of-the-art methods, validating its effectiveness in achieving superior crowd-counting accuracy. The source code is available at https://github.com/Naughty725.
Yi-Kuan Hsieh, Jun-Wei Hsieh, Xin Li 0005, Yu-Ming Zhang, Yu-Chee Tseng, Ming-Ching Chang
IEEE Trans. Image Process.3
2025 Incomplete Modalities Restoration via Hierarchical Adaptation for Robust Multimodal Segmentation
abstract
Multimodal semantic segmentation has significantly advanced the field of semantic segmentation by integrating data from multiple sources. However, this task often encounters missing modality scenarios due to challenges such as sensor failures or data transmission errors, which can result in substantial performance degradation. Existing approaches to addressing missing modalities predominantly involve training separate models tailored to specific missing scenarios, typically requiring considerable computational resources. In this paper, we propose a Hierarchical Adaptation framework to Restore Missing Modalities for Multimodal segmentation (HARM3), which enables frozen pretrained multimodal models to be directly applied to missing-modality semantic segmentation tasks with minimal parameter updates. Central to HARM3 is a text-instructed missing modality prompt module, which learns multimodal semantic knowledge by utilizing available modalities and textual instructions to generate prompts for the missing modalities. By incorporating a small set of trainable parameters, this module effectively facilitates knowledge transfer between high-resource domains and low-resource domains where missing modalities are more prevalent. Besides, to further enhance the model's robustness and adaptability, we introduce adaptive perturbation training and an affine modality adapter. Extensive experimental results demonstrate the effectiveness and robustness of HARM3 across a variety of missing modality scenarios.
Weisheng Dong, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi
IEEE Trans. Image Process.6
2025 Local Uncertainty Energy Transfer for Active Domain Adaptation
abstract
Active Domain Adaptation (ADA) improves knowledge transfer efficiency from the labeled source domain to the unlabeled target domain by selecting a few target sample labels. However, most existing active sampling methods ignore the local uncertainty of neighbors in the target domain,making it easier to pick out anomalous samples that are detrimental to the model. To address this problem, we present a new approach to active domain adaptation called Local Uncertainty Energy Transfer (LUET), which integrates active learning of local uncertainty confusion and energy transfer alignment constraints into a unified framework. First, in the active learning module, the uncertainty difficult and representative samples from the target domain are selected through local uncertainty energy selection and entropy-weighted class confusion selection. And the active learning strategy based on local uncertainty energy will avoid selecting anomalous samples in the target domain. Second, for the discrimination issue caused by domain shift, we use a global and local energy-transfer alignment constraint module to eliminate the domain gap and improve accuracy. Finally, we used negative log-likelihood loss for supervised learning of source domains and query samples. With the introduction of sample-based energy metrics, the active learning strategy is more closely with the domain alignment. Experiments on multiple domain-adaptive datasets have demonstrated that our LUET can achieve outstanding results and outperform existing state-of-the-art approaches.
Guangming Shi, Weisheng Dong, Xin Li 0005, Xuemei Xie
IEEE Trans. Image Process.4
2024 Pushing the Limit of Fine-Tuning for Few-Shot Learning: Where Feature Reusing Meets Cross-Scale Attention
abstract
Due to the scarcity of training samples, Few-Shot Learning (FSL) poses a significant challenge to capture discriminative object features effectively. The combination of transfer learning and meta-learning has recently been explored by pre-training the backbone features using labeled base data and subsequently fine-tuning the model with target data. However, existing meta-learning methods, which use embedding networks, suffer from scaling limitations when dealing with a few labeled samples, resulting in suboptimal results. Inspired by the latest advances in FSL, we further advance the approach of fine-tuning a pre-trained architecture by a strengthened hierarchical feature representation. The technical contributions of this work include: 1) a hybrid design named Intra-Block Fusion (IBF) to strengthen the extracted features within each convolution block; and 2) a novel Cross-Scale Attention (CSA) module to mitigate the scaling inconsistencies arising from the limited training samples, especially for cross-domain tasks. We conducted comprehensive evaluations on standard benchmarks, including three in-domain tasks (miniImageNet, CIFAR-FS, and FC100), as well as two cross-domain tasks (CDFSL and Meta-Dataset). The results have improved significantly over existing state-of-the-art approaches on all benchmark datasets. In particular, the FSL performance on the in-domain FC100 dataset is more than three points better than the latest PMF (Hu et al. 2022).
Ying-Yu Chen, Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang
AAAI3
2024 Inverse Weight-Balancing for Deep Long-Tailed Learning
abstract
The performance of deep learning models often degrades rapidly when faced with imbalanced data characterized by a long-tailed distribution. Researchers have found that the fully connected layer trained by cross-entropy loss has large weight-norms for classes with many samples, but not for classes with few samples. How to address the data imbalance problem with both the encoder and the classifier seems an under-researched problem. In this paper, we propose an inverse weight-balancing (IWB) approach to guide model training and alleviate the data imbalance problem in two stages. In the first stage, an encoder and classifier (the fully connected layer) are trained using conventional cross-entropy loss. In the second stage, with a fixed encoder, the classifier is finetuned through an adaptive distribution for IWB in the decision space. Unlike existing inverse image frequency that implements a multiplicative margin adjustment transformation in the classification layer, our approach can be interpreted as an adaptive distribution alignment strategy using not only the class-wise number distribution but also the sample-wise difficulty distribution in both encoder and classifier. Experiments show that our method can greatly improve performance on imbalanced datasets such as CIFAR100-LT with different imbalance factors, ImageNet-LT, and iNaturelists2018.
Wenqi Dang, Weisheng Dong, Xin Li 0005, Guangming Shi
AAAI4
2024 SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object Tracking
abstract
Despite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper introduces SMILEtrack, an innovative object tracker that effectively addresses these challenges by integrating an efficient object detector with a Siamese network-based Similarity Learning Module (SLM). The technical contributions of SMILETrack are twofold. First, we propose an SLM that calculates the appearance similarity between two objects, overcoming the limitations of feature descriptors in Separate Detection and Embedding (SDE) models. The SLM incorporates a Patch Self-Attention (PSA) block inspired by the vision Transformer, which generates reliable features for accurate similarity matching. Second, we develop a Similarity Matching Cascade (SMC) module with a novel GATE function for robust object matching across consecutive video frames, further enhancing MOT performance. Together, these innovations help SMILETrack achieve an improved trade-off between the cost (e.g., running speed) and performance (e.g., tracking accuracy) over several existing state-of-the-art benchmarks, including the popular BYTETrack method. SMILETrack outperforms BYTETrack by 0.4-0.8 MOTA and 2.1-2.2 HOTA points on MOT17 and MOT20 datasets. Code is available at http://github.com/pingyang1117/SMILEtrack_official.
Yu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Hung-Hin So, Xin Li 0005
AAAI6
2024 Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect Understanding
abstract
In precision agriculture, the detection and recognition of insects play an essential role in the ability of crops to grow healthy and produce a high-quality yield. The current machine vision model requires a large volume of data to achieve high performance. However, there are approximately 5.5 million different insect species in the world. None of the existing insect datasets can cover even a fraction of them due to varying geographic locations and acquisition costs. In this paper, we introduce a novel “Insect-1M” dataset, a game-changing resource poised to revolutionize insect-related foundation model training. Covering a vast spectrum of insect species, our dataset, including 1 million images with dense identification labels of taxonomy hierarchy and insect descriptions, offers a panoramic view of entomology, enabling foundation models to comprehend visual and semantic information about insects like never before. Then, to efficiently establish an Insect Foundation Model, we develop a micro-feature self-supervised learning method with a Patch-wise Relevant Attention mechanism capable of discerning the subtle differences among insect images. In addition, we introduce Description Consistency loss to improve micro-feature modeling via insect descriptions. Through our experiments, we illustrate the effectiveness of our proposed approach in insect modeling and achieve State-of-the-Art performance on standard benchmarks of insect-related tasks. Our Insect Foundation Model and Dataset promise to empower the next generation of insect-related vision models, bringing them closer to the ultimate goal of precision agriculture.
Hoang-Quan Nguyen, Thanh-Dat Truong, Xuan-Bac Nguyen, Ashley Dowling, Xin Li 0005, Khoa Luu
CVPR5
2024 Image Manipulation Detection with Implicit Neural Representation and Limited Supervision
Zhenfei Zhang, Mingyang Li 0007, Xin Li 0005, Ming-Ching Chang, Jun-Wei Hsieh
ECCV (88)3
2024 Set-Nas: Sample-Efficient Training For Neural Architecture Search With Strong Predictor And Stratified Sampling
abstract
Sample-efficient neural architecture search (NAS) techniques have advanced rapidly. Two lines of methods, namely neural predictor and sequential search, have shown promising performance in improving the sample efficiency of NAS. However, as far as we know, little attention has been paid to the middle ground between these two lines. Inspired by the analogy between NAS and evolutionary optimization, we propose a new Sample-Efficient Training for NAS (SETNAS) based on strategies that improve fitness scores and sampling mechanisms. We develop a strong neural predictor called the Fully Bidirectional Graph Convolutional Network evolutionary (Fully-BiGCN) that significantly enhances the predictor capability of the features in each layer. The developed predictor is embedded into an iterative stratified sampling process to retain only a subset of best-fit architectures using the same training budget. SET-NAS achieves remarkable results compared to the state-of-the-art in predictor-based NAS. Using NASBench-201 as the benchmark, SET-NAS takes only $27.1 \%$ (CIFAR-10), $49.0 \%$ (CIFAR-100), and $51.75 \%$ (ImageNet-16) of training cost of other state-of-the-art predictor-based methods to find the promising network architecture.
Yu-Ming Zhang, Jun-Wei Hsieh, Yu-Hsiu Chang, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan
ICIP4
2024 A Linearly Continous False Alarm Removing Method in a Multichannel SAR-GMTI System
abstract
In this paper, a method aimed at removing the linearly continuous false alarm caused by the artificial building with unignored elevation in a multichannel synthetic aperture radar (SAR) system with ground moving target indication (GMTI) mode is proposed. In the proposed method, the Hough transform is firstly utilized to obtain the potential false alarms during the iteration procedure. After that, the linearly continuous false alarms are finally obtained and removed based on the phase similarity characteristic of those building scatters. Real-measured SAR data processing results are presented to verity the effectiveness and feasibility of the proposed method.
Lingyu Wang 0004, Xin Li 0005, Penghui Huang, Haojuan Yuan, Changhong He, Muyang Zhan, Fengyuan Hu
IGARSS3
2024 Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models
abstract
The need to analyze graphs is ubiquitous across various fields, from social networks to biological research and recommendation systems. Therefore, enabling the ability of large language models (LLMs) to process graphs is an important step toward more advanced general intelligence. However, current LLM benchmarks on graph analysis require models to directly reason over the prompts describing graphtopology, and are thus limited to small graphs with only a few dozens of nodes. In contrast, human experts typically write programs based on popular libraries for task solving, and can thus handle graphs with different scales. To this end, a question naturally arises: can LLMs analyze graphs like professionals? In this paper, we introduce ProGraph, a manually crafted benchmark containing 3 categories of graph tasks. The benchmark expects solutions based on programming instead of directly reasoning over raw inputs. Our findings reveal that the performance of current LLMs is unsatisfactory, with the best model achieving only 36% accuracy. To bridge this gap, we propose LLM4Graph datasets, which include crawled documents and auto-generated codes based on 6 widely used graph libraries. By augmenting closed-source LLMs with document retrieval and fine-tuning open-source ones on the codes, we show 11-32% absolute improvements in their accuracies. Our results underscore that the capabilities of LLMs in handling structured data are still under-explored, and show the effectiveness of LLM4Graph in enhancing LLMs’ proficiency of graph analysis. The benchmark, datasets and enhanced open-sourcemodels are available at https://github.com/BUPT-GAMMA/ProGraph.
Xin Li 0005, Weize Chen, Qizhi Chu, Zhaojun Sun, Chuan Shi 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Cheng Yang 0002
NeurIPS1
2024 CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos
abstract
Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics in video sequences. In this paper, we introduce the new AeroEye dataset that focuses on multi-object relationship modeling in aerial videos. Our AeroEye dataset features various drone scenes and includes a visually comprehensive and precise collection of predicates that capture the intricate relationships and spatial arrangements among objects. To this end, we propose the novel Cyclic Graph Transformer (CYCLO) approach that allows the model to capture both direct and long-range temporal dependencies by continuously updating the history of interactions in a circular manner. The proposed approach also allows one to handle sequences with inherent cyclical patterns and process object relationships in the correct sequential order. Therefore, it can effectively capture periodic and overlapping relationships while minimizing information loss. The extensive experiments on the AeroEye dataset demonstrate the effectiveness of the proposed CYCLO model, demonstrating its potential to perform scene understanding on drone videos. Finally, the CYCLO method consistently achieves State-of-the-Art (SOTA) results on two in-the-wild scene graph generation benchmarks, i.e., PVSG and ASPIRe.
Pha A. Nguyen, Xin Li 0005, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu
NeurIPS3
2024 MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture Search
abstract
Neural Architecture Search (NAS) methods seek effective optimization toward performance metrics regarding model accuracy and generalization while facing challenges regarding search costs and GPU resources. Recent Neural Tangent Kernel (NTK) NAS methods achieve remarkable search efficiency based on a training-free model estimate; however, they overlook the non-convex nature of the DNNs in the search process. In this paper, we develop Multi-Objective Training-based Estimate (MOTE) for efficient NAS, retaining search effectiveness and achieving the new state-of-the-art in the accuracy and cost trade-off. To improve NTK and inspired by the Training Speed Estimation (TSE) method, MOTE is designed to model the actual performance of DNNs from macro to micro perspective by draw loss landscape and convergence speed simultaneously. Using two reduction strategies, the MOTE is generated based on a reduced architecture and a reduced dataset. Inspired by evolutionary search, our iterative ranking-based, coarse-to-fine architecture search is highly effective. Experiments on NASBench-201 show MOTE-NAS achieves 94.32% accuracy on CIFAR-10, 72.81% on CIFAR-100, and 46.38% on ImageNet-16-120, outperforming NTK-based NAS approaches. An evaluation-free (EF) version of MOTE-NAS delivers high efficiency in only 5 minutes, delivering a model more accurate than KNAS.
Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan
NeurIPS3
2024 Joint fuzzy background and adaptive foreground model for moving target detection
Yongfeng Dong, Linhao Li, Xin Li 0005
Frontiers Comput. Sci.5
2024 Knowledge-prompted ChatGPT: Enhancing drug trafficking detection on social media
Chuanbo Hu, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001, Minglei Yin
Inf. Manag.3
2024 TSUDepth: Exploring temporal symmetry-based uncertainty for unsupervised monocular depth estimation
Yufan Zhu, Weisheng Dong, Xin Li 0005, Guangming Shi
Neurocomputing4
2024 Robust ensemble person reidentification via orthogonal fusion with occlusion handling
Syeda Nyma Ferdous, Xin Li 0005
Image Vis. Comput.2
2024 Point Cloud Attacks in Graph Spectral Domain: When 3D Geometry Meets Graph Signal Processing
abstract
With the increasing attention in various 3D safety-critical applications, point cloud learning models have been shown to be vulnerable to adversarial attacks. Although existing 3D attack methods achieve high success rates, they delve into the data space with point-wise perturbation, which may neglect the geometric characteristics. Instead, we propose point cloud attacks from a new perspective-the graph spectral domain attack, aiming to perturb graph transform coefficients in the spectral domain that correspond to varying certain geometric structures. Specifically, leveraging on graph signal processing, we first adaptively transform the coordinates of points onto the spectral domain via graph Fourier transform (GFT) for compact representation. Then, we analyze the influence of different spectral bands on the geometric structure, based on which we propose to perturb the GFT coefficients via a learnable graph spectral filter. Considering the low-frequency components mainly contribute to the rough shape of the 3D object, we further introduce a low-frequency constraint to limit perturbations within imperceptible high-frequency components. Finally, the adversarial point cloud is generated by transforming the perturbed spectral representation back to the data domain via the inverse GFT. Experimental results demonstrate the effectiveness of the proposed attack in terms of both the imperceptibility and attack success rates.
Daizong Liu, Wei Hu 0003, Xin Li 0005
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Learning real-world heterogeneous noise models with a benchmark dataset
Jie Lin 0008, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
Pattern Recognit.4
2024 Abductive natural language inference by interactive model with structural loss
Linhao Li, Yongfeng Dong, Xin Li 0005
Pattern Recognit. Lett.5
2024 Uncertainty Modeling of the Transmission Map for Single Image Dehazing
abstract
Despite rapid progress of end-to-end optimization for single-image dehazing, a long-standing open problem is the non-homogenous haze, at the core of the differences between synthetic hazy images and real hazy images. The atmospheric scattering model (ASM) has been widely adopted to model the degradation process of haze images but based on the assumption of homogeneous haze. In realistic scenarios, non-homogeneous haze often makes it more difficult to estimate the transmission map in ASM, resulting in undesired artifacts in the restored images. To address the issue of non-homogeneous haze, we propose to model the uncertainty in the estimation of the transmission map and develop a spatially adaptive learning module for ASM correction. Specifically, we present an approach to enhancing the well-known Dark Channel prior (DCP) by relaxing the constraint with the transmission map in the DCP-net. Assuming the availability of paired training data, we have developed a strategy to address vulnerability in the DCP, leading to a more accurate estimation of the transmission map. Then, we explore the uncertainty between the estimated transmission map and target transmission map (Ground Truth) to reformulate the ASM for the presence of non-homogeneous haze. A robust and accurate estimated transmission map can boost the final dehazing performance of our DCP-net. Experiments on three popular synthetic and real non-homogeneous datasets show that our proposed approach has achieved better results on both synthetic scenes and real non-homogeneous scenes. The code is available athttps://see.xidian.edu.cn/faculty/wsdong/Projects/Projects/project_dehazing_TCSVT2024.htm
Bokang Wang, Qian Ning, Xin Li 0005, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.4
2024 Space Decoupled Prototype Learning for Few-Shot Attack Detection in Cyber-Physical Systems
abstract
Due to the lack of effective attack detection measures, cyberattacks may cause strong damage to industrial cyber–physical systems (CPSs). The embedding of attack categories learned by the existing attack detection methods is highly coupled to each other with fuzzy boundaries and overlapped neighborhood, leading to weak robustness and high false positive rates. To address these issues, in this article, we propose a few-shot attack detection method based on decoupled prototype learning (DPL-FSAD), aiming to enhance the detection accuracy and generalization capabilities for malicious attacks in CPS. Specifically, we first introduce feature contrastive learning to extract differentiated features from highly similar samples, achieving compact intraclass and sparse interclass feature embedding space. To solve the problem of fuzzy boundaries of different attack categories, prototype contrastive learning is then employed to reduce the coupling degree among prototypes and enhance their discriminability. A regularization term is exploited to mitigate the overfitting problem by reducing the gap between the feature embedding and prototypes. Furthermore, an orthogonal constraint is employed to separate prototypes of different attack types, generating a decoupled prototype embedding space. The experimental results on three public cyberattack datasets show that, compared with the suboptimal model a few-shot learning model with Siamese convolutional neural network (FSL-SCNN), the proposed DPL-FSAD can improve the precision by 5.53%,F1-score by 3.3%, and reduce the false positive rate by 2.37% in average, which proves that the space decoupled prototype learning is effective for improving the generalization and robustness of industrial CPS attack detection in few-shot scenario.
Haili Sun, Yan Huang 0026, Chunjie Zhou, Lansheng Han, Hongle Liu, Xin Li 0005
IEEE Trans. Ind. Informatics7
2024 TransVQA: Transferable Vector Quantization Alignment for Unsupervised Domain Adaptation
abstract
Unsupervised Domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain. Most existing domain adaptation methods are based on convolutional neural networks (CNNs) to learn cross-domain invariant features. Inspired by the success of transformer architectures and their superiority to CNNs, we propose to combine the transformer with UDA to improve their generalization properties. In this paper, we present a novel model named Trans ferable V ector Q uantization A lignment for Unsupervised Domain Adaptation (TransVQA), which integrates the Transferable transformer-based feature extractor (Trans), vector quantization domain alignment (VQA), and mutual information weighted maximization confusion matrix (MIMC) of intra-class discrimination into a unified domain adaptation framework. First, TransVQA uses the transformer to extract more accurate features in different domains for classification. Second, TransVQA, based on the vector quantization alignment module, uses a two-step alignment method to align the extracted cross-domain features and solve the domain shift problem. The two-step alignment includes global alignment via vector quantization and intra-class local alignment via pseudo-labels. Third, for intra-class feature discrimination problem caused by the fuzzy alignment of different domains, we use the MIMC module to constrain the target domain output and increase the accuracy of pseudo-labels. The experiments on several datasets of domain adaptation show that TransVQA can achieve excellent performance and outperform existing state-of-the-art methods.
Weisheng Dong, Xin Li 0005, Guangming Shi, Xuemei Xie
IEEE Trans. Image Process.3
2024 Multi-Scale Spatio-Temporal Memory Network for Lightweight Video Denoising
abstract
Deep learning-based video denoising methods have achieved great performance improvements in recent years. However, the expensive computational cost arising from sophisticated network design has severely limited their applications in real-world scenarios. To address this practical weakness, we propose a multiscale spatio-temporal memory network for fast video denoising, named MSTMN, aiming at striking an improved trade-off between cost and performance. To develop an efficient and effective algorithm for video denoising, we exploit a multiscale representation based on the Gaussian-Laplacian pyramid decomposition so that the reference frame can be restored in a coarse-to-fine manner. Guided by a model-based optimization approach, we design an effective variance estimation module, an alignment error estimation module and an adaptive fusion module for each scale of the pyramid representation. For the fusion module, we employ a reconstruction recurrence strategy to incorporate local temporal information. Moreover, we propose a memory enhancement module to exploit the global spatio-temporal information. Meanwhile, the similarity computation of the spatio-temporal memory network enables the proposed network to adaptively search the valuable information at the patch level, which avoids computationally expensive motion estimation and compensation operations. Experimental results on real-world raw video datasets have demonstrated that the proposed lightweight network outperforms current state-of-the-art fast video denoising algorithms such as FastDVDnet, EMVD, and ReMoNet with fewer computational costs.
Xin Li 0005, Jie Lin 0008, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.4
2024 Robust Geometry-Dependent Attack for 3D Point Clouds
abstract
Deep learning models for point clouds have shown to be vulnerable to adversarial attacks, which have received increasing attention in various safety-critical applications such as autonomous driving, robotics, and surveillance. Since existing 3D attack methods either modify the local points or perform global point-wise perturbations over the point cloud, they fail to capture the dependency between neighboring points for preserving the geometrical context and topological smoothness of the original 3D object. In this article, we propose a novel Geometry-Dependent Attack (GDA), which aims to generate more robust adversarial point clouds with lower perturbation costs by capturing and preserving the geometry-guided topology information. Specifically, we first analyze the geometric information of each benign point cloud following the graph signal processing and disentangle it into low-frequency (flat) and high-frequency (contour) components. Then, considering the varying characteristics of smoothness and sharpness after disentanglement, we design two collaborative patch-aware and point-aware attacks to perturb these two components separately to misclassify the 3D object. We test the proposed GDA attack using five popular point cloud networks (PointNet, PointNet++, DGCNN, PointTransformer, and PointMLP) on both ModelNet40 and ShapNetPart datasets. Experimental results show that our GDA attack achieves 100% success rates with the lowest perturbation cost. It also demonstrates the increased capability to defeat several existing defense models over other competing attacks.
Daizong Liu, Wei Hu 0003, Xin Li 0005
IEEE Trans. Multim.3
2024 MMGInpainting: Multi-Modality Guided Image Inpainting Based on Diffusion Models
abstract
Proper inference of semantics is necessary for realistic image inpainting. Most image inpainting methods use deep generative models, which require large image datasets to predict and generate content. However, predicting the missing regions and generating coherent content is difficult due to limited control. Existing approaches include image-guided or text-guided image inpainting, but none of them has taken both image and text as the guidance signals, as far as we know. To fill this gap, we propose a multi-modality guided (MMG) image inpainting approach based on the diffusion model. This MMGInpainting method uses both image and text as guidance for generating content within the target area for inpainting, effectively integrating the semantic information conveyed by the guiding image or text into the content of the inpainted region. To construct MMGInpainting, we start by enhancing the U-Net backbone with a customized Nonlinear Activation Free Network (NAFNet). This adapted NAFNet incorporates anAnchored Stripe Attentionmechanism, which utilizes anchor points to effectively model global contextual dependencies. To regulate inpainting, we use a Semantic Fusion Encoder to guide the inverse process of the diffusion model. The process is iteratively executed to denoise and generate the desired inpainting result. Additionally, we explore how different modes of meaning interact and coordinate to offer users useful guidance for a more manageable inpainting procedure. Experimental results demonstrate that our approach produces faithful results adhering to the guiding information, while significantly improving computational efficiency. Github Repository:https://github.com/skipper-zc/MMGInpainting/
Cong Zhang 0019, Wenxia Yang, Xin Li 0005, Huan Han
IEEE Trans. Multim.3
2023 Self-supervised Non-uniform Kernel Estimation with Flow-based Motion Prior for Blind Image Deblurring
abstract
Many deep learning-based solutions to blind image deblurring estimate the blur representation and reconstruct the target image from its blurry observation. However, these methods suffer from severe performance degradation in real-world scenarios because they ignore important prior information about motion blur (e.g., real-world motion blur is diverse and spatially varying). Some methods have attempted to explicitly estimate non-uniform blur kernels by CNNs, but accurate estimation is still challenging due to the lack of ground truth about spatially varying blur kernels in real-world images. To address these issues, we propose to represent the field of motion blur kernels in a latent space by normalizing flows, and design CNNs to predict the latent codes instead of motion kernels. To further improve the accuracy and robustness of non-uniform kernel estimation, we introduce uncertainty learning into the process of estimating latent codes and propose a multi-scale kernel attention module to better integrate image features with estimated kernels. Extensive experimental results, especially on real-world blur datasets, demonstrate that our method achieves state-of-the-art results in terms of both subjective and objective quality as well as excellent generalization performance for non-uniform image deblurring. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UFPNet.htm.
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
CVPR4
2023 Micron-BERT: BERT-Based Facial Micro-Expression Recognition
abstract
Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT ($(\mu$-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed$\mu$-BERT significantly outperforms all previous work in various micro-expression tasks.$\mu$-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show$\mu$-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT
Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li 0005, Susan Gauch, Han-Seok Seo, Khoa Luu
CVPR3
2023 Vector Quantization with Self-Attention for Quality-Independent Representation Learning
abstract
Recently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn invariant representations from multiple corrupted data through data augmentation or image-pair-based feature distillation to improve the robustness. Inspired by sparse representation in image restoration, we opt to address this issue by learning image-quality-independent feature representation in a simple plug-and-play manner, that is, to introduce discrete vector quantization (VQ) to remove redundancy in recognition models. Specifically, we first add a codebook module to the network to quantize deep features. Then we concatenate them and design a self-attention module to enhance the representation. During training, we enforce the quantization of features from clean and corrupted images in the same discrete embedding space so that an invariant quality-independentfeature representation can be learned to improve the recognition robustness of low-quality images. Qualitative and quantitative experimental results show that our method achieved this goal effectively, leading to a new state-of-the-art result of 43.1 % mCE on ImageNet-C with ResNet50 as the backbone. On other robustness benchmark datasets, such as ImageNet-R, our method also has an accuracy improvement of almost 2%. The source code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/VQSA.htm
Weisheng Dong, Xin Li 0005, Mengluan Huang, Guangming Shi
CVPR3
2023 Low-Light Image Enhancement with Multi-stage Residue Quantization and Brightness-aware Attention
abstract
Low-light image enhancement (LLIE) aims to recover illumination and improve the visibility of low-light images. Conventional LLIE methods often produce poor results because they neglect the effect of noise interference. Deep learning-based LLIE methods focus on learning a mapping function between low-light images and normal-light images that outperforms conventional LLIE methods. However, most deep learning-based LLIE methods cannot yet fully exploit the guidance of auxiliary priors provided by normal-light images in the training dataset. In this paper, we propose a brightness-aware network with normal-light priors based on brightness-aware attention and residual-quantized codebook. To achieve a more natural and realistic enhancement, we design a query module to obtain more reliable normal-light features and fuse them with low-light features by a fusion branch. In addition, we propose a brightness-aware attention module to further improve the robustness of the network to the brightness. Extensive experimental results on both real-captured and synthetic data show that our method outperforms existing state-of-the-art methods.
Weisheng Dong, Xin Li 0005, Guangming Shi
ICCV5
2023 GAFlow: Incorporating Gaussian Attention into Optical Flow
abstract
Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and smoothness, which yet is not fully explored by existing approaches. In this paper, we push Gaussian Attention (GA) into the optical flow models to accentuate local properties during representation learning and enforce the motion affinity during matching. Specifically, we introduce a novel Gaussian-Constrained Layer (GCL) which can be easily plugged into existing Transformer blocks to highlight the local neighborhood that contains fine-grained structural information. Moreover, for reliable motion analysis, we provide a new Gaussian-Guided Attention Module (GGAM) which not only inherits properties from Gaussian distribution to instinctively revolve around the neighbor fields of each point but also is empowered to put the emphasis on contextually related regions during matching. Our fully-equipped model, namely Gaussian Attention Flow network (GAFlow), naturally incorporates a series of novel Gaussian-based modules into the conventional optical flow framework for reliable motion analysis. Extensive experiments on standard optical flow datasets consistently demonstrate the exceptional performance of the proposed approach in terms of both generalization ability evaluation and online benchmark testing. Code is available at https://github.com/LA30/GAFlow.
Ao Luo, Fan Yang 0054, Xin Li 0005, Lang Nie, Chunyu Lin, Haoqiang Fan, Shuaicheng Liu
ICCV3
2023 Constraining Depth Map Geometry for Multi-View Stereo: A Dual-Depth Approach with Saddle-shaped Depth Cells
abstract
Learning-based multi-view stereo (MVS) methods deal with predicting accurate depth maps to achieve an accurate and complete 3D representation. Despite the excellent performance, existing methods ignore the fact that a suitable depth geometry is also critical in MVS. In this paper, we demonstrate that different depth geometries have significant performance gaps, even using the same depth prediction error. Therefore, we introduce an ideal depth geometry composed of Saddle-Shaped Cells, whose predicted depth map oscillates upward and downward around the ground-truth surface, rather than maintaining a continuous and smooth depth plane. To achieve it, we develop a coarse-to-fine framework called Dual-MVSNet (DMVSNet), which can produce an oscillating depth plane. Technically, we predict two depth values for each pixel (Dual-Depth), and propose a novel loss function and a checkerboard-shaped selecting strategy to constrain the predicted depth geometry. Compared to existing methods, DMVSNet achieves a high rank on the DTU benchmark and obtains the top performance on challenging scenes of Tanks and Temples, demonstrating its strong performance and generalization ability. Our method also points to a new research direction for considering depth geometry in MVS.
Weiyue Zhao, Tianqi Liu 0003, Zihao Huang 0001, Zhiguo Cao 0001, Xin Li 0005
ICCV6
2023 Fast Full-frame Video Stabilization with Iterative Optimization
abstract
Video stabilization refers to the problem of transforming a shaky video into a visually pleasing one. The question of how to strike a good trade-off between visual quality and computational speed has remained one of the open challenges in video stabilization. Inspired by the analogy between wobbly frames and jigsaw puzzles, we propose an iterative optimization-based learning approach using synthetic datasets for video stabilization, which consists of two interacting submodules: motion trajectory smoothing and full-frame outpainting. First, we develop a two-level (coarse-to-fine) stabilizing algorithm based on the probabilistic flow field. The confidence map associated with the estimated optical flow is exploited to guide the search for shared regions through backpropagation. Second, we take a divide-and-conquer approach and propose a novel multi-frame fusion strategy to render full-frame stabilized views. An important new insight brought about by our iterative optimization approach is that the target video can be interpreted as the fixed point of nonlinear mapping for video stabilization. We formulate video stabilization as a problem of minimizing the amount of jerkiness in motion trajectories, which guarantees convergence with the help of fixed-point theory. Extensive experimental results are reported to demonstrate the superiority of the proposed approach in terms of computational speed and visual quality. The code will be available on GitHub.
Weiyue Zhao, Xin Li 0005, Xianrui Luo, Hao Lu 0003, Zhiguo Cao 0001
ICCV2
2023 STAP Performance Evaluation for Spaceborne Radar Systems with Different Clutter Distribution Models
abstract
As one of the main statistical characteristics of clutter, the clutter amplitude distribution plays an important role in the accurate modeling of spaceborne multi-channel radar signal as well as the subsequent, maritime radar target detection. In this paper, considering that the space-time adaptive processing (STAP) technology is usually applied to accomplish the main-lobe clutter rejection in a space-borne surveillance radar, the influence of different clutter amplitude distributions on STAP in a spaceborne multichannel radar system are analyzed. Firstly, a spaceborne multi-channel clutter model is established based on radar equation and clutter space-time steering vector. Then, Rayleigh distribution, Weibull distribution, lognormal distribution, K distribution, generalized Pareto distribution, and IG-CG distribution are used to fit the clutter amplitude. Finally, the effects of these clutter amplitude distributions on STAP performance are analyzed, respectively. The simulation results show that in the case of the same clutter power, the influence of different clutter amplitude distributions on STAP performance is approximately the same.
Fan Yang 0054, Penghui Huang, Xin Li 0005, Bingliang Zhang, Junli Chen, Peili Xi, Guozhong Chen, Xingzhao Liu
IGARSS3
2023 Exploring Correlations in Degraded Spatial Identity Features for Blind Face Restoration
abstract
Blind face restoration aims to recover high-quality face images from low-quality ones with complex and unknown degradation. Existing approaches have achieved promising performance by leveraging pre-trained dictionaries or generative priors. However, these methods may fail to exploit the full potential of degraded inputs and facial identity features due to complex degradation. To address this issue, we propose a novel method that explores the correlation of degraded spatial identity features by learning a general representation using memory network. Specifically, our approach enhances degraded features with more identity by leveraging similar facial features retrieved from memory network. We also propose a fusion approach that fuses memorized spatial features with GAN prior features via affine transformation and blending fusion to improve fidelity and realism. Additionally, the memory network is updated online in an unsupervised manner along with other modules, which obviates the requirement for pre-training. Experimental results on synthetic and popular real-world datasets demonstrate the effectiveness of our proposed method, which achieves at least comparable and often better performance than other state-of-the-art approaches.
Qian Ning, Weisheng Dong, Xin Li 0005, Guangming Shi
ACM Multimedia4
2023 Video-based Contrastive Learning on Decision Trees: from Action Recognition to Autism Diagnosis
abstract
How can we teach a computer to recognize 10,000 different actions? Deep learning has evolved from supervised and unsupervised to self-supervised approaches. In this paper, we present a new contrastive learning-based framework for decision tree-based classification of actions, including human-human interactions (HHI) and human-object interactions (HOI). The key idea is to translate the original multi-class action recognition into a series of binary classification tasks on a pre-constructed decision tree. Under the new framework of contrastive learning, we present the design of an interaction adjacent matrix (IAM) with skeleton graphs as the backbone for modeling various action-related attributes such as periodicity and symmetry. Through the construction of various pretext tasks, we obtain a series of binary classification nodes on the decision tree that can be combined to support higher-level recognition tasks. Experimental justification for the potential of our approach in real-world applications ranges from interaction recognition to symmetry detection. In particular, we have demonstrated the promising performance of video-based autism spectrum disorder (ASD) diagnosis on the CalTech interview video database.
Mindi Ruan, Xiangxu Yu, Chuanbo Hu, Shuo Wang 0016, Xin Li 0005
MMSys6
2023 Fine-grained classification of drug trafficking based on Instagram hashtags
Chuanbo Hu, Bin Liu 0045, Yanfang Ye 0001, Xin Li 0005
Decis. Support Syst.4
2023 Memory Based Temporal Fusion Network for Video Deblurring
Chaohua Wang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
Int. J. Comput. Vis.3
2023 A2B: Anchor to Barycentric Coordinate for Robust Correspondence
Weiyue Zhao, Hao Lu 0003, Zhiguo Cao 0001, Xin Li 0005
Int. J. Comput. Vis.4
2023 Adaptive Search-and-Training for Robust and Efficient Network Pruning
abstract
Both network pruning and neural architecture search (NAS) can be interpreted as techniques to automate the design and optimization of artificial neural networks. In this paper, we challenge the conventional wisdom of training before pruning by proposing a joint search-and-training approach to learn a compact network directly from scratch. Using pruning as a search strategy, we advocate three new insights for network engineering: 1) to formulate adaptive search as a cold start strategy to find a compact subnetwork on the coarse scale; and 2) to automatically learn the threshold for network pruning; 3) to offer flexibility to choose between efficiency and robustness. More specifically, we propose an adaptive search algorithm in the cold start by exploiting the randomness and flexibility of filter pruning. The weights associated with the network filters will be updated by ThreshNet, a flexible coarse-to-fine pruning method inspired by reinforcement learning. In addition, we introduce a robust pruning strategy leveraging the technique of knowledge distillation through a teacher-student network. Extensive experiments on ResNet and VGGNet have shown that our proposed method can achieve a better balance in terms of efficiency and accuracy and notable advantages over current state-of-the-art pruning methods in several popular datasets, including CIFAR10, CIFAR100, and ImageNet. The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/AST-NP.htm.
Xiaotong Lu, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Learning Probabilistic Coordinate Fields for Robust Correspondences
abstract
We introduce Probabilistic Coordinate Fields (PCFs), a novel geometric-invariant coordinate representation for image correspondence problems. In contrast to standard Cartesian coordinates, PCFs encode coordinates in correspondence-specific barycentric coordinate systems (BCS) with affine invariance. To know when and where to trust the encoded coordinates, we implement PCFs in a probabilistic network termed PCF-Net, which parameterizes the distribution of coordinate fields as Gaussian mixture models. By jointly optimizing coordinate fields and their confidence conditioned on dense flows, PCF-Net can work with various feature descriptors when quantifying the reliability of PCFs by confidence maps. An interesting observation of this work is that the learned confidence map converges to geometrically coherent and semantically consistent regions, which facilitates robust coordinate representation. By delivering the confident coordinates to keypoint/feature descriptors, we show that PCF-Net can be used as a plug-in to existing correspondence-dependent approaches. Extensive experiments on both indoor and outdoor datasets suggest that accurate geometric invariant coordinates help to achieve the state of the art in several correspondence problems, such as sparse feature matching, dense image registration, camera pose estimation, and consistency filtering. Further, the interpretable confidence map predicted by PCF-Net can also be leveraged to other novel applications from texture transfer to multi-homography classification.
Weiyue Zhao, Hao Lu 0003, Zhiguo Cao 0001, Xin Li 0005
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Discriminative Few Shot Learning of Facial Dynamics in Interview Videos for Autism Trait Classification
abstract
Autism is a prevalent neurodevelopmental disorder characterized by impairments in social and communicative behaviors. Possible connections between autism and facial expression recognition have recently been studied in the literature. However, most works are based on facial images or short videos. Few works aim at Autism Diagnostic Observation Schedule (ADOS) videos due to their complexity (e.g., interaction between interviewer and interviewee) and length (e.g., usually last for hours). In this paper, we attempt to fill this gap by developing a novel discriminative few shot learning method to analyze hour-long video data and exploring the fusion of facial dynamics for the trait classification of ASD. Leveraging well-established computer vision tools from spatio-temporal feature extraction and marginal fisher analysis to few-shot learning and scene-level fusion, we have constructed a three-category system to classify an individual into Autism, Autism Spectrum, and Non-Spectrum. For the first time, we have shown that certain interview scenes carry more discriminative information for ASD trait classification than others. Experimental results are reported to demonstrate the potential of the proposed automatic ASD trait classification system (achieving 91.72% accuracy on the Caltech ADOS video dataset) and the benefits of few-shot learning and scene-level fusion strategy by extensive ablation studies.
Mindi Ruan, Shuo Wang 0016, Lynn K. Paul, Xin Li 0005
IEEE Trans. Affect. Comput.5
2023 Uncertainty-Driven Knowledge Distillation for Language Model Compression
abstract
Despite the remarkable performance on various Natural Language Processing (NLP) tasks, the parametric complexity of pretrained language models has remained a major obstacle due to limited computational resources in many practical applications. Techniques such as knowledge distillation, network pruning, and quantization have been developed for language model compression. However, it has remained challenging to achieve an optimal tradeoff between model size and inference accuracy. To address this issue, we propose a novel and efficient uncertainty-driven knowledge distillation compression method for transformer-based pretrained language models. Specifically, we design a method of parameter retention and feedforward network parameter distillation to compress N-stacked transformer modules into one module in the fine-tuning stage. A key innovation of our approach is to add the uncertainty estimation module (UEM) into the student network such that it can guide the student network's feature reconstruction in the latent space (similar to the teacher's). Across multiple datasets in the natural language inference tasks of GLUE, we have achieved more than 95% accuracy of the original BERT, while only using about 50% of the parameters.
Weisheng Dong, Xin Li 0005, Guangming Shi
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Deep Unfolding Network for Efficient Mixed Video Noise Removal
abstract
Existing image and video denoising algorithms have focused on removing homogeneous Gaussian noise. However, this assumption with noise modeling is often too simplistic for the characteristics of real-world noise. Moreover, the design of network architectures in most deep learning-based video denoising methods is heuristic, ignoring valuable domain knowledge. In this paper, we propose a model-guided deep unfolding network for the more challenging and realistic mixed noise video denoising problem, named DU-MVDnet. First, we develop a novel observation model/likelihood function based on the correlations among adjacent degraded frames. In the framework of Bayesian deep learning, we introduce a deep image denoiser prior and obtain an iterative optimization algorithm based on the maximum a posterior (MAP) estimation. To facilitate end-to-end optimization, the iterative algorithm is transformed into a deep convolutional neural network (DCNN)-based implementation. Furthermore, recognizing the limitations of traditional motion estimation and compensation methods, we propose an efficient multistage recursive fusion strategy to exploit temporal dependencies. Specifically, we divide video frames into several overlapping groups and progressively integrate these frames into one frame. Toward this objective, we implement a multiframe adaptive aggregation operation to integrate feature maps of intragroup with those of intergroup frames. Extensive experimental results on different video test datasets have demonstrated that the proposed model-guided deep network outperforms current state-of-the-art video denoising algorithms such as FastDVDnet and MAP-VDNet.
Xin Li 0005, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.4
2023 Searching Efficient Model-Guided Deep Network for Image Denoising
abstract
Unlike the success of neural architecture search (NAS) in high-level vision tasks, it remains challenging to find computationally efficient and memory-efficient solutions to low-level vision problems such as image restoration through NAS. One of the fundamental barriers to differential NAS-based image restoration is the optimization gap between the super-network and the sub-architectures, causing instability during the searching process. In this paper, we present a novel approach to fill this gap in image denoising application by connecting model-guided design (MoD) with NAS (MoD-NAS). Specifically, we propose to construct a new search space under a model-guided framework and develop more stable and efficient differential search strategies. MoD-NAS employs a highly reusable width search strategy and a densely connected search block to automatically select the operations of each layer as well as network width and depth via gradient descent. During the search process, the proposed MoD-NAS remains stable because of the smoother search space designed under the model-guided framework. Experimental results on several popular datasets show that our MoD-NAS method has achieved at least comparable even better PSNR performance than current state-of-the-art methods with fewer parameters, fewer flops, and less testing time. "The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/Mod-NAS.htm".
Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu
IEEE Trans. Image Process.3
2023 Spatially Varying Prior Learning for Blind Hyperspectral Image Fusion
abstract
In recent years, researchers have become more interested in hyperspectral image fusion (HIF) as a potential alternative to expensive high-resolution hyperspectral imaging systems, which aims to recover a high-resolution hyperspectral image (HR-HSI) from two images obtained from low-resolution hyperspectral (LR-HSI) and high-spatial-resolution multispectral (HR-MSI). It is generally assumed that degeneration in both the spatial and spectral domains is known in traditional model-based methods or that there existed paired HR-LR training data in deep learning-based methods. However, such an assumption is often invalid in practice. Furthermore, most existing works, either introducing hand-crafted priors or treating HIF as a black-box problem, cannot take full advantage of the physical model. To address those issues, we propose a deep blind HIF method by unfolding model-based maximum a posterior (MAP) estimation into a network implementation in this paper. Our method works with a Laplace distribution (LD) prior that does not need paired training data. Moreover, we have developed an observation module to directly learn degeneration in the spatial domain from LR-HSI data, addressing the challenge of spatially-varying degradation. We also propose to learn the uncertainty (mean and variance) of LD models using a novel Swin-Transformer-based denoiser and to estimate the variance of degraded images from residual errors (rather than treating them as global scalars). All parameters of the MAP estimation algorithm and the observation module can be jointly optimized through end-to-end training. Extensive experiments on both synthetic and real datasets show that the proposed method outperforms existing competing methods in terms of both objective evaluation indexes and visual qualities.
Xin Li 0005, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.3
2022 Robust Depth Completion with Uncertainty-Driven Loss Functions
abstract
Recovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey's prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics.
Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li 0005, Guangming Shi
AAAI5
2022 DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition
abstract
Human action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video action recognition with competitive results. However, these methods have suffered some fundamental limitations such as lack of robustness and generalization, e.g., how does the temporal ordering of video frames affect the recognition results? This work presents a novel end-to-end Transformer-based Directed Attention (Direc-Former) framework11The implementation of DirecFormer is available at https://github.com/uark-cviu/DirecFormer for robust action recognition. The method takes a simple but novel perspective of Transformer-based approach to understand the right order of sequence actions. Therefore, the contributions of this work are three-fold. Firstly, we introduce the problem of ordered temporal learning issues to the action recognition problem. Secondly, a new Directed Attention mechanism is introduced to understand and provide attentions to human actions in the right order. Thirdly, we introduce the conditional dependency in action sequence modeling that includes orders and classes. The proposed approach consistently achieves the state-of-the-art (SOTA) results compared with the recent action recognition methods [4, 18, 72, 74]. on three standard large-scale benchmarks, i.e. Jester, Kinetics-400 and Something-Something-V2.
Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li 0005, Khoa Luu
CVPR6
2022 Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (18)3
2022 Self-feature Distillation with Uncertainty Modeling for Degraded Image Recognition
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (24)3
2022 Uncertainty Aware Multitask Pyramid Vision Transformer for UAV-Based Object Re-Identification
abstract
Object Re-IDentification (ReID), one of the most significant problems in biometrics and surveillance systems, has been extensively studied by image processing and computer vision communities in the past decades. Learning a robust and discriminative feature representation is a crucial challenge for object ReID. The problem is even more challenging in ReID based on Unmanned Aerial Vehicle (UAV) as the images are characterized by continuously varying camera parameters (e.g., view angle, altitude, etc.) of a flying drone. To address this challenge, multiscale feature representation has been considered to characterize images captured from UAV flying at different altitudes. In this work, we propose a multitask learning approach, which employs a new multiscale architecture without convolution, Pyramid Vision Transformer (PVT), as the backbone for UAV-based object ReID. By uncertainty modeling of intraclass variations, our proposed model can be jointly optimized using both uncertainty-aware object ID and camera ID information. Experimental results are reported on PRAI and VRAI, two ReID data sets from aerial surveillance, to verify the effectiveness of our proposed approach.
Syeda Nyma Ferdous, Xin Li 0005, Siwei Lyu
ICIP2
2022 Model Attribution of Face-Swap Deepfake Videos
abstract
AI-created face-swap videos, commonly known as Deepfakes, have attracted wide attention as powerful impersonation attacks. Existing research on Deepfakes mostly focuses on binary detection to distinguish between real and fake videos. However, it is also important to determine the specific generation model for a fake video, which can help attribute it to the source for forensic investigation. In this paper, we fill this gap by studying the model attribution problem of Deepfake videos. We first introduce a new dataset with DeepFakes from Different Models (DFDM) based on several Autoencoder models. Specifically, five generation models with variations in encoder, decoder, intermediate layer, input resolution, and compression ratio have been used to generate a total of 6, 450 Deepfake videos based on the same input. Then we take Deepfakes model attribution as a multiclass classification task and propose a spatial and temporal attention based method to explore the differences among Deep-fakes in the new dataset. Experimental evaluation shows that most existing Deepfakes detection methods failed in Deep-fakes model attribution, while the proposed method achieved over 70% accuracy on the high-quality DFDM dataset1.
Shan Jia, Xin Li 0005, Siwei Lyu
ICIP2
2022 Dependency-Aware Traffic Management for Configuring On-demand in Service Meshes
abstract
Service mesh is a promising micro-services architecture due to its excellent governance capabilities. Unlike traditional service invocation, configurations for governance need to be issued in the service mesh. However, we find that the control-plane traffic of governance is distributed in full by default, i.e., each service in the data plane receives all configurations. The vast majority of the configurations are redundant for a specific service. Hence, it is important and challenging to make the control plane aware of the calling relationships between services. In this paper, we propose a traffic management mechanism named DATM. Using this mechanism, the entire cluster can be dynamically controlled and services can be configured on demand. It is implemented through a dependency-aware controller and monitors. The controller first processes the information listened to by the monitors and then analyzes the connection between the metrics and the service requests through intelligent algorithms. Finally, the control traffic for regulating the control plane is generated. Our proposed mechanism is experimentally compared with the default strategy and existing work across a wide set of load scenarios in a testbed based on Istio service mesh and Kubernetes. Experimental results demonstrate that our mechanism can save the storage resources of a single agent by 40% to 60%, and the number of cluster updates can be greatly reduced. From the perspective of the whole cluster, the optimization results are even better.
Lin Wang 0015, Xin Li 0005, Ning Wang 0018, Hao Li 0030, Xiaolin Qin, Jie Wu 0001
ICPADS2
2022 Learning Degradation Uncertainty for Unsupervised Real-world Image Super-resolution
abstract
Acquiring degraded images with paired high-resolution (HR) images is often challenging, impeding the advance of image super-resolution in real-world applications. By generating realistic low-resolution (LR) images with degradation similar to that in real-world scenarios, simulated paired LR-HR data can be constructed for supervised training. However, most of the existing work ignores the degradation uncertainty of the generated realistic LR images, since only one LR image has been generated given an HR image. To address this weakness, we propose learning the degradation uncertainty of generated LR images and sampling multiple LR images from the learned LR image (mean) and degradation uncertainty (variance) and construct LR-HR pairs to train the super-resolution (SR) networks. Specifically, uncertainty can be learned by minimizing the proposed loss based on Kullback-Leibler (KL) divergence. Furthermore, the uncertainty in the feature domain is exploited by a novel perceptual loss; and we propose to calculate the adversarial loss from the gradient information in the SR stage for stable training performance and better visual quality. Experimental results on popular real-world datasets show that our proposed method has performed better than other unsupervised approaches.
Qian Ning, Jingzhu Tang, Weisheng Dong, Xin Li 0005, Guangming Shi
IJCAI5
2022 Robust Dynamic Background Modeling for Foreground Estimation
abstract
Separating the background and foreground components from video frames is important to many tasks in computer vision and multimedia. As of today, robust principal component analysis (RPCA) has shown highly promising performance with the assumption that the background is low-rank and the foreground is sparse. However, existing RPCA-based methods have overlooked the uncertainty that some parts of the background (e.g., moving leaves in a dynamic background) or even the whole background (e.g., camera jittering) can be moving, which violates the low-rank assumption. To address this issue, we propose a novel enhanced RPCA framework (called ERPCA) by robustly modeling the dynamic background. Different from traditional RPCA framework, the background is decomposed into a low-rank component and a sparse component in the proposed ERPCA framework. Specifically, the sparse parts including foreground and dynamic parts of the background are modeled by Gaussian scale mixture (GSM) model. Moreover, those sparse components are further constrained by temporal consistency using nonzeromeans Gaussian models; the correspondences between sparse pixels in adjacent frames are explored by optical flow. Experimental results on 40 real videos demonstrate the superiority of our proposed method, with better average results than current state-of-the-art foreground estimation methods.
Qian Ning, Weisheng Dong, Jinjian Wu, Guangming Shi, Xin Li 0005
VCIP6
2022 Correlation filters based on spatial-temporal Gaussion scale mixture modelling for visual tracking
Guangming Shi, Weisheng Dong, Tianzhu Zhang 0001, Jinjian Wu, Xuemei Xie, Xin Li 0005
Neurocomputing7
2022 Rotation invariant point cloud analysis: Where local geometry meets global topology
Chen Zhao 0025, Jiaqi Yang 0002, Angfan Zhu, Zhiguo Cao 0001, Xin Li 0005
Pattern Recognit.6
2022 Bayesian Correlation Filter Learning With Gaussian Scale Mixture Model for Visual Tracking
abstract
Correlation filters (CF), a popular tool for visual tracking, suffer from unwanted boundary effects due to the periodic assumption needed for FFT implementation. To address this issue, spatially regularized discriminative correlation filters (SRDCF) have been proposed by introducing a weighting matrix to the regularization term. However, the existing design of spatial weighting matrix is often heuristic and non-adaptive. Inspired by recent advances in joint discrimination and reliability learning for correlation tracking, we propose a principled Bayesian correlation filter learning method using Gaussian scale mixture (GSM) model. The key idea is to decompose each CF coefficient into the product of a positive scalar multiplier and a Gaussian random variable. Treating positive multipliers as weighting coefficients, GSM-based modeling of CFs leads to a spatially adaptive regularization strategy with improved capability of handling various appearance-related uncertainty factors (e.g., scale variation, out-of-plane rotation, and motion blur). Moreover, by imposing a sparse prior over the multipliers, we can jointly learn multipliers and CFs under a unified Bayesian estimation framework. Structured GSM model allows us to better exploit the spatial correlations among CFs and further improve the tracking performance. Experimental results on OTB-2013, OTB-2015, Temple Color-128, VOT-2016, and VOT-2017 show that our tracking method performs favorably when compared with current state-of-the-art methods.
Guangming Shi, Tianzhu Zhang 0001, Weisheng Dong, Jinjian Wu, Xuemei Xie, Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.7
2021 Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach
abstract
Social media such as Instagram and Twitter have become important platforms for marketing and selling illicit drugs. Detection of online illicit drug trafficking has become critical to combat the online trade of illicit drugs. However, the legal status often varies spatially and temporally; even for the same drug, federal and state legislation can have different regulations about its legality. Meanwhile, more drug trafficking events are disguised as a novel form of advertising - commenting leading to information heterogeneity. Accordingly, accurate detection of illicit drug trafficking events (IDTEs) from social media has become even more challenging. In this work, we conduct the first systematic study on fine-grained detection of IDTEs on Instagram. We propose to take a deep multimodal multilabel learning (DMML) approach to detect IDTEs and demonstrate its effectiveness on a newly constructed dataset called multimodal IDTE (MM-IDTE). Specifically, our model takes text and image data as the input and combines multimodal information to predict multiple labels of illicit drugs. Inspired by the success of BERT, we have developed a self-supervised multimodal bidirectional transformer by jointly fine-tuning pretrained text and image encoders. We have constructed a large-scale dataset MM-IDTE with manually annotated multiple drug labels to support fine-grained detection of illicit drugs. Extensive experimental results on the MM-IDTE dataset show that the proposed DMML methodology can accurately detect IDTEs even in the presence of special characters and style changes attempting to evade detection.
Chuanbo Hu, Minglei Yin, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001
CIKM4
2021 Self-Contrastive Learning with Hard Negative Sampling for Self-supervised Point Cloud Learning
abstract
Point clouds have attracted increasing attention. Significant progress has been made in methods for point cloud analysis, which often requires costly human annotation as supervision. To address this issue, we propose a novel self-contrastive learning for self-supervised point cloud representation learning, aiming to capture both local geometric patterns and nonlocal semantic primitives based on the nonlocal self-similarity of point clouds. The contributions are two-fold: on the one hand, instead of contrasting among different point clouds as commonly employed in contrastive learning, we exploit self-similar point cloud patches within a single point cloud as positive samples and otherwise negative ones to facilitate the task of contrastive learning. On the other hand, we actively learn hard negative samples that are close to positive samples for discriminative feature learning, which are sampled conditional on each anchor patch leveraging on the degree of self-similarity. Experimental results show that the proposed method achieves state-of-the-art performance on widely used benchmark datasets for self-supervised point cloud segmentation and transfer learning for classification.
Bi'an Du, Xiang Gao 0014, Wei Hu 0003, Xin Li 0005
ACM Multimedia4
2021 Uncertainty-Driven Loss for Single Image Super-Resolution
abstract
In low-level vision such as single image super-resolution (SISR), traditional MSE or L1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual information than smooth areas in photographic images. How to achieve such spatial adaptation in a principled manner has been an open problem in both traditional model-based and modern learning-based approaches toward SISR. In this paper, we propose a new adaptive weighted loss for SISR to train deep networks focusing on challenging situations such as textured and edge pixels with high uncertainty. Specifically, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis into SISR solutions so the targeted pixels in a high-resolution image (mean) and their corresponding uncertainty (variance) can be learned simultaneously. Moreover, uncertainty estimation allows us to leverage conventional wisdom such as sparsity prior for regularizing SISR solutions. Ultimately, pixels with large certainty (e.g., texture and edge pixels) will be prioritized for SISR according to their importance to visual quality. For the first time, we demonstrate that such uncertainty-driven loss can achieve better results than MSE or L1 loss for a wide range of network architectures. Experimental results on three popular SISR networks show that our proposed uncertainty-driven loss has achieved better PSNR performance than traditional loss functions without any increased computation during testing. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UDL-SR.htm
Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
NeurIPS3
2021 Deep Maximum a Posterior Estimator for Video Denoising
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
Int. J. Comput. Vis.3
2021 Toward blind joint demosaicing and denoising of raw color filter array data
Weisheng Dong, Guangming Shi, Zhonglong Zheng, Xin Li 0005
Neurocomputing6
2021 Face spoofing detection under super-realistic 3D wax face attacks
Shan Jia, Chuanbo Hu, Xin Li 0005, Zhengquan Xu
Pattern Recognit. Lett.3
2021 Robust subspace clustering network with dual-domain regularization
Guangming Shi, Xin Li 0005, Weisheng Dong, Jinjian Wu
Pattern Recognit. Lett.4
2021 Hybrid sparsity learning for image restoration: An iterative and trainable approach
Weisheng Dong, Guangming Shi, Shaoyuan Cheng, Xin Li 0005
Signal Process.6
2021 3D Face Anti-Spoofing With Factorized Bilinear Coding
abstract
We have witnessed rapid advances in both face presentation attack models and presentation attack detection (PAD) in recent years. When compared with widely studied 2D face presentation attacks, 3D face spoofing attacks are more challenging because face recognition systems are more easily confused by the 3D characteristics of materials similar to real faces. In this work, we tackle the problem of detecting these realistic 3D face presentation attacks and propose a novel anti-spoofing method from the perspective of fine-grained classification. Our method, based on factorized bilinear coding of multiple color channels (namely MC_FBC), targets at learning subtle fine-grained differences between real and fake images. By extracting discriminative and fusing complementary information from RGB and YCbCr spaces, we have developed a principled solution to 3D face spoofing detection. A large-scale wax figure face database (WFFD) with both images and videos has also been collected as super realistic attacks to facilitate the study of 3D face presentation attack detection. Extensive experimental results show that our proposed method achieves the state-of-the-art performance on both our own WFFD and other face spoofing databases under various intra-database and inter-database testing scenarios.
Shan Jia, Xin Li 0005, Chuanbo Hu, Guodong Guo, Zhengquan Xu
IEEE Trans. Circuits Syst. Video Technol.2
2021 Model-Guided Deep Hyperspectral Image Super-Resolution
abstract
The trade-off between spatial and spectral resolution is one of the fundamental issues in hyperspectral images (HSI). Given the challenges of directly acquiring high-resolution hyperspectral images (HR-HSI), a compromised solution is to fuse a pair of images: one has high-resolution (HR) in the spatial domain but low-resolution (LR) in spectral-domain and the other vice versa. Model-based image fusion methods including pan-sharpening aim at reconstructing HR-HSI by solving manually designed objective functions. However, such hand-crafted prior often leads to inevitable performance degradation due to a lack of end-to-end optimization. Although several deep learning-based methods have been proposed for hyperspectral pan-sharpening, HR-HSI related domain knowledge has not been fully exploited, leaving room for further improvement. In this paper, we propose an iterative Hyperspectral Image Super-Resolution (HSISR) algorithm based on a deep HSI denoiser to leverage both domain knowledge likelihood and deep image prior. By taking the observation matrix of HSI into account during the end-to-end optimization, we show how to unfold an iterative HSISR algorithm into a novel model-guided deep convolutional network (MoG-DCN). The representation of the observation matrix by subnetworks also allows the unfolded deep HSISR network to work with different HSI situations, which enhances the flexibility of MoG-DCN. Extensive experimental results are reported to demonstrate that the proposed MoG-DCN outperforms several leading HSISR methods in terms of both implementation cost and visual quality. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/MoG-DCN.htm.
Weisheng Dong, Chen Zhou 0005, Jinjian Wu, Guangming Shi, Xin Li 0005
IEEE Trans. Image Process.6
2021 Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data Fusion
abstract
Illicit drug trafficking via social media sites such as Instagram have become a severe problem, thus drawing a great deal of attention from law enforcement and public health agencies. How to identify illicit drug dealers from social media data has remained a technical challenge for the following reasons. On the one hand, the available data are limited because of privacy concerns with crawling social media sites; on the other hand, the diversity of drug dealing patterns makes it difficult to reliably distinguish drug dealers from common drug users. Unlike existing methods that focus on posting-based detection, we propose to tackle the problem of illicit drug dealer identification by constructing a large-scale multimodal dataset named Identifying Drug Dealers on Instagram (IDDIG). Nearly 4,000 user accounts, of which more than 1,400 are drug dealers, have been collected from Instagram with multiple data sources including post comments, post images, homepage bio, and homepage images. We then design a quadruple-based multimodal fusion method to combine the multiple data sources associated with each user account for drug dealer identification. Experimental results on the constructed IDDIG dataset demonstrate the effectiveness of the proposed method in identifying drug dealers (almost 95% accuracy). Moreover, we have developed a hashtag-based community detection technique for discovering evolving patterns, especially those related to geography and drug types.
Chuanbo Hu, Minglei Yin, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001
ACM Trans. Intell. Syst. Technol.4
2020 Learning 3D Faces from Photo-Realistic Facial Synthesis
abstract
We present an approach to efficiently learn an accurate and complete 3D face model from a single image. Previous methods heavily rely on 3D Morphable Models to populate the facial shape space as well as an over-simplified shading model for image formulation. By contrast, our method directly augments a large set of 3D faces from a compact collection of facial scans and employs a high-quality rendering engine to synthesize the corresponding photo-realistic facial images. We first use a deep neural network to regress vertex coordinates from the given image and then refine them by a non-rigid deformation process to more accurately capture local shape similarity. We have conducted extensive experiments to demonstrate the superiority of the proposed approach on 2D-to-3D facial shape inference, especially its excellent generalization property on real-world selfie images.
Ruizhe Wang 0004, Chih-Fan Chen, Hao Peng 0016, Xudong Liu 0006, Xin Li 0005
3DV5
2020 dStyle-GAN: Generative Adversarial Network based on Writing and Photography Styles for Drug Identification in Darknet Markets
abstract
Despite the persistent effort by law enforcement, illicit drug trafficking in darknet markets has shown great resilience with new markets rapidly appearing after old ones being shut down. In order to more effectively detect, disrupt and dismantle illicit drug trades, there’s an imminent need to gain a deeper understanding toward the operations and dynamics of illicit drug trading activities. To address this challenge, in this paper, we design and develop an intelligent system (named dSytle-GAN) to automate the analysis for drug identification in darknet markets, by considering both content-based and style-aware information. To determine whether a given pair of posted drugs are the same or not, in dStyle-GAN, based on the large-scale data collected from darknet markets, we first present an attributed heterogeneous information network (AHIN) to depict drugs, vendors, texts and writing styles, photos and photography styles, and the rich relations among them; and then we propose a novel generative adversarial network (GAN) based model over AHIN to capture the underlying distribution of posted drugs’ writing and photography styles to learn robust representations of drugs for their identifications. Unlike existing approaches, our proposed GAN-based model jointly considers the heterogeneity of network and relatedness over drugs formulated by domain-specific meta-paths for robust node (i.e., drug) representation learning. To the best of our knowledge, the proposed dStyle-GAN represents the first principled GAN-based solution over graphs to simultaneously consider writing and photography styles as well as their latent distributions for node representation learning. Extensive experimental results based on large-scale datasets collected from six darknet markets and the obtained ground-truth demonstrate that dStyle-GAN outperforms the state-of-the-art methods. Based on the identified drug pairs in the wild by dStyle-GAN, we perform further analysis to gain deeper insights into the dynamics and evolution of illicit drug trading activities in darknet markets, whose findings may facilitate law enforcement for proactive interventions.
Yiming Zhang 0002, Yiyue Qian, Yujie Fan, Yanfang Ye 0001, Xin Li 0005, Fudong Shao
ACSAC5
2020 Sparse-to-Dense Depth Completion Revisited: Sampling Strategy and Graph Construction
Haipeng Xiong, Ke Xian, Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005
ECCV (21)6
2020 Beyond Network Pruning: a Joint Search-and-Training Approach
abstract
Network pruning has been proposed as a remedy for alleviating the over-parameterization problem of deep neural networks. However, its value has been recently challenged especially from the perspective of neural architecture search (NAS). We challenge the conventional wisdom of pruning-after-training by proposing a joint search-and-training approach that directly learns a compact network from the scratch. By treating pruning as a search strategy, we present two new insights in this paper: 1) it is possible to expand the search space of networking pruning by associating each filter with a learnable weight; 2) joint search-and-training can be conducted iteratively to maximize the learning efficiency. More specifically, we propose a coarse-to-fine tuning strategy to iteratively sample and update compact sub-network to approximate the target network. The weights associated with network filters will be accordingly updated by joint search-and-training to reflect learned knowledge in NAS space. Moreover, we introduce strategies of random perturbation (inspired by Monte Carlo) and flexible thresholding (inspired by Reinforcement Learning) to adjust the weight and size of each layer. Extensive experiments on ResNet and VGGNet demonstrate the superior performance of our proposed method on popular datasets including CIFAR10, CIFAR100 and ImageNet.
Xiaotong Lu, Weisheng Dong, Xin Li 0005, Guangming Shi
IJCAI4
2020 Image Feature Correspondence Selection: A Comparative Study and a New Contribution
abstract
Image feature correspondence selection is pivotal to many computer vision tasks from object recognition to 3D reconstruction. Although many correspondence selection algorithms have been developed in the past decade, there still lacks an in-depth evaluation and comparison in the open literature, which makes it difficult to choose the appropriate algorithm for a specific application. This paper attempts to fill this gap by evaluating eight competing correspondence selection algorithms including both classical methods and current state-of-the-art ones. In addition to preselected correspondences, we have compared different combinations of detector and descriptor on four standard datasets. The diversity of those datasets cover a wide range of uncertainty factors including zoom, rotation, blur, viewpoint change, JPEG compression, light change, different rendering styles and multiple structures. We have measured the quality of competing correspondence selection algorithms in terms of four performance metrics -i.e., precision, recall, F-measure and efficiency. Moreover, we propose to combine the strengths of eight competing methods by combining their correspondence selection results. Extensive experimental results are reported to demonstrate the superiority of several fusion strategies to individual methods, which suggests the possibility of adaptively combining those methods for even better performance.
Chen Zhao 0025, Zhiguo Cao 0001, Jiaqi Yang 0002, Ke Xian, Xin Li 0005
IEEE Trans. Image Process.5
2019 NM-Net: Mining Reliable Neighbors for Robust Feature Correspondences
abstract
Feature correspondence selection is pivotal to many feature-matching based tasks in computer vision. Searching spatially k-nearest neighbors is a common strategy for extracting local information in many previous works. However, there is no guarantee that the spatially k-nearest neighbors of correspondences are consistent because the spatial distribution of false correspondences is often irregular. To address this issue, we present a compatibility-specific mining method to search for consistent neighbors. Moreover, in order to extract and aggregate more reliable features from neighbors, we propose a hierarchical network named NM-Net with a series of graph convolutions that is insensitive to the order of correspondences. Our experimental results have shown the proposed method achieves the state-of-the-art performance on four datasets with various inlier ratios and varying numbers of feature consistencies.
Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005, Jiaqi Yang 0002
CVPR4
2019 Your Style Your Identity: Leveraging Writing and Photography Styles for Drug Trafficker Identification in Darknet Markets over Attributed Heterogeneous Information Network
abstract
Due to its anonymity, there has been a dramatic growth of underground drug markets hosted in the darknet (e.g., Dream Market and Valhalla). To combat drug trafficking (a.k.a. illicit drug trading) in the cyberspace, there is an urgent need for automatic analysis of participants in darknet markets. However, one of the key challenges is that drug traffickers (i.e., vendors) may maintain multiple accounts across different markets or within the same market. To address this issue, in this paper, we propose and develop an intelligent system named uStyle-uID leveraging both writing and photography styles for drug trafficker identification at the first attempt. At the core of uStyle-uID is an attributed heterogeneous information network (AHIN) which elegantly integrates both writing and photography styles along with the text and photo contents, as well as other supporting attributes (i.e., trafficker and drug information) and various kinds of relations. Built on the constructed AHIN, to efficiently measure the relatedness over nodes (i.e., traffickers) in the constructed AHIN, we propose a new network embedding model Vendor2Vec to learn the low-dimensional representations for the nodes in AHIN, which leverages complementary attribute information attached in the nodes to guide the meta-path based random walk for path instances sampling. After that, we devise a learning model named vIdentifier to classify if a given pair of traffickers are the same individual. Comprehensive experiments on the data collections from four different darknet markets are conducted to validate the effectiveness of uStyle-uID which integrates our proposed method in drug trafficker identification by comparisons with alternative approaches.
Yiming Zhang 0002, Yujie Fan, Shifu Hou, Yanfang Ye 0001, Xin Li 0005, Liang Zhao 0002, Chuan Shi 0001
WWW6
2019 Moving Object Detection in Video via Hierarchical Modeling and Alternating Optimization
abstract
In conventional wisdom of video modeling, background is often treated as the primary target and foreground is derived using the technique of background subtraction. Based on the observation that foreground and background are two sides of the same coin, we propose to treat them as peer unknown variables and formulate a joint estimation problem, called Hierarchical modeling and Alternating Optimization (HMAO). The motivation behind our hierarchical extensions of background and foreground models is to better incorporate a priori knowledge about the disparity between background and foreground. For background, we decompose it into temporally low-frequency and high-frequency components for the purpose of better characterizing the class of video with dynamic background; for foreground, we construct a Markov random field prior at a spatially low resolution as the pivot to facilitate noise-resilient refinement at higher resolutions. Built on hierarchical extensions of both models, we show how to successively refine their joint estimates under a unified framework known as alternating direction multipliers method. Experimental results have shown that our approach produces more discriminative background and demonstrates better robustness to noise than other competing methods. When compared against current state-of-the-art techniques, HMAO achieves at least comparable and often superior performance in terms of F-measure scores especially for video containing dynamic and complex background.
Linhao Li, Qinghua Hu, Xin Li 0005
IEEE Trans. Image Process.3
2018 ICSD: An Automatic System for Insecure Code Snippet Detection in Stack Overflow over Heterogeneous Information Network
abstract
As the popularity of modern social coding paradigm such as Stack Overflow grows, its potential security risks increase as well (e.g., insecure codes could be easily embedded and distributed). To address this largely overlooked issue, in this paper, we bring an important new insight to exploit social coding properties in addition to code content for automatic detection of insecure code snippets in Stack Overflow. To determine if the given code snippets are insecure, we not only analyze the code content, but also utilize various kinds of relations among users, badges, questions, answers, code snippets and keywords in Stack Overflow. To model the rich semantic relationships, we first introduce a structured heterogeneous information network (HIN) for representation and then use meta-path based approach to incorporate higher-level semantics to build up relatedness over code snippets. Later, we propose a novel network embedding model named snippet2vec for representation learning in HIN where both the HIN structures and semantics are maximally preserved. After that, a multi-view fusion classifier is constructed for insecure code snippet detection. To the best of our knowledge, this is the first work utilizing both code content and social coding properties to address the code security issues in modern software coding platforms. Comprehensive experiments on the data collections from Stack Overflow are conducted to validate the effectiveness of the developed system ICSD which integrates our proposed method in insecure code snippet detection by comparisons with alternative approaches.
Yanfang Ye 0001, Shifu Hou, Lingwei Chen, Xin Li 0005, Liang Zhao 0002, Shouhuai Xu
ACSAC4
2018 Lightweight Deep Residue Learning for Joint Color Image Demosaicking and Denoising
abstract
Color demosaicking and image denoising each plays an important role in digital cameras. Conventional model-based methods often fail around the areas of strong textures and produce disturbing visual artifacts such as aliasing and zippering. Recently developed deep learning based methods were capable of obtaining images of better qualities though at the price of high computational cost, which make them not suitable for real-time applications. In this paper, we propose a lightweight convolutional neural network for joint demosaicking and denoising (JDD) problem with the following salient features. First, the densely connected network is trained in an end-to-end manner to learn the mapping from the noisy low-resolution space (CFA image) to the clean high-resolution space (color image). Second, the concept of deep residue learning and aggregated residual transformations are extended from image denoising and classification to JDD supporting more efficient training. Third, the design of our end-to-end network architecture is inspired by a rigorous analysis of JDD using sparsity models. Experimental results conducted for both demosaicking-only and JDD tasks have shown that the proposed method performs much better than existing state-of-the-art methods (i.e., higher visual quality, smaller training set and lower computational cost).
Weisheng Dong, Guangming Shi, Xin Li 0005
ICPR5
2018 Automatic Opioid User Detection from Twitter: Transductive Ensemble Built on Different Meta-graph Based Similarities over Heterogeneous Information Network
abstract
Opioid (e.g., heroin and morphine) addiction has become one of the largest and deadliest epidemics in the United States. To combat such deadly epidemic, in this paper, we propose a novel framework named HinOPU to automatically detect opioid users from Twitter, which will assist in sharpening our understanding toward the behavioral process of opioid addiction and treatment. In HinOPU, to model the users and the posted tweets as well as their rich relationships, we introduce structured heterogeneous information network (HIN) for representation. Afterwards, we use meta-graph based approach to characterize the semantic relatedness over users; we then formulate different similarities over users based on different meta-graphs on HIN. To reduce the cost of acquiring labeled samples for supervised learning, we propose a transductive classification method to build the base classifiers based on different similarities formulated by different meta-graphs. Then, to further improve the detection accuracy, we construct an ensemble to combine different predictions from different base classifiers for opioid user detection. Comprehensive experiments on real sample collections from Twitter are conducted to validate the effectiveness of HinOPU in opioid user detection by comparisons with other alternate methods.
Yujie Fan, Yiming Zhang 0002, Yanfang Ye 0001, Xin Li 0005
IJCAI4
2018 Design of an Autonomous Precision Pollination Robot
abstract
Precision robotic pollination systems can not only fill the gap of declining natural pollinators, but can also surpass them in efficiency and uniformity, helping to feed the fast-growing human population on Earth. This paper presents the design and ongoing development of an autonomous robot named “BrambleBee”, which aims at pollinating bramble plants in a greenhouse environment. Partially inspired by the ecology and behavior of bees, BrambleBee employs state-of-the-art localization and mapping, visual perception, path planning, motion control, and manipulation techniques to create an efficient and robust autonomous pollination system.
Nicholas Ohi, Kyle Lassak, Ryan M. Watson, Jared Strader, Yixin Du, Chizhao Yang, Gabrielle Hedrick, Jennifer Nguyen, Scott Harper, Dylan Reynolds, Cagri Kilic, Jacob Hikes, Sarah Mills, Conner Castle, Benjamin Buzzo, Nicole Waterland, Jason N. Gross, Yong-Lak Park, Xin Li 0005, Yu Gu 0008
IROS19
2018 DeepAM: a heterogeneous deep learning framework for intelligent malware detection
Yanfang Ye 0001, Lingwei Chen, Shifu Hou, William Hardy, Xin Li 0005
Knowl. Inf. Syst.5
2018 Distribution Sensitive Product Quantization
abstract
Product quantization (PQ) seems to have become the most efficient framework of performing approximate nearest neighbor (ANN) search for high-dimensional data. However, almost all existing PQ-based ANN techniques uniformly allocate precious bit budget to each subspace. This is not optimal, because data are often not evenly distributed among different subspaces. A better strategy is to achieve an improved balance between data distribution and bit budget within each subspace. Motivated by this observation, we propose to develop an optimized PQ (OPQ) technique, named distribution sensitive PQ (DSPQ) in this paper. The DSPQ dynamically analyzes and compares the data distribution based on a newly defined aggregate degree for high-dimensional data; whenever further optimization is feasible, resources such as memory and bits can be dynamically rearranged from one subspace to another. Our experimental results have shown that the strategy of bit rearrangement based on aggregate degree achieves modest improvements on most datasets. Moreover, our approach is orthogonal to the existing optimization strategy for PQ; therefore, it has been found that distribution sensitive OPQ can even outperform previous OPQ in the literature.
Linhao Li, Qinghua Hu, Yahong Han, Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.4
2018 Image Super-Resolution With Parametric Sparse Model Learning
abstract
Recovering a high-resolution (HR) image from its low-resolution (LR) version is an ill-posed inverse problem. Learning accurate prior of HR images is of great importance to solve this inverse problem. Existing super-resolution (SR) methods either learn a non-parametric image prior from training data (a large set of LR/HR patch pairs) or estimate a parametric prior from the LR image analytically. Both methods have their limitations: the former lacks flexibility when dealing with different SR settings; while the latter often fails to adapt to spatially varying image structures. In this paper, we propose to take a hybrid approach toward image SR by combining those two lines of ideas - that is, a parametric sparse prior of HR images is learned from the training set as well as the input LR image. By exploiting the strengths of both worlds, we can more accurately recover the sparse codes and therefore HR image patches than conventional sparse coding approaches. Experimental results show that the proposed hybrid SR method significantly outperforms existing model-based SR methods and is highly competitive to current state-of-the-art learning-based SR methods in terms of both subjective and objective image qualities.
Weisheng Dong, Xuemei Xie, Guangming Shi, Jinjian Wu, Xin Li 0005
IEEE Trans. Image Process.6
2017 Social Media for Opioid Addiction Epidemiology: Automatic Detection of Opioid Addicts from Twitter and Case Studies
abstract
Opioid (e.g., heroin and morphine) addiction has become one of the largest and deadliest epidemics in the United States. To combat such deadly epidemic, there is an urgent need for novel tools and methodologies to gain new insights into the behavioral processes of opioid abuse and addiction. The role of social media in biomedical knowledge mining has turned into increasingly significant in recent years. In this paper, we propose a novel framework named AutoDOA to automatically detect the opioid addicts from Twitter, which can potentially assist in sharpening our understanding toward the behavioral process of opioid abuse and addiction. In AutoDOA, to model the users and posted tweets as well as their rich relationships, a structured heterogeneous information network (HIN) is first constructed. Then meta-path based approach is used to formulate similarity measures over users and different similarities are aggregated using Laplacian scores. Based on HIN and the combined meta-path, to reduce the cost of acquiring labeled examples for supervised learning, a transductive classification model is built for automatic opioid addict detection. To the best of our knowledge, this is the first work to apply transductive classification in HIN into drug-addiction domain. Comprehensive experiments on real sample collections from Twitter are conducted to validate the effectiveness of our developed system AutoDOA in opioid addict detection by comparisons with other alternate methods. The results and case studies also demonstrate that knowledge from daily-life social media data mining could support a better practice of opioid addiction prevention and treatment.
Yujie Fan, Yiming Zhang 0002, Yanfang Ye 0001, Xin Li 0005, Wanhong Zheng
CIKM4
2017 Guided deep network for depth map super-resolution: How much can color help?
abstract
Since the quality of depth maps produced by Time-of-Flight (TOF) cameras is low, color-guided recovery methods have been proposed to increase spatial resolution and suppress unwanted noise. Despite successful applications of deep neural networks in color image super-resolution (SR), their potential for depth map SR is largely unknown. In this paper, we present a deep neural network architecture to learn the end-to-end mapping between low-resolution and high-resolution depth maps. Furthermore, we introduce a novel color-guided deep Fully Convolutional Network (FCN) and propose to jointly learn two nonlinear mapping functions (color-to-depth and LR-to-HR) in the presence of noise. Experimental results on several benchmark data sets show that our method outperforms several existing state-of-the-art depth SR algorithms. Moreover, this work attempts to partially shed some light onto the fundamental question in color-guided depth recovery - how much can color help in depth SR?
Wentian Zhou, Xin Li 0005, Daryl S. Reynolds
ICASSP2
2017 Color-Guided Depth Recovery via Joint Local Structural and Nonlocal Low-Rank Regularization
abstract
High-quality depth recovery from RGB-D data has received increasingly more attention in recent years due to their wide applications from depth-based image rendering to three-dimensional imaging and video. Sharp contrast between high-quality color images and low-quality depth maps presents severe challenges to the development of color-guided depth recovery techniques. Previous works have emphasized either locally varying characteristics of color-depth dependence or nonlocal similarities around the discontinuities of the scene geometry. Therefore, it is desirable to exploit both local and nonlocal structural constraints for optimizing the performance of color-guided depth recovery. In this work, we propose a unified variational approach via joint local and nonlocal regularization. The local regularization term consists of two complementary parts-one characterizing the color-depth dependence in the gradient domain and the other in the spatial domain; nonlocal regularization involves a low-rank constraint suitable for large-scale depth discontinuities. Extensive experimental results are reported to show that our approach outperforms several existing state-of-the-art depth recovery methods on both synthetic and real-world data sets.
Weisheng Dong, Guangming Shi, Xin Li 0005, Kefan Peng, Jinjian Wu, Zhenhua Guo 0001
IEEE Trans. Multim.3
2016 Learning Parametric Sparse Models for Image Super-Resolution
abstract
Learning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency details are learned in the former methods. Though effective, they are heuristic and have limitations in dealing with blurred LR images; while the latter suffers from the limitations of frequency aliasing. In this paper, we propose to combine those two lines of ideas for image super-resolution. More specifically, the parametric sparse prior of the desirable high-resolution (HR) image patches are learned from both the input low-resolution (LR) image and a training image dataset. With the learned sparse priors, the sparse codes and thus the HR image patches can be accurately recovered by solving a sparse coding problem. Experimental results show that the proposed SR method outperforms existing state-of-the-art methods in terms of both subjective and objective image qualities.
Weisheng Dong, Xuemei Xie, Guangming Shi, Xin Li 0005, Donglai Xu
NIPS5
2016 Hyperspectral Image Super-Resolution via Non-Negative Structured Sparse Representation
abstract
Hyperspectral imaging has many applications from agriculture and astronomy to surveillance and mineralogy. However, it is often challenging to obtain high-resolution (HR) hyperspectral images using existing hyperspectral imaging techniques due to various hardware limitations. In this paper, we propose a new hyperspectral image super-resolution method from a low-resolution (LR) image and a HR reference image of the same scene. The estimation of the HR hyperspectral image is formulated as a joint estimation of the hyperspectral dictionary and the sparse codes based on the prior knowledge of the spatial-spectral sparsity of the hyperspectral image. The hyperspectral dictionary representing prototype reflectance spectra vectors of the scene is first learned from the input LR image. Specifically, an efficient non-negative dictionary learning algorithm using the block-coordinate descent optimization technique is proposed. Then, the sparse codes of the desired HR hyperspectral image with respect to learned hyperspectral basis are estimated from the pair of LR and HR reference images. To improve the accuracy of non-negative sparse coding, a clustering-based structured sparse coding method is proposed to exploit the spatial correlation among the learned sparse codes. The experimental results on both public datasets and real LR hypspectral images suggest that the proposed method substantially outperforms several existing HR hyperspectral image recovery techniques in the literature in terms of both objective quality metrics and computational efficiency.
Weisheng Dong, Fazuo Fu, Guangming Shi, Xun Cao, Jinjian Wu, Xin Li 0005
IEEE Trans. Image Process.7
2015 Low-Rank Tensor Approximation with Laplacian Scale Mixture Modeling for Multiframe Image Denoising
abstract
Patch-based low-rank models have shown effective in exploiting spatial redundancy of natural images especially for the application of image denoising. However, two-dimensional low-rank model can not fully exploit the spatio-temporal correlation in larger data sets such as multispectral images and 3D MRIs. In this work, we propose a novel low-rank tensor approximation framework with Laplacian Scale Mixture (LSM) modeling for multi-frame image denoising. First, similar 3D patches are grouped to form a tensor of d-order and high-order Singular Value Decomposition (HOSVD) is applied to the grouped tensor. Then the task of multiframe image denoising is formulated as a Maximum A Posterior (MAP) estimation problem with the LSM prior for tensor coefficients. Both unknown sparse coefficients and hidden LSM parameters can be efficiently estimated by the method of alternating optimization. Specifically, we have derived closed-form solutions for both subproblems. Experimental results on spectral and dynamic MRI images show that the proposed algorithm can better preserve the sharpness of important image structures and outperform several existing state-of-the-art multiframe denoising methods (e.g., BM4D and tensor dictionary learning).
Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001
ICCV4
2015 Image Restoration via Simultaneous Sparse Coding: Where Structured Sparsity Meets Gaussian Scale Mixture
Weisheng Dong, Guangming Shi, Yi Ma 0001, Xin Li 0005
Int. J. Comput. Vis.4
2015 Automated Depression Diagnosis Based on Facial Dynamic Analysis and Sparse Coding
abstract
Depression is a severe psychiatric disorder preventing a person from functioning normally in both work and daily lives. Currently, diagnosis of depression requires extensive participation from clinical experts. It has drawn much attention to develop an automatic system for efficient and reliable diagnosis of depression. Under the influence of depression, visual-based behavior disorder is readily observable. This paper presents a novel method of exploring facial region visual-based nonverbal behavior analysis for automatic depression diagnosis. Dynamic feature descriptors are extracted from facial region subvolumes, and sparse coding is employed to implicitly organize the extracted feature descriptors for depression diagnosis. Discriminative mapping and decision fusion are applied to further improve the accuracy of visual-based diagnosis. The integrated approach has been tested on the AVEC2013 depression database and the best visual-based mean absolute error/root mean square error results have been achieved.
Lingyun Wen, Xin Li 0005, Guodong Guo, Yu Zhu 0006
IEEE Trans. Inf. Forensics Secur.2
2014 A study on the influence of body weight changes on face recognition
abstract
Overweight and obesity is quite common in the modern society, which can result in many severe health problems. Thus weight loss has become a major event for many people to have a healthy living. A question is then raised for Biometrics or identity management: Is there any influence on face recognition when the facial shapes are varied, caused by body weight changes? No previous research has addressed this issue, to the best of our knowledge. In this paper, we study the influence of body weight changes on face recognition. Both synthesized and real face images are assembled as the databases to facilitate our study. Empirically, we found that large body weight alterations can significantly reduce the matching accuracy of the face recognition system. This is a new exploration to the biometrics society. Then we study if the influence of weight changes can be reduced to improve the face recognition performance. The partial least squares (PLS) method is applied for this purpose. Our preliminary results show that it is feasible to develop algorithms to address the influence of facial adiposity variation on face recognition, caused by weight changes.
Lingyun Wen, Guodong Guo, Xin Li 0005
IJCB3
2014 Image restoration via Bayesian structured sparse coding
abstract
In this work, we propose a Bayesian structured sparse coding (BSSC) framework containing a nonlocal extension of Gaussian scale mixture (GSM) model by exploiting structured sparsity. It is shown that the variances of sparse coefficients (the field of Gaussian scalars) - if treated as a latent variable - can besparse coefficients jointly estimated along with the unknown sparse coefficients via the the method of alternative optimization. When applied to image restoration, BSSC leads to closed-form solutions involving iterative shrinkage/filtering and therefore admits computationally efficient implementation. Our experimental results have shown that BSSC-based image restoration often delivers reconstructed images with higher subjective/objective qualities than other competing approaches including IDD-BM3D and NCSR.
Weisheng Dong, Xin Li 0005, Yi Ma 0001, Guangming Shi
ICIP2
2014 Graph-based joint denoising and super-resolution of generalized piecewise smooth images
abstract
Images are often decoded with noise at receiver due to capturing errors and/or signal quantization during compression. Further, it is often necessary to display a decoded image at a higher resolution than the captured one, given available high-resolution (HR) display or a need to zoom-in for detailed examination. In this paper, we address the problems of image denoising and super-resolution (SR) jointly in one unified graph-based framework, focusing on a special class of signals called generalized piecewise smooth (GPWS) images. GPWS images are composed mostly of smooth regions connected by transition regions, and represent an important subclass of images, including cartoon, sub-regions of video frames with captions, graphics images in video games, etc. Like our previous work on piecewise smooth (PWS) images, GPWS images also imply simple-enough graph representations in the pixel domain, so that suitable graph-based filtering techniques can be readily applied. Specifically, leveraging on previous work on graph spectral analysis, for a given pixel block in low-resolution (LR) we first use the second eigenvector of a computed graph Laplacian matrix to identify a hard boundary, and then use the third eigenvector to identify two piecewise smooth regions and a transition region that separates them. The LR hard boundary is then super-resolved into HR via a procedure based on local self-similarity, while graph weights of the LR transition region is mapped to those of the HR transition region via polynomial fitting. Using the computed HR boundary and weights in the transition region, we construct a suitable HR graph corresponding to the LR counterpart, and perform joint denoising / SR using a graph smoothness prior. Experimental results show that our proposed algorithm outperforms two representative separable denoising / SR schemes in both subjective and objective quality.
Wei Hu 0003, Gene Cheung, Xin Li 0005, Oscar C. Au
ICIP3
2014 Blind restoration of very-high-ISO photos via low-rank methods
abstract
We propose a new algorithm for blind restoration of very-high-ISO photos. Unlike previous methods that sequentially tackle the problem of noise estimation and image denoising, our approach alternatively refines the estimates of latent image and noise level function (NLF). We rigorously show how the existing low-rank based modeling of image prior can be extended to incorporate spatially inhomogeneous and signal-dependent noise. We develop a generalization of singular-value thresholding technique by making the thresh-old/regularization parameter doubly adaptive - adaptive to both local signal and noise variance estimates. Our experimental results have shown that the proposed auto-denoising algorithm is capable of achieving visually pleasant restoration of photos with ISO settings of above 6400 for a wide range of brand cameras and at a moderate computational cost.
Xin Li 0005
ICME1
2014 Compressive Sensing via Nonlocal Low-Rank Regularization
abstract
Sparsity has been widely exploited for exact reconstruction of a signal from a small number of random measurements. Recent advances have suggested that structured or group sparsity often leads to more powerful signal reconstruction techniques in various compressed sensing (CS) studies. In this paper, we propose a nonlocal low-rank regularization (NLR) approach toward exploiting structured sparsity and explore its application into CS of both photographic and MRI images. We also propose the use of a nonconvex log det ( X) as a smooth surrogate function for the rank instead of the convex nuclear norm and justify the benefit of such a strategy using extensive experiments. To further improve the computational efficiency of the proposed algorithm, we have developed a fast implementation using the alternative direction multiplier method technique. Experimental results have shown that the proposed NLR-CS algorithm can significantly outperform existing state-of-the-art CS techniques for image recovery.
Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001
IEEE Trans. Image Process.3
2013 Depth map denoising using graph-based transform and group sparsity
abstract
Depth maps, characterizing per-pixel physical distance between objects in a 3D scene and a capturing camera, can now be readily acquired using inexpensive active sensors such as Microsoft Kinect. However, the acquired depth maps are often corrupted due to surface reflection or sensor noise. In this paper, we build on two previously developed works in the image denoising literature to restore single depth maps-i.e., to jointly exploit local smoothness and nonlocal self-similarity of a depth map. Specifically, we propose to first cluster similar patches in a depth image and compute an average patch, from which we deduce a graph describing correlations among adjacent pixels. Then we transform similar patches to the same graph-based transform (GBT) domain, where the GBT basis vectors are learned from the derived correlation graph. Finally, we perform an iterative thresholding procedure in the GBT domain to enforce group sparsity. Experimental results show that for single depth maps corrupted with additive white Gaussian noise (AWGN), our proposed NLGBT denoising algorithm can outperform state-of-the-art image denoising methods such as BM3D by up to 2.37dB in terms of PSNR.
Wei Hu 0003, Xin Li 0005, Gene Cheung, Oscar C. Au
MMSP2
2013 Facial Expression Recognition Influenced by Human Aging
abstract
Facial expression recognition (FER) is an active research topic in computer vision. However, there is no study yet to discover whether FER is affected by human aging, from a computational perspective. We perform a computational study of FER within and across age groups and compare the FER accuracies. Two databases from the psychology society are introduced to the computer vision community and used for our study. We found that the FER is influenced significantly by human aging, and we analyze the influence and interpret it from a computational viewpoint. Next, we propose some schemes to reduce the influence of aging on FER and evaluate the effectiveness in dealing with lifespan FER.
Guodong Guo, Xin Li 0005
IEEE Trans. Affect. Comput.3
2013 Nonlocal Image Restoration With Bilateral Variance Estimation: A Low-Rank Approach
abstract
Simultaneous sparse coding (SSC) or nonlocal image representation has shown great potential in various low-level vision tasks, leading to several state-of-the-art image restoration techniques, including BM3D and LSSC. However, it still lacks a physically plausible explanation about why SSC is a better model than conventional sparse coding for the class of natural images. Meanwhile, the problem of sparsity optimization, especially when tangled with dictionary learning, is computationally difficult to solve. In this paper, we take a low-rank approach toward SSC and provide a conceptually simple interpretation from a bilateral variance estimation perspective, namely that singular-value decomposition of similar packed patches can be viewed as pooling both local and nonlocal information for estimating signal variances. Such perspective inspires us to develop a new class of image restoration algorithms called spatially adaptive iterative singular-value thresholding (SAIST). For noise data, SAIST generalizes the celebrated BayesShrink from local to nonlocal models; for incomplete data, SAIST extends previous deterministic annealing-based solution to sparsity optimization through incorporating the idea of dictionary learning. In addition to conceptual simplicity and computational efficiency, SAIST has achieved highly competent (often better) objective performance compared to several state-of-the-art methods in image denoising and completion experiments. Our subjective quality results compare favorably with those obtained by existing techniques, especially at high noise levels and with a large amount of missing data.
Weisheng Dong, Guangming Shi, Xin Li 0005
IEEE Trans. Image Process.3
2013 Nonlocally Centralized Sparse Representation for Image Restoration
abstract
Sparse representation models code an image patch as a linear combination of a few atoms chosen out from an over-complete dictionary, and they have shown promising results in various image restoration applications. However, due to the degradation of the observed image (e.g., noisy, blurred, and/or down-sampled), the sparse representations by conventional models may not be accurate enough for a faithful reconstruction of the original image. To improve the performance of sparse representation-based image restoration, in this paper the concept of sparse coding noise is introduced, and the goal of image restoration turns to how to suppress the sparse coding noise. To this end, we exploit the image nonlocal self-similarity to obtain good estimates of the sparse coding coefficients of the original image, and then centralize the sparse coding coefficients of the observed image to those estimates. The so-called nonlocally centralized sparse representation (NCSR) model is as simple as the standard sparse representation model, while our extensive experiments on various types of image restoration problems, including denoising, deblurring and super-resolution, validate the generality and state-of-the-art performance of the proposed NCSR algorithm.
Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xin Li 0005
IEEE Trans. Image Process.4
2012 Depth map compression using multi-resolution graph-based transform for depth-image-based rendering
abstract
Depth map compression is important for efficient network transmission of 3D visual data in texture-plus-depth format, where the observer can synthesize an image of a freely chosen viewpoint via depth-image-based rendering (DIBR) using received neighboring texture and depth maps as anchors. Unlike texture maps, depth maps exhibit unique characteristics like smooth interior surfaces and sharp edges that can be exploited for coding gain. In this paper, we propose a multi-resolution approach to depth map compression using previously proposed graph-based transform (GBT). The key idea is to treat smooth surfaces and sharp edges of large code blocks separately and encode them in different resolutions: encode edges in original high resolution (HR) to preserve sharpness, and encode smooth surfaces in low-pass-filtered and down-sampled low resolution (LR) to save coding bits. Because GBT does not filter across edges, it produces small or zero high-frequency components when coding smooth-surface depth maps and leads to a compact representation in the transform domain. By encoding down-sampled surface regions in LR GBT, we achieve representation compactness for a large block without the high computation complexity associated with an adaptive large-block GBT. At the decoder, encoded LR surfaces are up-sampled and interpolated while preserving encoded HR edges. Experimental results show that our proposed multi-resolution approach using GBT reduced bitrate by 68% compared to native H.264 intra with DCT encoding original HR depth maps, and by 55% compared to single-resolution GBT encoding small blocks.
Wei Hu 0003, Gene Cheung, Xin Li 0005, Oscar C. Au
ICIP3
2012 Image reconstruction with locally adaptive sparsity and nonlocal robust regularization
Weisheng Dong, Guangming Shi, Xin Li 0005, Lei Zhang 0006, Xiaolin Wu 0001
Signal Process. Image Commun.3
2012 Progressive Significance Map and Its Application to Error-Resilient Image Transmission
abstract
Set partition coding (SPC) has shown tremendous success in image compression. Despite its popularity, the lack of error resilience remains a significant challenge to the transmission of images in error-prone environments. In this paper, we propose a novel data representation called the progressive significance map (prog-sig-map) for error-resilient SPC. It structures the significance map (sig-map) into two parts: a high-level summation sig-map and a low-level complementary sig-map (comp-sig-map). Such a structured representation of the sig-map allows us to improve its error-resilient property at the price of only a slight sacrifice in compression efficiency. For example, we have found that a fixed-length coding of the comp-sig-map in the prog-sig-map renders 64% of the coded bitstream insensitive to bit errors, compared with 40% with that of the conventional sig-map. Simulation results have shown that the prog-sig-map can achieve highly competitive rate-distortion performance for binary symmetric channels while maintaining low computational complexity. Moreover, we note that prog-sig-map is complementary to existing independent packetization and channel-coding-based error-resilient approaches and readily lends itself to other source coding applications such as distributed video coding.
William A. Pearlman, Xin Li 0005
IEEE Trans. Image Process.3
2011 Sparsity-based image denoising via dictionary learning and structural clustering
abstract
Where does the sparsity in image signals come from? Local and nonlocal image models have supplied complementary views toward the regularity in natural images - the former attempts to construct or learn a dictionary of basis functions that promotes the sparsity; while the latter connects the sparsity with the self-similarity of the image source by clustering. In this paper, we present a variational framework for unifying the above two views and propose a new denoising algorithm built upon clustering-based sparse representation (CSR). Inspired by the success of l1-optimization, we have formulated a double-header l1-optimization problem where the regularization involves both dictionary learning and structural structuring. A surrogate-function based iterative shrinkage solution has been developed to solve the double-header l1-optimization problem and a probabilistic interpretation of CSR model is also included. Our experimental results have shown convincing improvements over state-of-the-art denoising technique BM3D on the class of regular texture images. The PSNR performance of CSR denoising is at least comparable and often superior to other competing schemes including BM3D on a collection of 12 generic natural images.
Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi
CVPR2
2011 Sparsity-based image deblurring with locally adaptive and nonlocally robust regularization
abstract
Important structures in photographic images such as edges and textures are jointly characterized by local variation and nonlocal invariance (similarity). Both of them provide valuable heuristics to the regularization of image restoration process. In this pa per, we propose to explore two sets of complementary ideas: 1) locally learn PCA-based dictionaries and estimate the sparsity regularization parameters for each coefficient; and 2) nonlocally enforce the invariance constraint by introducing a patch-similarity based term into the cost functional. The minimization of this new cost functional leads to an iterative thresholding-based image deblurring algorithm and its efficient implementation is discussed. Our experimental results have shown that the proposed scheme significantly outperforms several leading deblurring techniques in the literature on both objective and visual quality assessments.
Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi
ICIP2
2011 Inverse halftoning with nonlocal regularization
abstract
Conventional wisdom in inverse halftoning heavily relies on the assumption about the local smoothness of image signals. Motivated by the effectiveness of nonlocal denoising, we propose a new class of inverse halftoning techniques using nonlocal regularization in this paper. The continuous-tone image is characterized by the intersection of two constraint sets - one related to the quantization process of halftoning and the other specified by nonlocal similarity-based regularization. Our nonlocal inverse halftoning algorithms alternatively project onto these two constraint sets; since the nonlocal regularization constraint set is nonconvex, we have borrowed the idea of deterministic annealing to optimize the performance of the proposed technique. Our experimental results have shown that our nonlocal approach can significantly outperform several existing state-of-the-art techniques in terms of both subjective and objective qualities.
Xin Li 0005
ICIP1
2011 The Magic of Nonlocal Perona-Malik Diffusion
abstract
In this letter, we show that it is possible to obtain perfect reconstruction of the 256 X 256 phantom image from only 8 radial lines in the Fourier domain by a nonlocal extension of Perona-Malik diffusion.
Xin Li 0005
IEEE Signal Process. Lett.1
2011 Fine-Granularity and Spatially-Adaptive Regularization for Projection-Based Image Deblurring
abstract
This paper studies two classes of regularization strategies to achieve an improved tradeoff between image recovery and noise suppression in projection-based image deblurring. The first is based on a simple fact that r-times Landweber iteration leads to a fixed level of regularization, which allows us to achieve fine-granularity control of projection-based iterative deblurring by varying the value r. The regularization behavior is explained by using the theory of Lagrangian multiplier for variational schemes. The second class of regularization strategy is based on the observation that various regularized filters can be viewed as nonexpansive mappings in the metric space. A deeper understanding about different regularization filters can be gained by probing into their asymptotic behavior--the fixed point of nonexpansive mappings. By making an analogy to the states of matter in statistical physics, we can observe that different image structures (smooth regions, regular edges and textures) correspond to different fixed points of nonexpansive mappings when the temperature(regularization) parameter varies. Such an analogy motivates us to propose a deterministic annealing based approach toward spatial adaptation in projection-based image deblurring. Significant performance improvements over the current state-of-the-art schemes have been observed in our experiments, which substantiates the effectiveness of the proposed regularization strategies.
Xin Li 0005
IEEE Trans. Image Process.1
2010 Color rank and census transforms using perceptual color contrast
abstract
Rank and census transforms provide high resistance to radiometric distortion, vignette, and noise because they are based on the relative ordering of local pixel intensity values rather than the pixel values themselves. These transforms are widely used in many computer vision applications. An important step of computing these transforms is to compare or rank two grayscale values, which is very much like measuring color difference in color image. Color difference between two color points at any part of a uniform color space corresponds to the perceptual difference between the two colors by the human vision system. Based on this idea, we propose to use perceptual color contrast to implement color rank and census transforms and achieve this without significantly increasing the amount of data to process and without complicated computations. Furthermore, we demonstrate the feasibility of using these new transforms to find correspondences for stereo vision.
Guangming Xiong, Xin Li 0005, Jianwei Gong, Huiyan Chen, Dah-Jye Lee
ICARCV2
2010 Exemplar-Based EM-like image denoising via manifold reconstruction
abstract
Discovering local geometry of low-dimensional manifold embedded into a high-dimensional space has been widely studied in the literature of machine learning. Counter-intuitively, we will show for the class of signal-independent additive noise, noisy data do not destroy the manifold structure thanks to the blessing of dimensionality. Based on this observation, we propose to reconstruct the manifold for a collection of exemplars by alternating between image filtering and neighborhood search. The byproduct of such manifold reconstruction from noisy data is an exemplar-Based EM-like (EBEM) denoising algorithm with minimal number of control parameters. Despite its conceptual simplicity, EBEM can achieve comparable performance to other leading algorithms in the literature. Our results suggest the importance of understanding the physical origin of manifold constraint underlying natural images - the symmetry in natural scenes.
Xin Li 0005
ICIP1
2010 AN adaptive L1-L2 hybrid error model to super-resolution
abstract
A hybrid error model with L1and L2norm minimization criteria is proposed in this paper for image/video super-resolution. A membership function is defined to adaptively control the tradeoff between the L1and L2norm terms. Therefore, the proposed hybrid model can have the advantages of both L1norm minimization (i.e. edge preservation) and L2norm minimization (i.e. smoothing noise). In addition, an effective convergence criterion is proposed, which is able to terminate the iterative L1and L2norm minimization process efficiently. Experimental results on images corrupted with various types of noises demonstrate the robustness of the proposed algorithm and its superiority to representative algorithms.
Huihui Song 0003, Lei Zhang 0006, Peikang Wang, Kaihua Zhang 0001, Xin Li 0005
ICIP5
2010 Collective sensing: a fixed-point approach in the metric space
abstract
Conventional wisdom in signal processing heavily relies on the concept of inner product defined in the Hilbert space. Despite the popularity of Hilbert-space formulation, we argue it is overly-structured to account for the complexity of signals arising from the real-world. Inspired by the works on fractal image decoding and nonlocal image processing, we propose to view an image as the fixed-point of some nonexpansive mapping in the metric space in this paper. Recently proposed BM3D-based denoising and nonlocal TV filtering can be viewed as the special cases of nonexpansive mappings while differ on the choice of clustering techniques. The physical interpretation of clustering-based nonexpansive mappings is that they convey organizational principles of the dynamical system underlying the signals of interest. There is an interesting analogy between phases of matters in statistical physics and types of structures in image processing. From this perspective, image reconstruction can be solved by a deterministic-annealing based global optimization approach which collectively exploits the a priori information about unknown image. The potential of this new paradigm, which we call ollective sensing is demonstrated on the lossy compression application where significant gain over current state-of-the-art (SPIHT) coding scheme has been achieved.
Xin Li 0005
VCIP1
2009 Patch-Based Video Processing: A Variational Bayesian Approach
abstract
In this paper, we present a patch-based variational Bayesian framework for video processing and demonstrate its potential in denoising, inpainting and deinterlacing. Unlike previous methods based on explicit motion estimation, we propose to embed motion-related information into the relationship among video patches and develop a nonlocal sparsity-based prior for typical video sequences. Specifically, we first extend block matching (nearest neighbor search) into patch clustering (k-nearest-neighbor search), which represents motion in an implicit and distributed fashion. Then we show how to exploit the sparsity constraint by sorting and packing similar patches, which can be better understood from a manifold perspective. Under the Bayesian framework, we treat both patch clustering result and unobservable data as latent variables and solve the inference problem via variational EM algorithms. A weighted averaging strategy of fusing diverse inference results from overlapped patches is also developed. The effectiveness of patch-based video models is demonstrated by extensive experimental results on a wide range of video materials.
Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.1
2008 Variational Bayesian image processing on stochastic factor graphs
abstract
In this paper, we present a patch-based variational Bayesian framework of image processing using the language of factor graphs (FGs). The variable and factor nodes of FGs represent image patches and their clustering relationship respectively. Unlike previous probabilistic graphical models, we model the structure of FGs by a latent variable, which gives the name "stochastic factor graphs"(SFGs). A sparsity-based prior is enforced to the local distribution functions at factor nodes, which leads to a class of variational expectation-maximization (VEM) algorithms on SFGs. VEM algorithms allow us to infer graph structure along with the target of inference from the observation data. This new framework can systematically exploit nonlocal dependency in natural images as justified by the experimental results in image denoising and inpainting applications.
Xin Li 0005
ICIP1
2008 Directional interpolation of noisy images
abstract
Most of the existing image interpolation schemes assume that the image is noise free. This assumption is invalid in practice because noise will be corrupted in the image acquisition process. The conventional way is to denoise the image first and then interpolate the denoised image. The denoising process, however, may smooth much the image details and introduce some artifacts, which could be amplified in the interpolation process. This paper presents a directional estimation scheme to implement denoising and interpolation simultaneously. For each noisy sample, we compute multiple directional estimates of it and then fuse them for a more accurate output. The estimation parameters computed in the denoising process can be subsequently used for interpolation. Compared with the schemes that perform denoising and interpolation in tandem, the proposed method can better reproduce the image fine structures and reduce much the interpolation artifacts.
Lei Zhang 0006, Xin Li 0005
ICIP2
2008 Intra prediction using template matching with adaptive illumination compensation
abstract
Modern video coding standards such as H.264/AVC use intra prediction for efficient coding of Intra pictures. These usually exploit local directional signal correlations. More recently, intra prediction modes using non-local signal information have been introduced. A very popular approach is the so called template matching prediction (TMP), which uses template based texture synthesis for signal prediction. This, combined with regular directional prediction, significantly improves intra coding efficiency compared to H.264/AVC. However, current TMP techniques have trouble synthesizing picture data with non-uniform illumination characteristics. They assume that similar picture regions resemble at the same time in structure and illumination, which is often not the case. In order to solve this, we propose a template matching technique with locally adaptive illumination compensation. The proposed technique is based on a linear compensation model with a scaling and an offset parameters to compensate for contrast and brightness disparities respectively. The model parameters are calculated by solving an auto-regressive Least Square problem during the template search for TMP. This permits to synthesize signal structures while capturing the local characteristics of illumination without needing extra side information. The total improvement in intra coding efficiency with respect to H.264/AVC can be of up to 18%.
Peng Yin 0002, Òscar Divorra Escoda, Xin Li 0005, Cristina Gomila
ICIP4
2008 Automatic Construction of Dental Charts for Postmortem Identification
abstract
Identification of deceased individuals based on dental characteristics is receiving increased attention, especially with the large volume of victims encountered in mass disasters. An important problem in automated dental identification is automatic classification of teeth into four classes (molars, premolars, canines, and incisors). An equally important problem is the construction of a dental chart, which is a data structure that guides tooth-to-tooth matching. Dental charts are the key for avoiding illogical comparisons that inefficiently consume the limited computational resources and may mislead decision making. Labeling of the teeth is a challenging task which has received inadequate attention in the literature. We tackle this composite problem using a two-stage approach. The first stage utilizes low computational cost, appearance-based features for assigning an initial class. The second stage applies a string matching technique, based on teeth neighborhood rules, to validate initial teeth-classes and, hence, to assign each tooth a number corresponding to its location in the dental chart. Based on a large test dataset of 507 bitewing and periapical films that contain 2027 teeth, the proposed approach achieves classification accuracy of 87%. Experimental results indicate that the proposed approach works very fast, and achieves high performance compared to other approaches suggested in the literature.
Diaa Eldin M. Nassar, Ayman Abaza, Xin Li 0005, Hany H. Ammar
IEEE Trans. Inf. Forensics Secur.3
2007 Geometry-Adaptive Block Partitioning for Video Coding
abstract
Frame partitioning is a process of key importance in efficient video coding. Most recent video compression technologies, like H.264/AVC, use tree based frame partition. This reveals to be more efficient than simple uniform block partition, typically used in older video coding standards like MPEG-2 or H.263. However, tree based frame partition still does not code efficiently enough video information, as is unable to capture the geometric structure of 2D data. During last years, several works have been developed, mainly in the domain of still image representation and coding, in order to solve such limitations. An example is the use of wedge partitions. Based on these, in this paper, we study a way to better represent and code 2D video data by taking its 2D geometry into account. Our study is developed as an extension of H.264/AVC. Geometry-adaptive partitions are used to improve intra and inter prediction modes. Results obtained with the investigated method show that both better R-D and visual performance can be achieved.
Òscar Divorra Escoda, Peng Yin 0002, Congxia Dai, Xin Li 0005
ICASSP (1)4
2007 Geometry-Adaptive Block Partitioning for Intra Prediction in Image & Video Coding
abstract
Many modern video coding strategies, such as the H.264/AVC standard, use quadtree-based partition structures for coding intra macroblocks. Such a structure allows the coding algorithm to adapt to the complicated and non-stationary nature of natural images. Despite the adaptation flexibility of quadtree partitions, recent studies have shown that these are not efficient enough (in terms of rate-distortion performance) when images can be locally modeled as 2D piecewise-smooth signals. These observations motivate us to investigate the use of geometry based block partitioning for modeling intra data in video coding. In particular, in this paper, we study in detail the use of geometry-adaptive intra models, where wedgelet like discontinuities are used in order to define separate coding regions where different statistical/waveform modeling tools can be used. In order to implement this idea, we extend the existing H.264/AVC intra coding scheme by introducing two additional geometric modes: INTRA16X16GEO, and INTRA8X8GEO. Experimental results show that significantly improved R-D performance is achieved.
Congxia Dai, Òscar Divorra Escoda, Peng Yin 0002, Xin Li 0005, Cristina Gomila
ICIP (6)4
2007 Video Modeling by Spatio-Temporal Resampling and Bayesian Fusion
abstract
In this paper, we propose an empirical Bayesian approach toward video modeling and demonstrate its application in multiframe image restoration. Based on our previous work on spatio-temporall adaptive localized learning (STALL), we introduce a new concept of spatio-temporal resampling to facilitate the task of video modeling. Resampling produces a redundant representation of video signals with distributed spatio-temporal characteristics. When combined with STALL model, we show how to probabilistically combine the linear regression results of resampled video signals under a Bayesian framework. Such empirical Bayesian approach opens the door to develop a whole new class of video processing algorithms without explicit motion estimation or segmentation. The potential of our distributed video model is justified by considering its application into two multiframe image restoration tasks: repair damaged blocks and remove impulse noise.
Xin Li 0005
ICIP (6)2
2007 Parallel and Distributed Audio Concealment using Nonlocal Sparse Representations
abstract
We present a new class of parallel and distributed audio concealment (PDAC) algorithms which recover lost audio packets at the receiver to fight against channel impairment. The main contribution of this work is the proposal of using nonlocal sparse representations to characterize the prior constraint of undamaged audio. When combined with observation constraint, we obtain an alternating projection based audio concealment algorithm which recovers missing data in a parallel and distributed fashion. We also present two extensions of PDAC for more challenging situations: expectation-maximization PDAC (EM-PDAC) to handle consecutive packet loss and filter-bank PDAC (FB-PDAC) to repair complex music signals. Excellent preliminary experimental results are reported for a wide range of audio materials and loss conditions.
Xin Li 0005
ICME1
2007 Pedestrian detection and tracking in infrared imagery using shape and appearance
Congxia Dai, Xin Li 0005
Comput. Vis. Image Underst.3
2007 Video Processing Via Implicit and Mixture Motion Models
abstract
In this paper, we present an alternative framework for video processing without explicit motion estimation or segmentation. Motivated by the geometric constraint of motion trajectory, we propose an adaptive filtering-based model for video signals in which filter coefficients are locally estimated by the least-square method. Such localized estimation can be viewed as an implicit approach of exploiting motion-related temporal dependency. We also introduce the the concept of a virtual camera to further improve the modeling capability by exploiting the fundamental tradeoff between space and time. Using mixture models, we show how to probabilistically fuse the inference results obtained from virtual cameras in order to achieve spatio-temporal adaptation. Implicit and mixture motion model supplements the existing paradigm and provides a unified solution to a wide range of low-level vision problems including video dejittering, impulse removal, error concealment, video coding, and temporal interpolation.
Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.1
2006 Two-Dimensional Wiener Filters for Error Resilient Time Domain Lapped Transform
abstract
This paper presents the design of two-dimensional Wiener filters for error resilient time domain lapped transform. Two solutions are discussed, and a multi-pass approach is also proposed to make the algorithm adaptive to input statistics. Design examples and image coding experiments show that the adaptive 2-D Wiener filters provide significant improvement over the existing 1-D Wiener filtering method.
Jie Liang 0001, Xin Li 0005, Guoqian Sun, Trac D. Tran
ICASSP (3)2
2006 Subframe Video Synchronization via 3D Phase Correlation
abstract
This paper introduces an accurate approach for synchronization (temporal alignment) between two video sequences of the same dynamic scene captured by uncalibrated cameras. With the homography assumption in spatial domain, an iterative procedure that successively achieves the alignment in space and time is proposed and its convergence is experimentally verified. Subframe accuracy is achieved by extending the existing image subpixel registration scheme to subframe video synchronization. In order to demonstrate the accuracy of the proposed method, we adopt a novel use of audio signals for their high sampling rate to obtain the synchronization ground-truth. The proposed video synchronization technique has potential use in temporal super-resolution, image-based rendering and tele-immersion.
Congxia Dai, Xin Li 0005
ICIP3
2006 Symmetric Disparity Estimation in Distributed Coding of Stereo Images
abstract
Disparity estimation has been widely studied in the literature of stereo matching and coding. However, how to shift disparity estimation from encoder to decoder to support distributed coding applications poses a new challenge. In this paper, we present a symmetric distributed coding protocol in which interlaced representations of stereo pair are finely and coarsely quantized as primary and secondary channels respectively. At the decoder, side information is generated from the primary channel by an expectation maximization (EM)-like algorithm and a novel dual exploitation of the secondary channel is proposed to simultaneously resolve intensity uncertainty and refine disparity estimation. Preliminary experimental results for synthetic images are reported to demonstrate the potential of the proposed approach.
Xin Li 0005
ICIP1
2006 Accurate Video Alignment Using Phase Correlation
abstract
In this letter, we present an accurate technique for temporally aligning two video sequences of the same scene captured by nonsynchronized cameras. An iterative procedure is proposed to successively align the sequences in space and time; and the existing two-dimensional phase-correlation method is generalized into three dimensions to achieve subframe accuracy. The ground-truth of subframe temporal alignment is obtained by using supplementary audio signals sampled at a much higher rate. The accuracy of our technique is demonstrated by experimental results using real-world sequences
Congxia Dai, Xin Li 0005
IEEE Signal Process. Lett.3
2006 Edge-Directed Error Diffusion Halftoning
abstract
In this letter, we propose two simple extensions of the existing error diffusion (ED) halftoning technique: one stops the diffusion at edge pixels, and the other tunes the support of the diffusion kernel to match the local edge orientation. Experimental results are used to show that the proposed edge-directed ED schemes achieve noticeably better performance over existing ED with edge enhancement for a certain class of gray-scale images.
Xin Li 0005
IEEE Signal Process. Lett.1
2005 Contour Adaptive Image Coding
abstract
This paper presents a new image coding framework that separates singularities based on their topological dimension: point singularities (0D), line singularities (1D) and plane singularities (2D). Contours are smooth curves corresponding to line singularities and boundaries of plane singularities. We propose to directly code contour locations in the spatial domain and spatially adapt the bases to approximate various singularities conditioned on contour locations. The key to the success of our forward adaptive coding lies in the exploitation of contour geometry while achieving spatial adaptation. Preliminary experimental results are used to demonstrate the potential of our approach.
Xin Li 0005
ICASSP (2)1
2005 Improved wavelet decoding via set theoretic estimation
Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.1
2005 Demosaicing by successive approximation
abstract
In this paper, we present a fast and high-performance algorithm for color filter array (CFA) demosaicing. CFA demosaicing is formulated as a problem of reconstructing correlated signals from their downsampled versions with an opposite phase. The major contributions of this work include (1) a new iterative demosaicing algorithm in the color difference domain and (2) a spatially adaptive stopping criterion for suppressing color misregistration and zipper artifacts in the demosaiced images. We have compared the proposed demosaicing algorithm with two current state-of-the-art techniques reported in the literature. Ours outperforms both of them on demosaicing performance and computational cost.
Xin Li 0005
IEEE Trans. Image Process.1
2004 Distributed coding of multispectral images: a set theoretic approach
abstract
Distributed coding problem poses the challenge of how to shift the exploitation of the correlation structure of source from encoder to decoder with minimal degradation on coding efficiency. In this paper, we propose a novel convex-set theoretic framework for distributed coding of multispectral images. Alternating projection based decoding algorithms are developed to exploit the correlation among different spectral channels at the centralized decoder. Both asymmetric and symmetric protocols are studied and compared. Experiment results have shown that the proposed symmetric distributed coder only falls behind standard wavelet coders (e.g., JPEG2000) by less than 2 dB at the bit rate of 1-2 bpp.
Xin Li 0005
ICIP1
2004 Scalable video compression via overcomplete motion compensated wavelet coding
Xin Li 0005
Signal Process. Image Commun.1
2003 New directions in video coding
Xin Li 0005
VCIP1
2003 New results of phase shifting in the wavelet space
abstract
This paper investigates the relationship between even-phase coefficients and odd-phase coefficients in a two-channel perfect reconstruction filter bank. We demonstrate that they are linked to each other by a unique phase-shifting matrix. In the case of multilevel wavelet decomposition, we present an efficient recursive solution to directly perform phase shifting in the wavelet space. Our proposed solution can also be easily generalized into the case of two-dimensional wavelet transform. Direct phase-shifting methods in the wavelet space have potential applications in wavelet-based image/video coding and compressed domain processing.
Xin Li 0005
IEEE Signal Process. Lett.1
2003 On exploiting geometric constraint of image wavelet coefficients
abstract
In this paper, we investigate the problem of how to exploit geometric constraint of edges in wavelet-based image coding.The value of studying this problem is the potential coding gain brought by improved probabilistic models of wavelet high-band coefficients. Novel phase shifting and prediction algorithms are derived in the wavelet space. It is demonstrated that after resolving the phase uncertainty, high-band wavelet coefficients can be better modeled by biased-mean probability models rather than the existing zero-mean ones. In lossy coding, the coding gain brought by the biased-mean model is quantitatively analyzed within the conventional DPCM coding framework. Experiment results have shown the proposed phase shifting and prediction scheme improves both subjective and objective performance of wavelet-based image coders.
Xin Li 0005
IEEE Trans. Image Process.1
2002 Low Bit Rate Image Coding in the Scale Space
abstract
Scale-space representation has been extensively studied in the computer vision community for analyzing image structures at dierent scales. This paper borrows and develops useful mathematical tools from scale-space theory to facilitate the task of image compression. Instead of compressing the original image directly, we propose to compress its scale-space representation obtained by the forward diusion with a Gaussian kernel at the chosen scale. The major con tribution of this w ork is a no vel solution to the ill-posed inverse diusion problem. We analytically derive a nonlinear lter to deblur Gaussian blurring for 1D ideal step edges. The generalized 2D edge enhancing lter only requires the knowledge of local minimum/maximum and preserves the geometric constraint of edges. When combined with a standard wavelet-based image coder, the forward and inverse diusion can be viewed as a pair of pre-processing and post-processing stages used to select and preserve important image features at the given bit rate. Experiment results ha ve sho wn that the proposed diusion-based techniques can dramatically improve the visual quality of reconstructed images at low bit rate (below 0:25bpp).
Xin Li 0005
DCC1
2002 Embedded Coding of Palette Images in the Topological Space
abstract
Summary form only given. Most existing image coding techniques resolve the uncertainty of an image source on a pixel-by-pixel basis. We demonstrate the effectiveness of region-based image models for the class of palette images. We propose to represent the index map of a palette image by a collection of successively refined color regions, from which the original index map can be reconstructed without any error. Within the framework of region-based modeling, we present a conditional coding approach to avoid information leakage during the multiple passes. Motivated by the distinguished characteristics of palette images, we propose to exploit topological property of isolated uniform-color regions while resolving the uncertainty of region boundaries. Our region-based image model not only provides an embedded representation of palette images in the topological space but also achieves excellent compression performance.
Xin Li 0005
DCC1
2002 Blind image quality assessment
abstract
Blind image quality assessment refers to the problem of evaluating the visual quality of an image without any reference. It addresses a fundamental distinction between fidelity and quality, i.e. human vision system usually does not need any reference to determine the subjective quality of a target image. In this paper, we propose to appraise the image quality by three objective measures: edge sharpness level, random noise level and structural noise level. They jointly provide a heuristic approach of characterizing the most important aspects of visual quality. We investigate various mathematical tools (analytical, statistical and PDE-based) for accurately and robustly estimating those three levels. Extensive experiment results are used to justify the validity of our approach.
Xin Li 0005
ICIP (1)1
2002 On exploiting phase constraint with image wavelet coefficients
abstract
This paper investigates the potential of exploiting a phase-related constraint to improve the performance of wavelet-based image coders. The phase-related constraint originates from the fact that the FFT of the 1D intensity profile along an oblique edge is identical up to a fixed phase shift. However, due to the decimation operation of wavelet transform (WT), the linear phase shifting characteristic is destroyed in the wavelet domain. We propose to recover the odd-phase coefficients from the even-phase ones by a novel phase shifting filter and to interpolate the fractional-phase coefficients from the integer-phase ones by Lagrange filters. The fractional amount of phase shift can be estimated from the causal neighbors in the spatial and frequency domain. Our coding results have demonstrated the effectiveness of the proposed techniques for both synthetic and real images.
Xin Li 0005
ICIP (3)1
2002 Efficient motion field representation in the wavelet domain
abstract
It has been widely observed that improved motion field representation (e.g. fractional-pel and OBMC) is beneficial to video coding. In this paper, we investigate efficient motion field representation in the wavelet domain. We aim to demonstrate that motion estimation in the wavelet domain can: 1) effectively solve the occlusion problem; 2) transform the aperture from a problem into a property that can be exploited in video coding. Our study leads to a novel anisotropic motion field representation in the wavelet domain. Our new wavelet-based motion field offers more flexibility than traditional representations in the spatial domain and opens the door to an improved understanding of relationship between motion and intensity uncertainty model.
Xin Li 0005, Shawmin Lei
ICIP (3)1
2002 High-performance resolution-scalable video coding via all-phase motion-compensated prediction of wavelet coefficients
Xin Li 0005, Louis Kerofsky
VCIP1
2002 Novel sequential error-concealment techniques using orientation adaptive interpolation
abstract
This paper introduces a new framework for error concealment in block-based image coding systems: sequential recovery. Unlike previous approaches that simultaneously recover the pixels inside a missing block, we propose to recover them in a sequential fashion such that the previously-recovered pixels can be used in the recovery process afterwards. The principal advantage of the sequential approach is the improved capability of recovering important image features brought by the reduction in the complexity of statistical modeling, i.e., from blockwise to pixelwise. Under the framework of sequential recovery, we present an orientation adaptive interpolation scheme derived from the pixelwise statistical model. We also investigate the problem of error propagation with sequential recovery and propose a linear merge strategy to alleviate it. Extensive experimental results are used to demonstrate the improvement of the proposed sequential error-concealment technique over previous techniques in the literature.
Xin Li 0005, Michael T. Orchard
IEEE Trans. Circuits Syst. Video Technol.1
2001 All-phase motion compensated prediction in the wavelet domain for high performance video coding
abstract
This paper presents a novel framework of motion compensated prediction (MCP) techniques in the wavelet domain for high performance video coding. Our analysis reveals fundamental limitations with previous ad-hoc wavelet-based video coders from the motion accuracy point of view. We demonstrate that the phase associated with any wavelet transform carries critical information of the motion accuracy and we propose to restore the motion accuracy by considering the wavelet coefficients of the previous frame with all different phases. Our all-phase MCP approach can be viewed as predicting the wavelet coefficients from an over-complete expansion of the previous frame. Experimental results have shown that restoration of motion accuracy in the wavelet domain can dramatically improve the efficiency of MCP. The video coder (MCP-WT) built upon the MCP of wavelet coefficients has achieved 2 to 3 dB gain over existing the MPEG-2 coder at the bit rate of 1 to 9 Mbps. Moreover, MCP techniques in the wavelet domain offer a promising new ground for developing efficient scalable video coders.
Xin Li 0005, Louis Kerofsky, Shawmin Lei
ICIP (3)1
2001 On the study of lossless compression of computer generated compound images
abstract
This paper studies the problem of lossless compression of computer generated compound images that contain not only photographic images but also text and graphic images. We present a simple backward adaptive classification scheme to separate the image source into three classes: smooth regions, text regions and image regions. Different probability models are assigned within each class to maximize the compression performance. We also extend our scheme to exploit the interplane dependency for coding color images. The segmentation results of the reference color plane are used as the contexts for the classification and coding of the current color plane. Our new lossless coder significantly outperforms current state-of-the-art coders such as CALIC and JPEG-LS for compound images with modest computational complexity.
Xin Li 0005, Shawmin Lei
ICIP (3)1
2001 Block-based segmentation and adaptive coding for visually lossless compression of scanned documents
abstract
This paper presents a novel block-based segmentation and adaptive coding (BSAC) algorithm for visually lossless compression of scanned documents that contain not only photographic images but also text and graphic images. For such a compound image source, we structure the image into nonoverlapping blocks and classify each block into four different classes based on the empirical statistics within the block. Different coding strategies are applied to different classes in order to achieve the very best compression performance. Our new block-based image coder is able to provide visually lossless compression of scanned documents at the bit rate of around 1/spl sim/1.5 bpp with modest computational complexity and very low memory requirement.
Xin Li 0005, Shawmin Lei
ICIP (3)1
2001 Novel sequential error concealment techniques using orientation adaptive interpolation
Xin Li 0005, Michael T. Orchard
VCIP1
2001 Edge-directed prediction for lossless compression of natural images
abstract
This paper sheds light on the least-square (LS)-based adaptive prediction schemes for lossless compression of natural images. Our analysis shows that the superiority of the LS-based adaptation is due to its edge-directed property, which enables the predictor to adapt reasonably well from smooth regions to edge areas. Recognizing that LS-based adaptation improves the prediction mainly around the edge areas, we propose a novel approach to reduce its computational complexity with negligible performance sacrifice. The lossless image coder built upon the new prediction scheme has achieved noticeably better performance than the state-of-the-art coder CALIC with moderately increased computational complexity.
Xin Li 0005, Michael T. Orchard
IEEE Trans. Image Process.1
2001 New edge-directed interpolation
abstract
This paper proposes an edge-directed interpolation algorithm for natural images. The basic idea is to first estimate local covariance coefficients from a low-resolution image and then use these covariance estimates to adapt the interpolation at a higher resolution based on the geometric duality between the low-resolution covariance and the high-resolution covariance. The edge-directed property of covariance-based adaptation attributes to its capability of tuning the interpolation coefficients to match an arbitrarily oriented step edge. A hybrid approach of switching between bilinear interpolation and covariance-based adaptive interpolation is proposed to reduce the overall computational complexity. Two important applications of the new interpolation algorithm are studied: resolution enhancement of grayscale images and reconstruction of color images from CCD samples. Simulation results demonstrate that our new interpolation algorithm substantially improves the subjective quality of the interpolated images over conventional linear interpolation.
Xin Li 0005, Michael T. Orchard
IEEE Trans. Image Process.1
2000 New Edge Directed Interpolation
abstract
This paper presents a novel edge orientation adaptive interpolation scheme for resolution enhancement of still images. In order to achieve ideal orientation adaptation, we propose to estimate the local covariance characteristics at low resolution but cleverly use them to direct the interpolation at high resolution based on the resolution invariant property of edge orientation. The orientation adaptive property guarantees the interpolation always go along the edge orientation but not across it. Our new interpolation scheme can generate images with dramatically higher visual quality than linear interpolation techniques while keeping the computational complexity still modest.
Xin Li 0005, Michael T. Orchard
ICIP1
2000 Spatially Adaptive Image Denoising Under OverComplete Expansion
abstract
This paper presents a novel wavelet-based image denoising algorithm under overcomplete expansion. In order to optimize the denoising performance, we make a systematic study of both signal and noise characteristics under overcomplete expansion. High-band coefficients are viewed as the mixture of non-edge class and edge class observing different probability models. Based on improved statistical modeling of wavelet coefficients, we derive optimal MMSE estimation strategies to suppress noise for both non-edge and edge coefficients. We have achieved fairly better objective performance than most recently-published wavelet denoising schemes.
Xin Li 0005, Michael T. Orchard
ICIP1
1999 Edge Directed Prediction for Lossless Compression of Natural Images
abstract
Natural images are populated with edges characterized by abrupt changes of local statistics. They put severe challenges on probability modeling of image sources. This paper proposes to employ recursive least square (RLS)-based predictive modeling to characterize local statistics for edges. It can be viewed as estimating the covariance matrix from a local causal neighborhood and selecting the MMSE optimal predictor for the local covariance estimate. We demonstrate how the RLS-based adaptation can produce predictor with support ideally aligned along an arbitrarily-oriented edge and therefore we call it "Edge Directed Prediction"(EDP). When applied to lossless image compression, the EDP substantially outperforms former context-based prediction schemes for natural images. Based on our high-level understanding of EDP, we dramatically reduce its complexity with little sacrifice on the performance, thus facilitating its application in practice.
Xin Li 0005, Michael T. Orchard
ICIP (4)1
1998 On Implementing Transforms from Integers to Integers
Xin Li 0005, Michael T. Orchard
ICIP (3)1