Ruotao Xu

dblp:154/5633 · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0002-5277-9859ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021
YearPublicationVenuePosition
2026 When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
abstract
Large reasoning models (LRMs) have achieved remarkable performance in complex reasoning tasks, driven by their powerful inference-time scaling capability.However, LRMs often suffer from overthinking, which results in substantial computational redundancy and significantly reduces efficiency.Early-exit methods aim to mitigate this issue by terminating reasoning once sufficient evidence has been generated, yet existing approaches mostly rely on handcrafted or empirical indicators that are unreliable and impractical.In this work, we introduce Dynamic Thought Sufficiency in Reasoning (DTSR), a novel framework for efficient reasoning that enables the model to dynamically assess the sufficiency of its chain-of-thought (CoT) and determine the optimal point for early exit.Inspired by human metacognition, DTSR operates in two stages: (1) Reflection Signal Monitoring, which identifies reflection signals as potential cues for early exit, and (2) Thought Sufficiency Check, which evaluates whether the current CoT is sufficient to derive the final answer.Experimental results on the Qwen3 models show that DTSR reduces reasoning length by 28.9%-34.9%with minimal performance loss, effectively mitigating overthinking.We further discuss overconfidence in LRMs and self-evaluation paradigms, providing valuable insights for early-exit reasoning.
Yang Xiang 0003, Yixin Ji, Ruotao Xu, Zheming Yang, Juntao Li 0005, Min Zhang 0005
ACL (1)3
2026 ObjectDiff: An object-centric diffusion policy with modality-specific conditioning for robot manipulation
Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002
Knowl. Based Syst.3
2025 Self-Correcting Robot Manipulation via Gaussian-Splatted Foresight
abstract
Language-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the robotic arm entering an untrained state, potentially causing destructive results. Consequently, the ability to detect and self-correct failures is crucial for the development of practical robotic systems. To address this challenge, we propose a foresight-driven failure detection and self-correction module for robot manipulation. By leveraging 3D Gaussian Splatting, we represent the current scene with multiple Gaussians. Subsequently, we train a prediction network to forecast the Gaussian representation of future scenes conditioned on planned actions. Failure is detected when the predicted future significantly deviates from the real observation after action execution. In such cases, the end-effector rolls back to the previous action to avoid an untrained state. Integrating this approach with the PerACT framework, we develop a self-correcting robot manipulation policy. Evaluations on ten RLBench tasks with 166 variations demonstrate the superior performance of the proposed method, which outperforms state-of-the-art methods by 12.0% success rate on average.
Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu
AAAI3
2025 Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video Restoration
abstract
Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete Prior-based Temporal-Coherent content prediction transformer to address the challenge, and our model is referred to as DP-TempCoh. Specifically, we incorporate a spatial-temporal-aware content prediction module to synthesize high-quality content from discrete visual priors, conditioned on degraded video tokens. To further enhance the temporal coherence of the predicted content, a motion statistics modulation module is designed to adjust the content, based on discrete motion priors in terms of cross-frame mean and variance. As a result, the statistics of the predicted content can match with that of real videos over time. By performing extensive experiments, we verify the effectiveness of the design elements and demonstrate the superior performance of our DP-TempCoh in both synthetically and naturally degraded video restoration.
Lianxin Xie, Bingbing Zheng, Ruotao Xu, Si Wu 0002, Hau-San Wong
AAAI6
2025 RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting
abstract
Face retouching aims to remove facial imperfections from image and videos while at the same time preserving face attributes. The existing methods are designed to perform non-interactive end-to-end retouching, while the ability to interact with users is highly demanded in downstream applications. In this paper, we propose RetouchGPT, a novel framework that leverages Large Language Models (LLMs) to guide the interactive retouching process. Towards this end, we design an instruction-driven imperfection prediction module to accurately identify imperfections by integrating textual and visual features. To learn imperfection prompts, we further incorporate a LLM-based embedding module to fuse multi-modal conditioning information. The prompt-based feature modification is performed in each transformer block, such that the imperfection features are suppressed and replaced with the features of normal skin progressively. Extensive experiments have been performed to verify effectiveness of our design elements and demonstrate that RetouchGPT is a useful tool for interactive face retouching and achieves superior performance over state-of-the-arts.
Chun Ding, Ruotao Xu, Si Wu 0002, Yong Xu 0007, Hau-San Wong
AAAI3
2025 Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding
abstract
The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual region, we propose a Task-aware Crossmodal feature Refinement Transformer with LLMs for visual grounding, and our model is referred to as TCRT. To enable the LLM trained solely on text to understand images, we introduce an LLM adaptation module that extracts textrelated visual features to bridge the domain discrepancy between the textual and visual modalities. We feed the text and visual features into the LLM to obtain task-aware priors. To enable the priors to guide the feature fusion process, we further incorporate a cross-modal feature fusion module, which allows task-aware embeddings to refine visual features and facilitate information interaction between the Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES) tasks. We have performed extensive experiments to verify the effectiveness of the main components and demonstrate the superior performance of the proposed TCRT over state-of-the-art end-to-end visual grounding methods on RefCOCO, RefCOCOg, RefCOCO+ and ReferItGame.
Ruotao Xu, Si Wu 0002, Hau-San Wong
CVPR3
2025 A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse Artifacts
abstract
Structured artifacts are semi-regular, repetitive patterns that closely intertwine with genuine image content, making their removal highly challenging. In this paper, we introduce the Scale-Adaptive Deformable Transformer, an network architecture specifically designed to eliminate such artifacts from images. The proposed network features two key components: a scale-enhanced deformable convolution module for modeling scale-varying patterns with abundant orientations and potential distortions, and a scale-adaptive deformable attention mechanism for capturing long-range relationships among repetitive patterns with different sizes and non-uniform spatial distributions. Extensive experiments show that our network consistently outperforms state-of-the-art methods in diverse artifact removal tasks, including image deraining, image demoiréing, and image debanding.
Xuyi He, Yuhui Quan, Ruotao Xu, Hui Ji 0002
CVPR3
2025 Rethinking 3D Robotic Perception: Elastic Voxel Representation with Splatting Distillation
abstract
Language-guided robotic manipulation is advancing rapidly with Vision-Language-Action (VLA) models, yet faces fundamental challenges in 3D perception. This paper addresses two critical challenges: the scale elasticity requirement for simultaneously processing coarse environmental context and fine manipulation details, and the scarcity of action-annotated training data. We present Splat-Actor, a novel robotic manipulation framework that introduces two key innovations. First, we develop an elastic voxel encoder that combines multi-scale processing with selective tokenization, enabling efficient 3D spatial reasoning while adaptively focusing on informative regions. Second, we propose a depth-constrained feature distillation framework that leverages Gaussian Splatting to bridge 2D and 3D representations, transferring rich semantic features from pre-trained vision models to enhance 3D understanding. Extensive experiments across 10 manipulation tasks with 166 variations demonstrate that Splat-Actor achieves a 6.8% improvement over state-of-the-art methods while maintaining the computational efficiency.
Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu, Patrick Le Callet
ICME3
2025 Text to Trajectory: Enhancing and Evaluating LLMs for Embodied Task Planning
abstract
The increasing demand for effective human-machine interaction highlights the importance of integrating natural language processing with robotics technology. This paper addresses the challenges of using Large Language Models (LLMs) for embodied task planning in complex environments. We propose a comprehensive framework that combines environmental-aware LLM fine-tuning with a novel Stepwise Beam Search (SBS) strategy. In conjunction with the environmentally enhanced LLM, the SBS strategy facilitates comprehensive exploration of both token-level and step-level search spaces, overcoming the limitations of conventional greedy search methods. Additionally, to evaluate the effectiveness of embodied task planning, we introduce the Trajectory Match Score (TMS), a robust evaluation metric that leverages state-based simulation to assess plan success. Through extensive experiments on standard benchmarks, our framework demonstrates substantial improvements in both plan generation quality and task success rates, advancing the state-of-the-art in embodied task planning.
Yihan Tang, Yong Xu 0007, Ruotao Xu, Yan Huang 0031, Si Wu 0002, Patrick Le Callet
ICME3
2025 Image debanding using cross-scale invertible networks with banded deformable convolutions
Yuhui Quan, Xuyi He, Ruotao Xu, Yong Xu 0007, Hui Ji 0002
Neural Networks3
2024 PU-SDF: Arbitrary-Scale Uniformly Upsampling Point Clouds via Signed Distance Functions
abstract
Point cloud upsampling is a crucial technique for improving the performance of 3D data understanding, enabling the creation of dense and uniform point clouds from raw and sparse input data. However, most existing methods exhibit limitations in dealing with different scaling factors, either requiring multiple networks or producing unevenly distributed points. To tackle this issue, this paper proposes a deep learning method called PU-SDF, which consists of a local-feature-based signed distance function network (LSDF-network) and a 3D-grid query points generation module (GPGM). The proposed LSDF-network can perform point cloud upsampling at arbitrary rates and can be easily adapted to handle unseen datasets. Additionally, we propose the GPGM that generates uniform and unlimited query points in sparse voxel space with arbitrary resolution. Extensive qualitative and quantitative evaluations demonstrate the superior performance of the PU-SDF method, achieving state-of-the-art point cloud upsampling performance.
Shaohui Pan, Yong Xu 0007, Ruotao Xu
3DV3
2024 View-aligned pixel-level feature aggregation for 3D shape classification
Yong Xu 0007, Shaohui Pan, Ruotao Xu, Haibin Ling
Comput. Vis. Image Underst.3
2024 Deep Single Image Defocus Deblurring via Gaussian Kernel Mixture Learning
abstract
This paper proposes an end-to-end deep learning approach for removing defocus blur from a single defocused image. Defocus blur is a common issue in digital photography that poses a challenge due to its spatially-varying and large blurring effect. The proposed approach addresses this challenge by employing a pixel-wise Gaussian kernel mixture (GKM) model to accurately yet compactly parameterize spatially-varying defocus point spread functions (PSFs), which is motivated by the isotropy in defocus PSFs. We further propose a grouped GKM (GGKM) model that decouples the coefficients in GKM, so as to improve the modeling accuracy with an economic manner. Afterward, a deep neural network called GGKMNet is then developed by unrolling a fixed-point iteration process of GGKM-based image deblurring, which avoids the efficiency issues in existing unrolling DNNs. Using a lightweight scale-recurrent architecture with a coarse-to-fine estimation scheme to predict the coefficients in GGKM, the GGKMNet can efficiently recover an all-in-focus image from a defocused one. Such advantages are demonstrated with extensive experiments on five benchmark datasets, where the GGKMNet outperforms existing defocus deblurring methods in restoration quality, as well as showing advantages in terms of model complexity and computational efficiency.
Yuhui Quan, Zicong Wu, Ruotao Xu, Hui Ji 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Enhancing texture representation with deep tracing pattern encoding
abstract
Texture representation is a challenging problem due to the complex underlying physics of texture as well as the variations caused by changes in viewpoint. Recent progress in texture analysis has been made by the power of convolutional neural networks (CNNs) in feature learning . However, most current methods aggregate the features from the last convolutional layer of the CNN to obtain a global feature vector, which fails to leverage shallow low-level visual cues and cross-layer feature patterns, limiting their performance. In this paper, we propose to trace the features generated along the convolutional layers via a histogram of local 3D invariant binary patterns, called deep tracing patterns. This leads to a highly discriminative yet robust global feature representation module. Building such a module into a CNN backbone, we develop an effective approach for texture recognition. Extensive experiments on six benchmark datasets show that the proposed approach provides a discriminative and robust texture descriptor , with state-of-the-art performance achieved.
Zhile Chen, Yuhui Quan, Ruotao Xu, Yong Xu 0007
Pattern Recognit.3
2024 Image Smoothing via Multiscale Global Perception
abstract
Image smoothing provides a fundamental operation for image processing, with a broad spectrum of applications. It is a challenging task which requires global analysis on image patterns with scale awareness. Existing deep models for image smoothing are insufficiently efficient in global perception and multi-scale processing. This paper proposes a deep model with an efficient multi-scale fusion architecture and a series of global processing blocks. The architecture enhances multi-scale feature flow by incorporating features of different scales into both the encoder and decoder blocks of a U-shape network, with multi-scale feature fusion modules inserted between the encoder and the decoder. The global processing blocks leverage the multi-axis processing mechanism to achieve joint local and global perception. Benefiting from these two key designs, our proposed model enjoys superiority in both smoothing performance and computational complexity, as demonstrated in the experiments on two benchmark datasets.
Xuyi He, Yuhui Quan, Yong Xu 0007, Ruotao Xu
IEEE Signal Process. Lett.4
2024 Spatial Adaptive Filter Network With Scale-Sharing Convolution for Image Demoiréing
abstract
Removing moiré patterns is a challenging task as it is a spatially varying degradation that varies in shape, color and scale. Existing image restoration models often rely on static convolutional neural networks (CNNs)-based architectures, and hence potentially suboptimal for addressing the diverse manifestations of moiré patterns across different images and spatial positions. To this end, we propose a spatially adaptive neural network for image demoiréing. This network introduces a dual-branch filter prediction module engineered to predict pixel-wise adaptive filters that can process moiré patterns of varying orientations and color-shift issues. To further tackle the challenge presented by scale variability, a scale-sharing convolution module is proposed, utilizing pixel-wise adaptive filters with multiple dilations to handle moiré patterns of different sizes but similar shapes effectively. Upon extensive evaluations of three benchmark datasets, our model consistently outperforms existing methods, yielding a PSNR improvement of over 0.37dB across all evaluated datasets and providing additional benefits in terms of model size.
Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Zhu Liang Yu
IEEE Signal Process. Lett.3
2023 Deep Video Demoiréing via Compact Invertible Dyadic Decomposition
abstract
Removing moiré patterns from videos recorded on screens or complex textures is known as video demoiréing. It is a challenging task as both structures and textures of an image usually exhibit strong periodic patterns, which thus are easily confused with moiré patterns and can be significantly erased in the removal process. By interpreting video demoiréing as a multi-frame decomposition problem, we propose a compact invertible dyadic network called CIDNet that progressively decouples latent frames and the moiré patterns from an input video sequence. Using a dyadic cross-scale coupling structure with coupling layers tailored for multi-scale processing, CIDNet aims at disentangling the features of image patterns from that of moiré patterns at different scales, while retaining all latent image features to facilitate reconstruction. In addition, a compressed form for the network’s output is introduced to reduce computational complexity and alleviate overfitting. The experiments show that CIDNet outperforms existing methods and enjoys the advantages in model size and computational efficiency.
Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu
ICCV4
2023 Fingerprinting Deep Image Restoration Models
abstract
Fingerprinting is a promising non-invasive method for protecting the intellectual property rights (IPR) of deep neural network (DNN) models. It extracts a feature called a fingerprint from a DNN model to identify its ownership. Existing fingerprinting methods focus only on classification-related models that map images to labels, while inapplicable to models for image restoration that map images to images. This paper proposes a fingerprinting framework for DNN models of image restoration. The proposed framework defines the fingerprint using a critical image, which exhibits strongly discriminative patterns and is robust to modest model modifications. Model ownership is then verified by comparing the distance of color histograms and local gradient pattern histograms of critical images between the suspect and source models. We apply the proposed framework to two representative tasks, denoising and super-resolution. It outperforms the baselines of fingerprinting and competes against existing invasive model watermarking methods.
Yuhui Quan, Huan Teng, Ruotao Xu, Jun Huang 0007, Hui Ji 0002
ICCV3
2022 Deep Scale-Aware Image Smoothing
abstract
Image smoothing, a technique for smoothing out insignificant textures while preserving meaningful structures, is an important component in many vision and graphics applications. Scale-awareness plays a fundamental role in image smoothing, as insignificant textures and noise usually are at fine scales while meaningful boundary objects are at coarse scales. This paper proposes a deep-learning-based scale-aware image smoothing method, which is built on a downscaling-upscaling mechanism with attention. The downscaling mechanism is for predicting large-scale salient structures, and the upscaling mechanism is for identifying and inferring insignificant small-scale details from the salient structures. In the experiments, the proposed one provides a noticeable performance improvement over recent methods.
Kunkun Qin, Ruotao Xu, Hui Ji 0002
ICASSP3
2022 Deep Blind Image Quality Assessment Using Dual-Order Statistics
abstract
Deep convolutional neural networks (CNNs) have become a promising approach to blind image quality assessment (BIQA). Existing CNN-based BIQA methods often employ global average pooling (GAP) to aggregate feature maps into a fixed-size representation for regression, so as to handle input images with varying sizes. However, GAP is only capable of extracting the first-order statistics of feature distributions, which is ineffective for distinguishing complex distortions that cause local degradation or preserve global features. To tackle this problem, we introduce the second-order global covariance pooling (GCP) for aggregating feature maps, leading to a more distortion-sensitive and more discriminative global representation. By incorporating GCP and GAP into a ResNet backbone, we propose an effective deep model for BIQA. The experimental results on five BIQA benchmark datasets, including both the synthetic and authentic ones, have demon-strated the excellent performance of the proposed method.
Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Ruotao Xu
ICME4
2022 No-Reference Image Quality Assessment Using Dynamic Complex-Valued Neural Model
abstract
Deep convolutional neural networks (CNNs) have become a promising approach to no-reference image quality assessment (NR-IQA). This paper aims at improving the power of CNNs for NR-IQA in two aspects. Firstly, motivated by the deep connection between complex-valued transforms and human visual perception, we introduce complex-valued convolutions and phase-aware activations beyond traditional real-valued CNNs, which improves the accuracy of NR-IQA without bringing noticeable additional computational costs. Secondly, considering the content-awareness of visual quality perception, we include a dynamic filtering module for better extracting content-aware features, which predicts features based on both local content and global semantics. These two improvements lead to a complex-valued content-aware neural NR-IQA model with good generalization. Extensive experiments on both synthetically and authentically distorted data have demonstrated the state-of-the-art performance of the proposed approach.
Zihan Zhou 0007, Yong Xu 0007, Ruotao Xu, Yuhui Quan
ACM Multimedia3
2021 Structure-Texture Image Decomposition Using Discriminative Patch Recurrence
abstract
Morphology component analysis provides an effective framework for structure-texture image decomposition, which characterizes the structure and texture components by sparsifying them with certain transforms respectively. Due to the complexity and randomness of texture, it is challenging to design effective sparsifying transforms for texture components. This paper aims at exploiting the recurrence of texture patterns, one important property of texture, to develop a nonlocal transform for texture component sparsification. Since the plain patch recurrence holds for both cartoon contours and texture regions, the nonlocal sparsifying transform constructed based on such patch recurrence sparsifies both the structure and texture components well. As a result, cartoon contours could be wrongly assigned to the texture component, yielding ambiguity in decomposition. To address this issue, we introduce a discriminative prior on patch recurrence, that the spatial arrangement of recurrent patches in texture regions exhibits isotropic structure which differs from that of cartoon contours. Based on the prior, a nonlocal transform is constructed which only sparsifies texture regions well. Incorporating the constructed transform into morphology component analysis, we propose an effective approach for structure-texture decomposition. Extensive experiments have demonstrated the superior performance of our approach over existing ones.
Ruotao Xu, Yong Xu 0007, Yuhui Quan
IEEE Trans. Image Process.1
2021 Multi-View 3D Shape Recognition via Correspondence-Aware Deep Learning
abstract
In recent years, multi-view learning has emerged as a promising approach for 3D shape recognition, which identifies a 3D shape based on its 2D views taken from different viewpoints. Usually, the correspondences inside a view or across different views encode the spatial arrangement of object parts and the symmetry of the object, which provide useful geometric cues for recognition. However, such view correspondences have not been explicitly and fully exploited in existing work. In this paper, we propose a correspondence-aware representation (CAR) module, which explicitly finds potential intra-view correspondences and cross-view correspondences via k NN search in semantic space and then aggregates the shape features from the correspondences via learned transforms. Particularly, the spatial relations of correspondences in terms of their viewpoint positions and intra-view locations are taken into account for learning correspondence-aware features. Incorporating the CAR module into a ResNet-18 backbone, we propose an effective deep model called CAR-Net for 3D shape classification and retrieval. Extensive experiments have demonstrated the effectiveness of the CAR module as well as the excellent performance of the CAR-Net.
Yong Xu 0007, Chaoda Zheng, Ruotao Xu, Yuhui Quan, Haibin Ling
IEEE Trans. Image Process.3
2021 Factorized Tensor Dictionary Learning for Visual Tensor Data Completion
abstract
This paper aims at developing a dictionary-learning-based method for completing the visual tensor data with missing elements. Traditional dictionary learning approaches suffer from very high computational costs when processing high-dimensional tensor data. Some existing approaches for acceleration impose orthogonality constraints or rank-one decompositions on dictionary atoms; however, the expressibility of the resulting dictionary is rather limited. To address such issues, we propose a convolutional analysis model for tensor dictionary learning, where the update of sparse coefficients during dictionary learning is simple and fast. Furthermore, we propose an orthogonality-constrained convolutional factorization scheme for dictionary construction, in which each tensor dictionary atom is factorized by the convolution of two atoms selected from two orthogonal factor dictionaries respectively. This factorization scheme enables us to efficiently learn an expressive dictionary with over-completeness and non-rank-one atoms. Based on our convolutional analysis model and factorization scheme, an effective yet efficient dictionary learning method is proposed for visual tensor completion. Extensive experiments show that, our method not only outperforms existing dictionary-based approaches with relatively-low time cost, but also outperforms recent low-rank approaches.
Ruotao Xu, Yong Xu 0007, Yuhui Quan
IEEE Trans. Multim.1
2021 Image Quality Assessment Using Kernel Sparse Coding
abstract
One key in image quality assessment (IQA) is the design of image representations that can capture the changes of image structures caused by distortions. Recent studies show that sparse coding has emerged as a promising approach to analyzing image structures for IQA. However, existing sparse-coding-based IQA approaches use linear coding models, which ignore the nonlinearities of manifolds of image patches and thus cannot analyze complex image structures well. To overcome such a weakness, in this paper, we introduce nonlinear sparse coding to IQA. A kernel dictionary construction scheme is proposed, which combines analytic dictionaries and learnable dictionaries to guarantee both the stability and effectiveness of kernel sparse coding in the context of IQA. Built upon the kernel dictionary construction, an effective full-reference IQA metric is developed. Benefiting from the considerations on nonlinearities during sparse coding, the proposed IQA metric not only characterizes image distortions better, but also achieves improvement on the consistency with subjective perception, when compared to the metrics built upon linear sparse coding. Such benefits are demonstrated with the experimental results on eight benchmark datasets in terms of common criteria.
Zihan Zhou 0007, Jing Li 0026, Yuhui Quan, Ruotao Xu
IEEE Trans. Multim.4
2020 Cartoon-Texture Image Decomposition using Orientation Characteristics in Patch Recurrence
abstract
Cartoon-texture image decomposition is about decomposing an image into the linear sum of two layers: cartoon and texture, where the key challenge is how to resolve the ambiguity between two layers. It is observed that the recurrence of texture patches occurs along multiple orientations, and the recurrence of cartoon patches only occurs along certain orientations. This paper proposes to separate these two layers by exploiting their orientation characteristics of image patch recurrence, i.e., isotropy property of texture patch recurrence versus anisotropy property of cartoon patch recurrence. Together with the sparsity-based regularizations in the image domain, a variational method is then developed in this paper for cartoon-texture decomposition. The experiments show that the proposed method noticeably outperforms many well-established ones on test images.
Ruotao Xu, Yong Xu 0007, Yuhui Quan, Hui Ji 0002
SIAM J. Imaging Sci.1
2019 Multi-view Rank Pooling for 3D Object Recognition*
abstract
3D shape recognition via deep learning is drawing more and more attention due to huge industry interests. As 3D deep learning methods emerged, the view-based approaches have gained considerable success in object classification. Most of these methods focus on designing a pooling scheme to aggregate CNN features of multi-view images into a single compact one. However, these view-wise pooling techniques suffer from loss of visual information. To deal with this issue, an adaptive rank pooling layer is introduced in this paper. Unlike max-pooling which only considers the maximum or mean-pooling that treats each element indiscriminately, the proposed pooling layer takes all the elements into account and dynamically adjusts their importances during the training. Experiments conducted on ModelNet40 and ModelNet10 shows both efficiency and accuracy gain when inserting such a layer into a baseline CNN architecture.
Chaoda Zheng, Yong Xu 0007, Ruotao Xu, Hongyu Chi, Yuhui Quan
VCIP3
2019 Attention with structure regularization for action recognition
Yuhui Quan, Ruotao Xu, Hui Ji 0002
Comput. Vis. Image Underst.3