EDBT 2026 Demo / reviewers in the wild / expert
Tao Dai 0001
dblp:54/875-1
· DBLP profile ↗
112ranked-venue papers
16as first author
79since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 12 first-author · 55 since 2021Artificial intelligence and machine learning · 72 · 9 first-author · 59 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data AugmentationabstractThe generalization capability of deepfake detectors is crucial for real-world applications. Data augmentation to generate synthetic fake faces has served as an effective strategy to enhance generalization. Interestingly, current state-of-the-art (SoTA) methods rely on fixed augmentation strategies, raising a fundamental question: Can a single static augmentation approach suffice, or does the diversity of forgery features necessitate dynamic strategies? We argue that existing methods overlook the evolving complexity of real-world forgery patterns, such as facial warping, expression manipulation, and compression artifacts, which cannot be fully simulated by fixed policies. To bridge this gap, we propose CRDA (Curriculum Reinforcement-Learning Data Augmentation), a novel framework that guides the detector to progressively master multi-domain forgery features from simple to complex. CRDA synthesizes augmented samples using a configurable pool of forgery operations and dynamically generates adversarial samples tailored to the detector’s current learning state. Key to our approach is the integration of reinforcement learning (RL) and causal inference. To efficiently explore the vast augmentation space, an RL agent dynamically selects augmentation actions based on the detector’s performance, ensuring continuous adaptation to increasingly challenging forgeries. Simultaneously, the agent’s output is designed to introduce variations in action spaces, generating heterogeneous forgery patterns. These variations are guided by causal inference theory, which mitigates spurious correlations by suppressing task-irrelevant biases and enforcing the model to focus on causally invariant features. This integration ensures robust generalization by decoupling synthetic augmentation patterns from the model’s learned representations. Extensive experiments demonstrate that the proposed method significantly improves the generalizability of the detector, achieving superior performance compared to state-of-the-art methods on multiple cross-domain datasets. Yuxuan Chou, Tao Dai 0001, Shutao Xia |
AAAI | 5 |
| 2026 | CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly DetectionabstractDeep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly detection tasks, which limits their generalizability to other 3D tasks. In contrast, self-supervised point cloud models aim for general representation learning, yet our investigation reveals that these classical models are suboptimal at anomaly detection under the unified fine-tuning paradigm. This motivates us to develop a more generalizable 3D model that can effectively detect anomalies without relying on task-specific designs. Interestingly, we find that using only the curvature of each point as its anomaly score already outperforms several classical self-supervised and dedicated anomaly detection models, highlighting the critical role of curvature in 3D anomaly detection. In this paper, we propose a Curvature-Augmented Self-supervised Learning (CASL) framework based on a reconstruction paradigm. Built upon the classical U-Net architecture, our approach introduces multi-scale curvature prompts to guide the decoder in predicting the coordinates of each point. Without relying on any dedicated anomaly detection mechanisms, it achieves leading detection performance through straightforward anomaly classification fine-tuning. Moreover, the learned representations generalize well to standard 3D understanding tasks such as point cloud classification. Yaohua Zha, Xue Yuerong, Chunlin Fan, Yuansong Wang, Tao Dai 0001, Ke Chen 0004, Shutao Xia |
AAAI | 5 |
| 2026 | Learning gated experts for segment anything in the wild
Yizhen Guo, Hang Guo 0002, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
Pattern Recognit. | 3 |
| 2026 | Perceptual image compression with textual side information
Shiyu Qin, Bin Chen 0011, Yujun Huang, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
Pattern Recognit. | 5 |
| 2025 | DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object DetectionabstractKnowledge distillation (KD) has recently gained great success in the field of object detection. By transferring the knowledge of the spatial or channel domain from the teacher model to the student model, it allows for a more compact representation with minimal performance loss. Despite this progress, existing KD methods typically treat knowledge from spatial or channel domains independently, ignoring the exploitation of the mutual relationship between these domains. In this work, we first explore the connection between spatial and channel domains and find there exists a strong correlation between them, i.e. the salient channels tend to contain significant object regions in the spatial domain. Motivated by this observation, we propose DCSF-KD, a novel Dynamic Channel-wise Spatial Feature Knowledge Distillation framework for object detection by fully exploiting both spatial and channel knowledge. Specifically, we introduce channel-wise spatial feature distillation and global channel attention distillation, using information from both domains to improve the accuracy of the student network. Experiments demonstrate that our DCSF-KD outperforms existing detection methods on both homogeneous and heterogeneous teacher-student network pairs. For example, when using the MaskRCNN-Swin detector as the teacher, and based on RetinaNet and FCOS with ResNet-50 on MS COCO, our DCSF-KD can achieve 41.9% and 44.1% mAP, respectively. Tao Dai 0001, Hang Guo 0002, Jinbao Wang 0001, Zexuan Zhu 0001 |
AAAI | 1 |
| 2025 | GCD-Sampling: A General Cross-scale Decoupled Sampling for Point CloudabstractSampling strategy (e.g., fixed farthest point sampling) of point cloud has been an essential step for developing practical solutions in 3D computer vision tasks. Previous fixed sampling is simple, but suffer from suboptimal performance for downstream tasks. To adapt to target networks properly, adaptive sampling methods with trainable parameters have been recently developed to enhance the performance. However, existing adaptive sampling methods still suffer from the over-coupling problem of target network, and thus become model-specific, which limits their practical applications. To address this issue, we propose a novel general cross-scale decoupled sampling method (GCD-sampling) for point cloud, which consists of original feature cache, cross-scale feature fusion and convex combination learning for better feature extraction. To reduce the coupling relationship with the target task network, our method only utilizes the point cloud coordinates as the input and output of itself. Besides, we introduce an arbitrary scale structure to enable parameter sharing across multi-scale sampling in point cloud networks. Extensive experiments on different architectures demonstrate the effectiveness of our method over other existing adaptive sampling methods. Tao Dai 0001, Yanzi Wang, Jianyu Xiong, Yaohua Zha, Shutao Xia, Zexuan Zhu 0001 |
AAAI | 1 |
| 2025 | Adaptive Multi-Scale Decomposition Framework for Time Series ForecastingabstractTransformer-based and MLP-based methods have emerged as leading approaches in time series forecasting (TSF). However, real-world time series often show different patterns at different scales, and future changes are shaped by the interplay of these overlapping scales, requiring high-capacity models. While Transformer-based methods excel in capturing long-range dependencies, they suffer from high computational complexities and tend to overfit. Conversely, MLP-based methods offer computational efficiency and adeptness in modeling temporal dynamics, but they struggle with capturing temporal patterns with complex scales effectively. Based on the observation of multi-scale entanglement effect in time series, we propose a novel MLP-based Adaptive Multi-Scale Decomposition (AMD) framework for TSF. Our framework decomposes time series into distinct temporal patterns at multiple scales, leveraging the Multi-Scale Decomposable Mixing (MDM) block to dissect and aggregate these patterns. Complemented by the Dual Dependency Interaction (DDI) block and the Adaptive Multi-predictor Synthesis (AMS) block, our approach effectively models both temporal and channel dependencies and utilizes autocorrelation to refine multi-scale data integration. Comprehensive experiments demonstrate our AMD framework not only overcomes the limitations of existing methods but also consistently achieves state-of-the-art performance across various datasets. Yifan Hu 0006, Peiyuan Liu, Peng Zhu 0002, Dawei Cheng, Tao Dai 0001 |
AAAI | 5 |
| 2025 | CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-TuningabstractDeep learning (e.g., Transformer) has been widely and successfully used in multivariate time series forecasting (MTSF). Unlike existing methods that focus on training models from a single modal of time series input, large language models (LLMs) based MTSF methods with cross-modal text and time series input have recently shown great superiority, especially with limited temporal data. However, current LLM-based MTSF methods usually focus on adapting and fine-tuning LLMs, while neglecting the distribution discrepancy between textual and temporal input tokens, thus leading to sub-optimal performance. To address this issue, we propose a novel Cross-Modal LLM Fine-Tuning (CALF) framework for MTSF by reducing the distribution discrepancy between textual and temporal data, which mainly consists of the temporal target branch with temporal input and the textual source branch with aligned textual input. To reduce the distribution discrepancy, we develop the cross-modal match module to first align cross-modal input distributions. Additionally, to minimize the modality distribution gap in both feature and output spaces, feature regularization loss is developed to align the intermediate features between the two branches for better weight updates, while output consistency loss is introduced to allow the output representations of both branches to correspond effectively. Thanks to the modality alignment, CALF establishes state-of-the-art performance for both long-term and short-term forecasting tasks with low computational complexity, and exhibits favorable few-shot and zero-shot abilities similar to that in LLMs. Peiyuan Liu, Hang Guo 0002, Tao Dai 0001, Naiqi Li, Jigang Bao, Xudong Ren, Yong Jiang 0001, Shutao Xia |
AAAI | 3 |
| 2025 | EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language ModelsabstractVision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to incorrect retrieval results. To address this problem, we propose the Entity Visual Description enhanced CLIP (EvdCLIP), designed to leverage the visual knowledge of entities to enrich queries. Specifically, since humans recognize entities through visual cues, we employ a large language model (LLM) to generate Entity Visual Descriptions (EVDs) as alignment cues to complement textual data. These EVDs are then integrated into raw queries to create visually-rich, EVD-enhanced queries. Furthermore, recognizing that EVD-enhanced queries may introduce noise or low-quality expansions, we develop a novel, trainable EVD-aware Rewriter (EaRW) for vision-language retrieval tasks. EaRW utilizes EVD knowledge and the generative capabilities of the language model to effectively rewrite queries. With our specialized training strategy, EaRW can generate high-quality and low-noise EVD-enhanced queries. Extensive quantitative and qualitative experiments on image-text retrieval benchmarks validate the superiority of EvdCLIP on vision-language retrieval tasks. Guanghao Meng, Sunan He, Jinpeng Wang 0002, Tao Dai 0001, Jieming Zhu, Qing Li 0006, Rui Zhang 0003, Yong Jiang 0001 |
AAAI | 4 |
| 2025 | Diffusion Prior Interpolation for Flexibility Real-World Face Super-ResolutionabstractDiffusion models represent the state-of-the-art in generative modeling. Due to their high training costs, many works leverage pre-trained diffusion models' powerful representations for downstream tasks, such as face super-resolution (FSR), through fine-tuning or prior-based methods. However, relying solely on priors without supervised training makes it challenging to meet the pixel-level accuracy requirements of discrimination task. Although prior-based methods can achieve high fidelity and high-quality results, ensuring consistency remains a significant challenge. In this paper, we propose a masking strategy with strong and weak constraints and iterative refinement for real-world FSR, termed Diffusion Prior Interpolation (DPI). We introduce conditions and constraints on consistency by masking different sampling stages based on the structural characteristics of the face. Furthermore, we propose a condition Corrector (CRT) to establish a reciprocal posterior sampling process. DPI can balance consistency and diversity and can be seamlessly integrated into pre-trained models. In extensive experiments conducted on synthetic and real datasets, along with consistency validation in face recognition, DPI demonstrates superiority over SOTA FSR methods. Tao Dai 0001, Naiqi Li, Jinmin Li, Shutao Xia |
AAAI | 2 |
| 2025 | Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningabstractVisual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which in practice may not hold true. Multiple suitable prompts may exist, but individually they often fall short, leading to difficulties in selection and the exclusion of useful context. To address this, we propose a new perspective: prompt condensation. Rather than relying on a single prompt, candidate prompts collaborate to efficiently integrate informative contexts without sacrificing resolution. We devise Condenser, a lightweight external plugin that compresses relevant fine-grained context across multiple prompts. Optimized end-to-end with the backbone, Condenser ensures accurate integration of contextual cues. Experiments demonstrate Condenser outperforms state-of-the-arts across benchmark tasks, showing superior context compression, scalability with more prompts, and enhanced computational efficiency compared to ensemble methods, positioning it as a highly competitive solution for VICL. Code is open-sourced at https://github.com/gimpong/CVPR25-Condenser. Jinpeng Wang 0002, Tianci Luo, Yaohua Zha, Ruisheng Luo, Bin Chen 0011, Tao Dai 0001, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
CVPR | 7 |
| 2025 | MambaIRv2: Attentive State Space RestorationabstractThe Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts the full utilization of pixels across the image and thus presents new challenges in image restoration. In this work, we propose MambaIRv2, which equips Mamba with the non-causal modeling ability similar to ViTs to reach the attentive state space restoration model. Specifically, the proposed attentive state-space equation allows to attend beyond the scanned sequence and facilitate image unfolding with just one single scan. Moreover, we further introduce a semantic-guided neighboring mechanism to encourage interaction between distant but similar pixels. Extensive experiments show our MambaIRv2 outperforms SRFormer by even 0.35dB PSNR for lightweight SR even with 9.3% less parameters and suppresses HAT on classic SR by up to 0.29dB. Code is available at https://github.com/csguoh/MambaIR. Hang Guo 0002, Yaohua Zha, Yulun Zhang 0001, Wenbo Li 0002, Tao Dai 0001, Shutao Xia, Yawei Li 0001 |
CVPR | 6 |
| 2025 | Protecting Your Video Content: Disrupting Automated Video-based LLM AnnotationsabstractRecently, video-based large language models (video-based LLMs) have achieved impressive performance across various video comprehension tasks. However, this rapid advancement raises significant privacy and security concerns, particularly regarding the unauthorized use of personal video data in automated annotation by video-based LLMs. These unauthorized annotated video-text pairs can then be used to improve the performance of downstream tasks, such as text-to-video generation. To safeguard personal videos from unauthorized use, we propose two series of protective video watermarks with imperceptible adversarial perturbations, named Ramblings and Mutes. Concretely, Ramblings aim to mislead video-based LLMs into generating inaccurate captions for the videos, thereby degrading the quality of video annotations through inconsistencies between video content and captions. Mutes, on the other hand, are designed to prompt video-based LLMs to produce exceptionally brief captions, lacking descriptive detail. Extensive experiments demonstrate that our video watermarking methods effectively protect video data by significantly reducing video annotation performance across various video-based LLMs, showcasing both stealthiness and robustness in protecting personal video content. Our code is available at https://github.com/ttthhl/Protecting_Your_Video_Content. Haitong Liu, Kuofeng Gao, Yang Bai 0011, Jinmin Li, Jinxiao Shan, Tao Dai 0001, Shutao Xia |
CVPR | 6 |
| 2025 | Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionabstractPoint cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video understanding, the limited availability of 4D data and the high computational cost of training 4D-specific models remain significant obstacles. In this paper, we investigate the potential of transferring pre-trained static 3D point cloud models to the 4D domain, pointing out the limitations of static models that capture only spatial information while neglecting temporal dynamics. To address this, we propose a novel Cross-frame Spatio-temporal Adaptation (CSA) strategy by introducing the Point Tube Adapter as the embedding layer and the Geometric Constraint Temporal Adapter (GCTA) to enforce temporal consistency across frames. This strategy extracts both short-term and long-term temporal dynamics, effectively integrating them with spatial features and enriching the model’s understanding of temporal changes in point cloud videos. Extensive experiments on 3D action and gesture recognition tasks demonstrate that our method achieves state-of-the-art performance, establishing its effectiveness for point cloud video understanding. Code is available at: https://github.com/LvBaixuan/Point-CSA. Baixuan Lv, Yaohua Zha, Tao Dai 0001, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 3 |
| 2025 | PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterabstractApplying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementary information in the intermediate layer, thereby failing to fully unlock the potential of pre-trained models. To overcome this limitation, we propose an orthogonal solution: Point Mamba Adapter (PMA), which constructs an ordered feature sequence from all layers of the pre-trained model and leverages Mamba to fuse all complementary semantics, thereby promoting comprehensive point cloud understanding. Constructing this ordered sequence is non-trivial due to the inherent isotropy of 3D space. Therefore, we further propose a geometry-constrained gate prompt generator (G2PG) shared across different layers, which applies shared geometric constraints to the output gates of the Mamba and dynamically optimizes the spatial order, thus enabling more effective integration of multi-layer information. Extensive experiments conducted on challenging point cloud datasets across various tasks demonstrate that our PMA elevates the capability for point cloud understanding to a new level by fusing diverse complementary intermediate features. Code is available at https://github.com/zyh16143998882/PMA. Yaohua Zha, Yanzi Wang, Hang Guo 0002, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhihao Ouyang, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 5 |
| 2025 | LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video CodecabstractExisting Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However, NeRV implicitly stores all video information in the network, requiring post network compression techniques such as pruning. To integrate explicit compression and implicit representation into an end-to-end framework, this study proposes a novel neural representation-based video compression paradigm called Latent code based Neural representation video compression (LNeRV). Specifically, LNeRV consists of hierarchy feature grids, synthesis network and entropy coding network. With a single-stage training process, LNeRV achieves video compression and dynamically allocates bits according to video complexity, better fitting dynamic videos. We provide a comprehensive compression-to-decompression workflow for our approach. Extensive experimental results verify the effectiveness of our LNeRV. Jiahong Chen, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICASSP | 5 |
| 2025 | FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningabstractVisual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale step, leading to the complexity and runtime scaling dramatically with image resolution. To address this challenge, we propose FastVAR, a post-training acceleration method for efficient resolution scaling with VARs. Our key finding is that the majority of latency arises from the large-scale step where most tokens have already converged. Leveraging this observation, we develop the cached token pruning strategy that only forwards pivotal tokens for scale-specific modeling while using cached tokens from previous scale steps to restore the pruned slots. This significantly reduces the number of forwarded tokens and improves the efficiency at larger resolutions. Experiments show the proposed FastVAR can further speedup FlashAttention-accelerated VAR by 2.7$\times$ with negligible performance drop of <1%. We further extend FastVAR to zero-shot generation of higher resolution images. In particular, FastVAR can generate one 2K image with 15GB memory footprints in 1.5s on a single NVIDIA 3090 GPU. Code is available at https://github.com/csguoh/FastVAR. Hang Guo 0002, Yawei Li 0001, Taolin Zhang 0003, Jiangshan Wang, Tao Dai 0001, Shutao Xia, Luca Benini |
ICCV | 5 |
| 2025 | Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression
Shiyu Qin, Jinpeng Wang 0002, Yimin Zhou 0011, Bin Chen 0011, Tianci Luo, Baoyi An 0002, Tao Dai 0001, Shutao Xia, Yaowei Wang 0001 |
ICCV | 7 |
| 2025 | IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion ModelsabstractFine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation recipes are not inference-efficient. Specifically, additional post-training quantization (PTQ) on tuned weights is needed during deployment, which results in noticeable performance drop when the bit-width is low. Based on this observation, we introduce IntLoRA, which adapts quantized diffusion models with integer-type low-rank parameters, to include inference efficiency during tuning. Specifically, IntLoRA enables pre-trained weights to remain quantized during training, facilitating fine-tuning on consumer-level GPUs. During inference, IntLoRA weights can be seamlessly merged into pre-trained weights to directly obtain quantized downstream weights without PTQ. Extensive experiments show our IntLoRA achieves significant speedup on both training and inference without losing performance. Hang Guo 0002, Yawei Li 0001, Tao Dai 0001, Shutao Xia, Luca Benini |
ICML | 3 |
| 2025 | TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series ForecastingabstractTime series forecasting methods generally fall into two main categories: Channel Independent (CI) and Channel Dependent (CD) strategies. While CI overlooks important covariate relationships, CD captures all dependencies without distinction, introducing noise and reducing generalization. Recent advances in Channel Clustering (CC) aim to refine dependency modeling by grouping channels with similar characteristics and applying tailored modeling techniques. However, coarse-grained clustering struggles to capture complex, time-varying interactions effectively. To address these challenges, we propose TimeFilter, a GNN-based framework for adaptive and fine-grained dependency modeling. After constructing the graph from the input sequence, TimeFilter refines the learned spatial-temporal dependencies by filtering out irrelevant correlations while preserving the most critical ones in a patch-specific manner. Extensive experiments on 13 real-world datasets from diverse application domains demonstrate the state-of-the-art performance of TimeFilter. The code is available at https://github.com/TROUBADOUR000/TimeFilter. Yifan Hu 0006, Guibin Zhang, Peiyuan Liu, Disen Lan, Naiqi Li, Dawei Cheng, Tao Dai 0001, Shutao Xia, Shirui Pan |
ICML | 7 |
| 2025 | 3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric PriorsabstractExisting multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camera systems, commonly found in virtual reality and autonomous driving applications. To address this limitation, we propose 3D-LMVIC, a novel learning-based multi-view image compression framework that leverages 3D Gaussian Splatting to derive geometric priors for accurate disparity estimation. Furthermore, we introduce a depth map compression model to minimize geometric redundancy across views, along with a multi-view sequence ordering strategy based on a defined distance measure between views to enhance correlations between adjacent views. Experimental results demonstrate that 3D-LMVIC achieves superior performance compared to both traditional and learning-based methods. Additionally, it significantly improves disparity estimation accuracy over existing two-view approaches. Yujun Huang, Bin Chen 0011, Niu Lian, Xin Wang 0001, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICML | 6 |
| 2025 | TimeBridge: Non-Stationarity Matters for Long-term Time Series ForecastingabstractNon-stationarity poses significant challenges for multivariate time series forecasting due to the inherent short-term fluctuations and long-term trends that can lead to spurious regressions or obscure essential long-term relationships. Most existing methods either eliminate or retain non-stationarity without adequately addressing its distinct impacts on short-term and long-term modeling. Eliminating non-stationarity is essential for avoiding spurious regressions and capturing local dependencies in short-term modeling, while preserving it is crucial for revealing long-term cointegration across variates. In this paper, we propose TimeBridge, a novel framework designed to bridge the gap between non-stationarity and dependency modeling in long-term time series forecasting. By segmenting input series into smaller patches, TimeBridge applies Integrated Attention to mitigate short-term non-stationarity and capture stable dependencies within each variate, while Cointegrated Attention preserves non-stationarity to model long-term cointegration across variates. Extensive experiments show that TimeBridge consistently achieves state-of-the-art performance in both short-term and long-term forecasting. Additionally, TimeBridge demonstrates exceptional performance in financial forecasting on the CSI 500 and S&P 500 indices, further validating its robustness and effectiveness. Code is available at https://github.com/Hank0626/TimeBridge. Peiyuan Liu, Beiliang Wu, Yifan Hu 0006, Naiqi Li, Tao Dai 0001, Jigang Bao, Shutao Xia |
ICML | 5 |
| 2025 | Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised LearningabstractRecently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer from the problem of negative transfer, due to the disparity in semantic distribution between upstream data and downstream data. To address this issue, we propose an expert enhancement strategy for existing MPM-based methods. Specifically, we insert a Sparse Mixture of Experts (SMoE) layer after each block of the backbone network, which utilizes a multi-branch expert architecture with routers that allocate data of different semantics to the appropriate experts for analysis. During the pre-training phase, our expert-enhanced model not only learns universal 3D representations for the backbone network but also acquires powerful semantic routing capabilities for all expert layers. In the fine-tuning phase, we freeze all backbones and conduct end-to-end fine-tuning solely on our expert layers to adaptively select multiple experts most relevant to the semantics of each downstream data for analysis. Extensive downstream experiments demonstrate the superiority of our method, especially outperforming baseline (Point-MAE) by 5.16%, 5.86%, and 4.62% in three variants of ScanObjectNN while utilizing only 12% of its trainable parameters. Our code is released at https://github.com/chenchen1104/point_e2mae. Yaohua Zha, Naiqi Li, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
ICRA | 4 |
| 2025 | DIIN: Diffusion Iterative Implicit Networks for Arbitrary-scale Super-resolutionabstractImplicit neural representation (INR) aims to represent continuous domain signals via implicit neural functions and has achieved great success in arbitrary-scale image super-resolution (SR). However, most existing INR-based SR methods focus on learning implicit features from independent coordinate, while neglecting interactions of neighborhood coordinates, thus resulting in limited contextual awareness. In this paper, we rethink the forward process of implicit neural functions as a signal diffusion process, we propose a novel Diffusion Iterative Implicit Network (DIIN) for arbitrary-scale SR to promote global signal flow with neighborhood interactions. The DIIN framework mainly consists of stacked Diffusion Iteration Layers with dictionary cross-attention block to enrich the iterative update process with supplementary information. Besides, we develop the Position-Aware Embedding Block to strengthen spatial dependencies between consecutive input samples.Extensive experiments on public datasets demonstrate that our method achieves state-of-the-art or competitive performance, highlighting its effectiveness and efficiency for arbitrary-scale SR. Our code is available at https://github.com/Song-1205/DIIN. Tao Dai 0001, Hang Guo 0002, Zexuan Zhu 0001 |
IJCAI | 1 |
| 2025 | Efficient Differentiable Approximation of Generalized Low-rank RegularizationabstractLow-rank regularization (LRR) has been widely applied in various machine learning tasks, but the associated optimization is challenging. Directly optimizing the rank function under constraints is NP-hard in general. To overcome this difficulty, various relaxations of the rank function were studied. However, optimization of these relaxed LRRs typically depends on singular value decomposition, which is a time-consuming and nondifferentiable operator that cannot be optimized with gradient-based techniques. To address these challenges, in this paper we propose an efficient differentiable approximation of the generalized LRR. The considered LRR form subsumes many popular choices like the nuclear norm, the Schatten-p norm, and various nonconvex relaxations. Our method enables LRR terms to be appended to loss functions in a plug-and-play fashion, and the GPU-friendly operations enable efficient and convenient implementation. Furthermore, convergence analysis is presented, which rigorously shows that both the bias and the variance of our rank estimator rapidly reduce with increased sample size and iteration steps. In the experimental study, the proposed method is applied to various tasks, which demonstrates its versatility and efficiency. Code is available at https://github.com/naiqili/EDLRR. Naiqi Li, Yuqiu Xie, Peiyuan Liu, Tao Dai 0001, Yong Jiang 0001, Shutao Xia |
IJCAI | 4 |
| 2025 | Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised LearningabstractPoint clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on learning domain-specific 3D representations within a single domain, neglecting the complementary nature of cross-domain knowledge, which limits the learning of 3D representations. In this paper, we propose to learn a comprehensive Point cloud Mixture-of-Domain-Experts model (Point-MoDE) via a block-to-scene pre-training strategy. Specifically, We first propose a mixture-of-domain-expert model consisting of scene domain experts and multiple shared object domain experts. Furthermore, we propose a block-to-scene pretraining strategy, which leverages the features of point blocks in the object domain to regress their initial positions in the scene domain through object-level block mask reconstruction and scene-level block position regression. By integrating the complementary knowledge between object and scene, this strategy simultaneously facilitates the learning of both object-domain and scene-domain representations, leading to a more comprehensive 3D representation. Extensive experiments in downstream tasks demonstrate the superiority of our model. Yaohua Zha, Tao Dai 0001, Hang Guo 0002, Yanzi Wang, Bin Chen 0011, Ke Chen 0004, Shutao Xia |
IJCAI | 2 |
| 2025 | Taming Anomalies with Down-Up Sampling Networks: Group Center Preserving Reconstruction for 3D Anomaly DetectionabstractReconstruction-based methods have demonstrated very promising results for 3D anomaly detection. However, these methods face great challenges in handling high-precision point clouds due to the large scale and complex structure. In this study, a Down-Up Sampling Networks (DUS-Net) is proposed to reconstruct high-precision point clouds for 3D anomaly detection by preserving the group center geometric structure. The DUS-Net first introduces a Noise Generation module to generate noisy patches, which facilitates the diversity of training data and strengthens the feature representation for reconstruction. Then, a Down-sampling Network (Down-Net) is developed to learn an anomaly-free center point cloud from patches with noise injection. Subsequently, an Up-sampling Network (Up-Net) is designed to reconstruct high-precision point clouds by fusing multi-scale up-sampling features. Our method leverages group centers for construction, enabling the preservation of geometric structure and providing a more precise point cloud. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art (SOTA) performance, with an Object-level AUROC of 79.9% and 79.5% and a Point-level AUROC of 71.2% and 84.7% on the Real3D-AD and Anomaly-ShapeNet datasets, respectively. Hanzhe Liang, Jie Zhang 0090, Tao Dai 0001, LinLin Shen, Jinbao Wang 0001, Can Gao |
ACM Multimedia | 3 |
| 2024 | Procedural Level Generation with Diffusion Models from a Single ExampleabstractLevel generation is a central focus of Procedural Content Generation (PCG), yet deep learning-based approaches are limited by scarce training data, i.e., human-designed levels. Despite being a dominant framework, Generative Adversarial Networks (GANs) exhibit a substantial quality gap between generated and human-authored levels, alongside rising training costs, particularly with increasing token complexity. In this paper, we introduce a diffusion-based generative model that learns from just one example. Our approach involves two core components: 1) an efficient yet expressive level representation, and 2) a latent denoising network with constrained receptive fields. To start with, our method utilizes token semantic labels, similar to word embeddings, to provide dense representations. This strategy not only surpasses one-hot encoding in representing larger game levels but also improves stability and accelerates convergence in latent diffusion. In addition, we adapt the denoising network architecture to confine the receptive field to localized patches of the data, aiming to facilitate single-example learning. Extensive experiments demonstrate that our model is capable of generating stylistically congruent samples of arbitrary sizes compared to manually designed levels. It suits a wide range of level structures with fewer artifacts than GAN-based approaches. The source code is available at https://github.com/shiqi-dai/diffusioncraft. Shiqi Dai, Naiqi Li, Tao Dai 0001, Zhi Wang 0001 |
AAAI | 4 |
| 2024 | Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersabstractLearning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they often suffer from heavy computational burdens and heavily rely on massive cross-modal data pairs that are often unavailable, which hinders their applications in practice. Instead, single-modal methods with solely point clouds as input are preferred in real applications due to their simplicity and efficiency. However, such methods easily suffer from limited 3D representations with global random mask input. To learn compact 3D representations, we propose a simple yet effective Point Feature Enhancement Masked Autoencoders (Point-FEMAE), which mainly consists of a global branch and a local branch to capture latent semantic features. Specifically, to learn more compact features, a share-parameter Transformer encoder is introduced to extract point features from the global and local unmasked patches obtained by global random and local block mask strategies, followed by a specific decoder to reconstruct. Meanwhile, to further enhance features in the local branch, we propose a Local Enhancement Module with local patch convolution to perceive fine-grained local context at larger scales. Our method significantly improves the pre-training efficiency compared to cross-modal alternatives, and extensive downstream experiments underscore the state-of-the-art effectiveness, particularly outperforming our baseline (Point-MAE) by 5.16%, 5.00%, and 5.04% in three variants of ScanObjectNN, respectively. Code is available at https://github.com/zyh16143998882/AAAI24-PointFEMAE. Yaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 5 |
| 2024 | Vision-Language Pre-training with Object Contrastive Learning for 3D Scene UnderstandingabstractIn recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building task-specific models, and fail to extract universal 3D vision-language embedding that generalize well. We carefully investigate three common tasks in semantic 3D scene understanding, and derive key insights into the development of a pre-training model. Motivated by these observations, we propose a vision-language pre-training framework 3DVLP (3D vision-language pre-training with object contrastive learning), which transfers flexibly on 3D vision-language downstream tasks. 3DVLP takes visual grounding as the proxy task and introduces Object-level IoU-guided Detection (OID) loss to obtain high-quality proposals in the scene. Moreover, we design Object-level Cross-Contrastive alignment (OCC) task and Object-level Self-Contrastive learning (OSC) task to align the objects with descriptions and distinguish different objects in the scene, respectively. Extensive experiments verify the excellent performance of 3DVLP on three 3D vision-language tasks, reflecting its superiority in semantic 3D scene understanding. Code is available at https://github.com/iridescentttt/3DVLP. Taolin Zhang 0003, Sunan He, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
AAAI | 3 |
| 2024 | CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
Hao Fang 0011, Jiawei Kong 0001, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ECCV (28) | 4 |
| 2024 | MambaIR: A Simple Baseline for Image Restoration with State-Space Model
Hang Guo 0002, Jinmin Li, Tao Dai 0001, Zhihao Ouyang, Xudong Ren, Shutao Xia |
ECCV (18) | 3 |
| 2024 | WFTNet: Exploiting Global and Local Periodicity in Long-Term Time Series ForecastingabstractRecent CNN and Transformer-based models tried to utilize frequency and periodicity information for long-term time series forecasting. However, most existing work is based on Fourier transform, which cannot capture fine-grained and local frequency structure. In this paper, we propose a Wavelet-Fourier Transform Network (WFTNet) for long-term time series forecasting. WFTNet utilizes both Fourier and wavelet transforms to extract comprehensive temporal-frequency information from the signal, where Fourier transform captures the global periodic patterns and wavelet transform captures the local ones. Furthermore, we introduce a Periodicity-Weighted Coefficient (PWC) to adaptively balance the importance of global and local frequency patterns. Extensive experiments on various time series datasets show that WFTNet consistently outperforms other state-of-the-art baseline. Code is available at https://github.com/Hank0626/WFTNet. Peiyuan Liu, Beiliang Wu, Naiqi Li, Tao Dai 0001, Fengmao Lei, Jigang Bao, Yong Jiang 0001, Shutao Xia |
ICASSP | 4 |
| 2024 | Progressive Learning with Visual Prompt Tuning for Variable-Rate Image CompressionabstractIn this paper, we propose a progressive learning paradigm for transformer-based variable-rate image compression. Our approach covers a wide range of compression rates with the assistance of the Layer-adaptive Prompt Module (LPM). Inspired by visual prompt tuning, we use LPM to extract prompts for input images and hidden features at the encoder side and decoder side, respectively, which are fed as additional information into the swin transformer layer of a pre-trained transformer-based image compression model to affect the allocation of attention region and the bits, which in turn changes the target compression ratio of the model. To ensure the network is more lightweight, we involves the integration of prompt networks with less convolutional layers. Exhaustive experiments show that compared to methods based on multiple models, which are optimized separately for different target rates, the proposed method arrives at the same performance with 80% savings in parameter storage and 90% savings in datasets. Meanwhile, our model outperforms all current variable bitrate image methods in terms of rate-distortion performance and approaches the state-of-the-art fixed bitrate image compression methods trained from scratch. Shiyu Qin, Yimin Zhou 0011, Jin-Peng Wang, Bin Chen 0011, Baoyi An 0002, Tao Dai 0001, Shutao Xia |
ICIP | 6 |
| 2024 | Periodicity Decoupling Framework for Long-term Series ForecastingabstractConvolutional neural network (CNN)-based and Transformer-based methods have recently made significant strides in time series forecasting, which excel at modeling local temporal variations or capturing long-term dependencies. However, real-world time series usually contain intricate temporal patterns, thus making it challenging for existing methods that mainly focus on temporal variations modeling from the 1D time series directly. Based on the intrinsic periodicity of time series, we propose a novel Periodicity Decoupling Framework (PDF) to capture 2D temporal variations of decoupled series for long-term series forecasting. Our PDF mainly consists of three components: multi-periodic decoupling block (MDB), dual variations modeling block (DVMB), and variations aggregation block (VAB). Unlike the previous methods that model 1D temporal variations, our PDF mainly models 2D temporal variations, decoupled from 1D time series by MDB. After that, DVMB attempts to further capture short-term and long-term variations, followed by VAB to make final predictions. Extensive experimental results across seven real-world long-term time series datasets demonstrate the superiority of our method over other state-of-the-art methods, in terms of both forecasting performance and computational efficiency. Code is available at https://github.com/Hank0626/PDF. Tao Dai 0001, Beiliang Wu, Peiyuan Liu, Naiqi Li, Jigang Bao, Yong Jiang 0001, Shutao Xia |
ICLR | 1 |
| 2024 | Towards Faithful XAI Evaluation via Generalization-Limited Backdoor WatermarkabstractSaliency-based representation visualization (SRV) ($e.g.$, Grad-CAM) is one of the most classical and widely adopted explainable artificial intelligence (XAI) methods for its simplicity and efficiency. It can be used to interpret deep neural networks by locating saliency areas contributing the most to their predictions. However, it is difficult to automatically measure and evaluate the performance of SRV methods due to the lack of ground-truth salience areas of samples. In this paper, we revisit the backdoor-based SRV evaluation, which is currently the only feasible method to alleviate the previous problem. We first reveal its \emph{implementation limitations} and \emph{unreliable nature} due to the trigger generalization of existing backdoor watermarks. Given these findings, we propose a generalization-limited backdoor watermark (GLBW), based on which we design a more faithful XAI evaluation. Specifically, we formulate the training of watermarked DNNs as a min-max problem, where we find the `worst' potential trigger (with the highest attack effectiveness and differences from the ground-truth trigger) via inner maximization and minimize its effects and the loss over benign and poisoned samples via outer minimization in each iteration. In particular, we design an adaptive optimization method to find desired potential triggers in each inner maximization. Extensive experiments on benchmark datasets are conducted, verifying the effectiveness of our generalization-limited watermark. Our codes are available at \url{https://github.com/yamengxi/GLBW}. Mengxi Ya, Yiming Li 0004, Tao Dai 0001, Bin Wang 0034, Yong Jiang 0001, Shutao Xia |
ICLR | 3 |
| 2024 | FreqFormer: Frequency-aware Transformer for Lightweight Image Super-resolution
Tao Dai 0001, Hang Guo 0002, Jinmin Li, Jinbao Wang 0001, Zexuan Zhu 0001 |
IJCAI | 1 |
| 2024 | Invertible Residual Rescaling Models
Jinmin Li, Tao Dai 0001, Yaohua Zha, Yilu Luo, Longfei Lu, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
IJCAI | 2 |
| 2024 | Boundary-aware Decoupled Flow Networks for Realistic Extreme Rescaling
Jinmin Li, Tao Dai 0001, Shaoming Wang, Shutao Xia, Rizen Guo |
IJCAI | 2 |
| 2024 | GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process
Yuqiu Xie, Bolin Jiang, Jiawei Li 0006, Naiqi Li, Bin Chen 0011, Tao Dai 0001, Yuang Peng, Shutao Xia |
IJCAI | 6 |
| 2024 | Large Point-to-Gaussian Model for Image-to-3D Generation
Longfei Lu, Huachen Gao, Tao Dai 0001, Yaohua Zha, Zhi Hou, Junta Wu, Shutao Xia |
ACM Multimedia | 3 |
| 2024 | Towards High-resolution 3D Anomaly Detection via Group-Level Feature Contrastive LearningabstractHigh-resolution point clouds (HRPCD) anomaly detection (AD) plays a critical role in precision machining and high-end equipment manufacturing. Despite considerable 3D-AD methods that have been proposed recently, they still cannot meet the requirements of the HRPCD-AD task. There are several challenges: i) It is difficult to directly capture HRPCD information due to large amounts of points at the sample level; ii) The advanced transformer-based methods usually obtain anisotropic features, leading to degradation of the representation; iii) The proportion of abnormal areas is very small, which makes it difficult to characterize. To address these challenges, we propose a novel group-level feature-based network, called Group3AD, which has a significantly efficient representation ability. First, we design an Intercluster Uniformity Network (IUN) to present the mapping of different groups in the feature space as several clusters, and obtain a more uniform distribution between clusters representing different parts of the point clouds in the feature space. Then, an Intracluster Alignment Network (IAN) is designed to encourage groups within the cluster to be distributed tightly in the feature space. In addition, we propose an Adaptive Group-Center Selection (AGCS) based on geometric information to improve the pixel density of potential anomalous regions during inference. The experimental results verify the effectiveness of our proposed Group3AD, which surpasses Reg3D-AD by the margin of 5% in terms of object-level AUROC on Real3D-AD. We provide the code and supplementary information on our website: https://github.com/M-3LAB/Group3AD. Hongze Zhu, Guoyang Xie, Chengbin Hou, Tao Dai 0001, Can Gao, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 4 |
| 2024 | DDN: Dual-domain Dynamic Normalization for Non-stationary Time Series ForecastingabstractDeep neural networks (DNNs) have recently achieved remarkable advancements in time series forecasting (TSF) due to their powerful ability of sequence dependence modeling. To date, existing DNN-based TSF methods still suffer from unreliable predictions for real-world data due to its non-stationarity characteristics, i.e., data distribution varies quickly over time. To mitigate this issue, several normalization methods (e.g., SAN) have recently been specifically designed by normalization in a fixed period/window in the time domain. However, these methods still struggle to capture distribution variations, due to the complex time patterns of time series in the time domain. Based on the fact that wavelet transform can decompose time series into a linear combination of different frequencies, which exhibits distribution variations with time-varying periods, we propose a novel Dual-domain Dynamic Normalization (DDN) to dynamically capture distribution variations in both time and frequency domains. Specifically, our DDN tries to eliminate the non-stationarity of time series via both frequency and time domain normalization in a sliding window way. Besides, our DDN can serve as a plug-in-play module, and thus can be easily incorporated into other forecasting models. Extensive experiments on public benchmark datasets under different forecasting models demonstrate the superiority of our DDN over other normalization methods. Code will be made available following the review process. Tao Dai 0001, Beiliang Wu, Peiyuan Liu, Naiqi Li, Xue Yuerong, Shutao Xia, Zexuan Zhu 0001 |
NeurIPS | 1 |
| 2024 | BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingabstractAdaptation of
pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches.
Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain.
However, existing training-required TTA approaches like TPT necessitate entropy minimization that involves large computational overhead, while training-free methods like TDA overlook the potential for information mining from the test samples themselves.
In this paper, we break down the design of existing popular training-required and training-free TTA methods and bridge the gap between them within our framework.
Specifically, we maintain a light-weight key-value memory for feature retrieval from instance-agnostic historical samples and instance-aware boosting samples.
The historical samples are filtered from the testing data stream and serve to extract useful information from the target distribution, while the boosting samples are drawn from regional bootstrapping and capture the knowledge of the test sample itself.
We theoretically justify the rationality behind our method and empirically verify its effectiveness on both the out-of-distribution and the cross-domain datasets, showcasing its applicability in real-world situations. Taolin Zhang 0003, Jinpeng Wang 0002, Hang Guo 0002, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
NeurIPS | 4 |
| 2024 | Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-ExpertsabstractDesigning single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising results, the existing all-in-one paradigm still suffers from high computational costs as well as limited generalization on unseen degradations. In this work, we introduce an alternative solution to improve the generalization of image restoration models. Drawing inspiration from recent advancements in Parameter Efficient Transfer Learning (PETL), we aim to tune only a small number of parameters to adapt pre-trained restoration models to various tasks. However, current PETL methods fail to generalize across varied restoration tasks due to their homogeneous representation nature. To this end, we propose AdaptIR, a Mixture-of-Experts (MoE) with orthogonal multi-branch design to capture local spatial, global spatial, and channel representation bases, followed by adaptive base combination to obtain heterogeneous representation for different degradations. Extensive experiments demonstrate that our AdaptIR achieves stable performance on single-degradation tasks, and excels in hybrid-degradation tasks, with training only 0.6% parameters for 8 hours. Hang Guo 0002, Tao Dai 0001, Yuanchao Bai, Bin Chen 0011, Xudong Ren, Zexuan Zhu 0001, Shutao Xia |
NeurIPS | 2 |
| 2024 | ReFIR: Grounding Large Restoration Models with Retrieval AugmentationabstractRecent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs. Hang Guo 0002, Tao Dai 0001, Zhihao Ouyang, Taolin Zhang 0003, Yaohua Zha, Bin Chen 0011, Shutao Xia |
NeurIPS | 2 |
| 2024 | LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingabstractThe pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limitation, we first conduct a comprehensive analysis of existing Transformer-based MPM, emphasizing the idea that redundancy reduction is crucial for point cloud analysis. To this end, we propose a Locally constrained Compact point cloud Model (LCM) consisting of a locally constrained compact encoder and a locally constrained Mamba-based decoder. Our encoder replaces self-attention with our local aggregation layers to achieve an elegant balance between performance and efficiency. Considering the varying information density between masked and unmasked patches in the decoder inputs of MPM, we introduce a locally constrained Mamba-based decoder. This decoder ensures linear complexity while maximizing the perception of point cloud geometry information from unmasked patches with higher information density. Extensive experimental results show that our compact model significantly surpasses existing Transformer-based models in both performance and efficiency, especially our LCM-based Point-MAE model, compared to the Transformer-based model, achieved an improvement of 1.84%, 0.67%, and 0.60% in performance on the three variants of ScanObjectNN while reducing parameters by 88% and computation by 73%. The code is available at https://github.com/zyh16143998882/LCM. Yaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai 0001, Hang Guo 0002, Bin Chen 0011, Zhi Wang 0001, Zhihao Ouyang, Shutao Xia |
NeurIPS | 4 |
| 2024 | Pyramid hybrid pooling quantization for efficient fine-grained image retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia, Zhi Wang 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | SpirDet: Toward Efficient, Accurate, and Lightweight Infrared Small-Target DetectorabstractIn recent years, the detection of infrared small targets using deep learning methods has garnered substantial attention due to notable advancements. To improve the detection capability of small targets, these methods commonly maintain a pathway that preserves high-resolution (HR) features of sparse and tiny targets. However, it can result in redundant and expensive computations. To tackle this challenge, we propose a sparse infrared-target detector (SpirDet) for the efficient detection of infrared small targets. Specifically, to cope with the computational redundancy issue, we employ a new dual-branch sparse decoder (DBSD) to restore the feature map. First, the fast branch directly predicts a sparse map indicating potential small-target locations. Second, the slow branch conducts fine-grained adjustments at the positions indicated by the sparse map. In addition, we design a lightweight DO-RepEncoder based on reparameterization with the downsampling orthogonality (DO), which can effectively reduce memory consumption and inference latency. Extensive experiments show that the proposed SpirDet significantly outperforms state-of-the-art (SOTA) models while achieving faster inference speed and fewer parameters. For example, on the NUDT-SIRST dataset, SpirDet improves mean intersection over union (MIoU) by 2.09 and has a$3.1\times $frames/s acceleration compared to the previous SOTA model. The code is available athttps://github.com/laCorse/SpirDet-Pytorch. Qianchen Mao, Qiang Li 0055, Bingshu Wang, Yongjun Zhang 0007, Tao Dai 0001, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge TransferabstractReal-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets. Sunan He, Taian Guo, Tao Dai 0001, Ruizhi Qiao, Xiujun Shu, Bo Ren 0002, Shutao Xia |
AAAI | 3 |
| 2023 | Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature DomainabstractBeyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed compression scenario, the existing method only implements patch matching at the image domain to solve the parallax problem caused by the difference in viewing points. However, the patch matching at the image domain is not robust to the variance of scale, shape, and illumination caused by the different viewing angles, and can not make full use of the rich texture information of the side information image. To resolve this issue, we propose Multi-Scale Feature Domain Patch Matching (MSFDPM) to fully utilizes side information at the decoder of the distributed image compression model. Specifically, MSFDPM consists of a side information feature extractor, a multi-scale feature domain patch matching module, and a multi-scale feature fusion network. Furthermore, we reuse inter-patch correlation from the shallow layer to accelerate the patch matching of the deep layer. Finally, we find that our patch matching in a multi-scale feature domain further improves compression rate by about 20% compared with the patch matching method at image domain. Yujun Huang, Bin Chen 0011, Shiyu Qin, Jiawei Li 0006, Yaowei Wang 0001, Tao Dai 0001, Shutao Xia |
AAAI | 6 |
| 2023 | FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution NetworksabstractDeep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To address this issue, we propose a general frequency-oriented framework (FSR) to accelerate SR networks by considering data characteristics in frequency domain. Our FSR mainly contains dual feature aggregation module (DFAM) to extract informative features in both spatial and transform domains, followed by a four-path SR-Module with different capacities to super-resolve in the frequency domain. Specifically, DFAM further consists of a transform attention block (TABlock) and a spatial context block (SCBlock) to extract global spectral information and local spatial information, respectively, while SR-Module is a parallel network container that contains four to-be-accelerated branches. Furthermore, we propose an adaptive weight strategy for a trade-off between image details recovery and visual quality. Extensive experiments show that our FSR can save FLOPs by almost 40% while reducing inference time by 50% for other SR methods (e.g., FSRCNN, CARN, SRResNet and RCAN). Code is available at https://github.com/THU-Kingmin/FSR. Jinmin Li, Tao Dai 0001, Mingyan Zhu 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 2 |
| 2023 | Difficulty-Aware Data Augmentor for Scene Text RecognitionabstractDeep neural network (DNN) based scene text recognition (STR) methods usually require a large amount of annotated data for training, which is time-consuming and cost-expensive in practice. To address this issue, many data augmentation methods have been developed to train recognizers by improving the diversity of training samples. However, most existing methods neglect the difficulty inherent in samples, and easily suffer from the problem of over-diversity, i.e., the distribution of the augmented data significantly deviates from that of clean data. In this paper, we propose a novel difficulty-aware data augmentation framework for scene text recognition, which jointly considers the difficulty of samples and the strength of augmentations. Specifically, our framework first predicts the sample difficulty, followed by an adaptive data augmentation strategy. Furthermore, we build a more diverse set of augmentation methods for STR and integrate it into our augmentation framework. Extensive experiments on scene text recognition benchmarks show that our augmentation framework significantly improves the performance of recognizers. Guanghao Meng, Tao Dai 0001, Bin Chen 0011, Naiqi Li, Yong Jiang 0001, Shutao Xia |
ICASSP | 2 |
| 2023 | Semantic Preserving Learning for Task-Oriented Point Cloud DownsamplingabstractRecent years have witnessed a tremendous growth in the scale and resolution of point clouds. To facilitate the applications of point cloud in downsampling tasks (e.g., point cloud classification), several task-oriented downsampling works have been developed by training with the task-specific loss with one-hot encoded label. However, these methods still suffer from performance degradation at high downsampling scales. In this paper, we propose a general semantic-preserved downsampling framework (SPDF) for point clouds by exploiting the rich knowledge inherent in the task network. Specifically, we firstly refine the previous pipeline to generate richer semantic supervised information. Then, the semantic feature learning is subdivided into label-level and feature-level to guide the training of downsampling network, which can better limit the semantic loss during downsampling. Extensive experiments on the benchmark dataset show that SPDF outperforms state-of-the-art downsampling methods. Jianyu Xiong, Tao Dai 0001, Yaohua Zha, Xin Wang 0001, Shutao Xia |
ICASSP | 2 |
| 2023 | SFR: Semantic-Aware Feature Rendering of Point CloudabstractMulti-view projection methods have demonstrated their ability to reach state-of-the-art performance in point cloud downstream tasks(e.g., classification and retrieval). These methods first require rendering the point cloud into 2D multi-view images. However, conventional methods only project the geometry of the point cloud, and such projections inevitably suffer from a loss of point cloud semantic information due to dimensionality reduction. We propose a semantic-aware and task-oriented differentiable feature rendering (SFR), which reduces the information loss during projection by generating rendered images with more point cloud semantic information for downstream tasks. Our SFR method can be applied as a plug-and-play module added to any multi-view-based backbone network for end-to-end training. Extensive experiments on benchmark datasets show that our SFR method reaches state-of-the-art performance and brings general improvements to point cloud classification and retrieval tasks. Yaohua Zha, Rongsheng Li, Tao Dai 0001, Jianyu Xiong, Xin Wang 0001, Shutao Xia |
ICASSP | 3 |
| 2023 | Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud ModelsabstractPre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the efficiency when applying large-scale pre-trained models. Inspired by the recent success of visual prompt tuning (VPT), this paper attempts to explore prompt tuning on pre-trained point cloud models, to pursue an elegant balance between performance and parameter efficiency. We find while instance-agnostic static prompting, e.g. VPT, shows some efficacy in downstream transfer, it is vulnerable to the distribution diversity caused by various types of noises in real-world point cloud data. To conquer this limitation, we propose a novel Instance-aware Dynamic Prompt Tuning (IDPT) strategy for pre-trained point cloud models. The essence of IDPT is to develop a dynamic prompt generation module to perceive semantic prior features of each point cloud instance and generate adaptive prompt tokens to enhance the model's robustness. Notably, extensive experiments demonstrate that IDPT outperforms full finetuning in most tasks with a mere 7% of the trainable parameters, providing a promising solution to parameter-efficient learning for pre-trained point cloud models. Code is available at https://github.com/zyh16143998882/ICCV23-IDPT. Yaohua Zha, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ICCV | 3 |
| 2023 | Unsupervised Surface Anomaly Detection with Diffusion Probabilistic ModelabstractUnsupervised surface anomaly detection aims at discovering and localizing anomalous patterns using only anomaly-free training samples. Reconstruction-based models are among the most popular and successful methods, which rely on the assumption that anomaly regions are more difficult to reconstruct. However, there are three major challenges to the practical application of this approach: 1) the reconstruction quality needs to be further improved since it has a great impact on the final result, especially for images with structural changes; 2) it is observed that for many neural networks, the anomalies can also be well reconstructed, which severely violates the underlying assumption; 3) since reconstruction is an ill-conditioned problem, a test instance may correspond to multiple normal patterns, but most current reconstruction-based methods have ignored this critical fact. In this paper, we propose DiffAD, a method for unsupervised anomaly detection based on the latent diffusion model, inspired by its ability to generate high-quality and diverse images. We further propose noisy condition embedding and interpolated channels to address the aforementioned challenges in the general reconstruction-based pipeline. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MVTec dataset, especially in localization accuracy. Xinyi Zhang 0008, Naiqi Li, Jiawei Li 0006, Tao Dai 0001, Yong Jiang 0001, Shutao Xia |
ICCV | 4 |
| 2023 | Towards Robust Scene Text Image Super-resolution via Explicit Location EnhancementabstractScene text image super-resolution (STISR), aiming to improve image quality while boosting downstream scene text recognition accuracy, has recently achieved great success. However, most existing methods treat the foreground (character regions) and background (non-character regions) equally in the forward process, and neglect the disturbance from the complex background, thus limiting the performance. To address these issues, in this paper, we propose a novel method LEMMA that explicitly models character regions to produce high-level text-specific guidance for super-resolution. To model the location of characters effectively, we propose the location enhancement module to extract character region features based on the attention map sequence. Besides, we propose the multi-modal alignment module to perform bidirectional visual-semantic alignment to generate high-quality prior guidance, which is then incorporated into the super-resolution branch in an adaptive manner using the proposed adaptive fusion module. Experiments on TextZoom and four scene text recognition benchmarks demonstrate the superiority of our method over other state-of-the-art methods. Code is available at https://github.com/csguoh/LEMMA. Hang Guo 0002, Tao Dai 0001, Guanghao Meng, Shutao Xia |
IJCAI | 2 |
| 2023 | DDA: A Dynamic Difficulty-aware Data Augmenter for Image Super-resolutionabstractDeep neural networks (DNNs) have been recently widely used in image super-resolution (SR) and have achieved remarkable performance. However, most existing methods focus on elaborate network design, while rarely considering the training strategy, which affects the model performance and training efficiency. In practice, most SR methods still train the networks with the commonly-used data augmentation (e.g., random crop and sampling), which is shown to converge slowly for deep SR networks. To address this issue, in this paper, we propose a dynamic difficulty-aware data augmenter, named DDA, by considering the restoration difficulty and distribution of input patches. Our DDA mainly consists of difficulty-aware divider, dynamic sampler, and adaptive re-weighter. Specifically, our DDA first uses the difficulty-aware divider to divide the input image into small over-lapping patches, followed by classification into$N$different classes based on the restoration difficulty. Next, dynamic sampler samples the training patches from each class with probability based on training loss. Furthermore, to remedy the imbalance of training patches between different classes, adaptive re-weighter updates the weight of each training patch according to the accumulated training loss. Extensive experiments demonstrate the effectiveness of our DDA on different SR methods by improving training efficiency and model performance across a wide range of scenarios. Xinyi Zhang 0008, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
IJCNN | 2 |
| 2023 | One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferabstractRecognizing characters from low-resolution (LR) text images poses a significant challenge due to the information deficiency as well as the noise and blur in low-quality images. Current solutions for low-resolution text recognition (LTR) typically rely on a two-stage pipeline that involves super-resolution as the first stage followed by the second-stage recognition. Although this pipeline is straightforward and intuitive, it has to use an additional super-resolution network, which causes inefficiencies during training and testing. Moreover, the recognition accuracy of the second stage heavily depends on the reconstruction quality of the first stage, causing ineffectiveness.In this work, we attempt to address these challenges from a novel perspective: adapting the recognizer to low-resolution inputs by transferring the knowledge from the high-resolution. Guided by this idea, we propose an efficient and effective knowledge distillation framework to achieve multi-level knowledge transfer.Specifically, the visual focus loss is proposed to extract the character position knowledge with resolution gap reduction and character region focus, the semantic contrastive loss is employed to exploit the contextual semantic knowledge with contrastive learning, and the soft logits loss facilitates both local word-level and global sequence-level learning from the soft teacher label.Extensive experiments show that the proposed one-stage pipeline significantly outperforms super-resolution based two-stage frameworks in terms of effectiveness and efficiency, accompanied by favorable robustness.Code is available at https://github.com/csguoh/KD-LTR. Hang Guo 0002, Tao Dai 0001, Mingyan Zhu 0001, Guanghao Meng, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ACM Multimedia | 2 |
| 2023 | Adversarial Examples Generation for Deep Product Quantization Networks on Image RetrievalabstractDeep product quantization networks (DPQNs) have been successfully used in image retrieval tasks, due to their powerful feature extraction ability and high efficiency of encoding high-dimensional visual features. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples) for classification. However, little effort has been devoted to investigating how adversarial examples affect DPQNs, which raises the potential safety hazard when deploying DPQNs in a commercial search engine. To this end, we propose an adversarial example generation framework by generating adversarial query images for DPQN-based retrieval systems. Unlike the adversarial generation for the classic image classification task that heavily relies on ground-truth labels, we alternatively perturb the probability distribution of centroids assignments for a clean query, then we can induce effective non-targeted attacks on DPQNs in white-box and black-box settings. Moreover, we further extend the non-targeted attack to a targeted attack by a novel sample space averaging scheme ([Formula: see text]AS), whose theoretical guarantee is also obtained. Extensive experiments show that our methods can create adversarial examples to successfully mislead the target DPQNs. Besides, we found that our methods both significantly degrade the retrieval performance under a wide variety of experimental settings. The source code is available at https://github.com/Kira0096/PQAG. Bin Chen 0011, Tao Dai 0001, Jiawang Bai, Yong Jiang 0001, Shutao Xia, Xuan Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Toward Effective Image Manipulation Detection With Proposal Contrastive LearningabstractDeep models have been widely and successfully used in image manipulation detection, which aims to classify tampered images and localize tampered regions. Most existing methods mainly focus on extracting global features from tampered images, while neglecting the relationships of local features between tampered and authentic regions within a single tampered image. To exploit such spatial relationships, we propose Proposal Contrastive Learning (PCL) for effective image manipulation detection. Our PCL consists of a two-stream architecture by extracting two types of global features from RGB and noise views respectively. To further improve the discriminative power, we exploit the relationships of local features through a proxy proposal contrastive learning task by attracting/repelling proposal-based positive/negative sample pairs. Moreover, we show that our PCL can be easily adapted to unlabeled data in practice, which can reduce manual labeling costs and promote more generalizable features. Extensive experiments among several standard datasets demonstrate that our PCL can be a general module to obtain consistent improvement. The code is available athttps://github.com/Sandy-Zeng/PCL. Yuyuan Zeng, Bowen Zhao 0003, Shanzhao Qiu, Tao Dai 0001, Shutao Xia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Contrastive Quantization with Code Memory for Unsupervised Image RetrievalabstractThe high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a novel solution to unsupervised deep quantization, namely Contrastive Quantization with Code Memory (MeCoQ). Different from existing reconstruction-based strategies, we learn unsupervised binary descriptors by contrastive learning, which can better capture discriminative visual semantics. Besides, we uncover that codeword diversity regularization is critical to prevent contrastive learning-based quantization from model degeneration. Moreover, we introduce a novel quantization code memory module that boosts contrastive learning with lower feature drift than conventional feature memories. Extensive experiments on benchmark datasets show that MeCoQ outperforms state-of-the-art methods. Code and configurations are publicly released. Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
AAAI | 4 |
| 2022 | NeXT: Towards High Quality Neural Radiance Fields via Multi-skip Transformer
Peidong Liu 0003, Tao Dai 0001, Shutao Xia |
ECCV (32) | 4 |
| 2022 | Self-Ensemble Variance Regularization for Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge from a label-rich source domain to a different yet related fully-unlabeled target domain. Existing approaches utilize self-training scheme to learn discriminative target features and thus enforce class-level distribution alignment implicitly across the source and target domains. However, inherent noise of the pseudo labels due to domain shift could compromise the training process to negatively affect the adapted model performance. In this paper, we propose Self-Ensemble Variance Regularization for Domain Adaptaton (VRDA) method to rectify the learning with pseudo labels. To be specific, we regard the prediction distinction between the student and its self-ensemble teacher model as prediction variance, to regularize target domain prediction bias from pseudo labels. The experimental results reveal that the proposed VRDA achieves the state-of-the-art performance on several standard UDA datasets. Tao Dai 0001, Shutao Xia, Yong Jiang 0001 |
ICASSP | 2 |
| 2022 | Adaptive Local Implicit Image Function for Arbitrary-Scale Super-ResolutionabstractImage representation is critical for many visual tasks. Instead of representing images discretely with 2D arrays of pixels, a recent study, namely local implicit image function (LIIF), denotes images as a continuous function where pixel values are expansion by using the corresponding coordinates as inputs. Due to its continuous nature, LIIF can be adopted for arbitrary-scale image super-resolution tasks, resulting in a single effective and efficient model for various up-scaling factors. However, LIIF often suffers from structural distortions and ringing artifacts around edges, mostly because all pixels share the same model, thus ignoring the local properties of the image. In this paper, we propose a novel adaptive local image function (A-LIIF) to alleviate this problem. Specifically, our A-LIIF consists of two main components: an encoder and a expansion network. The former captures cross-scale image features, while the latter models the continuous up-scaling function by a weighted combination of multiple local implicit image functions. Accordingly, our A-LIIF can reconstruct the high-frequency textures and structures more accurately. Experiments on multiple benchmark datasets verify the effectiveness of our method. Our codes are available at https://github.com/LeeHW-THU/A-LIIF. Hongwei Li 0001, Tao Dai 0001, Yiming Li 0004, Xueyi Zou, Shutao Xia |
ICIP | 2 |
| 2022 | Deep image prior based defense against adversarial examples
Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia |
Pattern Recognit. | 1 |
| 2021 | Knowledge Refinery: Learning from Decoupled LabelabstractRecently, a variety of regularization techniques have been widely applied in deep neural networks, which mainly focus on the regularization of weight parameters to encourage generalization effectively. Label regularization techniques are also proposed with the motivation of softening the labels while neglecting the relation of classes. Among them, the technique of knowledge distillation proposes to distill the soft label, which contains the knowledge of class relations. However, this technique needs to pre-train an extra cumbersome teacher model. In this paper, we propose a method called Knowledge Refinery (KR), which enables the neural network to learn the relation of classes on-the-fly without the teacher-student training strategy. We propose the definition of decoupled labels, which consist of the original hard label and the residual label. To exhibit the generalization of KR, we evaluate our method in both fields of computer vision and natural language processing. Our empirical results show consistent performance gains under all experimental settings. Qianggang Ding, Tao Dai 0001, Jiadong Guo, Zhang-Hua Fu, Shutao Xia |
AAAI | 3 |
| 2021 | Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration NetabstractDeep neural networks (DNNs) have been widely used in facial manipulation. Existing methods focus on training deeper networks in indirect supervision ways (e.g., feature constraint), or in unsupervised ways (e.g., cycle-consistency loss) due to the lack of ground-truth face images for manipulated outputs. However, such methods can not synthesize realistic face images well and suffer from very high training overhead. To address this issue, we propose a novel Feature Disentanglement and Reintegraion network (FDRNet), which employs ground-truth images as informative supervision and dynamically adapts the fusion of informative features of the ground-truth images effectively and efficiently. FDRNet consists of a Feature Disentanglement (FD) Network and a Feature Reintegration (FR) Network, which encodes informative disentangled representations from the ground-truth images and fuses the disentangled representations to reconstruct the face images. By learning disentangled representations, our method can generate plausible faces conditioned on both landmarks and identities, which can be used for a variety of face manipulation tasks. Experiments on the CelebA-HQ and FFHQ datasets are conducted to demonstrate the superiority of our method over state-of-the-art methods in terms of effectiveness and efficiency. Tao Dai 0001, Bin Chen 0011, Shutao Xia, Xiu Li 0001 |
ICASSP | 2 |
| 2021 | HOCA: Higher-Order Channel Attention for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich feature statistics higher than second-order, thus hindering the representation ability of CNNs. To address this issue, we propose a higher-order channel attention (HOCA) module to enhance the representation ability of CNNs. In our HOCA module, to capture different types of semantic information, we first compute k-order of feature statistics, followed by channel attention to learn the feature interdependencies. Considering the diversity of input contents, we design a gate mechanism to adaptively select a specific k-order channel attention. Besides, our HOCA module serves as a plug-and-play module and can be easily plugged into existing state-of-art CNN-based SR methods. Extensive experiments on public benchmarks show that our HOCA module effectively improves the performance of various CNN-based SR methods. Yalei Lv, Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia, Jingchao Cao |
ICASSP | 2 |
| 2021 | Webly Supervised Deep Attentive QuantizationabstractLearning to hash has been widely applied in large-scale image retrieval. Although current deep hashing methods yield state-of-the-art performance, their heavy dependence on groundtruth information actually makes it difficult to deploy in practical applications such as social media. To solve this problem, we propose a novel method termed Webly Supervised Deep Attentive Quantization (WSDAQ), where deep quantization is trained on web images associated with some userprovided weak tags, without consulting any ground-truth labels. Specifically, we design a tag processing module to leverage semantic information of tags so as to better supervised quantization learning. Besides, we propose an end-to-end trainable Attentive Product Quantization Module (APQM) to quantize deep features of images. Furthermore, we use a noise-contrastive estimation loss to train the model from the perspective of contrastive learning. Experiments validate that WSDAQ is superior to state-of-the-art baselines in compact coding trained on weakly-tagged web images. Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ICASSP | 3 |
| 2021 | Class Aware Robust TrainingabstractAdversarial training (AT) has been one of the most effective ways to defend adversarial attack. However, existing AT variants exhibit a large imbalanced robust accuracies among different classes, which might harm the robustness of some important class(es) in some real-world applications. For instance, diseased cells are much more important than healthy ones in medical image recognition. Given a certain task, the important class is often a priori. To improve robust accuracy of the important class(es), we are the first to propose a novel adversarial training method with class imbalance taken into account. We term it Class-Aware Robust Training (CART). CART can significantly increase the robustness of the important class(es) by an optional weighted combination of original adversarial example generation and that of the important class. Extensive experiments on three benchmark datasets verify the efficacy of CART for enhancing the robust accuracy of important classes while keeping competitive average robust accuracy. Zhikang Xia, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
ICASSP | 3 |
| 2021 | Bishift-Net for Image InpaintingabstractImage inpainting remains a challenging task in computer vision which aims to fill in the missing area of a corrupted image with proper contents and generate photorealistic images by using the information from the existing area. Most existing methods always generate contents with blurry texture caused by propogating the convolutional feature through a fully connected layer. To address this problem, Shift-Net[1] is proposed to shift the encoder feature from existing area to serve as an estimation of the missing parts. However, the decoder feature with new encoding information is ingnored. Inspired by this, we propose a new inpainting model, which is called BiShift-Net. BiShift-Net adopts the structure of U-Net, and we introduce a BiShift layer to it. We use the BiShift layer to capture the information from both encoder and decoder features, rearranging the features to generate sharp texture. Experiments show that Bishift-Net outperforms the other state-of-the-art CNN-based methods, while produce more faithful results at the same time. Tao Dai 0001, Yong Jiang 0001, Shutao Xia |
ICASSP | 2 |
| 2021 | Attribute Structured Knowledge DistillationabstractKnowledge distillation aims at transferring sufficient knowledge from one cumbersome teacher network to another compressed student network. Most previous knowledge distillation methods mainly focus on mimicking the teacher's output of each instance or inter-instance relations from the teacher to the student, while neglecting the attribute structured relations at an instance level. In this paper, we propose a novel Attribute Structured Knowledge Distillation (ASKD) to transfer attribute-level structured relations. It models two types of structured relations, including inter-region structure and inter-class structure, which captures informative knowledge from the teacher model. Inter-region structure captures local feature relations, while interclass structure focuses on cross-class relations. Transferring such attribute-level structure from the teacher to the student would make the student focus on more informative knowledge of the teacher. Extensive experiments on public datasets, including CIFAR-100 and TinyImageNet, demonstrate the superiority of our method over the state-of-the-art methods under similar architecture and different architecture teacher-student pairs. Tao Dai 0001, Bin Chen 0011, Shutao Xia |
IJCNN | 2 |
| 2021 | D2Defend: Dual-Domain based Defense against Adversarial ExamplesabstractConvolutional neural networks (CNNs) have recently been widely applied in computer vision tasks, yet they are seriously vulnerable to imperceptible adversarial perturbations. Such phenomena have caused great attention on the adversary topic. Existing adversarial defense methods mainly focus on improving the robustness of models (e.g., adversarial training) or removing adversarial perturbations (e.g., input-transformation based methods) directly, while rarely considering the accurate recovery of image structures of the input, which also play a vital role in making predictions for CNNs. To this end, we propose a Dual-Domain based Defense (D2Defend) method by recovering low-frequency and high-frequency image structures in both spatial and transform domains, while removing adversarial perturbations simultaneously. Unlike the existing input-transformation based methods, our method can decompose the input image into edge feature and texture feature layers, accompanied with bilateral filtering and short-time fourier transform (STFT) filtering. Experimental results demonstrate the effectiveness of our method against various adversarial attacks, and show the superiority of our method over other adversarial defense methods especially at strong adversarial strength. Tao Dai 0001, Yang Bai 0011, Shutao Xia |
IJCNN | 3 |
| 2021 | EDKE: Encoder-Decoder based Kernel Estimation for Blind Image Super-resolutionabstractDeep neural networks (DNNs) have recently achieved wonderful results in single image super-resolution (SISR). These DNN-based SR methods are usually trained on supervised training datasets downscaled with a fixed kernel (e.g., bicubic), and suffer from the generalization problem on real world images due to the complexity of the degradation process. More recent works (e.g., KernelGAN [1]) use a generator network to realize kernel estimation from a single image. However, such methods neglect the anisotropy of kernel, thus limiting the estimation accuracy of kernel in practice. To address this issue, in this paper, we propose a novel encoder-decoder based kernel estimation (EDKE) framework by considering the anisotropy of kernel. In our EDKE framework, the input image is first fed into an encoder network to learn the degradation process, followed by a decoder network to map the encoder image into the original image itself. In this way, the encoder image contains more compact information for the degradation kernel. Besides, our EDKE can be plugged into other blind SR methods. Extensive experiments demonstrate the effectiveness of our method in kernel estimation, and show the superiority of our method over other state-of-the-art methods. Mingyan Zhu 0001, Tao Dai 0001, Shutao Xia, Maowei Hu |
IJCNN | 2 |
| 2021 | Knowledge Distillation via Channel Correlation Structure
Bin Chen 0011, Tao Dai 0001, Maowei Hu, Yong Jiang 0001, Shutao Xia |
KSEM | 4 |
| 2021 | Mix-order Attention Networks for Image RestorationabstractConvolutional neural networks (CNNs) have obtained great success in image restoration tasks, like single image denoising, demosaicing, and super-resolution. However, most existing CNN-based methods neglect the diversity of image contents and degradations in the corrupted images and treat channel-wise features equally, thus hindering the representation ability of CNNs. To address this issue, we propose deep mix-order attention networks (MAN) to extract features that capture rich feature statistics within networks. Our MAN is mainly built on simple residual blocks and our mix-order channel attention (MOCA) module, which further consists of feature gating and feature pooling blocks to capture different types of semantic information. With our MOCA, our MAN can be flexible to handle various types of image contents and degradations. Besides, our MAN can be generalized to different image restoration tasks, like image denoising, super-resolution, and demosaicing. Extensive experiments demonstrate that our method obtains favorably against state-of-the-art methods in terms of quantitative and qualitative metrics. Tao Dai 0001, Yalei Lv, Bin Chen 0011, Zhi Wang 0001, Zexuan Zhu 0001, Shutao Xia |
ACM Multimedia | 1 |
| 2021 | Correlation-based structural dropout for convolutional neural networks
Yuyuan Zeng, Tao Dai 0001, Bin Chen 0011, Shutao Xia, Jian Lu 0002 |
Pattern Recognit. | 2 |
| 2020 | Adversarial Attack on Deep Product Quantization Network for Image RetrievalabstractDeep product quantization network (DPQN) has recently received much attention in fast image retrieval tasks due to its efficiency of encoding high-dimensional visual features especially when dealing with large-scale datasets. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples). This phenomenon raises the concern of security issues for DPQN in the testing/deploying stage as well. However, little effort has been devoted to investigating how adversarial examples affect DPQN. To this end, we propose product quantization adversarial generation (PQ-AG), a simple yet effective method to generate adversarial examples for product quantization based retrieval systems. PQ-AG aims to generate imperceptible adversarial perturbations for query images to form adversarial queries, whose nearest neighbors from a targeted product quantizaiton model are not semantically related to those from the original queries. Extensive experiments show that our PQ-AQ successfully creates adversarial examples to mislead targeted product quantization retrieval models. Besides, we found that our PQ-AG significantly degrades retrieval performance in both white-box and black-box settings. Bin Chen 0011, Tao Dai 0001, Shutao Xia |
AAAI | 3 |
| 2020 | Corrdrop: Correlation Based Dropout for Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) can be easily over-fitted when they are over-parametered. The popular dropout that drops feature units randomly can't always work well for CNNs, due to the problem of under-dropping. To eliminate this problem, some structural dropout methods such as SpatialDropout, Cutout and DropBlock have been proposed. However, these methods that drop feature units in continuous regions randomly, may have the risk of over-dropping, thus leading to degradation of performance. To address these issues, we propose a novel structural dropout method, Correlation based Dropout (CorrDrop), to regularize CNNs by dropping feature units based on feature correlation, which reflects the discriminative information in feature maps. Specifically, the proposed method first obtains correlation map based on the activation in the feature maps, and then adaptively masks out those regions with small average correlation. Thus, the proposed method can regularize CNNs well by discarding part of contextual regions. Extensive experiments on image classification demonstrate the superiority of our method compared with other counterparts. Yuyuan Zeng, Tao Dai 0001, Shutao Xia |
ICASSP | 2 |
| 2020 | Enhanced Image Restoration Via Supervised Target Feature TransferabstractDeep learning has obtained remarkable success for image restoration. However, most existing deep image restoration models are trained by minimizing the pixel-level reconstruction error between restored images and target images (ground truth), while neglecting the rich information from the intermediate feature layers, thus hindering the representational power of networks. To address this problem, we propose a Supervised Target Feature Transfer (STFT) framework to enhance the power of feature expression of the deep image restoration models. Specifically, we introduce a self-supervised antoencoder-based target feature extractor to extract compact feature representation of target images, which serves as supervision signals to train the deep backbone models at the same time. With such feature-level supervised information, deep backbone model can be enhanced by transfer learning of such target features. Moreover, we theoretically analyze our STFT training strategies and demonstrate that it imposes learnable prior information on the backbone restoration model. Extensive experiments demonstrate the effectiveness of our proposed framework compared with the state-of-the-art image restoration models. Yuzhao Chen, Tao Dai 0001, Xi Xiao 0001, Jian Lu 0002, Shutao Xia |
ICIP | 2 |
| 2020 | Hrnet: Hamiltonian Rescaling Network for Image DownscalingabstractImage downscaling has become a classical problem in image processing and has recently connected to image super-resolution (SR), which restores high-quality images from low-resolution ones generated by predetermined downscaling kernels (e.g., bicubic). However, most existing image downscaling methods are deterministic and lose information during the downscaling process, while rarely designing specific downscaling methods for image SR. In this paper, we propose a novel learning-based image downscaling method, Hamiltonian Rescaling Network (HRNet). The design of HRNet is based on the discretization of Hamiltonian System, a pair of iterative updating equations, which formulate a mechanism of iterative correction of the error caused by information missing during image or feature downscaling. Extensive experiments demonstrate the effectiveness of our proposed method in terms of both quantitative and qualitative results. Yuzhao Chen, Xi Xiao 0001, Tao Dai 0001, Shutao Xia |
ICIP | 3 |
| 2020 | Fakd: Feature-Affinity Based Knowledge Distillation for Efficient Image Super-ResolutionabstractConvolutional neural networks (CNNs) have been widely used in image super-resolution (SR). Most existing CNN-based methods focus on achieving better performance by designing deeper/wider networks, while suffering from heavy computational cost problem, thus hindering the deployment of such models in mobile devices with limited resources. To relieve such problem, we propose a novel and efficient SR model, named Feature Affinity-based Knowledge Distillation (FAKD), by transferring the structural knowledge of a heavy teacher model to a lightweight student model. To transfer the structural knowledge effectively, FAKD aims to distill the second-order statistical information from feature maps and trains a lightweight student network with low computational and memory cost. Experimental results demonstrate the efficacy of our method and the effectiveness over other knowledge distillation based methods in terms of both quantitative and visual metrics. Zibin He, Tao Dai 0001, Jian Lu 0002, Yong Jiang 0001, Shutao Xia |
ICIP | 2 |
| 2020 | Progressive Splitting and Upscaling Structure for Super-ResolutionabstractRecently, very deep convolutional neural networks (CNNs) have shown great success in single image super-resolution (SISR). Most of these methods focus on the design of network architecture and adopt a sub-pixel convolution layer at the end of network, but few have paid attention to exploring potential representation ability of upscaling layer. Sub-pixel convolution layer aggregates several low resolution (LR) feature maps and builds super-resolution (SR) images in a single step. However, those LR feature maps share similar patterns as they are extracted from a single trunk network. Inspired by this, we propose a novel progressive splitting and upscaling structure, termed PSUS, which generates decoupled feature maps for upscaling layer to get better SR image. Experiments show that our method can not only speed up the convergence, but also achieve considerable improvement on image quality with fewer parameters and lower computational cost. Qiang Li 0055, Tao Dai 0001, Shutao Xia |
ICPR | 2 |
| 2020 | Sample-aware Data Augmentor for Scene Text RecognitionabstractDeep neural networks (DNNs) have been widely used in scene text recognition, and achieved remarkable performance. Such DNN-based scene text recognizers usually require plenty of training data for training, but data collection and annotation is usually cost-expensive in practice. To alleviate this issue, data augmentation is often applied to train the scene text recognizers. However, existing data augmentation methods including affine transformation and elastic transformation methods suffer from the problems of under- and over-diversity, due to the complexity of text contents and shapes. In this paper, we propose a sample-aware data augmentor to transform samples adaptively based on the contents of samples. Specifically, our data augmentor consists of three parts: gated module, affine transformation module, and elastic transformation module. In our data augmentor, affine transformation module focuses on keeping the affinity of samples, while elastic transformation module aims to improve the diversity of samples. With the gated module, our data augmentor determines transformation type adaptively based on the properties of training samples and the recognizer capability during the training process. Besides, our framework introduces an adversarial learning strategy to optimize the augmentor and the recognizer jointly. Extensive experiments on scene text recognition benchmarks show that our sample-aware data augmentor significantly improves the performance of state-of-the-art scene text recognizer. Guanghao Meng, Tao Dai 0001, Shudeng Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia |
ICPR | 2 |
| 2020 | Transferable Adversarial Attacks for Deep Scene Text DetectionabstractScene text detection (STD) aims to locate text in images and plays an important role in many computer vision tasks including automatic driving and text recognition systems. Recently, deep neural networks (DNNs) have been widely and successfully used in scene text detection, leading to plenty of DNN-based STD methods including regression-based and segmentation-based STD methods. However, recent studies have also shown that DNN is vulnerable to adversarial attacks, which can significantly degrade the performance of DNN models. In this paper, we investigate the robustness of DNN-based STD methods against adversarial attacks. To this end, we propose a generic and efficient attack method to generate adversarial examples, which are produced by adding small but imperceptible adversarial perturbation to the input images. Experiments on attacking four various models and a real-world STD engine of Google optical character recognition (OCR) show that the state-of-the-art DNN-based STD methods including regression-based and segmentation-based methods are vulnerable to adversarial attacks. Shudeng Wu, Tao Dai 0001, Guanghao Meng, Bin Chen 0011, Jian Lu 0002, Shutao Xia |
ICPR | 2 |
| 2020 | Generalized Local Aggregation for Large Scale Gaussian Process RegressionabstractDespite being one of the most popular nonparametric approaches, Gaussian process regression (GPR) suffers from O(n3) computational burden and the computation is infeasible for large-scale scenarios. To reduce the computational complexity, many Shannon-mutual-information-based aggregation methods were proposed, whereas these methods can not effectively identify the importance of experts in some cases. To address this problem, we generalize the traditional mutual information-based methods (GPoE, RBCM, GRBCM) based on Tsallis mutual information. Accordingly, the generated weight distribution is more sparse tending to focus on those experts with good performance. To obtain adaptive and data-dependent entropic-index in Tsallis entropy, we propose three heuristic algorithms to solve our model. Extensive experiments show that, the proposed method can improve the prediction of both the mean and variance, and the improvement of variance prediction is significant in many cases. Yinghua Gao, Naiqi Li, Ning Ding 0002, Yiming Li 0004, Tao Dai 0001, Shutao Xia |
IJCNN | 5 |
| 2020 | DIPDefend: Deep Image Prior Driven Defense against Adversarial ExamplesabstractDeep neural networks (DNNs) have shown serious vulnerability to adversarial examples with imperceptible perturbation to clean images. Most existing input-transformation based defense methods (e.g., ComDefend) rely heavily on the learned external priors from an external large training dataset, while neglecting the rich image internal priors of the input itself, thus limiting the generalization of the defense models against the adversarial examples with biased image statistics from the external training dataset. Motivated by deep image prior that can capture rich image statistics from a single image, we propose an effective Deep Image Prior Driven Defense (DIPDefend) method against adversarial examples. With a DIP generator to fit the target/adversarial input, we find that our image reconstruction exhibits quite interesting learning preference from a feature learning perspectives, i.e., the early stage primarily learns the robust features resistant to adversarial perturbation, followed by learning non-robust features that are sensitive to adversarial perturbation. Besides, we develop an adaptive stopping strategy that adapts our method to diverse images. In this way, the proposed model obtains a unique defender for each individual adversarial input, thus being robust to various attackers. Experimental results demonstrate the superiority of our method over the state-of-the-art defense methods against white-box and black-box adversarial attacks. Tao Dai 0001, Dongxian Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia |
ACM Multimedia | 1 |
| 2019 | Second-Order Attention Network for Single Image Super-ResolutionabstractRecently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature correlations of intermediate layers, hence hindering the representational power of CNNs. To address this issue, in this paper, we propose a second-order attention network (SAN) for more powerful feature expression and feature correlation learning. Specifically, a novel train- able second-order channel attention (SOCA) module is developed to adaptively rescale the channel-wise features by using second-order feature statistics for more discriminative representations. Furthermore, we present a non-locally enhanced residual group (NLRG) structure, which not only incorporates non-local operations to capture long-distance spatial contextual information, but also contains repeated local-source residual attention groups (LSRAG) to learn increasingly abstract feature representations. Experimental results demonstrate the superiority of our SAN network over state-of-the-art SISR methods in terms of both quantitative metrics and visual quality. Tao Dai 0001, Jianrui Cai, Yongbing Zhang 0002, Shutao Xia, Lei Zhang 0006 |
CVPR | 1 |
| 2019 | Hilbert-Based Generative Defense for Adversarial ExamplesabstractAdversarial perturbations of clean images are usually imperceptible for human eyes, but can confidently fool deep neural networks (DNNs) to make incorrect predictions. Such vulnerability of DNNs raises serious security concerns about their practicability in security-sensitive applications. To defend against such adversarial perturbations, recently developed PixelDefend purifies a perturbed image based on PixelCNN in a raster scan order (row/column by row/column). However, such scan mode insufficiently exploits the correlations between pixels, which further limits its robustness performance. Therefore, we propose a more advanced Hilbert curve scan order to model the pixel dependencies in this paper. Hilbert curve could well preserve local consistency when mapping from 2-D image to 1-D vector, thus the local features in neighboring pixels can be more effectively modeled. Moreover, the defensive power can be further improved via ensembles of Hilbert curve with different orientations. Experimental results demonstrate the superiority of our method over the state-of-the-art defenses against various adversarial attacks. Yang Bai 0011, Yisen Wang 0001, Tao Dai 0001, Shutao Xia, Yong Jiang 0001 |
ICCV | 4 |
| 2019 | Attentiondrop for Convolutional Neural NetworksabstractDropout has been widely used in fully connected networks but becomes less effective for convolutional neural networks (CNNs), since the spatially correlated features still allow dropped information to flow through the network. To make dropout more practical for CNNs, structured dropout methods have been recently proposed by dropping regions with fixed shapes and random positions, which nonetheless may lead to unexpected discarding of information. To address this problem, in this paper, we propose a novel dropout variant based on attention information named AttentionDrop that drops features adaptively. Specifically, it precisely localizes masks that have irregular shapes according to the values of activation units. In addition, the use of soft values in adaptive masks lowers the risk of a complete loss of indispensable information. Experimental results demonstrate the effectiveness of our AttentionDrop on public datasets for image classification. Zhihao Ouyang, Tianbo Hao, Tao Dai 0001, Shutao Xia |
ICME | 5 |
| 2019 | Residual Frame for Noisy Video Classification According to Perceptual Quality in Convolutional Neural NetworksabstractPerceptual quality of a video describes the quality consistent with human perception. The growing popularity of short video sharing on mobile platforms such as Tik Tok and WeSee makes the video assessment system based on perceptual quality a necessity. In practice, short videos captured by mobile devices often contain different types of distortions incurred by sensor noise or compression noise, which potentially makes the videos visually unpleasing to users and may degrade the performance of deep neural networks when applied to these noisy videos. Thus, it is necessary to identify noisy videos based on video perceptual quality. However, traditional video/image noise estimation methods are designed to estimate the variance of homogeneously distributed synthetic noise, not real noise. In this paper, we propose a simple yet effective method to recognize the noisy videos using their residual frames. Since the original video frame contains rich content information, which may result in under-or over-estimation of the noise, we construct residual frames to reduce the influence of the content information while maintaining the main noise information in the video. We also create a new data set with more than 30 thousand images captured from videos with real noise. Experimental results demonstrate the effectiveness of our proposed method. Huaixuan Zhang, Yuhai Lan, Tao Dai 0001, Ruizhi Qiao, Yao Yao 0006, Shutao Xia |
ICME | 3 |
| 2019 | Self-attentive Pyramid Network for Single Image De-raining
Taian Guo, Tao Dai 0001, Jiawei Li 0006, Shutao Xia |
ICONIP (1) | 2 |
| 2019 | Automatic Grassland Degradation Estimation Using Deep LearningabstractGrassland degradation estimation is essential to prevent global land desertification and sandstorms. Typically, the key to such estimation is to measure the coverage of indicator plants. However, traditional methods of estimation rely heavily on human eyes and manual labor, thus inevitably leading to subjective results and high labor costs. In contrast, deep learning-based image segmentation algorithms are potentially capable of automatic assessment of the coverage of indicator plants. Nevertheless, a suitable image dataset comprising grassland images is not publicly available. To this end, we build an original Automatic Grassland Degradation Estimation Dataset (AGDE-Dataset), with a large number of grassland images captured from the wild. Based on AGDE-Dataset, we are able to propose a brand new scheme to automatically estimate grassland degradation, which mainly consists of two components. 1) Semantic segmentation: we design a deep neural network with an improved encoder-decoder structure to implement semantic segmentation of grassland images. In addition, we propose a novel Focal-Hinge Loss to alleviate the class imbalance of semantics in the training stage. 2) Degradation estimation: we provide the estimation of grassland degradation based on the results of semantic segmentation. Experimental results show that the proposed method achieves satisfactory accuracy in grassland degradation estimation. Xiyu Yan, Yong Jiang 0001, Shutao Xia, Tao Dai 0001, Shuo Dong, Feng Zheng 0001 |
IJCAI | 7 |
| 2019 | Making Large Ensemble of Convolutional Neural Networks via Bootstrap Re-samplingabstractThe ensemble of Convolutional Neural Networks (CNNs) is known to be more accurate and robust than the component CNNs models. Along with the development of a fast training method, current research has managed to make an effective ensemble of several CNNs models and require no additional training cost. However, when the ensemble size of CNNs is further increased, it is hard to observe a corresponding performance enhancement. According to the generalization capability analysis of CNNs, this phenomenon can be explained by the oversaturation of model capacity and the close correlation among the component CNNs, especially when the CNNs are trained within the same dataset. To address this problem, we propose to train CNNs on re-sampled bootstrap datasets. Extensive experiments demonstrate the bootstrap re-sampling is effective for a large ensemble size (up to 80). Besides, benefiting from the usage of the bootstrap re-sampling technique, we can also have an unbiased estimate of the standard deviation of the ensemble output. Jiawei Li 0006, Xingchun Xiang, Tao Dai 0001, Shutao Xia |
VCIP | 3 |
| 2018 | Sure-Based Dual Domain Image DenoisingabstractRecently developed Dual Domain Image Denoising (DDID) algorithm is a simple version of block-matching 3D filtering (BM3D) by combining bilateral filter and frequency-based method. DDID and its invariants have achieved competitive results compared with state-of-the-art methods. However, this kind of methods share a common drawback: there are a few parameters of the algorithms that are data- and noise-dependent, and difficult to tune. In this paper, we propose to use Stein's unbiased risk estimate (SURE) to measure the mean square error (MSE) of the DDID algorithm for restoration of an image contaminated with additive white Gaussian noise. We derive an explicit expression for SURE value to optimize parameters without access to the noise-free signal. Experimental results demonstrate the effectiveness of the proposed parameter selection in term of both quantitative and qualitative metrics. Zhiya Xu, Tao Dai 0001, Li Niu 0002, Jiawei Li 0006, Qingtao Tang, Shutao Xia |
ICASSP | 2 |
| 2018 | Self -Paced Mixture of T Distribution ModelabstractGaussian mixture model (GMM) is a powerful probabilistic model for representing the probability distribution of observations in the population. However, the fitness of Gaussian mixture model can be significantly degraded when the data contain a certain amount of outliers. Although there are certain variants of GMM (e.g., mixture of Laplace, mixture of t distribution) attempting to handle outliers, none of them can sufficiently mitigate the effect of outliers if the outliers are far from the centroids. Aiming to remove the effect of outliers further, this paper introduces a Self-Paced Learning mechanism into mixture of t distribution, which leads to Self-Paced Mixture of t distribution model (SPTMM). We derive an Expectation-Maximization based algorithm to train SPTMM and show SPTMM is able to screen the outliers. To demonstrate the effectiveness of SPTMM, we apply the model to density estimation and clustering. Finally, the results indicate that SPTMM outperforms other methods, especially on the data with outliers. Yang Zhang 0016, Qingtao Tang, Li Niu 0002, Tao Dai 0001, Xi Xiao 0001, Shutao Xia |
ICASSP | 4 |
| 2018 | Cyclic Annealing Training Convolutional Neural Networks for Image Classification with Noisy LabelsabstractNoisy labels modeling makes a convolutional neural network (CNN) more robust for the image classification problem. However, current noisy labels modeling methods usually require an expectation-maximization (EM) based procedure to optimize the parameters, which is computationally expensive. In this paper, we utilize a fast annealing training method to speed up the CNN training in every M-step. Since the training is repeated executed along the entire EM optimization path and obtain many local minimal CNN models from every training cycle, we name it as the Cyclic Annealing Training (CAT) approach. In addition to reducing the training time, CAT can further bagging all the local minimal CNN models at the test time to improve the performance of classification. We evaluate the proposed method on several image classification datasets with different noisy labels patterns, and the results show that our CAT approach outperforms state-of-the-art noisy labels modeling methods. Jiawei Li 0006, Tao Dai 0001, Qingtao Tang, Yeli Xing, Shutao Xia |
ICIP | 2 |
| 2018 | Portrait-Aware Artistic Style TransferabstractThe goal of artistic style transfer is to transfer the style of artistic works into photos. However, the performances of existing style transfer algorithms on portraits are not very satisfactory, because the synthetic photo is either not sufficiently stylized or distorted severely in the portrait domain (i.e., foreground), which limits the use of style transfer for portraits. In this paper, we propose a novel portrait-aware artistic style transfer algorithm, which treats foreground and background differently. Particularly, we separate the foreground from the background, and apply fine-grained style transfer to the background and coarse-grained style transfer to the entire image at the same time, so that the artistic style of entire image can be transferred with the details of the portrait well preserved. Extensive experiments demonstrate the effectiveness of our proposed method. Yeli Xing, Jiawei Li 0006, Tao Dai 0001, Qingtao Tang, Li Niu 0002, Shutao Xia |
ICIP | 3 |
| 2018 | Color Image Noise Covariance Estimation with Cross-Channel Image Noise ModelingabstractNoise estimation is crucial in many image processing tasks such as denoising. Most of the existing noise estimation methods are specially developed for grayscale images. For color images, these methods simply handle each color channel independently, without considering the correlation across channels. In this work, we propose a multivariate Gaussian approach to model the noise in color images, in which we explicitly consider the inter-dependence among color channels. We design a practical method for estimating the noise covariance matrix within the proposed model. Specifically, a patch selection scheme is first introduced to select weakly textured patches through thresholding the texture strength indicators. Noticing that the patch selection actually depends on the unknown noise covariance, we present an iterative noise covariance estimation algorithm, where the patch selection and the covariance estimation are conducted alternately. Experimental results show that our method can effectively estimate the noise covariance. The practical usage is demonstrated with color image denoising. Li Dong 0006, Jiantao Zhou 0001, Tao Dai 0001 |
ICME | 3 |
| 2018 | Referenceless quality metric of multiply-distorted images based on structural degradation
Tao Dai 0001, Ke Gu 0001, Li Niu 0002, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia |
Neurocomputing | 1 |
| 2017 | Compressed Sensing Performance of Binary Matrices with Binary Column CorrelationsabstractThis paper studies a class of binary matrices with correlations between distinct columnsequal to zero or one, which has reported comparable performance with random matrices inrecent studies of compressed sensing. For such matrix, we analyze its structure propertyand provide an improved performance estimation. Weizhi Lu, Tao Dai 0001, Shutao Xia |
DCC | 2 |
| 2017 | Foveated nonlocal dual denoisingabstractRecently developed dual domain image denoising (DDID) algorithm and its variants, such as dual domain filter (DDF), achieve remarkable results by combining bilateral filter with frequency-based method. However, this kind of algorithms require large patches to guarantee the denoising performance and most of them produce ringing artifacts due to the Gibbs phenomenon induced by high-contrast details. To address these issues, we propose a Foveated Nonlocal Dual Denoising (FNDD) algorithm by unifying foveated nonlocal means and frequency-based methods. In this way, the ability to preserve the high-contrast details is noticeably improved by exploiting foveated self-similarity (patch similarity) instead of pixel similarity, thus leading to void of artifacts. Moreover, we propose an entropy-based back projection step for compensating the detail loss to further improve the performance. Experimental results validate that FNDD significantly outperforms DDID in terms of both quantitative metrics and subjective visual quality under much smaller patches, and even achieves comparable results against state-of-the-art competitors. Tao Dai 0001, Ke Gu 0001, Qingtao Tang, Kwok-Wai Hung, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia |
ICIP | 1 |
| 2017 | Blind quality assessment of multiply-distorted images based on structural degradationabstractIt is known that images available usually undergo some stages of processing (e.g., acquisition, compression, transmission and display), and each stage may introduce certain type of distortion. Hence, images distorted by multiple types of distortions are common in real applications. Research in human visual perception has evidenced that the human visual system (HVS) is sensitive to image structural information. This fact inspires us to design a new blind/no-reference (NR) image quality assessment (IQA) method to evaluate the visual quality of multiply-distorted images based on structural degradation. Specifically, quality-aware features are extracted from both the first- and high-order image structures by local binary pattern (LBP) operators. Experimental results on two well-known multiply-distorted image databases demonstrate the outstanding performance of the proposed method. Tao Dai 0001, Ke Gu 0001, Zhiya Xu, Qingtao Tang, Haoyi Liang, Yongbing Zhang 0002, Shutao Xia |
ICIP | 1 |
| 2017 | Robust Survey Aggregation with Student-t Distribution and Sparse RepresentationabstractMost existing survey aggregation methods assume that the sample data follow Gaussian distribution. However, these methods are sensitive to outliers, due to the thin-tailed property of the Gaussian distribution. To address this issue, we propose a robust survey aggregation method based on Student-t distribution and sparse representation. Specifically, we assume that the samples follow Student-$t$ distribution, instead of the common Gaussian distribution. Due to the Student-t distribution, our method is robust to outliers, which can be explained from both Bayesian point of view and non-Bayesian point of view. In addition, inspired by James-Stain estimator (JS) and Compressive Averaging (CAvg), we propose to sparsely represent the global mean vector by an adaptive basis comprising both data-specific basis and combined generic bases. Theoretically, we prove that JS and CAvg are special cases of our method. Extensive experiments demonstrate that our proposed method achieves significant improvement over the state-of-the-art methods on both synthetic and real datasets. Qingtao Tang, Tao Dai 0001, Li Niu 0002, Yisen Wang 0001, Shutao Xia, Jianfei Cai 0001 |
IJCAI | 2 |
| 2017 | Student-t Process Regression with Student-t LikelihoodabstractGaussian Process Regression (GPR) is a powerful Bayesian method. However, the performance of GPR can be significantly degraded when the training data are contaminated by outliers, including target outliers and input outliers. Although there are some variants of GPR (e.g., GPR with Student-t likelihood (GPRT)) aiming to handle outliers, most of the variants focus on handling the target outliers while little effort has been done to deal with the input outliers. In contrast, in this work, we aim to handle both the target outliers and the input outliers at the same time. Specifically, we replace the Gaussian noise in GPR with independent Student-t noise to cope with the target outliers. Moreover, to enhance the robustness w.r.t. the input outliers, we use a Student-t Process prior instead of the common Gaussian Process prior, leading to Student-t Process Regression with Student-t Likelihood (TPRT). We theoretically show that TPRT is more robust to both input and target outliers than GPR and GPRT, and prove that both GPR and GPRT are special cases of TPRT. Various experiments demonstrate that TPRT outperforms GPR and its variants on both synthetic and real datasets. Qingtao Tang, Li Niu 0002, Yisen Wang 0001, Tao Dai 0001, Wangpeng An, Jianfei Cai 0001, Shutao Xia |
IJCAI | 4 |
| 2017 | A generic denoising framework via guided principal component analysis
Tao Dai 0001, Zhiya Xu, Haoyi Liang, Ke Gu 0001, Qingtao Tang, Yisen Wang 0001, Weizhi Lu, Shutao Xia |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Entropy-based bilateral filtering with a new range kernel
Tao Dai 0001, Weizhi Lu, Wei Wang 0138, Jilei Wang, Shutao Xia |
Signal Process. | 1 |
| 2015 | Deterministic constructions of binary measurement matrices with various sizesabstractWe introduce a general framework to deterministically construct binary measurement matrices for compressed sensing. The proposed matrices are composed of (circulant) permutation submatrix blocks and zero submatrix blocks, thus making their hardware realization convenient and easy. Firstly, using the famous Johnson bound for binary constant weight codes, we derive a new lower bound for the coherence of binary matrices with uniform column weights. Afterwards, a large class of binary base matrices with coherence asymptotically achieving this new bound are presented. Finally, by choosing proper rows and columns from these base matrices, we construct the desired measurement matrices with various sizes and they show empirically comparable performance to that of the corresponding Gaussian matrices. Xin-Ji Liu, Shutao Xia, Tao Dai 0001 |
ICASSP | 3 |
| 2015 | PMPA: A patch-based multiscale products algorithm for image denoisingabstractPatch-based algorithms for image denoising have been widely used in recent years. Most of patch-based methods just exploit patch redundancy in spatial or frequency domain without considering inter-scale dependencies. In this paper, we propose a novel patch-based multiscale products algorithm (PMPA) for image denoising. It is based on patch similarity in spatial domain and multiscale products in wavelet domain. PMPA is divided into two stages to process the smooth areas and non smooth areas (such as edges) individually. The first stage is in the wavelet domain, then a locally adaptive window-based denoising method (LAWML) based on multiscale products is applied to process those wavelet coefficients corresponding to the non smooth areas, then obtain one initial denoised image. The second stage is in the spatial domain, then a non local means algorithm is used to process those pixels in the smooth areas to obtain another initial denoised image. The final denoised image is obtained by a weighted averaging of all common pixels in both initial denoised images. Experiments show that the proposed algorithm can have competitive performance compared with the state-of-the-art patch-based denoising algorithms for most of images. Tao Dai 0001, Chaobing Song, Shutao Xia |
ICIP | 1 |
| 2015 | Single-Frame Super-Resolution via Compressive Sampling on Hybrid Reconstructions
Tao Dai 0001, Shutao Xia |
ICONIP (3) | 2 |