Longquan Dai

dblp:144/1390 · DBLP profile ↗
← Back
34ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0001-7652-5135ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 9 first-author · 11 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 PC-Flow: Preference Alignment in Flow Matching via Classifier
abstract
Flow Matching (FM) is an efficient generative modeling framework, but aligning it with human preferences remains underexplored.~Although applying Direct Preference Optimization (DPO) to diffusion models has yielded improvements, directly extending DPO-like methods to FM poses three challenges: 1) Incompatibility with ODE-based models, 2) Heavy computational cost from full model fine-tuning, and 3) Reliance on reference model quality. To address these limitations, we propose Preference Classifier for Flow Matching (PC-Flow), a novel reference-free preference alignment framework. Specifically, we reinterpret FM’s deterministic ODE as an equivalent SDE to enable DPO-style learning. Then, we introduce a lightweight classifier to model relative preferences exclusively. This approach decouples alignment from the generative model, eliminating the need for costly fine-tuning or a reference model. Theoretically, PC-Flow guarantees consistent preference-guided distribution evolution, achieves a DPO-equivalent objective without a reference model, and progressively steers generation toward preferred outputs. Experiments show that PC-Flow achieves DPO-level alignment with significantly lower training costs.
Shaomeng Wang, He Wang 0054, Longquan Dai, Jinhui Tang 0001
AAAI3
2026 Flow-guided cascaded transformer for consistent video colorization
Yan Zhai, Zishan Li, Zhulin Tao, Longquan Dai, Xianglin Huang
Pattern Recognit.4
2025 EMControl: Adding Conditional Control to Text-to-Image Diffusion Models via Expectation-Maximization
abstract
Recent advances in diffusion models focus on efficiently handling conditional generative tasks without extra training. The process involves decomposing the result into two components: 1. unconditional sample, generated in the absence of conditions; 2. condition correction, adjusting unconditional sample to include the guidance image. This adjustment is quantified by the pixel-level measure, where the latent is decoded back into a pixel image, and the forward operator translates the noisy image into the guidance domain for comparison with the guidance image. To enhance the fidelity of condition correction, we propose a learnable latent forward operator, focusing on latent-space consistency with the expectation that this latent-space consistency approximates the pixel-level fidelity measure. The encoder translates the guidance image into the latent space, and a correctional operator is proposed to rectify model mismatching in the latent guidance model. The determination of the condition term and the correction estimation is akin to solving a blind inverse problem. Our EMControl employs the Expectation-Maximization (EM) algorithm to solve the blind inverse problem during the reverse sampling process. This technique ensures that samples, once consistent with the guidance, are accurately mapped back onto the noisy data manifold, adhering to the data's inherent distribution. The EMControl has proven its effectiveness by delivering superior performance in conditional diffusion generation tasks compared to previous approaches. Moreover, its application to multiple-condition scenarios underscores its versatility and robustness across a range of generative tasks.
He Wang 0054, Longquan Dai, Jinhui Tang 0001
AAAI2
2025 NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion Models
abstract
In training-free conditional generative tasks, diffusion models utilize differentiable loss functions to steer the generative reverse process, necessitating modifications to sampling algorithms like DDPM and DDIM. However, such adjustments likely reduce flexibility and reliability. In this paper, we propose NoiseCtrl, a sampling-algorithm-agnostic technique for controlled image generation. Essentially, diffusion models generate denoised results zt−1by adding a predicted meanµtwith random noise ϵt. NoiseCtrl specifically adjusts the random noise while leaving the underlying sampling algorithms unchanged. At each step t, NoiseCtrl converts the unconditional Gaussian noise into conditional noise $\varepsilon _t^\prime $ by substituting the isotropic Gaussian distribution with the von Mises–Fisher distribution. This substitution introduces a directional focus while preserving the randomness required for conditional image generation. Thanks to this non-intrusive design, NoiseCtrl is straightforward to integrate and has been extensively validated through experiments, demonstrating its adaptability for different diffusion algorithms and superior performance across various conditional generation tasks.
Longquan Dai, He Wang 0054, Jinhui Tang 0001
CVPR1
2025 AccCtr: Accelerating Training-Free Conditional Control For Diffusion Models
abstract
In current training-free Conditional Diffusion Models (CDM), the sampling process is steered by the gradient, which measures the discrepancy between the guidance and the condition extracted by a pre-trained condition extraction network. These methods necessitate small guidance steps, resulting in longer sampling times. To address the issue of slow sampling, we introduce AccCtr, a method that simplifies the conditional sampling algorithm by maximizing the sum of two objectives. The local maximum set of one objective is contained within the local maximum set of the other. Leveraging this relationship, we decompose the joint optimization into two parts, alternately maximizing each objective. By analyzing the steps involved in optimizing these objectives, we identify the most time-consuming steps and recommend retraining condition extraction network—a relatively simple task—to reduce its computational cost. Integrating AccCtr into current CDMs is a seamless task that does not impose a significant computational burden. Extensive testing has demonstrated that AccCtr offers superior sample quality and faster generation times.
Longquan Dai, He Wang 0054, Shaomeng Wang, Jinhui Tang 0001
IJCAI1
2025 Conducting Conditional Diffusion by Estimating the Mean Vector of von Mises-Fisher Distribution
abstract
Recent diffusion model advancements aim to handle conditional generative tasks without extra training. Existing training-free methods add a correction term at each denoising step, but they often face computational instability and lack controllability, especially with limited samples and large noise. We propose a new approach using the von Mises-Fisher (vMF) distribution to model the denoised result, turning the conditional generation task into an estimation problem for vMF parameters. We formulate the conditional diffusion model as a mean vector estimation problem for the Gaussian distribution, noting that this can be seen as an estimation problem from noisy observations. When the sampling number is small, the estimation is unstable. To address this, we optimize the mean vector of the vMF distribution by minimizing the KL divergence between the prior and posterior distributions. This approach not only addresses the computational instability but also improves the controllability and quality of the generated results. Once these parameters are determined, the denoised result can be sampled directly from the vMF distribution. Estimating the parameters requires minimal additional code and incurs negligible computational overhead while significantly improving performance. Extensive experiments across various conditional generation tasks, including depth maps, edge detection, segmentation, and style guidance, demonstrate the superiority and versatility of our method. Our approach consistently outperforms existing training-free methods and even surpasses some training-required methods in terms of visual quality and controllability.
Longquan Dai, He Wang 0054, Xiaolu Wei, Shaomeng Wang, Jinhui Tang 0001
ACM Multimedia1
2025 Generative Semantic Probing for Vision-Language Models via Hierarchical Feature Optimization
abstract
Vision-language models (VLMs) has demonstrated impressive cross-modal alignment. However, their internal mechanisms of associating text concepts with visual patterns remain opaque. This opacity raises a critical question: What visual patterns do VLMs inherently associate with text concepts? Current methods for decoding representations of VLMs often produce suboptimal outputs, hindering to probe the clear visual patterns. To address this, we introduce Generative Semantic Probing (GSP), a novel training-free framework that synthesizes images to probe the implicit semantic preferences of VLMs. Our method generates visual patterns that maximize the similarity to the target text embeddings, through three core components: (1) Hierarchical Feature Decomposition, which decomposes the image generation across multi-scale feature levels; (2) Feature Space Constraint, which constrains the optimization within semantically meaningful feature subspace; (3) Quality Assessment Module, which ensures the generation of visually plausible outputs. Experiments validate our method's strengths in high-fidelity image generation and interpretable model analysis. Beyond text-to-image generation, style transfer and image editing applications, our framework enables unprecedented visualization of VLMs' decision boundaries. By exposing implicit preferences and systematic biases in the cross-modal association, our work provides a valuable insight for both understanding and improvement of the vision-language alignment.
He Wang 0054, Longquan Dai, Shihao Pu, Shaomeng Wang, Jinhui Tang 0001
ACM Multimedia2
2025 Aligning Text-to-Image Diffusion Models to Human Preference by Classification
abstract
Text-to-image diffusion models are typically trained on large-scale web data, often resulting in outputs that misalign with human preferences. Inspired by preference learning in large language models, we propose ABC (Alignment by Classification), a simple yet effective framework for aligning diffusion models with human preferences. In contrast to prior DPO-based methods that depend on suboptimal supervised fine-tuned (SFT) reference models, ABC assumes access to an ideal reference model perfectly aligned with human intent and reformulates alignment as a classification problem. Under this view, we recognize that preference data naturally forms a semi-supervised classification setting. To address this, we propose a data augmentation strategy that transforms preference comparisons into fully supervised training signals. We then introduce a classification-based ABC loss to guide alignment. Our alignment by classification approach could effectively steer the diffusion model toward the behavior of the ideal reference. Experiments on various diffusion models show that our ABC consistently outperforms existing baselines, offering a scalable and robust solution for preference-based text-to-image fine-tuning.
Longquan Dai, Xiaolu Wei, He Wang 0054, Shaomeng Wang, Jinhui Tang 0001
NeurIPS1
2025 DISCO: DISCrete nOise for Conditional Control in Text-to-Image Diffusion Models
abstract
A major challenge in using diffusion models is aligning outputs with user-defined conditions. Existing conditional generation methods fall into two major categories: classifier-based guidance, which requires differentiable target models and gradient-based correction; and classifier-free guidance, which embeds conditions directly into the diffusion model but demands expensive joint training and architectural coupling. In this work, we introduce a third paradigm: DISCrete nOise (DISCO) guidance, which replaces the continuous conditional correction term with a finite codebook of discrete noise vectors sampled from a Gaussian prior. Conditional generation is reformulated as a code selection task, and we train prediction network to choose the optimal code given the intermediate diffusion state and the conditioning input. Our approach is differentiability-free, and training-efficient, avoiding the gradient computation and architectural redundancy of prior methods. Empirical results demonstrate that DISCO achieves competitive controllability while substantially reducing resource demands, positioning it as a scalable and effective alternative for conditional diffusion generation.
Longquan Dai, Dejiao Xue, He Wang 0054, Jinhui Tang 0001
NeurIPS1
2025 Physical-Layer Secure Optical Transmission Based on Randomized Quantization Noise
abstract
In this paper, we propose a novel physical layer security optical transmission scheme utilizing randomized quantization noise. The proposed approach encrypts low-order plaintext signals into ultra-high-order ciphertext using principles similar to quantum noise stream cipher (QNSC). While, the ultra-dense quadrature amplitude modulation (QAM) ciphertext waveforms are masked by intrinsic quantization noise. The combination of digital delta-sigma quantization and analog chaotic random scrambling not only produces randomized quantization noise, but also naturally supports one-time processing of masking ciphertext and generating keystream. Whereby, the in-band quantization noise (IBN) conceals nearby ciphertext levels bolstering security, and the out-of-band quantization noise (OOBN) is digitized to generate keystream using the Toeplitz hashing extractor. Experimental results show that 256-QAM plaintext signals were securely transmitted over standard single-mode fiber: 163-Gbps signals over 400 km, and 79.7-Gbps signals over 1800 km. To evaluate system performance, theoretical models for signal-to-noise ratio (SNR), bit error rate (BER), and number of masked signals (NMS) are derived as functions of the oversampling ratio (OSR). Our findings reveal a trade-off between transmission performance and security performance. Our results confirm that this scrambling effectively eliminates the correlation between the masking noise components. Toeplitz hashing extractor can effectively reduce the complexity of keystream generation and obtain a source-independent random keystream.
Longquan Dai, Qi Yang 0008, Lei Deng 0006, Deming Liu, Xiaoxiao Dai, Mengfan Cheng
IEEE Trans. Inf. Forensics Secur.3
2025 Causal Inference Hashing for Long-Tailed Image Retrieval
abstract
In hashing-based long-tailed image retrieval, the dominance of data-rich head classes often hinders the learning of effective hash codes for data-poor tail classes due to inherent long-tailed bias. Interestingly, this bias also contains valuable prior knowledge by revealing inter-class dependencies, which can be beneficial for hash learning. However, previous methods have not thoroughly analyzed this tangled negative and positive effects of long-tailed bias from a causal inference perspective. In this paper, we propose a novel hash framework that employs causal inference to disentangle detrimental bias effects from beneficial ones. To capture good bias in long-tailed datasets, we construct hash mediators that conserve valuable prior knowledge from class centers. Furthermore, we propose a de-biased hash loss To enhance the beneficial bias effects while mitigating adverse ones, leading to more discriminative hash codes. Specifically, this loss function leverages the beneficial bias captured by hash mediators to support accurate class label prediction, while mitigating harmful bias by blocking its causal path to the hash codes and refining predictions through backdoor adjustment. Extensive experimental results on four widely used datasets demonstrate that the proposed method improves retrieval performance against the state-of-the-art methods by large margins. The source code is available at https://github.com/IMAG-LuJin/CIH.
Lu Jin 0001, Zhengyun Lu, Zechao Li, Yonghua Pan, Longquan Dai, Jinhui Tang 0001, Ramesh Jain 0001
IEEE Trans. Image Process.5
2024 OmniFusion: Exemplar-Based Video Colorization Using OmniMotion and DifFusion Priors
Xiaoyuan Fang, Longquan Dai, Jinhui Tang 0001
ACCV (5)2
2024 The diversified-equal loss for image translation tasks
Qianhao Wu, Longquan Dai, Jinhui Tang 0001
Pattern Recognit. Lett.2
2023 Odd: One-Class Anomaly Detection Via The Diffusion Model
abstract
Anomaly detection identifies instances that deviate the distribution of the normal class. Recently, the diffusion models have shown great promise. Our research revealed that by training the diffusion model solely on normal data, it is able to transform both normal and anomalous samples into normal images. Employing this discovery, we propose ODD (One-Class Anomaly Detection via the Diffusion model), which consists of: a diffusion model to convert both normal and anomalous data into normal data, and a similarity network enhanced with outlier exposure to measure the semantic distance between the input and output of the diffusion model. If the score is low, the input is considered as an anomaly instance. The ODD is evaluated on a variety of datasets. Both qualitative and quantitative results demonstrate that our method outperforms existing state-of-the-art techniques.
He Wang 0054, Longquan Dai, Jinglin Tong, Yan Zhai
ICIP2
2023 Flow-Guided Transformer for Video Colorization
abstract
Video colorization aims to add color to black-and-white films. However, propagating color information to the whole video clip accurately is a challenging task. In this paper, we propose Flow-Guided Transformer for Video Colorization (FGTVC), consisting of a Global Motion Aggregation (GMA) module, Residual modules, Flow-Guided Attention blocks (FGAB) based on encoder and decoder, to exploit the information from the neighbor patch with high similarity for each video patch colorization. Specifically, we employ Transformer to capture the long-distance dependencies between frames and learn non-local self-similarity in the frame. To overcome the shortcomings of previous optical flow-based methods, FGAB enjoys the guidance of optical flow to sample elements from spatio-temporal adjacent frames when calculating self-attention. Experiments show that the proposed FGTVC has an outstanding performance than the state-of-the-art methods. In addition, comprehensive findings demonstrate the superiority of our framework in real-world video colorization tasks.
Yan Zhai, Zhulin Tao, Longquan Dai, He Wang 0054, Xianglin Huang, Lifang Yang
ICIP3
2023 Out-of-Distribution with Text-to-Image Diffusion Models
Jinglin Tong, Longquan Dai
PRCV (11)2
2022 A selection function for pitched instrument source separation
Yukai Gong, Longquan Dai, Jinhui Tang 0001
Multim. Syst.2
2022 RMVAE: one-class classification via divergence regularization and maximization mutual information
Longquan Dai
Multim. Syst.2
2022 iFlowGAN: An Invertible Flow-Based Generative Adversarial Network for Unsupervised Image-to-Image Translation
abstract
We propose iFlowGAN that learns an invertible flow (a sequence of invertible mappings) via adversarial learning and exploit it to transform a source distribution into a target distribution for unsupervised image-to-image translation. Existing GAN-based generative model such as CycleGAN [1], StarGAN [2], AGGAN [3] and CyCADA [4] needs to learn a highly under-constraint forward mapping F: X → Y from a source domain X to a target domain Y. Researchers do this by assuming there is a backward mapping B: Y → X such that x and y are fixed points of the composite functions B °F and F °B. Inspired by zero-order reverse filtering [5], we (1) understand F via contraction mappings on a metric space; (2) provide a simple yet effective algorithm to present B via the parameters of F in light of Banach fixed point theorem; (3) provide a Lipschitz-regularized network which indicates a general approach to compose the inverse for arbitrary Lipschitz-regularized networks via Banach fixed point theorem. This network is useful for image-to-image translation tasks because it could save the memory for the weights of B. Although memory can also be saved by directly coupling the weights of the forward and backward mappings, the performance of the image-to-image translation network degrades significantly. This explains why current GAN-based generative models including CycleGAN must take different parameters to compose the forward and backward mappings instead of employing the same weights to build both mappings. Taking advantage of the Lipschitz-regularized network, we not only build iFlowGAN to solve the redundancy shortcoming of CycleGAN but also assemble the corresponding iFlowGAN versions of StarGAN, AGGAN and CyCADA without breaking their network architectures. Extensive experiments show that the iFlowGAN version could produce comparable results of the original implementation while saving half parameters.
Longquan Dai, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 MAT-Net: Representing Appearance-Irrelevant Warp Field by Multiple Affine Transformations
abstract
Warp-based methods for image animation estimate a warp field what do a rearrangement on the pixels of the input image to roughly align with the target image. Current methods predict accurate warp field by using manually annotated data. In this paper, we propose a simple method (MAT-net) to predict more precise warp field in self-supervised way. MAT-net decomposes complex spatial object movement between two images into multiple simple local motions (i.e. affine transformation) occurring in different areas of images. Sequentially, our model calculates a warp field depicting complex object movement by combining all local motions. MAT-net encodes appearance-irrelevant object movement accurately. Compared to the state-of-the-art method, MAT-net generates more realistic images with faster inference speed. We published the source code of our project online.
Longquan Dai
ICIP2
2021 Coupled Patch Similarity Network FOR One-Shot Fine-Grained Image Recognition
abstract
One-shot fine-grained image recognition (OSFG) aims to distinguish different fine-grained categories with only one training sample per category. Previous works mainly focus on learning a global feature representation through only a using single similarity metric branch, which is unsuitable for OSFG to effectively capture subtle and local differences under limited supervision. In this work, we propose a Coupled Patch Similarity Network (CPSN) for OSFG. Firstly, we propose a Feature Enhancement Module (FEM) to extract more discriminative features of the fine-grained samples. Then, we develop two coupled and symmetrical branches to capture discriminative parts of the samples and reduce the deviation of the distance metric. For each branch, we design a Patch Similarity Module (PSM) to calculate the patch similarity map for the sample pair. Especially, a Patch Weight Generator (PWG) is proposed to generate the patch weight map, which indicates the degree of importance for each position in the patch similarity map, so that the model can focus on diverse and informative parts. We analyze the effect of the different components in the proposed network, and extensive experimental results demonstrate the effectiveness and superiority of the proposed method on two fine-grained benchmark datasets.
Hao Tang 0007, Longquan Dai
ICIP3
2021 Robust Kernelized Multiview Self-Representation for Subspace Clustering
abstract
In this article, we propose a multiview self-representation model for nonlinear subspaces clustering. By assuming that the heterogeneous features lie within the union of multiple linear subspaces, the recent multiview subspace learning methods aim to capture the complementary and consensus from multiple views to boost the performance. However, in real-world applications, data feature usually resides in multiple nonlinear subspaces, leading to undesirable results. To this end, we propose a kernelized version of tensor-based multiview subspace clustering, which is referred to as Kt-SVD-MSC, to jointly learn self-representation coefficients in mapped high-dimensional spaces and multiple views correlation in unified tensor space. In view-specific feature space, a kernel-induced mapping is introduced for each view to ensure the separability of self-representation coefficients. In unified tensor space, a new kind of tensor low-rank regularizer is employed on the rotated self-representation coefficient tensor to preserve the global consistency across different views. We also derive an algorithm to efficiently solve the optimization problem with all the subproblems having closed-form solutions. Furthermore, by incorporating the nonnegative and sparsity constraints, the proposed method can be easily extended to a useful variant, meaning that several useful variants can be easily constructed in a similar way. Extensive experiments of the proposed method are tested on eight challenging data sets, in which a significant (even a breakthrough) advance over state-of-the-art multiview clustering is achieved.
Yuan Xie 0006, Yanyun Qu, Dacheng Tao, Wensheng Zhang 0002, Longquan Dai, Lizhuang Ma
IEEE Trans. Neural Networks Learn. Syst.6
2020 X-NET For Single Image Raindrop Removal
abstract
Photos taken on rainy days are likely degraded by raindrops adhered to camera lenses. Removing raindrops from images is a tough task. Its difficulties lie in restoring high frequency information from corrupted images while keeping the color of restored images consistent with human perception. To solve these problems, we propose an end-to-end convolutional neural network consisting of X-Net and RAD-Net (Raindrop Automatic Detection Net). X-Net takes advantage of Long Skip Connections and Cross Branch Connections to generate raindrop-free image with enough details. RAD-Net assists X-Net to produce better results by yielding raindrop location. Extensive experiments show our approach outperforms state-of-the-art methods quantitatively and qualitatively.
Jiamin Lin, Longquan Dai
ICIP2
2020 Speed Up Bilateral Filtering via Sparse Approximation on a Learned Cosine Dictionary
abstract
The edge-preserving bilateral filter (BF) is a widely used smoothing tool in many applications. However, its brute-force implementation depends on the size of the box window. The shortcoming causes BF time-consuming for the image processing task with a large window. To make the computational complexity irrelevant to the window size, sparse approximations of the filtering kernels are calculated on a learned cosine dictionary by two steps. First, all possible frequencies are learned (estimated) from the filtering kernel to compose a cosine dictionary. Then, the sparse approximation is conducted on the learned dictionary to seek the optimal cosine approximation for both the range and spatial kernels. By making use of the one-dimensional cosine approximation for the range kernel, the BF is transformed into spatial convolutions. Subsequently, by employing the two-dimensional cosine approximation for the spatial kernel, spatial convolutions are decomposed into box filters of which the computational complexity is O(1). To the best of our knowledge, our approach is the first method that adaptively constructs a cosine dictionary according to the input kernel. This merit guarantees the best filtering accuracy and efficiency. These advantages are corroborated by several carefully designed experiments.
Longquan Dai, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Hyper-Laplacian Regularized Multilinear Multiview Self-Representations for Clustering and Semisupervised Learning
abstract
In this paper, we address the multiview nonlinear subspace representation problem. Traditional multiview subspace learning methods assume that the heterogeneous features of the data usually lie within the union of multiple linear subspaces. However, instead of linear subspaces, the data feature actually resides in multiple nonlinear subspaces in many real-world applications, resulting in unsatisfactory clustering performance. To overcome this, we propose a hyper-Laplacian regularized multilinear multiview self-representation model, which is referred to as HLR-M2VS, to jointly learn multiple views correlation and a local geometrical structure in a unified tensor space and view-specific self-representation feature spaces, respectively. In unified tensor space, a well-founded tensor low-rank regularization is adopted to impose on the self-representation coefficient tensor to ensure global consensus among different views. In view-specific feature space, hypergraph-induced hyper-Laplacian regularization is utilized to preserve the local geometrical structure embedded in a high-dimensional ambient space. An efficient algorithm is then derived to solve the optimization problem of the established model with theoretical convergence guarantee. Furthermore, the proposed model can be extended to semisupervised classification without introducing any additional parameters. An extensive experiment of our method is conducted on many challenging datasets, where a clear advance over state-of-the-art multiview clustering and multiview semisupervised classification approaches is achieved.
Yuan Xie 0006, Wensheng Zhang 0002, Yanyun Qu, Longquan Dai, Dacheng Tao
IEEE Trans. Cybern.4
2019 Fast and Error-Bounded Space-Variant Bilateral Filtering
Mengke Yuan, Longquan Dai, Dong-Ming Yan 0001, Liqiang Zhang 0001, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.2
2019 Interpreting and Extending the Guided Filter via Cyclic Coordinate Descent
abstract
The guided filter (GF) is a widely used smoothing tool in computer vision and image processing. However, to the best of our knowledge, few papers investigate the mathematical connection between this filter and the least-squares optimization. In this paper, we first interpret the guided filter as the cyclic coordinate descent (CCD) solver of a least-squares objective function. This discovery implies an extension approach to generalize the guided filter since we can change the least-squares objective function and define new filters as the first pass iteration of the CCD solver of modified objective functions. In addition, referring to the iterative minimizing procedure of the CCD, we can derive new rolling filtering schemes. So, we are reasonable to say that our discovery not only reveals an approach to design new GF-like filters adapting to specific requirements of applications but also offers thorough explanations for two rolling filtering schemes of the guided filter as well as the method to extend them. Experiments prove our new proposed filters and rolling filtering schemes could produce state-of-the-art results.
Longquan Dai, Mengke Yuan, Yuan Xie 0006, Xiaopeng Zhang 0001, Jinhui Tang 0001
IEEE Trans. Image Process.1
2018 Designing by Training: Acceleration Neural Network for Fast High-Dimensional Convolution
abstract
The high-dimensional convolution is widely used in various disciplines but has a serious performance problem due to its high computational complexity. Over the decades, people took a handmade approach to design fast algorithms for the Gaussian convolution. Recently, requirements for various non-Gaussian convolutions have emerged and are continuously getting higher. However, the handmade acceleration approach is no longer feasible for so many different convolutions since it is a time-consuming and painstaking job. Instead, we propose an Acceleration Network (AccNet) which turns the work of designing new fast algorithms to training the AccNet. This is done by: 1, interpreting splatting, blurring, slicing operations as convolutions; 2, turning these convolutions to $g$CP layers to build AccNet. After training, the activation function $g$ together with AccNet weights automatically define the new splatting, blurring and slicing operations. Experiments demonstrate AccNet is able to design acceleration algorithms for a ton of convolutions including Gaussian/non-Gaussian convolutions and produce state-of-the-art results.
Longquan Dai, Yuan Xie 0006, Jinhui Tang 0001
NeurIPS1
2017 Hardware-Efficient Guided Image Filtering for Multi-label Problem
abstract
The Guided Filter (GF) is well-known for its linear complexity. However, when filtering an image with an n-channel guidance, GF needs to invert an n × n matrix for each pixel. To the best of our knowledge existing matrix inverse algorithms are inefficient on current hardwares. This shortcoming limits applications of multichannel guidance in computation intensive system such as multi-label system. We need a new GF-like filter that can perform fast multichannel image guided filtering. Since the optimal linear complexity of GF cannot be minimized further, the only way thus is to bring all potentialities of current parallel computing hardwares into full play. In this paper we propose a hardware-efficient Guided Filter (HGF), which solves the efficiency problem of multichannel guided image filtering and yields competent results when applying it to multi-label problems with synthesized polynomial multichannel guidance. Specifically, in order to boost the filtering performance, HGF takes a new matrix inverse algorithm which only involves two hardware-efficient operations: element-wise arithmetic calculations and box filtering. In order to break the linear model restriction, HGF synthesizes a polynomial multichannel guidance to introduce nonlinearity. Benefiting from our polynomial guidance and hardware-efficient matrix inverse algorithm, HGF not only is more sensitive to the underlying structure of guidance but also achieves the fastest computing speed. Due to these merits, HGF obtains state-of-the-art results in terms of accuracy and efficiency in the computation intensive multi-label systems.
Longquan Dai, Mengke Yuan, Zechao Li, Xiaopeng Zhang 0001, Jinhui Tang 0001
CVPR1
2016 Speeding Up the Bilateral Filter: A Joint Acceleration Way
abstract
Computational complexity of the brute-force implementation of the bilateral filter (BF) depends on its filter kernel size. To achieve the constant-time BF whose complexity is irrelevant to the kernel size, many techniques have been proposed, such as 2D box filtering, dimension promotion, and shiftability property. Although each of the above techniques suffers from accuracy and efficiency problems, previous algorithm designers were used to take only one of them to assemble fast implementations due to the hardness of combining them together. Hence, no joint exploitation of these techniques has been proposed to construct a new cutting edge implementation that solves these problems. Jointly employing five techniques: kernel truncation, best N-term approximation as well as previous 2D box filtering, dimension promotion, and shiftability property, we propose a unified framework to transform BF with arbitrary spatial and range kernels into a set of 3D box filters that can be computed in linear time. To the best of our knowledge, our algorithm is the first method that can integrate all these acceleration techniques and, therefore, can draw upon one another's strong point to overcome deficiencies. The strength of our method has been corroborated by several carefully designed experiments. In particular, the filtering accuracy is significantly improved without sacrificing the efficiency at running time.
Longquan Dai, Mengke Yuan, Xiaopeng Zhang 0001
IEEE Trans. Image Process.1
2015 Fully Connected Guided Image Filtering
abstract
This paper presents a linear time fully connected guided filter by introducing the minimum spanning tree (MST) to the guided filter (GF). Since the intensity based filtering kernel of GF is apt to overly smooth edges and the fixed-shape local box support region adopted by GF is not geometric-adaptive, our filter introduces an extra spatial term, the tree similarity, to the filtering kernel of GF and substitutes the box window with the implicit support region by establishing all-pairs-connections among pixels in the image and assigning the spatial-intensity-aware similarity to these connections. The adaptive implicit support region composed by the pixels with large kernel weights in the entire image domain has a big advantage over the predefined local box window in presenting the structure of an image for the reason that: 1, MST can efficiently present the structure of an image, 2, the kernel weight of our filter considers the tree distance defined on the MST. Due to these reasons, our filter achieves better edge-preserving results. We demonstrate the strength of the proposed filter in several applications. Experimental results show that our method produces better results than state-of-the-art methods.
Longquan Dai, Mengke Yuan, Feihu Zhang, Xiaopeng Zhang 0001
ICCV1
2015 Segment Graph Based Image Filtering: Fast Structure-Preserving Smoothing
abstract
In this paper, we design a new edge-aware structure, named segment graph, to represent the image and we further develop a novel double weighted average image filter (SGF) based on the segment graph. In our SGF, we use the tree distance on the segment graph to define the internal weight function of the filtering kernel, which enables the filter to smooth out high-contrast details and textures while preserving major image structures very well. While for the external weight function, we introduce a user specified smoothing window to balance the smoothing effects from each node of the segment graph. Moreover, we also set a threshold to adjust the edge-preserving performance. These advantages make the SGF more flexible in various applications and overcome the "halo" and "leak" problems appearing in most of the state-of-the-art approaches. Finally and importantly, we develop a linear algorithm for the implementation of our SGF, which has an O(N) time complexity for both gray-scale and high dimensional images, regardless of the kernel size and the intensity range. Typically, as one of the fastest edge-preserving filters, our CPU implementation achieves 0.15s per megapixel when performing filtering for 3-channel color images. The strength of the proposed filter is demonstrated by various applications, including stereo matching, optical flow, joint depth map upsampling, edge-preserving smoothing, edges detection, image abstraction and texture editing.
Feihu Zhang, Longquan Dai, Shiming Xiang, Xiaopeng Zhang 0001
ICCV2
2015 Depth map upsampling using compressive sensing based model
Longquan Dai, Haoxing Wang, Xiaopeng Zhang 0001
Neurocomputing1
2015 Fast Minimax Path-Based Joint Depth Interpolation
abstract
We propose a fast minimax path-based depth interpolation method. The algorithm computes for each target pixel varying contributions from reliable depth seeds, and weighted averaging is used to interpolate missing depths. Compared with state-of-the-art joint geodesic upsampling method which selects the K nearest seeds to interpolate missing depths with O(Kn) complexity, our method does not need to limit the number of seeds to K and reduces the computational complexity to O(n). In addition, the minimax path chooses a path with the smallest maximum immediate pairwise pixel difference on it, so it tends to preserve sharp depth discontinuities better. In contrast to the results of previous depth upsampling algorithms, our approach can provide accurate depths with fewer artifacts.
Longquan Dai, Feihu Zhang, Xing Mei, Xiaopeng Zhang 0001
IEEE Signal Process. Lett.1