Xin Du 0005

dblp:18/4833-5 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
9since 2021 · last 2023
0000-0002-6215-9733ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Vision Transformers for Single Image Dehazing
abstract
Image dehazing is a representative low-level vision task that estimates latent haze-free images from hazy images. In recent years, convolutional neural network-based methods have dominated image dehazing. However, vision Transformers, which has recently made a breakthrough in high-level vision tasks, has not brought new dimensions to image dehazing. We start with the popular Swin Transformer and find that several of its key designs are unsuitable for image dehazing. To this end, we propose DehazeFormer, which consists of various improvements, such as the modified normalization layer, activation function, and spatial information aggregation scheme. We train multiple variants of DehazeFormer on various datasets to demonstrate its effectiveness. Specifically, on the most frequently used SOTS indoor set, our small model outperforms FFA-Net with only 25% #Param and 5% computational cost. To the best of our knowledge, our large model is the first method with the PSNR over 40 dB on the SOTS indoor set, dramatically outperforming the previous state-of-the-art methods. We also collect a large-scale realistic remote sensing dehazing dataset for evaluating the method's capability to remove highly non-homogeneous haze. We share our code and dataset at https://github.com/IDKiro/DehazeFormer.
Yuda Song 0002, Zhuqing He, Hui Qian 0001, Xin Du 0005
IEEE Trans. Image Process.4
2022 Modular Degradation Simulation and Restoration for Under-Display Camera
Yuda Song 0002, Xin Du 0005
ACCV (3)3
2022 Multi-Curve Translator for High-Resolution Photorealistic Image Translation
Yuda Song 0002, Hui Qian 0001, Xin Du 0005
ECCV (15)3
2022 FreSCo: Frequency-Domain Scan Context for LiDAR-based Place Recognition with Translation and Rotation Invariance
abstract
Place recognition plays a crucial role in relocalization and loop closure detection tasks for robots and vehicles. This paper seeks a well-defined global descriptor for LiDAR-based place recognition. Compared to local descriptors, global descriptors show remarkable performance in urban road scenes but are usually viewpoint-dependent. To this end, we propose a simple yet robust global descriptor dubbed FreSCo that decomposes the viewpoint difference during revisit and achieves both translation and rotation invariance by leveraging Fourier Transform and circular shift technique. Besides, a fast two-stage pose estimation method is proposed to estimate the relative pose after place retrieval by utilizing the compact 2D point clouds extracted from the original data. Experiments show that FreSCo exhibited superior performance than contemporaneous methods on sequences of different scenes from multiple datasets. Code will be publicly available at https://github.com/soytony/FreSCo.
Yongzhi Fan, Xin Du 0005, Lun Luo, Jizhong Shen
ICARCV2
2022 DPVI: A Dynamic-Weight Particle-Based Variational Inference Framework
abstract
The recently developed Particle-based Variational Inference (ParVI) methods drive the empirical distribution of a set of fixed-weight particles towards a given target distribution by iteratively updating particles' positions. However, the fixed weight restriction greatly confines the empirical distribution's approximation ability, especially when the particle number is limited. In this paper, we propose to dynamically adjust particles' weights according to a Fisher-Rao reaction flow. We develop a general Dynamic-weight Particle-based Variational Inference (DPVI) framework according to a novel continuous composite flow, which evolves the positions and weights of particles simultaneously. We show that the mean-field limit of our composite flow is actually a Wasserstein-Fisher-Rao gradient flow of the associated dissimilarity functional. By using different finite-particle approximations in our general framework, we derive several efficient DPVI algorithms. The empirical results demonstrate the superiority of our derived DPVI algorithms over their fixed-weight counterparts.
Chao Zhang 0029, Xin Du 0005, Hui Qian 0001
IJCAI3
2021 StarEnhancer: Learning Real-Time and Style-Aware Image Enhancement
abstract
Image enhancement is a subjective process whose targets vary with user preferences. In this paper, we propose a deep learning-based image enhancement method covering multiple tonal styles using only a single model dubbed StarEnhancer. It can transform an image from one tonal style to another, even if that style is unseen. With a simple one-time setting, users can customize the model to make the enhanced images more in line with their aesthetics. To make the method more practical, we propose a well-designed enhancer that can process a 4K-resolution image over 200 FPS but surpasses the contemporaneous single style image enhancement methods in terms of PSNR, SSIM, and LPIPS. Finally, our proposed enhancement method has good inter-actability, which allows the user to fine-tune the enhanced image using intuitive options.
Yuda Song 0002, Hui Qian 0001, Xin Du 0005
ICCV3
2021 SHPOS: A Theoretical Guaranteed Accelerated Particle Optimization Sampling Method
abstract
Recently, the Stochastic Particle Optimization Sampling (SPOS) method is proposed to solve the particle-collapsing pitfall of deterministic Particle Variational Inference methods by ultilizing the stochastic Overdamped Langevin dynamics to enhance exploration. In this paper, we propose an accelerated particle optimization sampling method called Stochastic Hamiltonian Particle Optimization Sampling (SHPOS). Compared to the first-order dynamics used in SPOS, SHPOS adopts an augmented second-order dynamics, which involves an extra momentum term to achieve acceleration. We establish a non-asymptotic convergence analysis for SHPOS, and show that it enjoys a faster convergence rate than SPOS. Besides, we also propose a variance-reduced stochastic gradient variant of SHPOS for tasks with large-scale datasets and complex models. Experiments on both synthetic and real data validate our theory and demonstrate the superiority of SHPOS over the state-of-the-art.
Chao Zhang 0029, Hui Qian 0001, Xin Du 0005, Lingwei Peng
IJCAI4
2021 Low-rate DoS attack detection method based on hybrid deep neural networks
Congyuan Xu, Jizhong Shen, Xin Du 0005
J. Inf. Secur. Appl.3
2021 Estimating Generalized Gaussian Blur Kernels for Out-of-Focus Image Deblurring
abstract
Out-of-focus blur is a common image degradation phenomenon that occurs in case of lens defocusing. The out-of-focus blur kernel is usually modeled as a Gaussian function or a uniform disk in previous work. In this paper, we propose that it can be more accurately depicted using the generalized Gaussian (GG) function. This is motivated by the theoretical analysis of the out-of-focus blur and the practical observation of real blur kernels. We show that as the out-of-focus blur kernels are of specific shapes, the GG function can be further simplified to a single-parameter model. We estimate the parameter of the GG blur kernel from image patches containing step edges, and obtain the clear image by non-blind image deblurring. Experimental results validate that the proposed GG blur kernel estimation algorithm outperforms the state-of-the-art ones deploying either parametric (disk and Gaussian) or nonparametric kernels, and consequently benefits the image deblurring process.
Yuqi Liu 0005, Xin Du 0005, Shujie Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Automatical Enhancement and Denoising of Extremely Low-light Images
abstract
Deep convolutional neural networks (DCNN) based methodologies have achieved remarkable performance on various low-level vision tasks recently. Restoring images captured at night is one of the trickiest low-level vision tasks due to its high-level noise and low-level intensity. We propose a DCNN-based methodology, Illumination and Noise Separation Network (INSNet), which performs both denoising and enhancement on these extremely low-light images. INSNet fully utilizes global-ware features and local-ware features using the modified network structure and image sampling scheme. Compared to well-designed complex neural networks, our proposed methodology only needs to add a bypass network to the existing network. However, it can boost the quality of recovered images dramatically but only increase the computational cost by less than 0.1%. Even without any manual settings, INSNet can stably restore the extremely low-light images to desired high-quality images.
Yuda Song 0002, Yunfang Zhu, Xin Du 0005
ICPR3
2020 Web page classification based on heterogeneous features and a combination of multiple classifiers
abstract
Precise web page classification can be achieved by evaluating features of web pages, and the structural features of web pages are effective complements to their textual features. Various classifiers have different characteristics, and multiple classifiers can be combined to allow classifiers to complement one another. In this study, a web page classification method based on heterogeneous features and a combination of multiple classifiers is proposed. Different from computing the frequency of HTML tags, we exploit the tree-like structure of HTML tags to characterize the structural features of a web page. Heterogeneous textual features and the proposed tree-like structural features are converted into vectors and fused. Confidence is proposed here as a criterion to compare the classification results of different classifiers by calculating the classification accuracy of a set of samples. Multiple classifiers are combined based on confidence with different decision strategies, such as voting, confidence comparison, and direct output, to give the final classification results. Experimental results demonstrate that on the Amazon dataset, 7-web-genres dataset, and DMOZ dataset, the accuracies are increased to 94.2%, 95.4%, and 95.7%, respectively. The fusion of the textual features with the proposed structural features is a comprehensive approach, and the accuracy is higher than that when using only textual features. At the same time, the accuracy of the web page classification is improved by combining multiple classifiers, and is higher than those of the related web page classification algorithms.
Xin Du 0005, Jizhong Shen
Frontiers Inf. Technol. Electron. Eng.2
2020 Model Parameter Learning for Real-Time High-Resolution Image Enhancement
abstract
Deep learning-based methods have achieved remarkable performance in image enhancement but generally require huge GPU memory, and computational cost when enhancing the high-resolution images. We explore to enhance high-resolution images like manual retouching by digital artists to obtain stable, and excellent performance in real-time. We extract the features from their thumbnails, and utilizes these features to guide the enhancement of the high-resolution images. Our designed light-weight convolution, and polynomial attention consume limited computational cost on the full-resolution image to retouch the image from both pixel-level, and global-level. Compared with the contemporary methods, our proposed method can work in real-time but surpass the state-of-the-art methods over 1.0 dB on the high-resolution image enhancement benchmark.
Yuda Song 0002, Yunfang Zhu, Xin Du 0005
IEEE Signal Process. Lett.3
2020 Grouped Multi-Scale Network for Real-World Image Denoising
abstract
Deep learning-based methods have surpassed the traditional methods in image denoising due to the prior knowledge accumulated on the large dataset. Because of the difference between additive Gaussian white noise (AWGN) and real noise, researchers have recently paid more attention to real-world image denoising. Based on the characteristics of real-world image denoising, we revise the method of noise synthesis and design a novel network to make full use of the multi-scale context. Our proposed approach dramatically surpasses all the contemporary approaches on the sRGB track of the DND benchmark [1] with PSNR over 40 dB and SSIM over 0.96.
Yuda Song 0002, Yunfang Zhu, Xin Du 0005
IEEE Signal Process. Lett.3
2020 A Method of Few-Shot Network Intrusion Detection Based on Meta-Learning Framework
abstract
Conventional intrusion detection systems based on supervised learning techniques require a large number of samples for training, while in some scenarios, such as zero-day attacks, security agencies can only intercept a limited number of shots of malicious samples. Therefore, there is a need for few-shot detection. In this paper, a detection method based on a meta-learning framework is proposed for this purpose. The proposed method can be used to distinguish and compare a pair of network traffic samples as a basic task of learning, including a normal unaffected sample and a malicious one. To accomplish this task, we design a deep neural network (DNN) named FC-Net, which mainly comprises two parts: feature extraction network and comparison network. FC-Net learns a pair of feature maps for classification from a pair of network traffic samples, then compares the obtained feature maps, and finally determines whether the pair of samples belongs to the same type. To evaluate the proposed detection method, we construct two datasets for few-shot network intrusion detection based on real network traffic data sources, using a specifically developed approach. The experimental results indicate that the proposed detection method is universal and is not limited to specific datasets or attack types. Training and testing on the same datasets demonstrate that the proposed method can achieve the average detection rate up to 98.88%. The outcome of training on one dataset and testing on the other one confirms that the proposed method can achieve better performance. In a few-shot scenario, malicious samples in an untrained dataset can be detected successfully, and the average detection rate is up to 99.62%.
Congyuan Xu, Jizhong Shen, Xin Du 0005
IEEE Trans. Inf. Forensics Secur.3
2019 Detection method of domain names generated by DGAs based on semantic representation and deep neural network
Congyuan Xu, Jizhong Shen, Xin Du 0005
Comput. Secur.3