VLDB 2026 Research / reviewers in the wild / expert
Fei Wang 0008
dblp:52/3194-8
· DBLP profile ↗
102ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0003-3462-8472ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 3 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 1 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-authorDatabases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient superpixel-guided global-local spectral clustering for large-scale HSI
Ben Yang, Xuetao Zhang 0001, Yongqiang Luo, Feiping Nie 0001, Fei Wang 0008, Badong Chen |
Neurocomputing | 5 |
| 2026 | Attribute graph adjusted trace ratio linear discriminant analysis for feature extraction
Fei Wang 0008, Xinpei Wen, Zhiping Lin 0001, Feiping Nie 0001 |
Pattern Recognit. | 3 |
| 2026 | Improve noise tolerance of robust feature selection via block-sparse projection learning
Jie Wang 0164, Zheng Wang 0037, Yu Guo 0006, Rong Wang 0001, Fei Wang 0008, Feiping Nie 0001 |
Pattern Recognit. | 5 |
| 2026 | Decision forest with fast-determined optimal parameter intervals of base learners
Fei Wang 0008, Badong Chen, Zhiping Lin 0001, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2026 | One-Step Multi-View Clustering With Adaptive Low-Rank Anchor-Graph LearningabstractIn light of their capability to capture structural information while reducing computing complexity, anchor graph-based multi-view clustering (AGMC) methods have attracted considerable attention in large-scale clustering problems. Nevertheless, existing AGMC methods still face the following two issues: 1) They directly embedded diverse anchor graphs into a consensus anchor graph (CAG), and hence ignore redundant information and numerous noises contained in these anchor graphs, leading to a decrease in clustering effectiveness; 2) They drop effectiveness and efficiency due to independent post-processing to acquire clustering indicators. To overcome the aforementioned issues, we deliver a novel one-step multi-view clustering method with adaptive low-rank anchor-graph learning (OMCAL). To construct a high-quality CAG, OMCAL provides a nuclear norm-based adaptive CAG learning model against information redundancy and noise interference. Then, to boost clustering effectiveness and efficiency substantially, we incorporate category indicator acquisition and CAG learning into a unified framework. Numerous studies conducted on ordinary and large-scale datasets indicate that OMCAL outperforms existing state-of-the-art methods in terms of clustering effectiveness and efficiency. Zhiyuan Xue, Ben Yang, Xuetao Zhang 0001, Fei Wang 0008, Zhiping Lin 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingabstractListening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeline that first generates intermediate 3D motion signals such as 3DMM coefficients, and then synthesizes the videos by deterministic rendering, suffering from limited motion expressiveness and low visual quality (e.g. 256×256). In this work, we propose a novel listening head generation method that harnesses the generative capabilities of the diffusion model for both motion generation and high-quality rendering. Crucially, we propose an effective hybrid motion modeling module that addresses training difficulties caused by the scarcity of listening head data while preserving the intricate details that may be lost in explicit motion representations. We further develop a tailored control guidance for head pose and facial expression, by integrating their intrinsic motion characteristics. Our method enables high-fidelity video generation with 512 × 512 resolution and delivers vivid listener motion feedback. We conduct comprehensive experiments and obtain superior performance in terms of both visual quality and motion expressiveness compared with existing methods. Yanbo Fan, Xuan Wang 0009, Yu Guo 0006, Fei Wang 0008 |
CVPR | 5 |
| 2025 | PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
Yong Liu 0031, Hang Dong 0001, Jinshan Pan, Qingji Dong, Kai Chen 0023, Rongxiang Zhang, Lean Fu, Fei Wang 0008 |
ICCV | 8 |
| 2025 | Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures Via Joint Reconstruction and Registration
Yuan Sun 0003, Xuan Wang 0009, WeiLi Zhang, Yanbo Fan, Yu Guo 0006, Fei Wang 0008 |
ICCV | 7 |
| 2025 | UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
Yong Liu 0031, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008 |
ACM Multimedia | 7 |
| 2025 | Revitalizing Image Dehazing in the Real World: A High-Quality Dataset and a Customized MethodabstractExisting dehazing methods face challenges in generalization due to the lack of paired real-world training data and tailored models. Recently, some semi-supervised/unsupervised schemes have been explored, achieving impressive performance. However, their performance still depends heavily on synthetic training data and the introduced prior-based strong constraints do not always hold. In this paper, we first introduce RealHQ-HAZE, a new dataset with 200 collected real-world hazy images, 200 corresponding carefully rendered haze-free images, and an additional 1000 varicolored hazy images transferred from the collected images. We also propose a prior-compensated multi-stage dehazing network, PMDN, which can learn different levels of real-world haze distribution through multi-stage progressive learning. To utilize prior knowledge effectively, we introduce a prior-based feature compensation module, guiding intermediate results with an adaptive weight. Additionally, we propose a MixCut consistent dehazing strategy to mix paired and derived images using a cross-cutting scheme, reinforcing dehazing through consistency principles. Extensive experiments demonstrate the effectiveness of our dataset and the superiority of PMDN compared to existing state-of-the-art dehazing methods. Yong Liu 0031, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008 |
Comput. Vis. Media | 5 |
| 2025 | Correntropy-Induced Hypergraph Spectral Clustering With Discrete OptimizationabstractHypergraph clustering has garnered considerable attention in complex learning tasks due to its powerful capacity for modeling high-order relationships among samples. Nevertheless, existing methods encounter two fundamental challenges: 1) The need for an additional discretization step following low-dimensional spectral embedding, which introduces a suboptimal mismatch between continuous embeddings and discrete cluster assignments, thereby impairing clustering performance; and 2) the susceptibility to diverse and complex noise are commonly present in real-world scenarios, which significantly compromises clustering robustness. To address these issues, we propose a novel correntropy-induced hypergraph spectral clustering (CIHSC) model. Different from current spectral clustering methods, CIHSC integrates a correntropy-based framework to enable direct discrete spectral decomposition on hypergraphs, eliminating the need for post discretization and thereby enhancing clustering fidelity and robustness. To effectively address the non-convex optimization arising from the correntropy-induced objective, we develop a half-quadratic optimization strategy tailored to the CIHSC model. Extensive experiments conducted on both real-world and noise-contaminated datasets demonstrate that CIHSC consistently outperforms state-of-the-art clustering methods in terms of performance and robustness. Jiaqi Nie, Ben Yang, Zhiyuan Xue, Xuetao Zhang 0001, Fei Wang 0008 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Scalable Min-Max Multi-View Spectral ClusteringabstractMulti-view spectral clustering has attracted considerable attention since it can explore common geometric structures from diverse views. Nevertheless, existing min-min framework-based models adopt internal minimization to find the view combination with the minimized within-cluster variance, which will lead to effectiveness loss since the real clusters often exhibit high within-cluster variance. To address this issue, we provide a novel scalable min-max multi-view spectral clustering (SMMSC) model to improve clustering performance. Besides, anchor graphs, rather than full sample graphs, are utilized to reduce the computational complexity of graph construction and singular value decomposition, thereby enhancing the applicability of SMMSC to large-scale applications. Then, we rewrite the min-max model as a minimized optimal value function, demonstrate its differentiability, and develop an efficient gradient descent-based algorithm to optimize it with linear computational complexity. Moreover, we demonstrate that the resultant solution of the proposed algorithm is the global optimum. Numerous experiments on different real-world datasets, including some large-scale datasets, demonstrate that SMMSC outperforms existing state-of-the-art multi-view clustering methods regarding clustering performance. Ben Yang, Xuetao Zhang 0001, Jinghan Wu, Feiping Nie 0001, Fei Wang 0008, Badong Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Fast Multiview Anchor-Graph ClusteringabstractDue to its high computational complexity, graph-based methods have limited applicability in large-scale multiview clustering tasks. To address this issue, many accelerated algorithms, especially anchor graph-based methods and indicator learning-based methods, have been developed and made a great success. Nevertheless, since the restrictions of the optimization strategy, these accelerated methods still need to approximate the discrete graph-cutting problem to a continuous spectral embedding problem and utilize different discretization strategies to obtain discrete sample categories. To avoid the loss of effectiveness and efficiency caused by the approximation and discretization, we establish a discrete fast multiview anchor graph clustering (FMAGC) model that first constructs an anchor graph of each view and then generates a discrete cluster indicator matrix by solving the discrete multiview graph-cutting problem directly. Since the gradient descent-based method makes it hard to solve this discrete model, we propose a fast coordinate descent-based optimization strategy with linear complexity to solve it without approximating it as a continuous one. Extensive experiments on widely used normal and large-scale multiview datasets show that FMAGC can improve clustering effectiveness and efficiency compared to other state-of-the-art baselines. Ben Yang, Xuetao Zhang 0001, Jinghan Wu, Feiping Nie 0001, Zhiping Lin 0001, Fei Wang 0008, Badong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Multi-View Subspace Clustering With Consensus Graph Contrastive LearningabstractA significant challenge in multi-view clustering lies in the comprehensive extraction of consistency and complementary information from heterogeneous multi-view data. Numerous methods employ contrastive learning techniques to explore the information between views. However, the basic contrastive learning strategy does not consider cluster information when constructing sample pairs, potentially leading to the emergence of false negative pairs (FNPs). To tackle this concern, we propose a Multi-view Subspace Clustering with Consensus Graph Contrastive Learning (CGCL) model. Specifically, a self-representation layer is designed to acquire a consensus graph that elucidates the overall data distribution. Furthermore, a contrastive learning layer utilizes the cluster information embedded in the consensus graph to yield reliable sample pairs, resulting in a reduction of the detrimental FNPs and the extraction of complementary information from the various views. Extensive experiments on public datasets demonstrate the effectiveness of CGCL. Jie Zhang 0090, Yuan Sun 0003, Yu Guo 0006, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008 |
ICASSP | 6 |
| 2024 | Spectral Aggregation Cross-Square Transformer for Hyperspectral Image Denoising
Yang Liu 0385, Yantao Ji, Jiahua Xiao, Yu Guo 0006, Peilin Jiang, Haiwei Yang, Fei Wang 0008 |
ICPR (15) | 7 |
| 2024 | Dual-path dehazing network with spatial-frequency feature fusion
Li Wang 0072, Hang Dong 0001, Chao Zhu 0007, Huibin Tao, Yu Guo 0006, Fei Wang 0008 |
Pattern Recognit. | 7 |
| 2024 | Joint learning of latent subspace and structured graph for multi-view clustering
Yu Guo 0006, Zheng Wang 0037, Fei Wang 0008 |
Pattern Recognit. | 4 |
| 2024 | Coordinate Descent Optimized Trace Difference Model for Joint Clustering and Feature Extraction
Fei Wang 0008, Zhongheng Li, Zheng Wang 0037, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2024 | Efficient Local Coherent Structure Learning via Self-Evolution Bipartite GraphabstractDimensionality reduction (DR) targets to learn low-dimensional representations for improving discriminability of data, which is essential for many downstream machine learning tasks, such as image classification, information clustering, etc. Non-Gaussian issue as a long-standing challenge brings many obstacles to the applications of DR methods that established on Gaussian assumption. The mainstream way to address above issue is to explore the local structure of data via graph learning technique, the methods based on which however suffer from a common weakness, that is, exploring locality through pairwise points causes the optimal graph and subspace are difficult to be found, degrades the performance of downstream tasks, and also increases the computation complexity. In this article, we first propose a novel self-evolution bipartite graph (SEBG) that uses anchor points as the landmark of subclasses, and learns anchor-based rather than pairwise relationships for improving the efficiency of locality exploration. In addition, we develop an efficient local coherent structure learning (ELCS) algorithm based on SEBG, which possesses the ability of updating the edges of graph in learned subspace automatically. Finally, we also provide a multivariable iterative optimization algorithm to solve proposed problem with strict theoretical proofs. Extensive experiments have verified the superiorities of the proposed method compared to related SOTA methods in terms of performance and efficiency on several real-world benchmarks and large-scale image datasets with deep features. Zheng Wang 0037, Qi Li 0045, Feiping Nie 0001, Rong Wang 0001, Fei Wang 0008, Xuelong Li 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Double-Structured Sparsity Guided Flexible Embedding Learning for Unsupervised Feature SelectionabstractIn this article, we propose a novel unsupervised feature selection model combined with clustering, named double-structured sparsity guided flexible embedding learning (DSFEL) for unsupervised feature selection. DSFEL includes a module for learning a block-diagonal structural sparse graph that represents the clustering structure and another module for learning a completely row-sparse projection matrix using the$\ell_{2,0}$-norm constraint to select distinctive features. Compared with the commonly used$\ell_{2,1}$-norm regularization term, the$\ell_{2,0}$-norm constraint can avoid the drawbacks of sparsity limitation and parameter tuning. The optimization of the$\ell_{2,0}$-norm constraint problem, which is a nonconvex and nonsmooth problem, is a formidable challenge, and previous optimization algorithms have only been able to provide approximate solutions. In order to address this issue, this article proposes an efficient optimization strategy that yields a closed-form solution. Eventually, through comprehensive experimentation on nine real-world datasets, it is demonstrated that the proposed method outperforms existing state-of-the-art unsupervised feature selection methods. Yu Guo 0006, Yuan Sun 0003, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Toward Robust Discriminative Projections Learning Against Adversarial Patch AttacksabstractAs one of the most popular supervised dimensionality reduction methods, linear discriminant analysis (LDA) has been widely studied in machine learning community and applied to many scientific applications. Traditional LDA minimizes the ratio of squared norms, which is vulnerable to the adversarial examples. In recent studies, many -norm-based robust dimensionality reduction methods are proposed to improve the robustness of model. However, due to the difficulty of -norm ratio optimization and weakness on defending a large number of adversarial examples, so far, scarce works have been proposed to utilize sparsity-inducing norms for LDA objective. In this article, we propose a novel robust discriminative projections learning (rDPL) method based on the -norm trace-ratio minimization optimization algorithm. Minimizing the -norm ratio problem directly is a much more challenging problem than the traditional methods, and there is no existing optimization algorithm to solve such nonsmooth terms ratio problem. We derive a new efficient algorithm to solve this challenging problem and provide a theoretical analysis on the convergence of our algorithm. The proposed algorithm is easy to implement and converges fast in practice. Extensive experiments on both synthetic data and several real benchmark datasets show the effectiveness of the proposed method on defending the adversarial patch attack by comparison with many state-of-the-art robust dimensionality reduction methods. Zheng Wang 0037, Feiping Nie 0001, Hua Wang 0007, Heng Huang 0001, Fei Wang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Joint Anchor Graph Embedding and Discrete Feature Scoring for Unsupervised Feature SelectionabstractThe success of existing unsupervised feature selection (UFS) methods heavily relies on the assumption that the intrinsic relationships among original high-dimensional (HD) data samples exist in the discriminative low-dimension (LD) subspace. However, previous UFS methods commonly construct pairwise graphs and employ$\ell_{2,1}$-norm regularization to severally preserve the local structure and calculate the score of features, which is computationally complex and easy to get stuck into local optimum, so that those approaches cannot be applied in dealing with large-scale datasets in practice. To overcome this challenge, we propose a novel UFS method, in which a novel anchor graph embedding paradigm is designed to extract the local neighborhood relationships among data samples by reducing the computational complexity of graph construction to be linear in the number of data. Moreover, to improve the optimality of selected features as well as the performance of downstream tasks, we propose a discrete feature scoring mechanism, which imposes orthogonal$\ell_{2,0}$-norm constraints on learned projections, in order to enhance the distinction of feature scores as well as reduce the probability of falling into local optimum. In addition, solving the proposed nonconvex and nonsmooth NP-hard problem is challenging, and we present an efficient optimization algorithm to address it and acquire a closed-form solution of the transformation matrix. Extensive experiments demonstrate the effectiveness and efficiency of the proposed UFS by comparison with several state-of-the-art approaches to clustering and image segmentation tasks. Zheng Wang 0037, Dongming Wu 0001, Rong Wang 0001, Feiping Nie 0001, Fei Wang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Local-to-Global Registration for Bundle-Adjusting Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural3D representations and registering camera frames exist, they are susceptible to suboptimal solutions if poorly initialized. We propose L2G-NeRF, a Local-to-Global registration method for bundle-adjusting Neural Radiance Fields: first, a pixel-wise flexible alignment, followed by a framewise constrained parametric alignment. Pixel-wise local alignment is learned in an unsupervised way via a deep network which optimizes photometric reconstruction errors. framewise global alignment is performed using differentiable parameter estimation solvers on the pixel-wise correspondences to find a global transformation. Experiments on synthetic and real-world data show that our method outperforms the current state-of-the-art in terms of high-fidelity reconstruction and resolving large camera pose misalignment. Our module is an easy-to-use plugin that can be applied to NeRF variants and other neural field applications. The Code and supplementary materials are available at https://rover-xingyu.github.io/L2G-NeRF/. Xuan Wang 0009, Qi Zhang 0029, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 7 |
| 2023 | UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceabstractNeural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach that can render an editable free-view video of a human performer in real-time. It separates the high-frequency (i.e., non-smooth) human appearance from the 3D volume, and encodes them into 2D neural texture stacks (NTS). The smooth UV volumes allow much smaller and shallower neural networks to obtain densities and texture coordinates in 3D while capturing detailed appearance in 2D NTS. For editability, the mapping between the parameterized human model and the smooth texture coordinates allows us a better generalization on novel poses and shapes. Furthermore, the use of NTS enables interesting applications, e.g., retexturing. Extensive experiments on CMU Panoptic, ZJU Mocap, and H36M datasets show that our model can render$960\times 540$images in 30FPS on average with comparable photo-realism to state-of-the-art methods. The project and supplementary materials are available at https://fanegg.github.io/UV-Volumes. Xuan Wang 0009, Qi Zhang 0029, Xiaoyu Li 0002, Yu Guo 0006, Jue Wang 0001, Fei Wang 0008 |
CVPR | 8 |
| 2023 | SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationabstractGenerating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other hand, explicitly using 3D information also suffers problems of stiff expression and incoherent video. We present SadTalker, which generates 3D motion coefficients (head pose, expression) of the 3DMM from audio and implicitly modulates a novel 3D-aware face render for talking head generation. To learn the realistic motion coefficients, we explicitly model the connections between audio and different types of motion coefficients individually. Precisely, we present ExpNet to learn the accurate facial expression from audio by distilling both coefficients and 3D-rendered faces. As for the head pose, we design PoseVAE via a conditional VAE to synthesize head motion in different styles. Finally, the generated 3D motion coefficients are mapped to the unsupervised 3D keypoints space of the proposed face render to synthesize the final video. We conducted extensive experiments to demonstrate the superiority of our method in terms of motion and video quality.11The code and demo videos are available at https://sadtalker.github.io. Xiaodong Cun, Xuan Wang 0009, Yong Zhang 0034, Xi Shen 0001, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 8 |
| 2023 | Boosting Video Super Resolution with Patch-Based Temporal Redundancy Optimization
Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Boyang Liang, Yu Guo 0006, Ding Liu 0001, Lean Fu, Fei Wang 0008 |
ICANN (7) | 9 |
| 2023 | Multi-Source Fusion for Voxel-Based 7-DoF Grasping Pose EstimationabstractIn this work, we tackle the problem of 7-DoF grasping pose estimation(6-DoF with the opening width of parallel-jaw gripper) from point cloud data, which is a fundamental task in robotic manipulation. Most existing methods adopt 3D voxel CNNs as the backbone for their efficiency in handling unordered point cloud data. However, we found that these approaches overlook detailed information of the point clouds, resulting in decreased performance. Through our analysis, we identified quantization loss and boundary information loss within 3D convolutional layers as the primary causes of this issue. To address these challenges, we introduced two novel branches: one adds an extra positional encoding operation to preserve details and unique features for each point, and the other uses a 2D CNN to operate on the range-based image, which better aggregates boundary information on a continuous 2D domain. To integrate these branches with the original branch, we introduced a novel multi-source fusion gated mechanism to aggregate features. Our approach achieved state-of-the-art performance on the Graspnet-1Billion benchmark and demonstrated high success rates in real robotic experiments across different scenes. Our work has the potential to improve the performance of robotic grasping systems and contribute to the field of robotics. Junning Qiu, Fei Wang 0008, Zheng Dang |
IROS | 2 |
| 2023 | Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-ResolutionabstractRecent years have witnessed a few attempts of vision transformers for single image super-resolution (SISR). Since the high resolution of intermediate features in SISR models increases memory and computational requirements, efficient SISR transformers are more favored. Based on some popular transformer backbone, many methods have explored reasonable schemes to reduce the computational complexity of the self-attention module while achieving impressive performance. However, these methods only focus on the performance on the training platform (e.g., Pytorch/Tensorflow) without further optimization for the deployment platform (e.g., TensorRT). Therefore, they inevitably contain some redundant operators, posing challenges for subsequent deployment in real-world applications. In this paper, we propose a deployment-friendly transformer unit, namely UFONE (i.e., UnFolding ONce is Enough), to alleviate these problems. In each UFONE, we introduce an Inner-patch Transformer Layer (ITL) to efficiently reconstruct the local structural information from patches and a Spatial-Aware Layer (SAL) to exploit the long-range dependencies between patches. Based on UFONE, we propose a Deployment-friendly Inner-patch Transformer Network (DITN) for the SISR task, which can achieve favorable performance with low latency and memory usage on both training and deployment platforms. Furthermore, to further boost the deployment efficiency of the proposed DITN on TensorRT, we also provide an efficient substitution for layer normalization and propose a fusion optimization strategy for specific operators. Extensive experiments show that our models can achieve competitive results in terms of qualitative and quantitative performance with high deployment efficiency. Yong Liu 0031, Hang Dong 0001, Boyang Liang, Songwei Liu, Qingji Dong, Kai Chen 0023, Fangmin Chen, Lean Fu, Fei Wang 0008 |
ACM Multimedia | 9 |
| 2023 | Multi-view semi-supervised learning with adaptive graph fusion
Qianyao Qiang, Bin Zhang 0022, Feiping Nie 0001, Fei Wang 0008 |
Neurocomputing | 4 |
| 2023 | Efficient random subspace decision forests with a simple probability dimensionality setting scheme
Fei Wang 0008, Zhongheng Li, Peilin Jiang, Fuji Ren, Feiping Nie 0001 |
Inf. Sci. | 2 |
| 2023 | Multi-View Discrete Clustering: A Concise ModelabstractIn most existing graph-based multi-view clustering methods, the eigen-decomposition of the graph Laplacian matrix followed by a post-processing step is a standard configuration to obtain the target discrete cluster indicator matrix. However, we can naturally realize that the results obtained by the two-stage process will deviate from that obtained by directly solving the primal clustering problem. In addition, it is essential to properly integrate the information from different views for the enhancement of the performance of multi-view clustering. To this end, we propose a concise model referred to as Multi-view Discrete Clustering (MDC), aiming at directly solving the primal problem of multi-view graph clustering. We automatically weigh the view-specific similarity matrix, and the discrete indicator matrix is directly obtained by performing clustering on the aggregated similarity matrix without any post-processing to best serve graph clustering. More importantly, our model does not introduce an additive, nor does it has any hyper-parameters to be tuned. An efficient optimization algorithm is designed to solve the resultant objective problem. Extensive experimental results on both synthetic and real benchmark datasets verify the superiority of the proposed model. Qianyao Qiang, Bin Zhang 0022, Fei Wang 0008, Feiping Nie 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | An Effective Clustering Optimization Method for Unsupervised Linear Discriminant AnalysisabstractThe recent work Unsupervised Linear Discriminant Analysis (Un-LDA) completes its clustering process during the alternating optimization by converting equivalently the objective and finally using the K-means algorithm. However, the K-means algorithm has its inherent drawbacks. It is hard for the K-means algorithm to deal well with some complex clustering cases where there are too many real clusters or non-convex clusters. In this paper, a novel clustering optimization method is presented to accomplish the clustering process in Un-LDA and the resulting method can be named Un-LDA(CD). Specifically, instead of the K-means algorithm, an elaborately designed coordinate descent algorithm is adopted to obtain the clusters after the objective function goes through a series of simple but deft equivalent conversions. Extensive experiments have demonstrated that the coordinate descent clustering solution for Un-LDA can outperform the original K-means based solution on the tested data sets especially those complex data sets with a pretty large number of real clusters. Fei Wang 0008, Fuji Ren, Zhongheng Li, Feiping Nie 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Efficient Multi-View K-Means Clustering With Multiple Anchor GraphsabstractMulti-view clustering has attracted a lot of attention due to its ability to integrate information from distinct views, but how to improve efficiency is still a hot research topic. Anchor graph-based methods and k-means-based methods are two current popular efficient methods, however, both have limitations. Clustering on the derived anchor graph takes a while for anchor graph-based methods, and the efficiency of k-means-based methods drops significantly when the data dimension is large. To emphasize these issues, we developed an efficient multi-view k-means clustering method with multiple anchor graphs (EMKMC). It first constructs anchor graphs for each view and then integrates these anchor graphs using an improved k-means strategy to obtain sample categories without any extra post-processing. Since EMKMC combines the high-efficiency portions of anchor graph-based methods and k-means-based methods, its efficiency is substantially higher than current fast methods, especially when dealing with large-scale high-dimensional multi-view data. Extensive experiments demonstrate that, compared to other state-of-the-art methods, EMKMC can boost clustering efficiency by several to thousands of times while maintaining comparable or even exceeding clustering effectiveness. Ben Yang, Xuetao Zhang 0001, Zhongheng Li, Feiping Nie 0001, Fei Wang 0008 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | ECCA: Efficient Correntropy-Based Clustering Algorithm With Orthogonal Concept FactorizationabstractOne of the hottest topics in unsupervised learning is how to efficiently and effectively cluster large amounts of unlabeled data. To address this issue, we propose an orthogonal conceptual factorization (OCF) model to increase clustering effectiveness by restricting the degree of freedom of matrix factorization. In addition, for the OCF model, a fast optimization algorithm containing only a few low-dimensional matrix operations is given to improve clustering efficiency, as opposed to the traditional CF optimization algorithm, which involves dense matrix multiplications. To further improve the clustering efficiency while suppressing the influence of the noises and outliers distributed in real-world data, an efficient correntropy-based clustering algorithm (ECCA) is proposed in this article. Compared with OCF, an anchor graph is constructed and then OCF is performed on the anchor graph instead of directly performing OCF on the original data, which can not only further improve the clustering efficiency but also inherit the advantages of the high performance of spectral clustering. In particular, the introduction of the anchor graph makes ECCA less sensitive to changes in data dimensions and still maintains high efficiency at higher data dimensions. Meanwhile, for various complex noises and outliers in real-world data, correntropy is introduced into ECCA to measure the similarity between the matrix before and after decomposition, which can greatly improve the clustering effectiveness and robustness. Subsequently, a novel and efficient half-quadratic optimization algorithm was proposed to quickly optimize the ECCA model. Finally, extensive experiments on different real-world datasets and noisy datasets show that ECCA can archive promising effectiveness and robustness while achieving tens to thousands of times the efficiency compared with other state-of-the-art baselines. Ben Yang, Xuetao Zhang 0001, Feiping Nie 0001, Badong Chen, Fei Wang 0008, Zhixiong Nan, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Deep Recurrent Neural Network with Multi-Scale Bi-directional Propagation for Video DeblurringabstractThe success of the state-of-the-art video deblurring methods stems mainly from implicit or explicit estimation of alignment among the adjacent frames for latent video restoration. However, due to the influence of the blur effect, estimating the alignment information from the blurry adjacent frames is not a trivial task. Inaccurate estimations will interfere the following frame restoration. Instead of estimating alignment information, we propose a simple and effective deep Recurrent Neural Network with Multi-scale Bi-directional Propagation (RNN-MBP) to effectively propagate and gather the information from unaligned neighboring frames for better video deblurring. Specifically, we build a Multi-scale Bi-directional Propagation (MBP) module with two U-Net RNN cells which can directly exploit the inter-frame information from unaligned neighboring hidden states by integrating them in different scales. Moreover, to better evaluate the proposed algorithm and existing state-of-the-art methods on real-world blurry scenes, we also create a Real-World Blurry Video Dataset (RBVD) by a well-designed Digital Video Acquisition System (DVAS) and use it as the training and evaluation dataset. Extensive experimental results demonstrate that the proposed RBVD dataset effectively improve the performance of existing algorithms on real-world blurry videos, and the proposed algorithm performs favorably against the state-of-the-art methods on three typical benchmarks. The code is available at https://github.com/XJTU-CVLAB-LOWLEVEL/RNN-MBP. Chao Zhu 0007, Hang Dong 0001, Jinshan Pan, Boyang Liang, Lean Fu, Fei Wang 0008 |
AAAI | 7 |
| 2022 | Frequency-aware Deep Dual-path Feature Enhancement Network for Image DehazingabstractSingle image dehazing is a challenging task due to the severe degradations caused by the particles in the air. Recently, various CNN-based methods have been proposed and they have achieved promising results on some dehazing tasks. However, the existing end-to-end dehazing networks process high-frequency information and low-frequency information at the same time. Therefore, most dehazing methods cannot restore dehazed image with satisfying high-frequency details. In this paper, we propose a Frequency-aware deep Dual-path Feature enhancement Network (FDF-Net) to better restore the high-frequency information while removing the haze. To achieve this, we introduce a Dual-path Feature Enhancement (DFE) block, which contains two branches: one path is to remedy the missing spatial information from high-resolution features, and the other one is to obtain new features to increase the variety of features. We believe the dual-path architecture can help the first path to focus on the recovering the high-frequency information. Furthermore, to reserve more detailed image information from the features with larger resolution, we adopt a wavelet transform module during the downsampling process of the encoder module to directly pass the high frequency information to the next level. The extensive experiments show the superiority of the proposed model over previous methods on the benchmark datasets as well as real-world hazy images. Hang Dong 0001, Li Wang 0072, Boyang Liang, Yu Guo 0006, Fei Wang 0008 |
ICPR | 6 |
| 2022 | Adaptive weighted robust iterative closest point
Yu Guo 0006, Luting Zhao, Xuetao Zhang 0001, Shaoyi Du, Fei Wang 0008 |
Neurocomputing | 6 |
| 2022 | Multi-view unsupervised dimensionality reduction with probabilistic neighbors
Qianyao Qiang, Bin Zhang 0022, Fei Wang 0008, Feiping Nie 0001 |
Neurocomputing | 3 |
| 2022 | Efficient and Robust MultiView Clustering With Anchor Graph RegularizationabstractMulti-view clustering has received widespread attention owing to its effectiveness by integrating multi-view data appropriately, but traditional algorithms have limited applicability to large-scale real-world data due to their high computational complexity and low robustness. Focusing on the aforementioned issues, we propose an efficient and robust multi-view clustering algorithm with anchor graph regularization (ERMC-AGR). In this work, a novel anchor graph regularization (ARG) is designed to improve the quality of the learned embedded anchor graph (EAG), and the obtained EAG is decomposed by nonnegative matrix factorization (NMF) under correntropy criterion to acquire clustering results directly. Different from the traditional graph regularization that needs to construct a large-scale Laplacian matrix pertaining to the all-sample graph, our lightweight AGR, constructed from the perspective of anchors, can reduce the computational complexity significantly while improving the EAG quality. Moreover, a factor matrix of NMF is constrained to be the cluster indicator matrix to omit additional k-means after optimization. Subsequently, correntropy is utilized to improve the effectiveness and robustness of ERMC-AGR owing to its promising performance to complex noises and outliers. Extensive experiments on real-world datasets and noisy datasets show that ERMC-ARG can improve the clustering efficiency and robustness while ensuring comparable or even better effectiveness. Ben Yang, Xuetao Zhang 0001, Zhiping Lin 0001, Feiping Nie 0001, Badong Chen, Fei Wang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Fast Multiview Clustering With Spectral EmbeddingabstractSpectral clustering has been a hot topic in unsupervised learning owing to its remarkable clustering effectiveness and well-defined framework. Despite this, due to its high computation complexity, it is unable of handling large-scale or high-dimensional data, particularly multi-view large-scale data. To address this issue, in this paper, we propose a fast multi-view clustering algorithm with spectral embedding (FMCSE), which speeds up both the spectral embedding and spectral analysis stages of multi-view spectral clustering. Furthermore, unlike conventional spectral clustering, FMCSE can acquire all sample categories directly after optimization without extra k-means, which can significantly enhance efficiency. Moreover, we also provide a fast optimization strategy for solving the FMCSE model, which divides the optimization problem into three decoupled small-scale sub-problems that can be solved in a few iteration steps. Finally, extensive experiments on a variety of real-world datasets (including large-scale and high-dimensional datasets) show that, when compared to other state-of-the-art fast multi-view clustering baselines, FMCSE can maintain comparable or even better clustering effectiveness while significantly improving clustering efficiency. Ben Yang, Xuetao Zhang 0001, Feiping Nie 0001, Fei Wang 0008 |
IEEE Trans. Image Process. | 4 |
| 2022 | Fast Multi-View Semi-Supervised Learning With Learned GraphabstractMulti-view semi-supervised learning (SSL) has attracted great attention due to its effectiveness in information utilization of multiple views and labeled and unlabeled data to solve practical problems. However, most existing methods exhibit high computational complexity. Effective integration of the information on different views to achieve enhanced performance remains a challenging task. In this study, we combine an anchor-based approach with multi-view semi-supervised learning to address these problems. A novel multi-view SSL method called fast multi-view SSL (FMSSL) based on learned graph is proposed. Starting from the affinity graphs constructed by using an anchor-based strategy, FMSSL learns an optimal multi-view consensus graph by using feature and label information. The learned graph can jointly consider the relation of multiple views to approximate the manifold structure. The learned graph is then introduced into the SSL model as the weight matrix of a bipartite graph to simultaneously perform separate classification on the original samples and anchors. Accordingly, multi-view SSL can be efficiently performed, and the computational complexity can be significantly reduced. We propose an effective algorithm to optimize the objective function. Extensive experimental results on different real-world datasets demonstrate the effectiveness and efficiency of the proposed algorithm. Bin Zhang 0022, Qianyao Qiang, Fei Wang 0008, Feiping Nie 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Fast Multi-view Discrete Clustering with Anchor GraphsabstractGenerally, the existing graph-based multi-view clustering models consists of two steps: (1) graph construction; (2) eigen-decomposition on the graph Laplacian matrix to compute a continuous cluster assignment matrix, followed by a post-processing algorithm to get the discrete one. However, both the graph construction and eigen-decomposition are time-consuming, and the two-stage process may deviate from directly solving the primal problem. To this end, we propose Fast Multi-view Discrete Clustering (FMDC) with anchor graphs, focusing on directly solving the spectral clustering problem with a small time cost. We efficiently generate representative anchors and construct anchor graphs on different views. The discrete cluster assignment matrix is directly obtained by performing clustering on the automatically aggregated graph. FMDC has a linear computational complexity with respect to the data scale, which is a significant improvement compared to the quadratic one. Extensive experiments on benchmark datasets demonstrate its efficiency and effectiveness. Qianyao Qiang, Bin Zhang 0022, Fei Wang 0008, Feiping Nie 0001 |
AAAI | 3 |
| 2021 | Learning To Restore Hazy Video: A New Real-World Dataset and a New MethodabstractMost of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, video dehazing algorithms, which can acquire more satisfying dehazing results by exploiting the temporal redundancy from neighborhood hazy frames, receive less attention due to the absence of the video dehazing datasets. Therefore, we propose the first REal-world VIdeo DEhazing (REVIDE) dataset which can be used for the supervised learning of the video dehazing algorithms. By utilizing a well-designed video acquisition system, we can capture paired real-world hazy and haze-free videos that are perfectly aligned by recording the same scene (with or without haze) twice. Considering the challenge of exploiting temporal redundancy among the hazy frames, we also develop a Confidence Guided and Improved Deformable Network (CG-IDN) for video dehazing. The experiments demonstrate that the hazy scenes in the REVIDE dataset are more realistic than the synthetic datasets and the proposed algorithm also performs favorably against state-of-the-art dehazing methods. Xinyi Zhang 0005, Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Fei Wang 0008 |
CVPR | 9 |
| 2021 | Photometric Stereo Based on Multiple Kernel Learning
Yu Guo 0006, Xiaoxiao Yang, Xuetao Zhang 0001, Fei Wang 0008 |
ICIG (3) | 5 |
| 2021 | Dynamic Hypergraph Regularized Broad Learning System for Image Classification
Xiaoxiao Yang, Yu Guo 0006, Peilin Jiang, Fei Wang 0008 |
ICIG (1) | 5 |
| 2021 | Monocular 3D multi-person pose estimation via predicting factorized correction factors
Yu Guo 0006, Lichen Ma, Zhi Li 0055, Xuan Wang 0009, Fei Wang 0008 |
Comput. Vis. Image Underst. | 5 |
| 2021 | Online robust echo state broad learning system
Yu Guo 0006, Xiaoxiao Yang, Fei Wang 0008, Badong Chen |
Neurocomputing | 4 |
| 2021 | Flexible multi-view semi-supervised learning with unified graph
Zhongheng Li, Qianyao Qiang, Bin Zhang 0022, Fei Wang 0008, Feiping Nie 0001 |
Neural Networks | 4 |
| 2021 | Eigendecomposition-Free Training of Deep Networks for Linear Least-Square ProblemsabstractMany classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be tackled by solving a linear least-square problem, which can be done by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning frameworks would allow us to explicitly encode known notions of geometry, instead of having the network implicitly learn them from data. However, performing eigendecomposition within a network requires the ability to differentiate this operation. While theoretically doable, this introduces numerical instability in the optimization process in practice. In this paper, we introduce an eigendecomposition-free approach to training a deep network whose loss depends on the eigenvector corresponding to a zero eigenvalue of a matrix predicted by the network. We demonstrate that our approach is much more robust than explicit differentiation of the eigendecomposition using two general tasks, outlier rejection and denoising, with several practical examples including wide-baseline stereo, the perspective-n-point problem, and ellipse fitting. Empirically, our method has better convergence properties and yields state-of-the-art results. Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Fast Multi-View Clustering via Nonnegative and Orthogonal FactorizationabstractThe rapid growth of the number of data brings great challenges to clustering, especially the introduction of multi-view data, which collected from multiple sources or represented by multiple features, makes these challenges more arduous. How to clustering large-scale data efficiently has become the hottest topic of current large-scale clustering tasks. Although several accelerated multi-view methods have been proposed to improve the efficiency of clustering large-scale data, they still cannot be applied to some scenarios that require high efficiency because of the high computational complexity. To cope with the issue of high computational complexity of existing multi-view methods when dealing with large-scale data, a fast multi-view clustering model via nonnegative and orthogonal factorization (FMCNOF) is proposed in this paper. Instead of constraining the factor matrices to be nonnegative as traditional nonnegative and orthogonal factorization (NOF), we constrain a factor matrix of this model to be cluster indicator matrix which can assign cluster labels to data directly without extra post-processing step to extract cluster structures from the factor matrix. Meanwhile, the F-norm instead of the L2-norm is utilized on the FMCNOF model, which makes the model very easy to optimize. Furthermore, an efficient optimization algorithm is proposed to solve the FMCNOF model. Different from the traditional NOF optimization algorithm requiring dense matrix multiplications, our algorithm can divide the optimization problem into three decoupled small size subproblems that can be solved by much less matrix multiplications. Combined with the FMCNOF model and the corresponding fast optimization method, the efficiency of the clustering process can be significantly improved, and the computational complexity is nearly O(n) . Extensive experiments on various benchmark data sets validate our approach can greatly improve the efficiency when achieve acceptable performance. Ben Yang, Xuetao Zhang 0001, Feiping Nie 0001, Fei Wang 0008, Weizhong Yu, Rong Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Flexible Multi-View Unsupervised Graph EmbeddingabstractFaced with the increasing data diversity and dimensionality, multi-view dimensionality reduction has been an important technique in computer vision, data mining and multi-media applications. Since collecting labeled data is difficult and costly, unsupervised learning is of great significance. Generally, it is crucial to explore the complementarity or independence of different feature spaces in multi-view learning. How to find a low-dimensional subspace to preserve the intrinsic structure of original unlabeled high-dimensional multi-view data is still challenging. In addition, noises and outliers always appear in real data. In this study, we propose a novel model called flexible multi-view unsupervised graph embedding (FMUGE). A flexible regression residual term is introduced so that the strict linear mapping is relaxed, new-coming data and noises are better handled, and the raw data negotiate with the learned low-dimensional representation in the procedure. To ensure the consistency among multiple views, FMUGE adaptively weights different features and fuses them to get an optimal multi-view consensus similarity graph, which assists high-quality graph embedding. We propose an efficient alternating iterative algorithm to optimize the proposed model. Finally, experimental results on synthetic and benchmark datasets show the significant improvement of FMUGE over the state-of-the-art methods. Bin Zhang 0022, Qianyao Qiang, Fei Wang 0008, Feiping Nie 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Unsupervised Linear Discriminant Analysis for Jointly Clustering and Subspace LearningabstractLinear discriminant analysis (LDA) is one of commonly used supervised subspace learning methods. However, LDA will be powerless faced with the no-label situation. In this paper, the unsupervised LDA (Un-LDA) is proposed and first formulated as a seamlessly unified objective optimization which guarantees convergence during the iteratively alternative solving process. The objective optimization is in both the ratio trace and the trace ratio forms, forming a complete framework of a new approach to jointly clustering and unsupervised subspace learning. The extension of LDA into Un-LDA enables to not only complete unsupervised subspace learning via the explicitly presented subspace projection matrix but also simultaneously finish clustering and even clustering out-of-sample data via the explicitly presented transformation matrix. To overcome the difficulty in solving the non-convex objective optimization, we mathematically prove that the Un-LDA optimization in both forms can be transformed into the simple K-means clustering optimization when the subspace is determined. The Un-LDA optimization is eventually completed by alternatively optimizing the clusters using K-means and the subspace using the supervised LDA methods and iterating this whole process until convergence or stopping criterion. The experiments demonstrate that our proposed Un-LDA algorithms are comparable or even much superior to the counterparts. Fei Wang 0008, Feiping Nie 0001, Zhongheng Li, Weizhong Yu, Rong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Multi-Scale Boosted Dehazing Network With Dense Feature FusionabstractIn this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strengthen-Operate-Subtract boosting strategy in the decoder of the proposed model, we develop a simple yet effective boosted decoder to progressively restore the haze-free image. To address the issue of preserving spatial information in the U-Net architecture, we design a dense feature fusion module using the back-projection feedback scheme. We show that the dense feature fusion module can simultaneously remedy the missing spatial information from high-resolution features and exploit the non-adjacent features. Extensive evaluations demonstrate that the proposed model performs favorably against the state-of-the-art approaches on the benchmark datasets as well as real-world hazy images. Hang Dong 0001, Jinshan Pan, Xinyi Zhang 0005, Fei Wang 0008, Ming-Hsuan Yang 0001 |
CVPR | 6 |
| 2020 | Deep Multi-Scale Gabor Wavelet Network for Image RestorationabstractDue to the limitations of the imaging processors and complex weather conditions, image degradation is often inevitable. Existing deep learning-based image restoration methods often rely on the powerful feature representation capacity of deep networks and pay less attention to the inherent properties of the degradation signal, e.g. variations in spatial scale and orientations across the image, which makes them ineffective for the image restoration tasks. In this paper, we propose a Multiscale Gabor Wavelet Network (MsGWN) for image restoration. We apply the multi-scale architecture to extract the contaminated feature from input at different spatial scales, and thus the contaminated feature can be effectively restored in a corse- to-fine manner. However, using multi-scale architecture alone cannot remove the degradations with different orientations. To overcome this problem, we introduce a Gabor Wavelet Module (GWM) to further extract the contaminated features from four orientations. By decomposing the features into four multi-orientation components, the restoration process can be facilitated by avoiding learning the mixed degradations all-in- one. We evaluate the proposed method on image demoirding, image deraining, and image dehazing. Experiments on these applications demonstrate that the proposed method can achieve favorable results against the state-of-the-art approaches. Hang Dong 0001, Xinyi Zhang 0005, Yu Guo 0006, Fei Wang 0008 |
ICASSP | 4 |
| 2020 | Recursive Maximum Correntropy Criterion Based Randomized Recurrent Broad Learning System
Yu Guo 0006, Fei Wang 0008 |
ICONIP (5) | 3 |
| 2020 | Detail Fusion GAN: High-Quality Translation for Unpaired Images with GAN-based Data AugmentationabstractImage-to-image translation, a task to learn the mapping relation between two different domains, is a rapid-growing research field in deep learning. Although existing Generative Adversarial Network (GAN)-based methods have achieved decent results in this field, there are still some limitations in generating high-quality images for practical applications (e.g., data augmentation and image inpainting). In this work, we aim to propose a GAN-based network for data augmentation which can generate translated images with more details and less artifacts. The proposed Detail Fusion Generative Adversarial Network (DFGAN) consists of a detail branch, a transfer branch, a filter module, and a reconstruction module. The detail branch is trained by a super-resolution loss and its intermediate features can be used to introduce more details to the transfer branch by the filter module. Extensive evaluations demonstrate that our model generates more satisfactory images against the state-of-the-art approaches for data augmentation. Yaochen Li, Hang Dong 0001, Peilin Jiang, Fei Wang 0008 |
ICPR | 6 |
| 2020 | Fair and cache blocking aware warp scheduling for concurrent kernel execution on GPU
Chen Zhao 0009, Wu Gao, Feiping Nie 0001, Fei Wang 0008, Huiyang Zhou |
Future Gener. Comput. Syst. | 4 |
| 2020 | Gated Fusion Network for Degraded Image Super Resolution
Xinyi Zhang 0005, Hang Dong 0001, Wei-Sheng Lai, Fei Wang 0008, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 5 |
| 2020 | A forest of trees with principal direction specified oblique split on random subspace
Fei Wang 0008, Feiping Nie 0001, Weizhong Yu, Rong Wang 0001, Zhongheng Li |
Neurocomputing | 1 |
| 2020 | A linear multivariate binary decision tree classifier based on K-means splitting
Fei Wang 0008, Feiping Nie 0001, Zhongheng Li, Weizhong Yu, Fuji Ren |
Pattern Recognit. | 1 |
| 2020 | Diffusion adaptation framework for compressive sensing reconstruction
Yicong He, Fei Wang 0008, Badong Chen |
Signal Process. | 2 |
| 2019 | On Boosting Single-Frame 3D Human Pose Estimation via Monocular VideosabstractThe premise of training an accurate 3D human pose estimation network is the possession of huge amount of richly annotated training data. Nonetheless, manually obtaining rich and accurate annotations is, even not impossible, tedious and slow. In this paper, we propose to exploit monocular videos to complement the training dataset for the single-image 3D human pose estimation tasks. At the beginning, a baseline model is trained with a small set of annotations. By fixing some reliable estimations produced by the resulting model, our method automatically collects the annotations across the entire video as solving the 3D trajectory completion problem. Then, the baseline model is further trained with the collected annotations to learn the new poses. We evaluate our method on the broadly-adopted Human3.6M and MPI-INF-3DHP datasets. As illustrated in experiments, given only a small set of annotations, our method successfully makes the model to learn new poses from unlabelled monocular videos, promoting the accuracies of the baseline model by about 10%. By contrast with previous approaches, our method does not rely on either multi-view imagery or any explicit 2D keypoint annotations. Zhi Li 0055, Xuan Wang 0009, Fei Wang 0008, Peilin Jiang |
ICCV | 3 |
| 2019 | Dice Loss in Siamese Network for Visual Object Tracking
Zhao Wei, Changhao Zhang, Kaiming Gu, Fei Wang 0008 |
ICIC (2) | 4 |
| 2019 | Gated Contiguous Memory U-Net for Single Image Dehazing
Hang Dong 0001, Fei Wang 0008, Yu Guo 0006, Kaisheng Ma |
ICONIP (2) | 3 |
| 2019 | Stacked Mixed-Scale Networks for Human Pose Estimation
Xuan Wang 0009, Zhi Li 0055, Peilin Jiang, Fei Wang 0008 |
PRICAI (1) | 5 |
| 2019 | Maximum correntropy adaptation approach for robust compressive sensing reconstruction
Yicong He, Fei Wang 0008, Jiuwen Cao, Badong Chen |
Inf. Sci. | 2 |
| 2018 | Gated Fusion Network for Joint Image Deblurring and Super-Resolution
Xinyi Zhang 0005, Hang Dong 0001, Wei-Sheng Lai, Fei Wang 0008, Ming-Hsuan Yang 0001 |
BMVC | 5 |
| 2018 | Online Multi-Object Tracking with Structural Invariance Constraint
Peilin Jiang, Zhao Wei, Hang Dong 0001, Fei Wang 0008 |
BMVC | 5 |
| 2018 | Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses
Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
ECCV (5) | 4 |
| 2018 | Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline StallsabstractFollowing the advances in technology scaling, graphics processing units (GPUs) incorporate an increasing amount of computing resources and it becomes difficult for a single GPU kernel to fully utilize the vast GPU resources. One solution to improve resource utilization is concurrent kernel execution (CKE). Early CKE mainly targets the leftover resources. However, it fails to optimize the resource utilization and does not provide fairness among concurrent kernels. Spatial multitasking assigns a subset of streaming multiprocessors (SMs) to each kernel. Although achieving better fairness, the resource underutilization within an SM is not addressed. Thus, intra-SM sharing has been proposed to issue thread blocks from different kernels to each SM. However, as shown in this study, the overall performance may be undermined in the intra-SM sharing schemes due to the severe interference among kernels. Specifically, as concurrent kernels share the memory subsystem, one kernel, even as computing-intensive, may starve from not being able to issue memory instructions in time. Besides, severe L1 D-cache thrashing and memory pipeline stalls caused by one kernel, especially a memory-intensive one, will impact other kernels, further hurting the overall performance. In this study, we investigate various approaches to overcome the aforementioned problems exposed in intra-SM sharing. We first highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose two approaches to reduce memory pipeline stalls. The first is to balance memory accesses of concurrent kernels. The second is to limit the number of inflight memory instructions issued from individual kernels. Our evaluation shows that the proposed schemes significantly improve the weighted speedup of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK, by 24.6% and 27.2% on average, respectively, with lightweight hardware overhead. Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou |
HPCA | 5 |
| 2018 | A Deep Encoder-Decoder Networks for Joint Deblurring and Super-ResolutionabstractIn this paper, we propose an end-to-end convolution neural network (CNN) to restore a clear high-resolution image from a severely blurry image. It's a highly ill-posed problem and brings tremendous challenges to state-of-art deblurring or super-resolution (SR) methods. A straightforward way to solve this problem is to concatenate two types of networks directly. However, experiments show that the concatenation of independent networks increases computation complexity instead of generating satisfying high-resolution images. Consequently, we focus on designing a single deep network to solve the deblurring and SR problems in parallel. Our method, called ED-DSRN, extends the traditional Super-Resolution network by adding a deblurring branch that shares the same feature maps extracted from an encoder-decoder module with the original SR branch. Extensive experiments show that our method produces remarkable deblurred and super-resolved images simultaneously with high efficiency. Xinyi Zhang 0005, Fei Wang 0008, Hang Dong 0001, Yu Guo 0006 |
ICASSP | 2 |
| 2018 | Generalized Maximum Correntropy-Based Echo State Network for Robust Nonlinear System IdentificationabstractIn this paper, we propose a robust method for non-linear system identification that incorporates robustness to echo state networks (ESNs). In particular, the ESNs utilize generalized correntropy as a loss function to get optimal solutions. Generalized correntropy is a more flexible extension of correntropy in information theoretic learning (ITL). Generalized correntropy induced metric (GCIM) is robust to outliers with a proper shape parameter. The ESNs with GCIM can provide the anti-noise capacity and are insensitive outliers which are prevalent in real-world tasks. They also inherit the basic architecture of echo state network but replaces the commonly used mean square error (MSE) criterion with GCIM. The stochastic gradient descent method is adopted to optimize the generalized correntropy-based cost function. Numerical simulations are given to show that the proposed algorithm is robust to the non-Gaussian noise and outliers. Changhao Zhang, Yu Guo 0006, Fei Wang 0008, Badong Chen |
IJCNN | 3 |
| 2018 | Efficient tree classifiers for large scale datasets
Fei Wang 0008, Feiping Nie 0001, Weizhong Yu, Rong Wang 0001 |
Neurocomputing | 1 |
| 2018 | Multi-view embedded clustering with unsupervised trace ratio LDA
Weizhong Yu, Rong Wang 0001, Feiping Nie 0001, Fei Wang 0008 |
Neurocomputing | 4 |
| 2018 | An improved locality preserving projection with ℓ1-norm minimization for dimensionality reduction
Weizhong Yu, Rong Wang 0001, Feiping Nie 0001, Fei Wang 0008 |
Neurocomputing | 4 |
| 2018 | Robust real-time visual object tracking via multi-scale fully convolutional Siamese networks
Longchao Yang, Peilin Jiang, Fei Wang 0008, Xuan Wang 0009 |
Multim. Tools Appl. | 3 |
| 2018 | Point-wise saliency detection on 3D point clouds via covariance descriptors
Yu Guo 0006, Fei Wang 0008, Jingmin Xin |
Vis. Comput. | 2 |
| 2017 | POSTER: Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline StallsabstractIn this study, we demonstrate that the performance may be undermined in the state-of-the-art intra-SM sharing schemes for concurrent kernel execution (CKE) on GPUs, due to the interference among concurrent kernels. We highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose to balance memory accesses and limit the number of inflight memory instructions issued from concurrent kernels to reduce memory pipeline stalls. Our proposed schemes significantly improve the performance of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK. Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou |
PACT | 5 |
| 2017 | A neural filter-based scheme for synchronizing chaotic systemsabstractSynchronization of chaotic systems and/or maps is a key step to implement secure communication schemes with chaos. If the process to synchronize chaotic systems is modeled stochastic, schemes based on extended Kalman filter (EKF) and unscented Kalman filter (UKF) have been studied in the past. However, such nonlinear filters are employed with assumptions of Gaussian noise processes and the Markov property. Further, EKF and UKF are suboptimal filtering methods, incurring unacceptable errors for high nonlinear systems. In this paper, neural filter (NF) is proposed for chaotic synchronization. This new approach requires no mentioned assumptions and achieves optimal filter. Numerical comparisons between the proposed approach and existing schemes are presented in this paper, showing the superiority of the proposed approach. Yu Guo 0006, Fei Wang 0008, James Ting-Ho Lo |
ICASSP | 2 |
| 2017 | Saliency-Guided Smoothing for 3D Point Clouds
Fei Wang 0008, Yu Guo 0006, Peilin Jiang |
ICIC (1) | 2 |
| 2017 | Recovering complex non-rigid 3D structures from monocular images by union of nonlinear subspacesabstractNon-rigid structure from motion (NRSfM) is a well-known challenging task due to its inherent ambiguities. Most existing approaches rely on kinds of low-rank linear subspaces assumption to make the problem well-constrained. In this paper, we make two contributions. First, we empirically present that relying on the assumption, 3D shapes lie on a union of non-linear subspaces, can better model the complex non-rigid motion than its linear counterparts. Second, we introduce the nonlinear low-rank representation as a regularizer to the objective function for NRSfM and show that it can be solved by alternating direction multiplier method (ADMM). Our experiments demonstrate that our method yields more accurate reconstruction and more reasonable clustering results than state-of-the-art methods, on CMU MoCap and UMPM datasets. Fei Wang 0008, Xuan Wang 0009 |
ICIP | 2 |
| 2017 | Region-based fully convolutional siamese networks for robust real-time visual trackingabstractPartial occlusions and deformations in visual object tracking are still very challenging. Existing Convolutional Neural Networks (CNNs) trackers either fail to handle these issues or can just run in low speed. In this paper, we present a real-time tracker which is robust to occlusions and deformations based on a Region-based, Fully Convolutional Siamese Network (R-FCSN). In the proposed R-FCSN, the information of regions is extracted separately by the proposition of position-sensitive score maps. Combining these score maps via adaptive weights leads to accurate location of the target on a new frame. The experiments illustrate that our method outperforms state-of-the-art approaches, and can handle the cases of object deformation and occlusion at about 51 FPS. Longchao Yang, Peilin Jiang, Fei Wang 0008, Xuan Wang 0009 |
ICIP | 3 |
| 2017 | Robust echo state networks based on correntropy induced loss function
Yu Guo 0006, Fei Wang 0008, Badong Chen, Jingmin Xin |
Neurocomputing | 2 |
| 2017 | Maximum total correntropy adaptive filtering against heavy-tailed noises
Fei Wang 0008, Yicong He, Badong Chen |
Signal Process. | 1 |
| 2017 | A Novel Method of Minimizing View Synthesis Distortion Based on Its Non-Monotonicity in 3D VideoabstractIn depth-based 3D video, the view synthesis distortion (VSD), is generally measured by modeling the effect of texture and depth errors separately. With such a development, it has been referred that the VSD changes monotonically with respect to to both the texture and depth distortions. In this paper, we find that the VSD does not always change monotonically with them by both theoretical analysis and experimental test, when the effect of the texture and depth errors is considered together. Specifically, first, we prove that the VSD is non-monotonic with the texture distortion. That is, the VSD increases with the increasing texture distortion at higher distortion range but conversely decreases with it at lower range. It is different from the general scenario that only considering the effect of the texture errors. We also analytically depict their relationship with low computational cost and identify the turning point at which the change of the VSD is converted. Second, we confirm that the VSD is always monotonic with the depth distortion, which is consistent with the general scenario that only considering the effect of the depth errors. The non-monotonicity property of the VSD can be utilized to improve the viewing performance of 3D video in relevant applications, since a minimal value of the VSD exists at the turning point. We conduct two applications for this purpose. First, it is used to generate the synthesis view of minimal distortion, which achieves 0.51-dB gain of PSNR on average for the tested scenarios. Second, it is used for lossy compression of texture videos in 3D video, which reduces the coding rate by 24% on average for the tested scenarios, meanwhile, keeps the VSD not increased simultaneously. Meng Yang 0002, Nanning Zheng 0001, Ce Zhu, Fei Wang 0008 |
IEEE Trans. Image Process. | 4 |
| 2016 | Template-Free 3D Reconstruction of Poorly-Textured Nonrigid Surfaces
Xuan Wang 0009, Mathieu Salzmann, Fei Wang 0008, Jizhong Zhao |
ECCV (7) | 3 |
| 2016 | A Novel Feature Point Detection Algorithm of Unstructured 3D Point Cloud
Bei Tian, Peilin Jiang, Xuetao Zhang 0001, Yulong Zhang 0003, Fei Wang 0008 |
ICIC (3) | 5 |
| 2016 | Natural Scene Digit Classification Using Convolutional Neural Networks
Ziqin Wang, Peilin Jiang, Xuetao Zhang 0001, Fei Wang 0008 |
ICIC (2) | 4 |
| 2016 | The Measurement of Human Height Based on Coordinate Transformation
Peilin Jiang, Xuetao Zhang 0001, Bin Zhang 0022, Fei Wang 0008 |
ICIC (3) | 5 |
| 2016 | Selectively GPU Cache Bypassing for Un-Coalesced LoadsabstractGPUs are widely used to accelerate general purpose applications, and could hide memory latency through massive multithreading. But multithreading can increase contention for the L1 data caches (L1D). This problem is exacerbated when an application contains irregular memory references which would lead to un-coalesced memory accesses. In this paper, we propose a simple yet effective GPU cache Bypassing scheme for Un-Coalesced Loads (BUCL). BUCL makes bypassing decisions at two granularities. At the instruction-level, when the number of memory accesses generated by a non-coalesced load instruction is bigger than a threshold, referred as the threshold of un-coalescing degree (TUCD), all the accesses generated from this load will bypass L1D. The reason is that the cache data filled by un-coalesced loads typically have low probabilities to be reused. At the level of each individual memory access, when the L1D is stalled, the accessed data is likely with low locality, and the utilization of the target memory sub-partition is not high, this memory access may also bypass L1D. Our experiments show that BUCL achieves 36% and 5% performance improvement over the baseline GPU for memory un-coalesced and memory coherent benchmarks, respectively, and also significantly outperforms prior GPU cache bypassing and warp throttling schemes. Chen Zhao 0009, Fei Wang 0008, Huiyang Zhou, Nanning Zheng 0001 |
ICPADS | 2 |
| 2016 | Kernel adaptive filtering under generalized Maximum Correntropy CriterionabstractOwing to their universal approximation capability and online learning manner, kernel adaptive filters have been widely used in nonlinear systems modeling. Under Gaussian assumption, traditional kernel adaptive algorithms utilize the well-known mean square error(MSE) as a cost function to get optimal solutions. For non-Gaussian situations, MSE will not properly represent the statistics of the error, and hence degrade the performance. In recent years, an information theoretic learning(ITL) based criterion called Maximum Correntropy Criterion(MCC) has been proposed and applied in robust adaptive filtering. The correntropy is a generalized correlation measure in kernel space, which uses Gaussian kernel as a default kernel function. Of course, Gaussian kernel is not always the best choice. Recently, a more flexible definition of correntropy, called generalized correntropy, has been proposed. With a proper shape parameter, the generalized correntropy may get better performance than original correntropy with Gaussian kernel. In this paper, we take advantages of both kernel methods and generalized correntropy to develop a new kernel adaptive algorithm called Generalized Kernel Maximum Correntropy(GKMC) algorithm. We analyze theoretically the stability and steady-state performance of the new algorithm. In addition, we propose a Quantized GKMC(QGKMC) algorithm to curb the growth of the network size in GKMC while maintaining the performance. Simulation results confirm the theoretical expectations and show superior performance compared with existing methods. Yicong He, Fei Wang 0008, Jing Yang 0014, Hai-Jun Rong, Badong Chen |
IJCNN | 2 |
| 2015 | Pose Estimation for Vehicles Based on Binocular Stereo Vision in Urban Traffic
Fei Wang 0008, Yicong He, Hang Dong 0001, Haiwei Yang, Yang Yang 0066 |
ICIC (1) | 2 |
| 2014 | Error Tracing and Analysis of Vision Measurement System
Fei Wang 0008, Haiwei Yang, Yongjian He |
ICIC (1) | 2 |
| 2014 | Monocular 3D Shape Recovery of Inextensibility Deformable Surface by Using DE-Based Niching Algorithm with Partial Reinitialization
Xuan Wang 0009, Fei Wang 0008 |
ICIC (2) | 2 |
| 2014 | Clustering-Based Latent Variable Models for Monocular Non-rigid 3D Shape Recovery
Fei Wang 0008, Xuan Wang 0009 |
ICIC (2) | 2 |
| 2014 | Robust Pose Estimation Algorithm for Approximate Coplanar Targets
Haiwei Yang, Fei Wang 0008, Yicong He, Yongjian He |
ICIC (2) | 2 |
| 2014 | Overtaking vehicle detection using a spatio-temporal CRFabstractOvertaking vehicle detection is vital for road safety, as the dangerous behavior of that vehicle may affect the safety of ego-vehicle and the time is not enough for the driver to attend and react. Therefore, it is one of the key components of the Advanced Driver Assistance Systems. Mostly, traditional methods only use local information, appearance or motion. In this paper, we build a novel CRF model to make use of the interaction between local regions, and the motion features from multiple scales as well. The whole model is based on the low-level optical flows. In order to increase the robustness to the noise in the flow, we divided the motion field into small blocks, and learned Mixture of Probabilistic Principle Analysis models for the common motion patterns of the background. Moreover, we also adopted an online scheme for updating the parameters. Results of testing on the real road images demonstrated the capability of the proposed algorithm. Xuetao Zhang 0001, Peilin Jiang, Fei Wang 0008 |
Intelligent Vehicles Symposium | 3 |
| 2012 | Emotion Ontology Construction from Chinese Knowledge
Peilin Jiang, Fei Wang 0008, Fuji Ren, Nanning Zheng 0001 |
CICLing (1) | 2 |
| 2011 | Linear Pose Estimation Algorithm Based on Quaternion
Yongjian He, Caigui Jiang, Chengwei Hu, Jingmin Xin, Fei Wang 0008 |
ICIC (1) | 6 |
| 2011 | Research on Dynamic Human Object Tracking Algorithm
Yongjian He, Shoupeng Feng, Rongkun Zhou, Yonghua Xing, Fei Wang 0008 |
ICIC (1) | 6 |
| 2011 | Part-based on-road vehicle detection using hidden random field
Xuetao Zhang 0001, Yongjian He, Fei Wang 0008 |
Sci. China Inf. Sci. | 3 |
| 2007 | Automatic Region-of-Interest Coding in JPEG2000 Based on Morphology Segmentation and LLn Subband Analysis
Fei Wang 0008, Nanning Zheng 0001, Yuehu Liu |
ICIC (3) | 1 |