EDBT 2026 Demo / reviewers in the wild / expert
Fei Wen 0005
dblp:21/8350-5
· DBLP profile ↗
43ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0002-3083-9611ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 14 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Fully Test Time Adaptation via Information Preserving Source Training
Zecheng Wang, Fei Wen 0005 |
ICIC (8) | 4 |
| 2026 | NeuroUNI: A Unified Event-Driven Multi-Core Architecture Optimizing Neuromorphic Primitives for Brain-Inspired ComputingabstractThe hardware convergence of Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) is hindered by conflicting computational paradigms: dense tensor parallelism versus asynchronous sparse dynamics. Existing unifications typically rely on inefficient spatial partitioning or mode-reconfigurable datapaths, limiting the flexibility needed by heterogeneous ANN-SNN hybrid models requiring frequent cross-domain interaction. To resolve this, we present NeuroUNI, a unified event-driven multi-core architecture. Unlike partitioned designs, NeuroUNI unifies computation at the primitive level using a novel Five-Tuple Event Model, abstracting both continuous activations and discrete spikes to naturally leverage dynamic sparsity. The architecture features a co-optimized hierarchical Macro-Micro-$\mu $OP ISA, a superscalar SIMD-based microarchitecture, and a decentralized multi-core synchronization protocol. Validated in TSMC 28nm technology via post-synthesis simulation and on a Xilinx VCU129 FPGA prototype, NeuroUNI demonstrates competitive cross-paradigm efficiency. It achieves$35.0\times $and$1.21\times $the ANN energy efficiency of the NVIDIA V100 and EyerissV2, respectively, while delivering$5.7\times $the SNN throughput of TrueNorth. In a unified mapless navigation workload, NeuroUNI attains 422.6 GOPS/W (ANN) and 190.8GSOPS/W (SNN), outperforming Loihi 1 with$56.7\times $the throughput and$3.85\times $the energy efficiency, proving the viability of a primitive/ISA-level unified silicon substrate. Faquan Chen, Qingyang Tian, Ziren Wu, Xiangcheng Shi, Rendong Ying, Fei Wen 0005 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-PropagationabstractFine-tuning pre-trained speech foundation models for Automatic Speech Recognition (ASR) is prevalent, yet constrained by substantial GPU memory requirements. We introduce ZO-ASR, a memory-efficient Zeroth-Order (ZO) method that avoids Back-Propagation (BP) and activation memory by estimating gradients via forward passes. When combined with SGD optimizer, ZO-ASR-SGD fine-tunes ASR models using only inference memory. Our evaluation spans supervised and unsupervised tasks. For Supervised Domain Adaptation on Whisper-Large-V3, ZO-ASR’s multiple query mechanism enhances robustness and achieves up to an 18.9% relative Word Error Rate reduction over zero-shot baselines, outperforming existing ZO methods. For unsupervised Test-Time Adaptation on Wav2Vec2-Base, ZO-ASR exhibits moderately lower performance compared to first-order optimizer Adam. Our BP-free approach provides a viable solution for fine-tuning ASR models in computationally resource-constrained or gradient-inaccessible scenarios. Yuezhang Peng, Fei Wen 0005, Xie Chen 0001 |
ASRU | 5 |
| 2025 | MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsabstractFine-tuning pre-trained large language models (LLMs) on downstream tasks has achieved significant success across various domains.However, as model sizes grow, traditional firstorder fine-tuning algorithms incur substantial memory overhead due to the need for activation storage for back-propagation (BP).The BP-free Memory-Efficient Zeroth-Order Optimization (MeZO) method estimates gradients through finite differences, avoiding the storage of activation values, and has been demonstrated as a viable approach for fine-tuning large language models.This work proposes the MUltiple-query Memory Efficient Zeroth-Order (MUZO) method, which is based on variance-reduced multiple queries to obtain the average of gradient estimates.When combined with Adam optimizer, MUZO-Adam demonstrates superior performance in fine-tuning various LLMs.Furthermore, we provide theoretical guarantees for the convergence of the MUZO-Adam optimizer.Extensive experiments empirically demonstrate that MUZO-Adam converges better than MeZO-SGD and achieves near first-order optimizer performance on downstream classification, multiple-choice, and generation tasks. Yuezhang Peng, Fei Wen 0005, Xie Chen 0001 |
EMNLP | 3 |
| 2025 | Explicit Mutual Information Maximization for Self-Supervised LearningabstractRecently, self-supervised learning (SSL) has been extensively studied. Theoretically, mutual information maximization (MIM) is an optimal criterion for SSL, with a strong theoretical foundation in information theory. However, it is difficult to directly apply MIM in SSL since the data distribution is not analytically available in applications. In practice, many existing methods can be viewed as approximate implementations of the MIM criterion. This work shows that, based on the invariance property of MI, explicit MI maximization can be applied to SSL under a generic distribution assumption, i.e., a relaxed condition of the data distribution. We further illustrate this by analyzing the generalized Gaussian distribution. Based on this result, we derive a loss function based on the MIM criterion using only second-order statistics. We implement the new loss for SSL and demonstrate its effectiveness via extensive experiments. Lele Chang, Qinghai Guo, Fei Wen 0005 |
ICASSP | 4 |
| 2025 | A Skeleton-Based Topological Planner for Exploration in Complex Unknown EnvironmentsabstractThe capability of autonomous exploration in complex, unknown environments is important in many robotic applications. While recent research on autonomous exploration have achieved much progress, there are still limitations, e.g., existing methods relying on greedy heuristics or optimal path planning are often hindered by repetitive paths and high computational demands. To address such limitations, we propose a novel exploration framework that utilizes the global topology information of observed environment to improve exploration efficiency while reducing computational overhead. Specifically, global information is utilized based on a skeletal topological graph representation of the environment geometry. We first propose an incremental skeleton extraction method based on wavefront propagation, based on which we then design an approach to generate a lightweight topological graph that can effectively capture the environment's structural characteristics. Building upon this, we introduce a finite state machine that leverages the topological structure to efficiently plan coverage paths, which can substantially mitigate the back-and-forth maneuvers (BFMs) problem. Experimental results demonstrate the superiority of our method in comparison with state-of-theart methods. The source code will be made publicly available at: https://github.com/Haochen-Niu/STGPlanner. Haochen Niu, Xingwu Ji, Lantao Zhang, Fei Wen 0005, Rendong Ying |
ICRA | 4 |
| 2025 | Self-supervised End-to-end ToF Imaging Based on RGB-D Cross-modal DependencyabstractTime-of-Flight (ToF) imaging systems are susceptible to various noise and degradation, which can severely affect image quality. Traditional sequential imaging pipelines often suffer from error accumulation due to separate multi-stage processing. Existing end-to-end methods typically rely on noisy-clean depth image pairs for supervised learning. However, acquiring ground-truth is challenging in real-world scenarios due to factors such as Multi-Path Interference (MPI), phase wrapping, and complex noise patterns. In this paper, we propose a self-supervised learning framework for end-to-end ToF imaging, which does not require any noisy-clean pairs yet generalizes well across various off-the-shelf cameras. Our framework leverages the cross-modal dependencies between RGB and depth data as implicit supervision to effectively suppress noise and maintain image fidelity. Additionally, the loss function integrates the statistical characteristics of raw measurement data, enhancing robustness against noise and artifacts. Extensive experiments on both synthetic and real-world data demonstrate that our approach achieves performance comparable to supervised methods, without requiring paired noisy-clean data for training. Furthermore, our method consistently delivers strong performance across all evaluated cameras, highlighting its generalization capabilities. The code is available at https://github.com/WeihangWANG/RGBD_imaging. Weihang Wang 0012, Jun Wang 0137, Fei Wen 0005 |
IJCAI | 3 |
| 2025 | Probing Implicit Bias in Semi-gradient Q-learning: Visualizing the Non-equilibrium Loss Landscapes via the Fokker-Planck EquationabstractSemi-gradient Q-learning is widely applied across various fields; however, due to the absence of an explicit loss function, understanding its preference for critical points in the parameter space—namely, the implicit bias—remains challenging. This paper leverages the Fokker–Planck equation to construct and visualize the non-equilibrium loss landscape in a two-dimensional parameter space. Our visualization reveals that the semi-gradient method transforms stable critical points into unstable ones, resulting in training dynamics biased against unstable critical points. Furthermore, we show that this phenomenon extends to high-dimensional parameter spaces and neural network settings, such as DQN. This paper provides valuable insights into the implicit bias of semi-gradient Q-learning. Shuyu Yin, Fei Wen 0005, Tao Luo 0012 |
IJCNN | 2 |
| 2025 | Lifelong Test-Time Adaptation via Online Learning in Tracked Low-Dimensional SubspaceabstractTest-time adaptation (TTA) aims to adapt a source model to a target domain using only test data. Existing methods predominantly rely on unsupervised entropy minimization or its variants, which suffer from degeneration, leading to trivial solutions with low-entropy but inaccurate predictions. In this work, we identify *entropy-deceptive* (ED) samples, instances where the model makes highly confident yet incorrect predictions, as the underlying cause of degeneration. Further, we reveal that the gradients of entropy minimization in TTA have an intrinsic low-dimensional structure, driven primarily by *entropy-truthful* (ET) samples whose gradients are highly correlated. In contrast, ED samples have scattered, less correlated gradients. Leveraging this observation, we show that the detrimental impact of ED samples can be suppressed by constraining model updates within the principal subspace of backward gradients. Building on this insight, we propose LCoTTA, a lifelong continual TTA method that tracks the principal subspace of gradients online and utilizes their projections onto this subspace for adaptation. Further, we provide theoretical analysis to show that the proposed subspace-based method can enhance the robustness against detrimental ED samples. Extensive experiments demonstrate that LCoTTA effectively overcomes degeneration and significantly outperforms existing methods in long-term continual adaptation scenarios. Code is available online. Dexin Duan, Fei Wen 0005 |
NeurIPS | 4 |
| 2025 | Self-ensemble for test time adaptation
Liujia Ma, Mingyue Qin, Fei Wen 0005 |
Neurocomputing | 5 |
| 2025 | MOS-GAN: Mean Opinion Score GAN for Unsupervised Speech Enhancement
Wenbin Jiang 0003, Fei Wen 0005, Kai Yu 0004 |
IEEE Signal Process. Lett. | 2 |
| 2025 | A Learning Framework for Perceptual Lossy Compression With Stochastic CodingabstractRecent studies in perceptual lossy compression have highlighted the advantage of stochastic coding with shared randomness between encoder and decoder, over deterministic encoding in the regime of “high perceptual quality”. While the theoretical benefits of stochastic coding have been well-established and demonstrated by analytic examples, its practical realization remains challenging and largely unexplored. In this work, we propose a practical learning framework for stochastic coding that effectively realizes its theoretical advantages. Starting with a theoretically optimal scheme, we develop an implementation closely approximates it through a two-stage training process: learning a stochastic encoder, followed by a stochastic decoder which is modeled as an optimal transport problem conditioned on minimum mean square error (MMSE) decoding. Additionally, for training stochastic coding models, we prove the equivalence between quantized representation and “noisy” representation. Based on this insight, we introduce a quantization-free training method that effectively addresses the non-differentiability challenge posed by quantization. Experiments on a circular distribution example and the MNIST dataset validate our findings and demonstrate the effectiveness of the proposed method. Fei Wen 0005 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Brain-Inspired Online Adaptation for Remote Sensing With Spiking Neural NetworkabstractOn-device computing, or edge computing, is becoming increasingly important for remote sensing, particularly in applications like deep network-based perception on on-orbit satellites and unmanned aerial vehicles (UAVs). In these scenarios, two brain-like capabilities are crucial for remote sensing models: (1)high energy efficiency, allowing the model to operate on edge devices with limited computing resources, and (2)online adaptation, enabling the model to quickly adapt to environmental variations, weather changes, and sensor drift. This work addresses these needs by proposing an online adaptation framework based on spiking neural networks (SNNs) for remote sensing. Starting with a pretrained SNN model, we design an efficient, unsupervised online adaptation algorithm, which adopts an approximation of the BPTT algorithm and only involves forward-in-time computation that significantly reduces the computational complexity of SNN adaptation learning. Besides, we propose an adaptive activation scaling scheme to boost online SNN adaptation performance, particularly in low time-steps. Furthermore, for the more challenging remote sensing detection task, we propose a confidence-based instance weighting scheme, which substantially improves adaptation performance in the detection task. To our knowledge, this work is the first to address the online adaptation of SNNs. Extensive experiments on seven benchmark datasets across classification, segmentation, and detection tasks demonstrate that our proposed method significantly outperforms existing domain adaptation and domain generalization approaches under varying weather conditions. The proposed method enables energy-efficient and fast online adaptation on edge devices, and has much potential in applications such as remote perception on on-orbit satellites and UAV. Code is available at https://github.com/ThunderDavid/OASNN. Dexin Duan, BingWei Hui, Fei Wen 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Unsupervised Speech Enhancement Using Optimal Transport and Speech Presence ProbabilityabstractSpeech enhancement models based on deep learning are typically trained in a supervised manner, requiring a substantial amount of paired noisy-to-clean speech data for training. However, synthetically generated training data can only capture a limited range of realistic environments, and it is often challenging or even impractical to gather real-world pairs of noisy and ground-truth clean speech. To overcome this limitation, we propose an unsupervised learning approach for speech enhancement that eliminates the need for paired noisy-to-clean training data. Specifically, our method utilizes the optimal transport criterion to train the speech enhancement model in an unsupervised manner. It employs a fidelity loss based on noisy speech and a distribution divergence loss to minimize the difference between the distribution of the model's output and that of unpaired clean speech. Further, we use the speech presence probability as an additional optimization objective and incorporate the short-time Fourier transform (STFT) domain loss as an extra term for the unsupervised learning loss. We also apply the multi-resolution STFT loss as the validation loss to enhance the stability of the training process and improve the algorithm's performance. Experimental results on the VCTK + DEMAND benchmark demonstrate that the proposed method achieves competitive performance compared to the supervised methods. Furthermore, the speech recognition results on the CHiME4 benchmark show the superiority of the proposed method over its supervised counterpart. Wenbin Jiang 0003, Kai Yu 0004, Fei Wen 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | ParallelNN: A Parallel Octree-based Nearest Neighbor Search Accelerator for 3D Point CloudsabstractAs Light Detection And Ranging (LiDAR) increasingly becomes an essential component in robotic navigation and autonomous driving, the processing of high throughput 3D point clouds in real time is widely required. This work considers the point cloud k-Nearest Neighbor (kNN) search, which is an important 3D processing kernel. Although applying fine-grained parallelism optimization on internal processing, e.g., using multiple workers, has demonstrated high efficiency, previous accelerators with DDR external memory are fundamentally limited by the external bandwidth bottleneck. To break this bottleneck, this work proposes a highly parallel architecture, namely ParallelNN, for highly efficient kNN search processing of high throughput point clouds. First, we optimize the multichannel cache based on High Bandwidth Memory (HBM) and on-chip memory to provide large external bandwidth. Then, a novel parallel depth-first octree construction algorithm is proposed and mapped onto multiple construction branches with trace-coded construction queues, which can regularize random accesses and perform multi-branch octree construction efficiently. Furthermore, in the search stage, we present algorithm-architecture co-optimization, including parallel keyframe-based scheduling and multi-branch flexible search engines, to provide conflict-free access and maximum reuse opportunities for reference points, which achieves more than 27.0× speedup compared with baseline architectures. We prototype ParallelNN on Virtex HBM FPGA and perform extensive benchmarking on the KITTI dataset. The results demonstrate that ParallelNN achieves up to 107.7× and 12.1× speedup over CPU and GPU implementations, while being more energy efficient, e.g., outperforming CPU and GPU implementations by 73.6× and 31.1×, respectively. Besides, with the proposed algorithm-architecture co-optimization, ParallelNN achieves 11.4× speedup over state-of-the-art architecture. Moreover, ParallelNN is configurable and can be easily generalized to similar octree-based applications. Faquan Chen, Rendong Ying, Jianwei Xue, Fei Wen 0005 |
HPCA | 4 |
| 2023 | UnSE: Unsupervised Speech Enhancement Using Optimal Transport
Wenbin Jiang 0003, Fei Wen 0005, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2023 | SPAR: An efficient self-attention network using Switching Partition Strategy for skeleton-based action recognition
Rendong Ying, Fei Wen 0005 |
Neurocomputing | 3 |
| 2023 | Optimal Transport for Unsupervised Denoising LearningabstractRecently, much progress has been made in unsupervised denoising learning. However, existing methods more or less rely on some assumptions on the signal and/or degradation model, which limits their practical performance. How to construct an optimal criterion for unsupervised denoising learning without any prior knowledge on the degradation model is still an open question. Toward answering this question, this work proposes a criterion for unsupervised denoising learning based on the optimal transport theory. This criterion has favorable properties, e.g., approximately maximal preservation of the information of the signal, whilst achieving perceptual reconstruction. Furthermore, though a relaxed unconstrained formulation is used in practical implementation, we prove that the relaxed formulation in theory has the same solution as the original constrained formulation. Experiments on synthetic and real-world data, including realistic photographic, microscopy, depth, and raw depth images, demonstrate that the proposed method even compares favorably with supervised methods, e.g., approaching the PSNR of supervised methods while having better perceptual quality. Particularly, for spatially correlated noise and realistic microscopy images, the proposed method not only achieves better perceptual quality but also has higher PSNR than supervised methods. Besides, it shows remarkable superiority in harsh practical conditions with complex noise, e.g., raw depth images. Code is available at https://github.com/wangweiSJTU/OTUR. Fei Wen 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Domain Generalization for Face Anti-Spoofing via Negative Data AugmentationabstractIn practical applications, the generalization capability of face anti-spoofing (FAS) models on unseen domains is of paramount importance to adapt to diverse camera sensors, device drift, environmental variation, and unpredictable attack types. Recently, various domain generalization (DG) methods have been developed to improve the generalization capability of FAS models via training on multiple source domains. These DG methods commonly require collecting sufficient real-world attack samples of different attack types for each source domain. This work aims to learn a FAS model without using any real-world attack sample in any source domain but can generalize well to the unseen domain, which can significantly reduce the learning cost. Toward this goal, we draw inspiration from the theoretical error bound of domain generalization to use negative data augmentation instead of real-world attack samples for training. We show that using only a few types of simple synthesized negative samples, e.g., color jitter and color mask, the learned model can achieve competitive performance over state-of-the-art DG methods trained using real-world attack samples. Moreover, a dynamic global common loss and a local contrast loss are proposed to prompt the model to learn a compact and common feature representation for real face samples from different source domains, which can further improve the generalization capability. Experimental results of extensive cross-dataset testing demonstrate that our method can even outperform state-of-the-art DG methods using real-world attack samples for training. The code for reproducing the results of our method is available at https://github.com/WeihangWANG/NDA-FAS. Weihang Wang 0003, Rendong Ying, Fei Wen 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Self-Supervised Learning for RGB-Guided Depth Enhancement by Exploiting the Dependency Between RGB and DepthabstractDue to the imaging mechanism of time-of-flight (ToF) sensors, the captured depth images usually suffer from severe noise and degradation. Though many RGB-guided methods have been proposed for depth image enhancement in the past few years, yet the enhancement performance on real-world depth images is still largely unsatisfactory. Two main reasons are the complexity of realistic noise and degradation in depth images, and the difficulty in collecting noise-clean pairs for supervised enhancement learning. This work aims to develop a self-supervised learning method for RGB-guided depth image enhancement, which does not require any noisy-clean pairs but can significantly boost the enhancement performance on real-world noisy depth images. To this end, we exploit the dependency between RGB and depth images to self-supervise the learning of the enhancement model. It is achieved by maximizing the cross-modal dependency between RGB and depth to promote the enhanced depth having dependency with the RGB of the same scene as much as possible. Furthermore, we augment the cross-modal dependency maximization formulation based on the optimal transport theory to achieve further performance improvement. Experimental results on both synthetic and real-world data demonstrate that our method can significantly outperform existing state-of-the-art methods on depth denoising, multi-path interference suppression, and hole filling. Particularly, our method shows remarkable superiority over existing ones on real-world data in handling various realistic complex degradation. Code is available at https://github.com/wjcyt/SRDE. Jun Wang 0137, Fei Wen 0005 |
IEEE Trans. Image Process. | 3 |
| 2022 | Speech Enhancement with Neural Homomorphic SynthesisabstractMost deep learning-based speech enhancement methods operate directly on time-frequency representations or learned features without making use of the model of speech production. This work proposes a new speech enhancement method based on neural homomorphic synthesis. The speech signal is firstly decomposed into excitation and vocal tract with complex cepstrum analysis. Then, two complex-valued neural networks are applied to estimate the target complex spectrum of the decomposed components. Finally, the time-domain speech signal is synthesized from the estimated excitation and vocal tract. Furthermore, we investigated numerous loss functions and found that the multi-resolution STFT loss, commonly used in the TTS vocoder, benefits speech enhancement. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art complex-valued neural network-based methods in terms of both PESQ and eSTOI. Wenbin Jiang 0003, Kai Yu 0004, Fei Wen 0005 |
ICASSP | 4 |
| 2022 | Optimally Controllable Perceptual Lossy CompressionabstractRecent studies in lossy compression show that distortion and perceptual quality are at odds with each other, which put forward the tradeoff between distortion and perception (D-P). Intuitively, to attain different perceptual quality, different decoders have to be trained. In this paper, we present a nontrivial finding that only two decoders are sufficient for optimally achieving arbitrary (an infinite number of different) D-P tradeoff. We prove that arbitrary points of the D-P tradeoff bound can be achieved by a simple linear interpolation between the outputs of a minimum MSE decoder and a specifically constructed perfect perceptual decoder. Meanwhile, the perceptual quality (in terms of the squared Wasserstein-2 distance metric) can be quantitatively controlled by the interpolation factor. Furthermore, to construct a perfect perceptual decoder, we propose two theoretically optimal training frameworks. The new frameworks are different from the distortion-plus-adversarial loss based heuristic framework widely used in existing methods, which are not only theoretically optimal but also can yield state-of-the-art performance in practical perceptual decoding. Finally, we validate our theoretical finding and demonstrate the superiority of our frameworks via experiments. Code is available at: https://github.com/ZeyuYan/Controllable-Perceptual-Compression Fei Wen 0005 |
ICML | 2 |
| 2022 | A Complementary Fusion Strategy for RGB-D Face Recognition
Weihang Wang 0003, Fei Wen 0005 |
MMM (1) | 3 |
| 2022 | On optimality of multidimensional scaling for time differences of arrival/frequency differences of arrival based moving source localisationabstractAbstract Multidimensional scaling (MDS) is an attractive technique for a moving source localisation from time and frequency difference of arrival (time differences of arrival (TDOA)/frequency differences of arrival (FDOA)) measurements. However, its optimality has not yet been proven theoretically because of the difficult Moore–Penrose pesudo‐inverse operation. In addition to the theoretical incompleteness of the MDS technique, the sensor uncertainties are not considered in the MDS framework for the moving source localisation either. A closed‐form estimator is proposed for the TDOA/FDOA‐based localisation with senor uncertainties by exploiting the MDS technique. Furthermore, based on the fundamental corollaries in the MDS analysis, an elegant and detailed analytical proof is also presented for the optimality of the MDS estimator thoroughly in the presence and absence of sensor uncertainties. The theoretical derivation is corroborated by numerical examples. He-Wen Wei, Fei Wen 0005 |
IET Signal Process. | 2 |
| 2022 | Conv-MLP: A Convolution and MLP Mixed Model for Multimodal Face Anti-SpoofingabstractLocal features contain crucial clues for face anti-spoofing. Convolutional neural networks (CNNs) are powerful in extracting local features, but the intrinsic inductive bias of CNNs limits the ability to capture long-range dependencies. This paper aims to develop a simple yet effective framework that is versatile in extracting both local information and long-range dependencies for face anti-spoofing. To this end, we propose a novel architecture, namely Conv-MLP, which incorporates local patch convolution with global multi-layer perceptrons (MLP). Conv-MLP breaks the inductive bias limitation of traditional full CNNs and can be expected to better exploit long-range dependencies. Furthermore, we design a new loss specifically for the face anti-spoofing task, namely moat loss. The moat loss benefits discriminative representations learning and can improve the generalization capability on unseen presentation attacks. In this work, multi-modal data are directly fused at the signal level to extract complementary features. Extensive experiments on single and multi-modal datasets demonstrate that Conv-MLP outperforms existing state-of-the-art methods while being more computationally efficient. The code is available at https://github.com/WeihangWANG/Conv-MLP. Weihang Wang 0003, Fei Wen 0005, Rendong Ying |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Cross-modal Graph Matching Network for Image-text RetrievalabstractImage-text retrieval is a fundamental cross-modal task whose main idea is to learn image-text matching. Generally, according to whether there exist interactions during the retrieval process, existing image-text retrieval methods can be classified into independent representation matching methods and cross-interaction matching methods. The independent representation matching methods generate the embeddings of images and sentences independently and thus are convenient for retrieval with hand-crafted matching measures (e.g., cosine or Euclidean distance). As to the cross-interaction matching methods, they achieve improvement by introducing the interaction-based networks for inter-relation reasoning, yet suffer the low retrieval efficiency. This article aims to develop a method that takes the advantages of cross-modal inter-relation reasoning of cross-interaction methods while being as efficient as the independent methods. To this end, we propose a graph-based Cross-modal Graph Matching Network (CGMN) , which explores both intra- and inter-relations without introducing network interaction. In CGMN, graphs are used for both visual and textual representation to achieve intra-relation reasoning across regions and words, respectively. Furthermore, we propose a novel graph node matching loss to learn fine-grained cross-modal correspondence and to achieve inter-relation reasoning. Experiments on benchmark datasets MS-COCO, Flickr8K, and Flickr30K show that CGMN outperforms state-of-the-art methods in image retrieval. Moreover, CGMM is much more efficient than state-of-the-art methods using interactive matching. The code is available at https://github.com/cyh-sj/CGMN . Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | R-SDSO: Robust stereo direct sparse odometry
Ruihang Miao, Fei Wen 0005, Wuyang Xue, Rendong Ying |
Vis. Comput. | 3 |
| 2021 | On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training FrameworkabstractLossy compression algorithms are typically designed to achieve the lowest possible distortion at a given bit rate. However, recent studies show that pursuing high perceptual quality would lead to increase of the lowest achievable distortion (e.g., MSE). This paper provides nontrivial results theoretically revealing that, 1) the cost of achieving perfect perception quality is exactly a doubling of the lowest achievable MSE distortion, 2) an optimal encoder for the “classic” rate-distortion problem is also optimal for the perceptual compression problem, 3) distortion loss is unnecessary for training a perceptual decoder. Further, we propose a novel training framework to achieve the lowest MSE distortion under perfect perception constraint at a given bit rate. This framework uses a GAN with discriminator conditioned on an MSE-optimized encoder, which is superior over the traditional framework using distortion plus adversarial loss. Experiments are provided to verify the theoretical finding and demonstrate the superiority of the proposed training framework. Fei Wen 0005, Rendong Ying |
ICML | 2 |
| 2021 | Fast and Positive Definite Estimation of Large Covariance Matrix for High-Dimensional Data AnalysisabstractLarge covariance matrix estimation is a fundamental problem in many high-dimensional statistical analysis applications arises in economics and finance, bioinformatics, social networks, and climate studies. To achieve reliable estimation in the high-dimensional setting, an effective technique is to exploit the intrinsic structure of the covariance matrix, e.g., by sparsity regularization. For sparsity regularization, the lasso penalty is popular and convenient due to its convexity but has a bias problem. A nonconvex penalty can alleviate the bias problem, but the involved nonconvex problem under positive-definiteness constraint is generally difficult to solve. In this work, we propose an efficient algorithm for positive-definiteness constrained covariance estimation by combining the iteratively reweighted method and the alternative direction method of multipliers (ADMM). The iterative reweighting scheme can achieve better sparsity regularization than the lasso method. Meanwhile, the proposed algorithm solves convex subproblems in each iteration and hence is easy to converge. The efficiency and effectiveness of the proposed algorithm has been demonstrated by both simulation study and a gene clustering example for tumor tissues. Code for reproducing the results is available at https://github.com/FWen/pdlc.git. Fei Wen 0005, Rendong Ying |
IEEE Trans. Big Data | 1 |
| 2021 | A Simple Local Minimal Intensity Prior and an Improved Algorithm for Blind Image DeblurringabstractBlind image deblurring is a long standing challenging problem in image processing and low-level vision. Recently, sophisticated priors such as dark channel prior, extreme channel prior, and local maximum gradient prior, have shown promising effectiveness. However, these methods are computationally expensive. Meanwhile, since these priors involved subproblems cannot be solved explicitly, approximate solution is commonly used, which limits the best exploitation of their capability. To address these problems, this work firstly proposes a simplified sparsity prior of local minimal pixels, namely patch-wise minimal pixels (PMP). The PMP of clear images is much more sparse than that of blurred ones, and hence is very effective in discriminating between clear and blurred images. Then, a novel algorithm is designed to efficiently exploit the sparsity of PMP in deblurring. The new algorithm flexibly imposes sparsity inducing on the PMP under the maximum a posterior (MAP) framework rather than directly uses the half quadratic splitting algorithm. By this, it avoids non-rigorous approximation solution in existing algorithms, while being much more computationally efficient. Extensive experiments demonstrate that the proposed algorithm can achieve better practical stability compared with state-of-the-arts. In terms of deblurring quality, robustness and computational efficiency, the new algorithm is superior to state-of-the-arts. Code for reproducing the results of the new method is available at https://github.com/FWen/deblur-pmp.git. Fei Wen 0005, Rendong Ying, Yipeng Liu 0001, Trieu-Kien Truong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | AMP-Net: Denoising-Based Deep Unfolding for Compressive Image SensingabstractMost compressive sensing (CS) reconstruction methods can be divided into two categories, i.e. model-based methods and classical deep network methods. By unfolding the iterative optimization algorithm for model-based methods onto networks, deep unfolding methods have the good interpretation of model-based methods and the high speed of classical deep network methods. In this article, to solve the visual image CS problem, we propose a deep unfolding model dubbed AMP-Net. Rather than learning regularization terms, it is established by unfolding the iterative denoising process of the well-known approximate message passing algorithm. Furthermore, AMP-Net integrates deblocking modules in order to eliminate the blocking artifacts that usually appear in CS of visual images. In addition, the sampling matrix is jointly trained with other network parameters to enhance the reconstruction performance. Experimental results show that the proposed AMP-Net has better reconstruction accuracy than other state-of-the-art methods with high reconstruction speed and a small number of network parameters. Yipeng Liu 0001, Jiani Liu 0002, Fei Wen 0005, Ce Zhu |
IEEE Trans. Image Process. | 4 |
| 2020 | Restoration of Motion Blur in Time-of-Flight Depth Image Using Data AlignmentabstractTime-of-flight (ToF) sensors are vulnerable to motion blur in the presence of moving objects. This is due to the principle of ToF camera that it estimates depth from the phase-shift between emitted and received modulated signals. And the phase-shift is measured by four sequential phase-shifted images, which is assumed to be consistent in an integration time. However, object motion would give rise to disparity among the four phase-shifted images, contributing to unreliable depth measurement. In this paper, we propose a novel method that is capable of aligning the four phase-shifted images through investigating the electronic value of each pixel in the phase images. It consists of two steps, motion detecting and deblurring. Furthermore, a refinement utilizing an additional group of phase-shifted images is adopted to further improve the accuracy of depth measurement. Experiment results on a new elaborated dataset with ground-truth demonstrate that the proposed method compares favorably over existing methods in both accuracy and runtime. Particularly, the new method can achieve the best accuracy while being computationally efficient that can support real-time running. Fei Wen 0005, Jun Wang 0137, Rendong Ying |
3DV | 3 |
| 2020 | Robust PCA Using Generalized Nonconvex RegularizationabstractRecently, the robustification of principal component analysis (PCA) has attracted much research attention in numerous areas of science and engineering. The most popular and successful approach is to model the robust PCA problem as a low-rank matrix recovery problem in the presence of sparse corruption. With this model, the nuclear norm and 11-norm penalties are usually used for low-rank and sparsity promotion. Although the nuclear norm and 11-norm are favorable due to their convexity, they have a bias problem. In comparison, nonconvex penalties can be expected to yield better recovery performance. In this paper, we consider a formulation for robust PCA using generalized nonconvex penalties for low-rank and sparsity inducing. This nonconvex formulation is efficiently solved by a multi-block alternative direction method of multipliers (ADMM) algorithm. A sufficient condition for the convergence of this new ADMM algorithm has been derived. Furthermore, to address the important issue of nonconvex penalty selection, we have evaluated the new algorithm via numerical experiments in various low-rank and sparsity conditions. The results indicate that, “exact” recovery of the low-rank principle component can be achieved only by nonconvex regularization. MATLAB code is available at https://github.com/FWen/RPCA.git. Fei Wen 0005, Rendong Ying, Robert C. Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Efficient Algorithms for Maximum Consensus Robust FittingabstractMaximum consensus robust fitting is a fundamental problem in many computer vision applications, such as vision-based robotic navigation and mapping. While exact search algorithms are computationally demanding, randomized algorithms are cheap but the solution quality is not guaranteed. Deterministic algorithms fill the gap between these two kinds of algorithms, which have better solution quality than randomized algorithms while being much faster than exact algorithms. In this article, we develop two highly efficient deterministic algorithms based on the alternating direction method of multipliers (ADMM) and proximal block coordinate descent (BCD) frameworks. Particularly, the proposed BCD algorithm is guaranteed convergent. Furthermore, on the slack variable in the BCD algorithm, which indicates the inliers and outliers, we establish some meaningful properties, such as support convergence within finite iterations and convergence to restricted strictly local minimizer. Compared with state-of-the-art algorithms, the new algorithms with initialization from a randomized or convex relaxed algorithm can achieve improved solution quality while being much more efficient (e.g., more than an order of magnitude faster). An application of the new ADMM algorithm in simultaneous localization and mapping (SLAM) has also been provided to demonstrate its effectiveness. Code for reproducing the results is available online. Fei Wen 0005, Rendong Ying |
IEEE Trans. Robotics | 1 |
| 2019 | Robust Precoding Design for Coarsely Quantized MU-MIMO Under Channel Uncertainties-V0abstractRecently, multi-user multiple input multiple output (MU-MIMO) systems with low-resolution digital-to-analog converters (DACs) has received considerable attention, owing to the capability of dramatically reducing the hardware cost. Besides, it has been shown that the use of low-resolution DACs enable great reduction in power consumption while maintain the performance loss within acceptable margin, under the assumption of perfect knowledge of channel state information (CSI). In this paper, we investigate the precoding problem for the coarsely quantized MU-MIMO system without such an assumption. The channel uncertainties are modeled to be a random matrix with finite second-order statistics. By leveraging a favorable relation between the multi-bit DACs outputs and the single-bit ones, we first reformulate the original complex precoding problem into a nonconvex binary optimization problem. Then, using the S-procedure lemma, the nonconvex problem is recast into a tractable formulation with convex constraints and finally solved by the semidefinite relaxation (SDR) method. Compared with existing representative methods, the proposed precoder is robust to various channel uncertainties and is able to support a MU-MIMO system with higher-order modulations, e.g., 16QAM. Fei Wen 0005, Robert C. Qiu |
ICC | 2 |
| 2019 | Action Recognition Based on 3D Skeleton and RGB Frame FusionabstractAction recognition has wide applications in assisted living, health monitoring, surveillance, and human-computer interaction. In traditional action recognition methods, RGB video-based ones are effective but computationally inefficient, while skeleton-based ones are computationally efficient but do not make use of low-level detail information. This work considers action recognition based on a multimodal fusion between the 3D skeleton and the RGB image. We design a neural network that uses a 3D skeleton sequence and a single middle frame from an RGB video as input. Specifically, our method picks up one frame in a video and extracts spatial features from it using two attention modules, a self-attention module and a skeleton-attention module. Further, temporal features are extracted from the skeleton sequence via a BI-LSTM sub-network. Finally, the spatial features and the temporal features are combined via a feature fusion network for action classification. A distinct feature of our method is that it uses only a single RGB frame rather than an RGB video. Accordingly, it has a light-weighted architecture and is more efficient than RGB video-based methods. Comparative evaluation on two public datasets, NTU-RGBD and SYSU, demonstrates that, our method can achieve competitive performance compared with state-of-the-art methods. Guiyu Liu, Jiuchao Qian, Fei Wen 0005, Xiaoguang Zhu, Rendong Ying |
IROS | 3 |
| 2019 | CD-ABM: Curriculum Design with Attention Branch Model for Person Re-identification
Jiuchao Qian, Xiaoguang Zhu, Fei Wen 0005 |
PRICAI (3) | 4 |
| 2019 | 3DTI-Net: Learn 3D Transform-Invariant Feature Using Hierarchical Graph CNN
Guanghua Pan, Jun Wang 0137, Rendong Ying, Fei Wen 0005 |
PRICAI (2) | 5 |
| 2019 | Efficient Nonlinear Precoding for Massive MIMO Downlink Systems With 1-Bit DACsabstractThe power consumption of digital-to-analog converters (DACs) constitutes a significant proportion of the total power consumption in a massive multiuser multiple-input multiple-output (MU-MIMO) base station (BS). Using 1-bit DACs can significantly reduce the power consumption. This paper addresses the precoding problem for the massive narrow-band MU-MIMO downlink system equipped with 1-bit DACs at each BS. In such a system, the precoding problem plays a central role as the precoded symbols are affected by extra distortion introduced by 1-bit DACs. In this paper, we develop a highly efficient nonlinear precoding algorithm based on the alternative direction method framework. Unlike the classic algorithms, such as the semidefinite relaxation (SDR) and squared-infinity norm Douglas-Rachford splitting (SQUID) algorithms, which solve convex relaxed versions of the original precoding problem, the new algorithm solves the original nonconvex problem directly. The new algorithm is guaranteed to globally converge under some mild conditions. A sufficient condition for its convergence has been derived. The experimental results in various conditions demonstrated that the new algorithm can achieve the state-of-the-art performance comparable with the SDR algorithm while being much more efficient (e.g., more than 300 times faster than the SDR algorithm). Fei Wen 0005, Lily Li 0003, Robert C. Qiu |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | Robust sparse recovery for compressive sensing in impulsive noise using ℓp-norm model fittingabstractThis work considers the robust sparse recovery problem in compressive sensing (CS) in the presence of impulsive measurement noise. We propose a robust formulation for sparse recovery using the generalized lp-norm with 0 < p < 2 as the metric for the residual error under l1-norm regularization. An alternative direction method (ADM) has been proposed to solve this formulation efficiently. Moreover, a smoothing strategy has been used to derive a convergent method for the nonconvex case of p < 1. The convergence conditions of the proposed algorithm for both the convex and nonconvex cases have been provided. Numerical simulations demonstrated that the new algorithm can achieve state-of-the-art robust performance in highly impulsive noise. Fei Wen 0005, Yipeng Liu 0001, Robert C. Qiu, Wenxian Yu |
ICASSP | 1 |
| 2016 | An improved sparse reconstruction algorithm for speech compressive sensing using structured priorsabstractThis work addresses the issue of sparse reconstruction in compressive sensing (CS) for speech signals. We propose a novel sparse reconstruction algorithm based on the approximate message passing (AMP) framework, via exploiting the intrinsic structures of real-life speech signals in the modified discrete cosine transform (MDCT) domain. We use a Gaussian mixture model to characterize the marginal distribution of the MDCT coefficients, and employ a first order Markov chain model to capture the inter-dependencies between neighboring MDCT coefficients. The parameters of these two models are adaptively learned using an expectation-maximization (EM) learning procedure. Compared with several state-of-the-art algorithms, the new algorithm showed significantly better performance in reconstruction experiments on real speech signals. Xiaobo Jiang, Rendong Ying, Fei Wen 0005, Sumxin Jiang |
ICME | 3 |
| 2014 | Robust Capon beamforming exploiting the second-order noncircularity of signals
Fei Wen 0005, Qun Wan, He-Wen Wei, Yong-Jie Luo |
Signal Process. | 1 |
| 2014 | Improved MUSIC Algorithm for Multiple Noncoherent SubarraysabstractThis work addresses the direction-of-arrival (DOA) estimation issue with multiple noncoherent subarrays. We use a maximum likelihood approach to derive a weighted MUSIC (w-MUSIC) algorithm for such arrays, which obtains the overall spatial spectrum via combining the weighted MUSIC spectrum of the subarrays. Theoretical analysis and numerical examples demonstrate that the w-MUSIC algorithm has a better performance compared to a previously introduced MUSIC algorithm for noncoherent subarrays. Fei Wen 0005, Qun Wan, He-Wen Wei |
IEEE Signal Process. Lett. | 1 |