Lu Gan 0002

dblp:45/3353-2 · DBLP profile ↗
← Back
51ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0003-1056-7660ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 4 · 2 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Joint Transceiver Beamforming and IRS Phase Shift Design for mmWave MIMO with Modulo ADCs
abstract
Millimeter-wave (mmWave) multiple-input multiple-output (MIMO) has accelerated the efficiency of modern wireless systems, yet its performance is constrained by the propagation blocking effect and the receiver analog-to-digital converters (ADC) saturation. Motivated by these, this paper first attempted to study an intelligent reflecting surface (IRS)-aided downlink mmWave MIMO system equipped with a modulo ADC front end at the receiver and explicitly models receiver-side quantization error. We propose spectral efficiency (SE) maximization under total-power and constant-modulus constraints and develop a real-valued genetic algorithm-based scheme that, in each generation, (i) explores the IRS reflection matrix and (ii) updates the transceiver beamformers by performing singular value decomposition on the cascaded channel, updated with the IRS reflection matrix. Experimental results demonstrate that the modulo ADC achieves higher SE and lower bit error rates with fewer bit resolution, effectively mitigating saturation-induced clipping and outperforming conventional ADCs.
Wenyi Yan, Lu Gan 0002, Qilie Liu, John Cosmas
ICC3
2026 Line spectral estimation with unlimited sensing
Hongwei Wang 0005, Jun Fang 0001, Hongbin Li 0001, Geert Leus, Ruixiang Zhu, Lu Gan 0002
Signal Process.6
2026 A novel sparse adaptive filter for suppressing impulsive disturbance in audio signals
Hongqing Liu 0002, Lu Gan 0002, Yi Zhou 0014, Maciej Niedzwiecki, Trieu-Kien Truong
Signal Process.3
2026 Integrating Robust CRT With Signal Unwrapping for Modulo Sampling
abstract
Two-channel modulo sampling based on the robust Chinese remainder theorem (RCRT) enables closed-form, point wise reconstruction, yet its recovery remains fundamentally limited by the effective RCRT threshold. This letter addresses this limitation through a hybrid architecture that integrates RCRT based recovery with difference-based unwrapping. The RCRT stage is reinterpreted as a coarse modulo operator with an expanded threshold, while the difference stage resolves the residual ambiguity by exploiting temporal structure. The proposed scheme extends recovery beyond the intrinsic RCRT range and simultaneously relaxes the oversampling requirement for a fixed difference order. Theoretical analysis establishes the resulting sampling condition, and the effectiveness of the proposed method is validated through both simulations and hardware experiments.
Chutong Shen, Wenyi Yan, Yimin Zhang 0001, Lu Gan 0002
IEEE Signal Process. Lett.4
2026 Bit-Efficient Quantization for Two-Channel Modulo-Sampling Systems
abstract
Two-channel modulo analog-to-digital converters (ADCs) enable high-dynamic-range signal sensing at the Nyquist rate per channel, but existing designs quantise both channel outputs independently, incurring redundant bitrate costs. This paper proposes a bit-efficient quantisation scheme that exploits the integer-valued structure of inter-channel differences, transmitting one quantised channel output together with a compact difference index. We prove that this approach requires only 1-2 bits per signal sample overhead relative to conventional ADCs, despite operating with a much smaller per-channel dynamic range. Simulations confirm the theoretical error bounds and bitrate analysis, while hardware experiments demonstrate substantial bitrate savings compared with existing modulo sampling schemes, while maintaining comparable reconstruction accuracy. These results highlight a practical path towards high-resolution, bandwidth-efficient modulo ADCs for bitrate-constrained systems.
Wenyi Yan, Lu Gan 0002, Honqing Liu, Guoquan Li 0001
IEEE Signal Process. Lett.3
2025 Threshold Sensitivity in Two-Channel Modulo ADCs: Analysis and Robust Reconstruction
abstract
This paper presents a comprehensive analysis of two-channel modulo analog-to-digital converters (ADCs) systems, focusing on the sensitivity of ADC thresholds. By exploiting analytic number theory, we first investigate the relationship among ADC threshold precision, maximum signal dynamic range, and error tolerance. Our analysis reveals that even slight deviations in ADC thresholds can substantially impact the maximum reconstructed signal dynamic range and error tolerance. To address these sensitivity issues, we propose a novel approach that strategically sacrifices signal dynamic range to stabilise error tolerance in the presence of slight ADC threshold variations. We also introduce a low-complexity reconstruction algorithm that exploits this trade-off, thereby enhancing system robustness. Simulation results validate the theoretical framework and confirm the efficiency of our proposed algorithm.
Wenyi Yan, Lu Gan 0002, Yimin Zhang 0001
ICASSP2
2025 A Robust Hybrid ACC-PM Approach for Personal Sound Zones
abstract
The performance of personal sound systems is often degraded by inaccurate acoustic measurements. To achieve robust control while balancing acoustic contrast and signal distortion, this work proposes a robust hybrid optimization method that exploits both acoustic contrast control and pressure matching (ACC-PM). The method addresses perturbations caused by uncertainties in the acoustic transfer functions such as temperature changes, head movement, etc, modeled as norm-bounded uncertainties. Although the resulting worst-case optimization is inherently non-convex, it is reformulated as a second-order cone programming problem, which can be efficiently solved. Numerical simulations demonstrate the effectiveness of the proposed robust ACC-PM algorithm, showing an improvement over 18% in terms of AC compared to vanilla ACC-PM.
Yaqi Zhu, Hongqing Liu 0001, Liming Shi, Lu Gan 0002
INTERSPEECH5
2025 Live Demonstration: Real-Time High-Amplitude Signal Acquisition with 2-Channel Modulo ADC
abstract
Modulo analog-to-digital converters (ADCs) offer a potential solution to the clipping challenges in conventional ADCs by folding signals that exceed the threshold. This makes them suitable for high-amplitude signal acquisition in wide dynamic range applications. In this demonstration, we present a 2-channel modulo ADC system implemented on a field-programmable gate array (FPGA) with an integrated real-time recovery algorithm. By managing both signal folding and recovery entirely in hard-ware, the system ensures low-latency processing. The FPGA efficiently handles high-bandwidth signals, making it a promising option for applications that require robust performance in high-amplitude signal environments.
Wenyi Yan, Ruixiang Zhu, Lu Gan 0002, Hongqing Liu 0001
ISCAS4
2025 Generalized Score Matching: Bridging $f$-Divergence and Statistical Estimation Under Correlated Noise
abstract
Relative Fisher information, also known as score matching, is a recently introduced learning method for parameter estimation. Fundamental relations between relative entropy and score matching have been established in the literature for scalar and isotropic Gaussian channels. This paper demonstrates that such relations hold for a much larger class of observation models. We introduce the vector channel where the perturbation is non-isotropic Gaussian noise. For such channels, we derive new representations that connect the$f$-divergence between two distributions to the estimation loss induced by mismatch at the decoder. This approach not only unifies but also greatly extends existing results from both the isotropic Gaussian and classical relative entropy frameworks. Building on this generalization, we extend De Bruijn's identity to mismatched non-isotropic Gaussian models and demonstrate that the connections to generative models naturally follow as a consequence application of this new result.
Yirong Shen, Lu Gan 0002, Cong Ling 0001
ISIT2
2025 Information Theoretic Learning for Diffusion Models with Warm Start
abstract
Generative models that maximize model likelihood have gained traction in many practical settings. Among them, perturbation-based approaches underpin many state-of-the-art likelihood estimation models, yet they often face slow convergence and limited theoretical understanding. In this paper, we derive a tighter likelihood bound for noise-driven models to improve both the accuracy and efficiency of maximum likelihood learning. Our key insight extends the classical Kullback–Leibler (KL) divergence–Fisher information relationship to arbitrary noise perturbations, going beyond the Gaussian assumption and enabling structured noise distributions. This formulation allows flexible use of randomized noise distributions that naturally account for sensor artifacts, quantization effects, and data distribution smoothing, while remaining compatible with standard diffusion training. Treating the diffusion process as a Gaussian channel, we further express the mismatched entropy between data and model, showing that the proposed objective upper-bounds the negative log-likelihood (NLL). In experiments, our models achieve competitive NLL on CIFAR-10 and state-of-the-art results on ImageNet across multiple resolutions, all without data augmentation, and the framework extends naturally to discrete data.
Yirong Shen, Lu Gan 0002, Cong Ling 0001
NeurIPS2
2024 Understanding Gaussian Noise Mismatch: A Hellinger Distance Approach
abstract
This paper explores noise-mismatched models using the Hellinger distance. In many applications, the design/training stage often assumes an independent and identically distributed (i.i.d.) Gaussian prior noise, but the real world introduces Gaussian noise with arbitrary covariance, creating a mismatch. We analyze the impact on system output and study optimal injected noise intensity for training/design. While theory assumes Gaussian sources, it provides guidance for non-Gaussian settings too. Experiments with Cycle-GAN for image-to-image translation validate the theory, producing results consistenting with derivations. Overall, this work provides theoretical and empirical insights into designing systems robust to noise uncertainties beyond simplified assumptions.
Chaohua Shi, Lu Gan 0002, Hongqing Liu 0001
ICASSP3
2024 Towards Optimized Multi-Channel Modulo-ADCs: Moduli Selection Strategies and Bit Depth Analysis
abstract
This paper presents a theoretical analysis of multi-channel modulo analog-to-digital converters (ADCs) for high-dynamic range sampling under bounded noise. In particular, we derive the maximum error tolerance in terms of ADC dynamic range, signal dynamic range, and channel number. Additionally, we present closed-form expressions for ADC thresholds, ensuring near-optimal error resilience, and analyzing the minimal bit-depth needed for stable recovery. Compared to single-channel modulo ADCs, our approach achieves superior error tolerance with reduced sampling rates. Moreover, it demands a minor bit rate increase compared to conventional ADCs but operates with a significantly smaller ADC dynamic range.
Wenyi Yan, Lu Gan 0002, Shaoqing Hu, Hongqing Liu 0001
ICASSP2
2024 On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection
abstract
Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection during training has been utilized, its theoretical underpinnings have been unclear. This work provides a robust theoretical framework elucidating the role of Gaussian noise injection in I2I translation models. We address critical questions on the influence of noise variance on distribution divergence, resilience to unseen noise types, and optimal noise intensity selection. Our contributions include connecting $f$-divergence and score matching, unveiling insights into the impact of Gaussian noise on aligning probability distributions, and demonstrating generalized robustness implications. We also explore choosing an optimal training noise level for consistent performance in noisy environments. Extensive experiments validate our theoretical findings, showing substantial improvements over various I2I baseline models in noisy settings. Our research rigorously grounds Gaussian noise injection for I2I translation, offering a sophisticated theoretical understanding beyond heuristic applications.
Chaohua Shi, Lu Gan 0002, Hongqing Liu 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
ICLR3
2024 Cross Domain Optimization for Speech Enhancement: Parallel or Cascade?
abstract
This paper introduces five novel deep-learning architectures for speech enhancement. Existing methods typically use time-domain, time-frequency representations, or a hybrid approach. Recognizing the unique contributions of each domain to feature extraction and model design, this study investigates the integration of waveform and complex spectrogram models through cross-domain fusion to enhance speech feature learning and noise reduction, thereby improving speech quality. We examine both cascading and parallel configurations of waveform and complex spectrogram models to assess their effectiveness in speech enhancement. Additionally, we employ an orthogonal projection-based error decomposition technique and manage the inputs of individual sub-models to analyze factors affecting speech quality. The network is trained by optimizing three specific loss functions applied across all sub-models. Our experiments, using the DNS Challenge (ICASSP 2021) dataset, reveal that the proposed models surpass existing benchmarks in speech enhancement, offering superior speech quality and intelligibility. These results highlight the efficacy of our cross-domain fusion strategy.
Hongqing Liu 0001, Liming Shi, Yi Zhou 0014, Lu Gan 0002
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Robust Precoding for HF Skywave Massive MIMO With Slepian Transform
abstract
In this paper, we address robust precoding in high-frequency (HF) skywave massive multiple-input multiple-output (MIMO) systems with imperfect channel state information (CSI). We first employ a sparse beam baseda posteriorichannel model and demonstrate that robust precoding can be efficiently solved in the Slepian transform domain with a large number of base station (BS) antennas. Next, we introduce two Slepian transform based robust precoding methods, including a joint approach that leverages inverse fast Fourier transform (IFFT) for reduced complexity with a large number of user terminals (UTs). We then establish a local optimum for the Slepian transform domain robust precoder (STRP) design using the majorization minimization (MM) algorithm, taking advantages of HF skywave massive MIMO channel sparsity and Slepian sequence properties. Further, two distinct designs are presented: separate STRP (SSTRP) and joint STRP (JSTRP). Simulation results confirm the effectiveness of proposed robust precoders, showcasing their excellent ergodic sum-rate performance and low complexity.
Linfeng Song, Ding Shi, Lu Gan 0002, Xiqi Gao 0001
IEEE Trans. Commun.3
2023 mdctGAN: Taming transformer-based GAN for speech super-resolution with Modified DCT spectra
abstract
Annual Conference of the International Speech Communication Association
Chenhao Shuai, Chaohua Shi, Lu Gan 0002, Hongqing Liu 0001
INTERSPEECH3
2022 Acoustic Echo Cancellation and Noise Suppression with a Full Time-Frequency Cascaded Neural Network
abstract
With the developments of various multi-function communication services, acoustic echoes and background noises inevitably appear in hands-free calling occasions. Different from using the combination of neural network and traditional acoustic echo cancellation (AEC) method, this paper directly proposes a time-frequency complex cascaded neural network (TFCN) for echo cancellation and noise suppression. To that aim, in frequency domain, complex LSTM layers are employed to process the real and imaginary signals. After that, an end-to-end time domain network is designed using dilated convolution layers to further remove residual interferences. By adding rich delay information to the dataset and optimizing the model by a weighted loss function, the generalization ability of the model is also improved. The extensive experimental results show that the proposed frame-work is robust to blind test datasets, effectively removes echoes and noises, and achieves an excellent performance on AECMOS scores. The subjective mean score of the proposed method is 4.37, which is 0.50 higher than the INTERSPEECH2021 AEC-Challenge baseline.
Hongqing Liu 0001, Yi Zhou 0014, Lu Gan 0002
MMSP5
2021 Fast Binary Embedding of Deep Learning Image Features Using Golay-Hadamard Matrices
abstract
Convolutional neural networks (CNNs) have emerged as powerful tools for image retrieval and classification. In this paper, we study binary embedding of CNN-based image descriptors for resource-constrained devices. We propose two classes of fast computable and memory-efficient dimension reduction operators using Golay-Hadamard matrices (GHMs), which are constructed by multiplying the columns of a Hadamard matrix with a Golay sequence. Simulation results on CNN-based instance image retrieval and classification show that GHM-based operators can offer competitive performance to those of full random Gaussian matrices and Gaussian circulant matrices at much lower computational cost and storage space. This implies the potential of proposed approaches on devices with low memory, bandwidth and power restrictions.
Chanattra Ammatmanee, Lu Gan 0002, Hongqing Liu 0001
ICME2
2021 Multi-Channel Modulo Samplers Constructed From Gaussian Integers
abstract
Recently, there is an increased interest in the study of modulo analog to digital converters (ADCs). These new systems can reconstruct a signal whose amplitude is much higher than the conventional ADC's dynamic range. Modulo ADCs are characterized by their modulo threshold and in the current literature, all existing works are limited to real-valued moduli. In this paper, we propose multi-channel modulo samplers with complex-valued moduli to sample a band-limited complex signal. Specifically, we discuss the construction of complex divisors from Gaussian integers and propose their efficient implementations. A memory-efficient, closed-form recovery algorithm is also proposed. Simulation results demonstrate that the proposed systems can provide stable reconstruction of a high dynamic range complex-valued signal at low sampling rates.
Lu Gan 0002, Hongqing Liu 0001
IEEE Signal Process. Lett.2
2020 Spike Sorting Based On Low-Rank And Sparse Representation
abstract
As the first step to study the coding mechanism and synergistic behaviour of neurons, spike sorting plays an important role in the neurosciences research community. Despite many empirical successes in spike sorting models, there are still sufferings from the overlapping and noise corruption problems. To ease these situations, in this paper, we present an efficient and effective method with the help of optimization theory. Firstly, by introducing the low-rank strategy, the global structure underlying the spike data could be discovered. Secondly, by engaging the sparse coding to balance the noise, the proposed model is robust in the overlapping and noise spike sorting scenario. We have conducted experiments on the Wave-clus dataset compared with two state of the art models. The results verify the efficacy of our scheme and confirm the claims above.
Libo Huang 0001, Bingo Wing-Kuen Ling, Yan Zeng 0002, Lu Gan 0002
ICME4
2020 Clutter Reduction and Target Tracking in Through-the-Wall Radar
abstract
This article addresses the problem of tracking targets behind the wall using through-the-wall radar. To that end, the wall reflection, i.e., clutter, must be eliminated first because it interferes with the subsequent image formation operation. The low-rank of the clutter and sparseness of the useful signal are utilized to devise a joint low-rank and sparse framework to simultaneously suppress the clutter and recover the target returns, where alternating direction method of multipliers (ADMM) approach is developed to solve the corresponding optimization. Since then, an effective observation window scheme is proposed to locate the target and further to facilitate the tracking process. The tracking is finally provided by Kalman filter and particle filter. The numerical studies are provided to demonstrate that the performance of the proposed framework is superior to that of other methods in terms of clutter removal and tracking accuracy.
Hongqing Liu 0001, Lu Gan 0002, Yi Zhou 0014, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.3
2019 3D Coprime Arrays in Sparse Sensing
abstract
Coprime arrays are a class of sensor arrays that play a crucial role in various signal processing tasks because of their desirable properties such as sparsity and increased degrees of freedom (DOF) of coarrays. In this contribution, a new class of three-dimensional (3D) arrays is constructed from pure cubic fields. By studying the properties of cubic integers, we convert the problem of finding two coprime 3-by-3 integer matrices to that of two coprime integers in the ring of integers of a cubic field, which significantly reduces the design complexity and expands the design space of these matrices. The proposed construction offers naturally commutative matrices and includes generalized circulant matrices as a special case (under certain restriction of a parameter). The surged DOF is guaranteed by the generalized Chinese Remainder Theorem (CRT) for rings and ideals.
Conghui Li, Lu Gan 0002, Cong Ling 0001
ICASSP2
2019 RFI Suppression Based on Atomic Norm Minimization in SAR Signal Recovery
abstract
The recovery problem of synthetic aperture radar (SAR) signal in the presence of radio frequency interference (RFI) is studied. To perform RFI suppression, in this paper, the RFI is modeled as the combination of multiple complex sinusoids such that the RFI suppression problem becomes a frequency estimation one. To accurately estimate model parameters, by exploiting sparse representation of the RFI, a gridless approach based on atomic norm minimization is proposed, which completely removes the off-grid issue. Finally, to recover the SAR signal, a joint scheme is devised to simultaneously perform the RFI suppression and the SAR signal recovery under an optimization framework. The resultant optimization is efficiently solved by a two-step process based on the coordinate descent approach. Simulation results and real-world experiments are provided to show the superior performance of the proposed approach.
Hongqing Liu 0001, Lu Gan 0002, Dong Li 0007, Trieu-Kien Truong
ICIP2
2016 Compressive CSI Acquisition and Non-Orthogonal Pilot Design for Downlink Massive MIMO Systems
abstract
Channel state information (CSI) is usually necessary for downlink precoding, power allocation, etc. in multiple-input multiple-output (MIMO) systems. When the base stations (BS) are equipped with massive elements, the training overhead required by conventional CSI estimation methods becomes overwhelming, leading to unacceptable loss of spectrum efficiency. In this paper, we investigate pilot design and CSI acquisition issues for downlink massive MIMO transmission. By exploiting the sparsity of beam domain channel, we first derive the optimal pilot structure in case of non-orthogonal pilots with compressive sensing (CS) framework. A deterministic sensing matrix design method is then proposed that satisfies the restricted isometry property (RIP). As beam domain channels are usually approximately sparse, we propose a modified subspace pursuit (SP) algorithm to recover the signals with tradeoff between noise and approximation error. Numerical results demonstrate that the proposed sensing matrices have better performance than conventional random CS matrices, and the new channel estimation scheme achieves significant performance improvement with reduced pilots consumption over conventional least square (LS) method.
Wenjin Wang 0001, Lu Gan 0002, Xiqi Gao 0001
GLOBECOM3
2015 Atomic norm denoising-based channel estimation for massive multiuser MIMO systems
abstract
In this paper, we propose a novel channel estimation method for massive multiple-input multiple-output (MIMO) systems operating in time-division duplexing (TDD) mode. By exploiting the fact that the degrees of freedom of the physical channel matrix are smaller than the number of free parameters, the channel estimation is formulated as an atomic norm denoising problem and solved efficiently via the alternating direction method of multipliers (ADMM). Both theoretical analysis and numerical simulations demonstrate that our proposed method outperforms existing ones in terms of the channel estimation performance.
Peng Zhang 0020, Lu Gan 0002, Sumei Sun, Cong Ling 0001
ICC2
2013 Golay sequence for parital Fourier and Hadamard compressive imaging
abstract
This paper introduces Golay sequence for partial Fourier and Hadamard compressive imaging. In the proposed system, the signal is pre-modulated by a binary Golay sequence before applying a random subsampled Fourier or Hadamard transform. The correpsonding sampling operator has been proved to be incoherent with the (block) DCT or the Haar wavelet transform. Empirical results show that they are also incoherent with the Daubechies wavelets. It is well known that natural images are sparse in the DCT and the wavelet basis. Hence, the proposed sampling operators are promising in many Fourier or Hadamard compressive imaging applications. In fact, they can achieve near-optimal reconstruction performance with small memory requirement and simple hardware implementation. Some simulation results and proof-of-concept experimental results are included to demonstrate the validity of the theory and the potential of the proposed sampling operators.
Lu Gan 0002, Yaochun Shen
ICASSP1
2013 Lattice quantization noise revisited
abstract
Dithered quantization is widely used in signal processing and coding due to its many desirable properties in theory and in practice. In this paper, we show that dither is unnecessary for Gaussian sources when the flatness factor of the quantization lattice is small. This means that the quantization noise behaves much like that in dithered quantization. In particular, it tends to be uniformly distributed over any fundamental region of the lattice and be uncorrelated with the signal; further, for optimum lattice quantizers, it approaches the rate-distortion bound of Gaussian sources with minimum mean-square error (MMSE) estimation.
Cong Ling 0001, Lu Gan 0002
ITW2
2012 Golay meets Hadamard: Golay-paired Hadamard matrices for fast compressed sensing
abstract
This paper introduces Golay-paired Hadamard matrices for fast compressed sensing of sparse signals in the time or spectral domain. These sampling operators feature low-memory requirement, hardware-friendly implementation and fast computation in reconstruction. We show that they require a nearly optimal number of measurements for faithful reconstruction of a sparse signal in the time or frequency domain. Simulation results demonstrate that the proposed sensing matrices offer a reconstruction performance similar to that of fully random matrices.
Lu Gan 0002, Kezhi Li, Cong Ling 0001
ITW1
2011 Deterministic compressed-sensing matrices: Where Toeplitz meets Golay
abstract
Recently, the statistical restricted isometry property (STRIP) has been formulated to analyze the performance of deterministic sampling matrices for compressed sensing. In this paper, a class of deterministic matrices which satisfy STRIP with overwhelming probability are proposed, by taking advantage of concentration inequalities using Stein's method. These matrices, called orthogonal symmetric Toeplitz matrices (OSTM), guarantee successful recovery of all but an exponentially small fraction of K-sparse signals. Such matrices are deterministic, Toeplitz, and easy to generate. We derive the STRIP performance bound by exploiting the specific properties of OSTM, and obtain the near-optimal bound by setting the underlying sign sequence of OSTM as the Golay sequence. Simulation results show that these deterministic sensing matrices can offer reconstruction performance similar to that of random matrices.
Kezhi Li, Cong Ling 0001, Lu Gan 0002
ICASSP3
2010 Fast dimension reduction through random permutation
abstract
This paper studies permutation-based dimension reduction, which can be implemented by first scrambling the input data, then applying the FFT, DCT or Walsh-Hadamard transform and finally using either uniformly random sampling or sparse random projection. By exploiting concentration inequalities of random permutation, we show that this subclass of operators can offer (near) optimal theoretical guarantee. Besides, as random permutation of N elements can be implemented in O(N) time, the proposed algorithm has very low complexity. Some numerical examples are presented to demonstrate the validity of our theoretical development and their promising applications in image processing.
Lu Gan 0002, Thong T. Do, Trac D. Tran
ICIP1
2010 Secure private fragile watermarking scheme with improved tampering localisation accuracy
abstract
In some applications, such as surveillance cameras, it is essential that images acquired with the same device be deemed distinct. To fulfil this requirement, many fragile watermarking methods need to employ a different key for every single image. Keeping track of the correct key associated with each image may become a challenging task as the number of images increases. An existing image index-based scheme as a suitable solution for this problem is revisited, and its security limitations in applications are highlighted where higher localisation accuracy is required. Then, a new fragile watermarking scheme that enhances the localisation accuracy is proposed, while hindering brute force attacks. Experiments demonstrate that the proposed scheme outperforms the state-of-the-art private fragile watermarking methods.
Sergio Bravo-Solorio, Lu Gan 0002, Asoke K. Nandi, Maurice F. Aburdene
IET Inf. Secur.2
2009 A fast and efficient heuristic nuclear-norm algorithm for affine rank minimization
abstract
The problem of affine rank minimization seeks to find the minimum rank matrix that satisfies a set of linear equality constraints. Generally, since affine rank minimization is NP-hard, a popular heuristic method is to minimize the nuclear norm that is a sum of singular values of the matrix variable. A recent intriguing paper shows that if the linear transform that defines the set of equality constraints is nearly isometrically distributed and the number of constraints is at least O(r(m + n) logmn), where r and m times n are the rank and size of the minimum rank matrix, minimizing the nuclear norm yields exactly the minimum rank matrix solution. Unfortunately, it takes a large amount of computational complexity and memory buffering to solve the nuclear norm minimization problem with known nearly isometric transforms. This paper presents a fast and efficient algorithm for nuclear norm minimization that employs structurally random matrices for its linear transform and a projected subgradient method that exploits the unique features of structurally random matrices to substantially speed up the optimization process. Theoretically, we show that nuclear norm minimization using structurally random linear constraints guarantees the minimum rank matrix solution if the number of linear constraints is at least O(r(m+n) log3mn). Extensive simulations verify that structurally random transforms still retain optimal performance while their implementation complexity is just a fraction of that of completely random transforms, making them promising candidates for large scale applications.
Thong T. Do, Yi Chen 0014, Nam H. Nguyen, Lu Gan 0002, Trac D. Tran
ICASSP4
2009 Fast and efficient dimensionality reduction using Structurally Random Matrices
abstract
Structurally Random Matrices (SRM) are first proposed in [1] as fast and highly efficient measurement operators for large scale compressed sensing applications. Motivated by the bridge between compressed sensing and the Johnson-Lindenstrauss lemma [2] , this paper introduces a related application of SRMs regarding to realizing a fast and highly efficient embedding. In particular, it shows that a SRM is also a promising dimensionality reduction transform that preserves all pairwise distances of high dimensional vectors within an arbitrarily small factor ∈, provided that the projection dimension is on the order of O(∈−2log3N), where N denotes the number of d-dimensional vectors. In other words, SRM can be viewed as the sub-optimal Johnson-Lindenstrauss embedding that, however, owns very low computational complexity O(d log d) and highly efficient implementation that uses only O(d) random bits, making it a promising candidate for practical, large scale applications where efficiency and speed of computation are highly critical.
Thong T. Do, Lu Gan 0002, Yi Chen 0014, Nam P. Nguyen, Trac D. Tran
ICASSP2
2009 Distributed compressed video sensing
abstract
This paper proposes a novel framework called Distributed Compressed Video Sensing (DISCOS) - a solution for Distributed Video Coding (DVC) based on the recently emerging Compressed Sensing theory. The DISCOS framework compressively samples each video frame independently at the encoder. However, it recovers video frames jointly at the decoder by exploiting an interframe sparsity model and by performing sparse recovery with side information. In particular, along with global frame-based measurements, the DISCOS encoder also acquires local block-based measurements for block prediction at the decoder. Our interframe sparsity model mimics state-of-the-art video codecs: the sparsest representation of a block is a linear combination of a few temporal neighboring blocks that are in previously reconstructed frames or in nearby key frames. This model enables a block to be optimally predicted from its local measurements by l1-minimization. The DISCOS decoder also employs a sparse recovery with side information to jointly reconstruct a frame from its global measurements and its local block-based prediction. Simulation results show that the proposed framework outperforms the baseline compressed sensing-based scheme of intraframe-coding and intraframe-decoding by 8 – 10dB. Finally, unlike conventional DVC schemes, our DISCOS framework can perform most encoding operations in the analog domain with very low-complexity, making it be a promising candidate for real-time, practical applications where the analog to digital conversion is expensive, e.g., in Terahertz imaging.
Thong T. Do, Yi Chen 0014, Dzung T. Nguyen, Nam P. Nguyen, Lu Gan 0002, Trac D. Tran
ICIP5
2009 Robust video transmission using Layered Compressed Sensing
abstract
We propose a novel Layered Compressed Sensing (CS) approach for robust transmission of video signals over packet loss channels. In our proposed method, the encoder consists of a base layer and an enhancement layer. The base layer is a conventionally encoded bitstream and transmitted without any error protection. The additional enhancement layer is a stream of compressed measurements taken across slices of video signals for error-resilience. The decoder regards the corrupted base layer as the side information (SI) and employs a sparse recovery with SI to recover approximation of lost packets. By exploiting the SI at the decoder, the enhancement layer is required to transmit a minimal amount of compressed measurements for error protection that is only proportional to the amount of lost packets. Simulation results show that both compression efficiency and error-resilience capacity of the proposed scheme are competitive with those of other state-of-the-art robust transmission methods, in which Wyner-Ziv (WZ) coders often generate an enhancement layer. Thanks to the soft-decoding feature of sparse recovery algorithms, our CS-based scheme can avoid the cliff effect that often occurs with otherWyner-Ziv based schemes when the error rate is over the error correction capacity of the channel code. In addition, our result suggests that compressed sensing is actually closer to source coding with decoder side information than to conventional source coding.
Thong T. Do, Yi Chen 0014, Dzung T. Nguyen, Nam P. Nguyen, Lu Gan 0002, Trac D. Tran
MMSP5
2008 Fragile logo watermarking for public authentication
abstract
A new fragile logo watermarking scheme is proposed for public authentication and integrity verification of images. The security of the proposed block-wise scheme relies on a public encryption algorithm and a hash function. The encoding and decoding methods can provide public detection capabilities even in the absence of the image indices and the original logos. Furthermore, the detector automatically authenticates input images and extracts possible multiple logos and image indices, which can be used not only to localise tampered regions, but also to identify the original source of images used to generate counterfeit images. Results are reported to illustrate the effectiveness of the proposed method.
Sergio Bravo-Solorio, Lu Gan 0002, Asoke K. Nandi, Maurice F. Aburdene
ICASSP2
2008 Fast compressive sampling with structurally random matrices
abstract
This paper presents a novel framework of fast and efficient compressive sampling based on the new concept of structurally random matrices. The proposed framework provides four important features. (i) It is universal with a variety of sparse signals. (ii) The number of measurements required for exact reconstruction is nearly optimal. (iii) It has very low complexity and fast computation based on block processing and linear filtering. (iv) It is developed on the provable mathematical model from which we are able to quantify trade-offs among streaming capability, computation/memory requirement and quality of reconstruction. All currently existing methods only have at most three out of these four highly desired features. Simulation results with several interesting structurally random matrices under various practical settings are also presented to verify the validity of the theory as well as to illustrate the promising potential of the proposed framework.
Thong T. Do, Trac D. Tran, Lu Gan 0002
ICASSP3
2008 Lifting-based Laplacian Pyramid reconstruction schemes
abstract
Laplacian Pyramid (LP) provides a redundant signal representation and can be characterized as an oversampled filter bank (FB). In this paper, a generic lifting-based parameterization reconstruction algorithm is proposed to characterize all LP synthesis banks that can satisfy the perfect reconstruction property. Two typical lifting-based LP reconstruction schemes are then derived from this general representation. The first scheme presents the dual frame LP reconstruction and its closed-form solutions for any LP filters. The second LP reconstruction scheme leads to an efficient FB, which demonstrates improvements over the usual LP reconstruction in the presence of noise.
Lijie Liu, Lu Gan 0002, Trac D. Tran
ICIP2
2007 Expected Run-Time Distortion Based Scheduling for Scalable Video Transmission with Hybrid FEC/ARQ Error Control
abstract
The optimal packet scheduling for transmitting scalable media over the packet erasure networks has been extensively studied in the past. In the existing work, only retransmission is used for packet loss recovery. As a result, when the round-trip-time (RTT) of the network increases, the performance of the system degrades fast. In this paper, a scheduling scheme for hybrid FEC/ARQ error control is proposed. In our proposed scheme, the importance of both data packets and FEC packets are evaluated by considering several factors, such as data dependency structure of scalable video, transmission history of other packets, as well as the decoding deadline, and the most important packet is chosen to be sent out. In this way, our proposed scheme is able to achieve more stable playback video quality by taking advantage of both FEC and ARQ, as demonstrated by the experimental results.
Tong Gan, Lu Gan 0002, Kai-Kuang Ma
ICASSP (1)2
2007 Computation of the Dual Frame: Forward and Backward Greville Formulas
abstract
We study the computation of the dual frame for oversampled filter banks (OFBs) by exploiting Greville's formula, which was derived in 1960 to compute the pseudo inverse of a matrix when a new row is appended. In this paper, we first develop the backward Greville formula to handle the case of row deletion. Based on Greville's formula, we then study the dual frame computation of the Laplacian pyramid. Through the backward Greville formula, we investigate OFBs for robust transmission over erasure channels. The necessary and sufficient conditions for OFBs robust to one erasure channel are derived. A post-filtering structure is also presented to implement the dual frame when the transform coefficients in one subband are completely lost.
Lu Gan 0002, Cong Ling 0001
ICASSP (3)1
2007 Undersampled Boundary Pre-/Postfilters for Low Bit-Rate DCT-Based Block Coders
abstract
It has been well established that critically sampled boundary pre-/postfiltering operators can improve the coding efficiency and mitigate blocking artifacts in traditional discrete cosine transform-based block coders at low bit rates. In these systems, both the prefilter and the postfilter are square matrices. This paper proposes to use undersampled boundary pre- and postfiltering modules, where the pre-/postfilters are rectangular matrices. Specifically, the prefilter is a "fat" matrix, while the postfilter is a "tall" one. In this way, the size of the prefiltered image is smaller than that of the original input image, which leads to improved compression performance and reduced computational complexities at low bit rates. The design and VLSI-friendly implementation of the undersampled pre-/postfilters are derived. Their relations to lapped transforms and filter banks are also presented. Two design examples are also included to demonstrate the validity of the theory. Furthermore, image coding results indicate that the proposed undersampled pre-/postfiltering systems yield excellent and stable performance in low bit-rate image coding.
Lu Gan 0002, Chengjie Tu, Jie Liang 0001, Trac D. Tran, Kai-Kuang Ma
IEEE Trans. Image Process.1
2007 Wiener Filter-Based Error Resilient Time-Domain Lapped Transform
abstract
In this paper, the design of the error resilient time-domain lapped transform is formulated as a linear minimal mean-squared error problem. The optimal Wiener solution and several simplifications with different tradeoffs between complexity and performance are developed. We also prove the persymmetric structure of these Wiener filters. The existing mean reconstruction method is proven to be a special case of the proposed framework. Our method also includes as a special case the linear interpolation method used in DCT-based systems when there is no pre/postfiltering and when the quantization noise is ignored. The design criteria in our previous results are scrutinized and improved solutions are obtained. Various design examples and multiple description image coding experiments are reported to demonstrate the performance of the proposed method.
Jie Liang 0001, Chengjie Tu, Lu Gan 0002, Trac D. Tran, Kai-Kuang Ma
IEEE Trans. Image Process.3
2006 Reducing video-quality fluctuations for streaming scalable video using unequal error protection, retransmission, and interleaving
abstract
Forward error correction based multiple description (MD-FEC) transcoding for transmitting embedded bitstream over the packet erasure networks has been extensively studied in the past. In the existing work, a single embedded source bitstream, e.g., the bitstream of a group of pictures (GOP) encoded using three-dimensional set partitioning in hierarchical trees is optimally protected unequal error protection (UEP) in the rate-distortion sense. However, most of the previous work on transmitting embedded video using MD-FEC assumed that one GOP is transmitted only once, and did not consider the chance of retransmission. This may lead to noticeable video quality variations due to varying channel conditions. In this paper, a novel window-based packetization scheme is proposed, which combats bursty packet loss by combining the following three techniques: UEP, retransmission, and GOP-level interleaving. In particular, two retransmission mechanisms, namely segment-wise retransmission and byte-wise retransmission, are proposed based on different types of receiver feedback. Moreover, two levels of rate allocations are introduced: intra-GOP rate allocation minimizes the distortion of individual GOP; while inter-GOP rate allocation intends to reduce video quality fluctuations by adaptively allocating bandwidth according to video signal characteristics and client buffer status. In this way, more consistent video quality can be achieved under various packet loss probabilities, as demonstrated by our experimental results.
Tong Gan, Lu Gan 0002, Kai-Kuang Ma
IEEE Trans. Image Process.2
2005 Wiener Filtering for Generalized Error Resilient Time Domain Lapped Transform
abstract
In this paper, we revisit the design of the time-domain lapped transform for error resilient image transmission. A general structure is first proposed whose solution is given by a Wiener filter. Two simplified schemes with different tradeoffs between complexity and performance are then developed, for which Wiener filter solutions also exist. We show that the existing method is a special case of the general scheme. Design examples and image coding experiments verify that the performance of our new approach is significantly better than existing techniques.
Jie Liang 0001, Chengjie Tu, Trac D. Tran, Lu Gan 0002
ICASSP (2)4
2003 On efficient implementation of oversampled linear phase perfect reconstruction filter banks
abstract
In this paper, we first present an alternative way of generating oversampled linear phase perfect reconstruction filter banks (OSLPPRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed.
Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma
ICASSP (6)2
2003 Oversampled lapped transforms via time-domain pre- and post-processing
abstract
This paper introduces a large family of oversampled lapped transforms with symmetric basis functions. These new transforms are implemented by adding time-domain oversampled pre- and post- filters to the DCT and the IDCT, respectively. Structures and parameterizations of the corresponding pre-/postfilters are proposed. Two design examples along with some image coding results are presented to demonstrate the validity of the theory and the potential of the new transforms.
Lu Gan 0002, Kai-Kuang Ma
ICIP (3)1
2003 On efficient implementation of oversampled linear phase perfect reconstruction filter banks
abstract
In this paper, we first present an alternative way of generating over-sampled linear phase perfect reconstruction filter banks (OSLP-PRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed.
Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma
ICME2
2002 Theory and lattice factorization of oversampled linear-phase perfect reconstruction filter banks
abstract
This paper presents the theory and structure of a large family of oversampled linear-phase perfect reconstruction filter banks (OLPPRFBs). For such filter banks, we first derive the necessary existence conditions on the number of symmetric filters and antisymmetric filters. We then develop lattice factorizations of these OLPPRFBs, followed by two design examples to confirm the validity of the theory.
Lu Gan 0002, Kai-Kuang Ma
ICASSP1
2002 On lattice factorization of symmetric-antisymmetric multifilter banks
abstract
We introduce a new structure for symmetric-antisymmetric multiwavelets (SAMWTs) and symmetric-antisymmetric multifilter banks (SAMFBs). First, by exploring the connection between SAMFBs and traditional (scalar) linear phase perfect reconstruction filter banks (LPPRFBs), we show that the implementation and design of an SAMFB can be converted into that of a LPPRFB. Then, based on the lattice factorization for LPPRFBs, we propose a fast, modular, minimal structure for SAMFBs. To demonstrate the effectiveness of the proposed lattice structure, a multiplierless SAMWT design example is presented along with its application in image coding.
Lu Gan 0002, Kai-Kuang Ma
ICIP (1)1
2002 On the completeness of the lattice factorization for linear-phase perfect reconstruction filter banks
abstract
In this letter, we re-examine the completeness of the lattice factorization for M-channel linear-phase perfect reconstruction filter bank (LPPRFB) with filters of the same length L=KM as discussed by Tran et al. (see IEEE Trans. Signal Processing, vol.48, p.133-47, Jan. 2000). We point out that the assertion of completeness is incorrect. Examples are presented to show that the proposed lattice structure of Tran et al. is not complete when K>2. In addition, we verify that the lattice structure is complete only when K/spl les/2.
Lu Gan 0002, Kai-Kuang Ma, Truong Q. Nguyen, Trac D. Tran, Ricardo L. de Queiroz
IEEE Signal Process. Lett.1
2001 A simplified lattice factorization for linear-phase perfect reconstruction filter bank
abstract
We propose a simplified version of lattice factorization for linear-phase perfect reconstruction filter bank (LPPRFB) derived by T.D. Tran et al. (see IEEE Trans. Signal Processing. vol.48, no.1, p.133-47, Jan. 2000). The proposed new lattice structure spans the same class of LPPRFB, while substantially reducing free parameters in nonlinear optimization and saving computation cost in hardware implementation. To further address the importance of our proposed structure, we generalize our factorization to multidimensional LPPRFB (MD-LPPRFB), and show its effectiveness.
Lu Gan 0002, Kai-Kuang Ma
IEEE Signal Process. Lett.1