Guohe Zhang

dblp:127/9568 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0001-8092-8009ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DELTAT-SNN: An efficient dynamic execution latency tuning through adaptive timestep SNN for patient-specific multifunctional seizure monitoring
Ran Wang 0007, Jian Zhang 0091, Yumu Wang, Guohe Zhang
Neurocomputing6
2026 Spike-hammer: An efficient spike-driven hybrid architecture for multi-modal emotion recognition with physiological signals
Haisheng Fu, Yuchen Zou, Guohe Zhang, Jie Liang 0001
Neural Networks4
2026 A 61.4 Gb/s/mm Wireline Transceiver Using a 7 bit-Over-8 Lane Symmetric Correlated Coding for High-Density Interconnects
abstract
This paper presents a symmetric correlated coding (SCC) scheme and the corresponding high-density transceiver that deliver 7-bit data over 8 lanes. The SCC method implements the 7-bit data encoding jointly in 8 correlated channels which improves the pin efficiency up to 87.5%. Besides, as an advanced chord signaling, the SCC exhibits the immunity to crosstalk, common-mode noise (CMN) and simultaneous switching noise (SSN). The proposed SCC transceiver innovates a SCC source-series-terminated (SST) driver that maintains an unchanged bandwidth even with a low-voltage power supply, and the CTLE decoder adopts the active inductor to compensate the channel loss. Prototyped in 28-nm CMOS, the proposed wireline transceiver using SCC method supports a maximum data rate of$7\times 10$Gb/s with a bit error rate (BER) <1e−12, achieving a data rate density (DRD) up to 61.4 Gb/s/mm. The SCC transceiver dissipates 99.4 mW with the energy efficiency of 1.42 pJ/bit.
Geng Zhang 0001, Fangxu Lv, Liquan Xiao, Xuqiang Zheng, Heng Huang 0009, Kewei Xin, Liangyong Yuan, Ruixiao Kuai, Bolin Ren, Ruotian Yin, Guohe Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.12
2025 Resource-aware strategies for real-time multi-person pose estimation
Mohammed A. Esmail, Guoliang Zhu, Guohe Zhang
Image Vis. Comput.6
2025 Multiobjective Optimization of Class-F Oscillators
abstract
To address the complex nonlinear problem of determining class-F voltage-controlled oscillator (VCO) dimensions, this article introduces an electronic design automation (EDA) framework that rapidly optimizes multiple design objectives yielding superior outcomes. The framework incorporates fast frequency determination, harmonic alignment, and extremal optimization of multiobjective particle swarm optimization with crowding distance (FHE-MOPSO-CD), an efficient algorithm we developed specifically for class-F VCOs, which includes transformer-based tank circuit strategies and extremum optimization techniques. Using a 55-nm CMOS process, this algorithm optimized various class-F VCO topologies, achieving excellent metrics and confirming its versatility. Optimization results indicate that at a 10-MHz offset, the figure of merit (FoM) is at least 8.81 dBc/Hz higher than values reported in the literature. Compared with other analog/RF dimension optimization methods, our approach yielded a higher hypervolume, indicating better convergence and greater diversity of solutions.
Zhenjiao Chen, Xingqiang Shi, Guohe Zhang, Feng Liang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Learned Image Compression with Dual-Branch Encoder and Conditional Information Coding
abstract
Recent advancements in deep learning-based image compression are notable. However, prevalent schemes that employ a serial context-adaptive entropy model to enhance rate-distortion (R-D) performance are markedly slow. Furthermore, the complexities of the encoding and decoding networks are substantially high, rendering them unsuitable for some practical applications. In this paper, we propose two techniques to balance the trade-off between complexity and performance. First, we introduce two branching coding networks to independently learn a low-resolution latent representation and a high-resolution latent representation of the input image, discriminatively representing the global and local information therein. Second, we utilize the high-resolution latent representation as conditional information for the low-resolution latent representation, furnishing it with global information, thus aiding in the reduction of redundancy between low-resolution information. We do not utilize any serial entropy models. Instead, we employ a parallel channel-wise auto-regressive entropy model for encoding and decoding low-resolution and high-resolution latent representations. Experiments demonstrate that our method is approximately twice as fast in both encoding and decoding compared to the parallelizable checkerboard context model, and it also achieves a 1.2% improvement in R-D performance compared to state-of-the-art learned image compression schemes. Our method also outperforms classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods in rate-distortion performance, as validated by both PSNR and MS-SSIM metrics on the Kodak dataset.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
DCC5
2024 WeConvene: Learned Image Compression with Wavelet-Domain Convolution and Entropy Model
Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han, Feng Liang 0001, Guohe Zhang
ECCV (50)6
2024 Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context Model
abstract
Recently, learning-based image compression approaches have achieved superior performance over classical image compression methods. However, their complexities remain quite high. In this paper, we propose two efficient modules to reduce the complexity. First, we introduce a selective kernel residual module into the core network, which effectively expands the receptive field and captures global information. Second, we present an improved channel-wise causal context model, designed to not only reduce encoding and decoding time but also ensure rate-distortion performance. Experimental results demonstrate that our proposed method achieves better tradeoff than recent leading learned image compression methods, and also outperforms the latest H.266/VVC (4:4:4) in terms of PSNR and MS-SSIM metrics.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
ICASSP5
2024 Fast and High-Performance Learned Image Compression With Improved Checkerboard Context Model, Deformable Residual Module, and Knowledge Distillation
abstract
Deep learning-based image compression has made great progresses recently. However, some leading schemes use serial context-adaptive entropy model to improve the rate-distortion (R-D) performance, which is very slow. In addition, the complexities of the encoding and decoding networks are quite high and not suitable for many practical applications. In this paper, we propose four techniques to balance the trade-off between the complexity and performance. We first introduce the deformable residual module to remove more redundancies in the input image, thereby enhancing compression performance. Second, we design an improved checkerboard context model with two separate distribution parameter estimation networks and different probability models, which enables parallel decoding without sacrificing the performance compared to the sequential context-adaptive model. Third, we develop a three-pass knowledge distillation scheme to retrain the decoder and entropy coding, and reduce the complexity of the core decoder network, which transfers both the final and intermediate results of the teacher network to the student network to improve its performance. Fourth, we introduce$L_{1}$regularization to make the numerical values of the latent representation more sparse, and we only encode non-zero channels in the encoding and decoding process to reduce the bit rate. This also reduces the encoding and decoding time. Experiments show that compared to the state-of-the-art learned image coding scheme, our method can be about 20 times faster in encoding and 70-90 times faster in decoding, and our R-D performance is also 2.3% higher. Our method achieves better rate-distortion performance than classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods, as measured by both PSNR and MS-SSIM metrics on the Kodak and Tecnick-40 datasets.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
IEEE Trans. Image Process.6
2024 An Improved DEM for Multibit DT ΣΔMs Based on Poles Splitting Technique and Segmented VQ
abstract
This brief presents a higher-order vector DEM for multibit discrete-time (DT)$\boldsymbol \Sigma \boldsymbol \Delta $modulators ($\boldsymbol \Sigma \boldsymbol \Delta $Ms) to achieve higher linearity. By using the proposed vector filter (VF) with a poles-splitting technique, the root locus outside the unit circle of higher-order DEM can be eliminated, leading to the DAC mismatch-shaping stability will not be limited by the closed-loop poles of the vector DEM, thus a self-stabilizing vector DEM at higher order can be achieved. Besides, a complexity-reduced vector quantizer (VQ) using the segmented architecture is proposed. Compared to recent works, the proposed technique can achieve higher-order mismatch shaping and decouple the correlation between the mismatch errors and the input pattern, leading to stronger suppression of DAC mismatch and canceling in-band tones.
Ke Chang, Qian Xing, Guoliang Jia, Yang Pu, Yan Wang 0119, Yanlong Zhang, Guohe Zhang
IEEE Trans. Very Large Scale Integr. Syst.8
2023 Asymmetric Learned Image Compression With Multi-Scale Residual Block, Importance Scaling, and Post-Quantization Filtering
abstract
Recently, deep learning-based image compression has made significant progresses, and has achieved better rate-distortion (R-D) performance than the latest traditional method, H.266/VVC, in both MS-SSIM metric and the more challenging PSNR metric. However, a major problem is that the complexities of many leading learned schemes are too high. In this paper, we propose an efficient and effective image coding framework, which achieves similar R-D performance with lower complexity than the state of the art. First, we develop an improved multi-scale residual block (MSRB) that can expand the receptive field and capture global information more efficiently, which further reduces the spatial correlation of the latent representations. Second, an importance scaling network is introduced to directly scale the latents to achieve content-adaptive bit allocation without sending side information, which is more flexible than previous importance map methods. Third, we apply a post-quantization filter (PQF) to reduce the quantization error, motivated by the Sample Adaptive Offset (SAO) filter in video coding. Moreover, our experiments show that the performance of the system is less sensitive to the complexity of the decoder. Therefore, we design an asymmetric paradigm, in which the encoder employs three stages of MSRBs to improve the learning capacity, whereas the decoder only uses one stage of MSRB, which reduces the decoder complexity and still yields satisfactory performance. Experimental results show that compared to the state-of-the-art method, the encoding and decoding time of the proposed method are about 17 times faster, and the R-D performance is only reduced by about 1% on both Kodak and Tecnick-40 datasets, which is still better than H.266/VVC(4:4:4) and other leading learning-based methods. Our source code is publicly available athttps://github.com/fengyurenpingsheng.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jingning Han
IEEE Trans. Circuits Syst. Video Technol.5
2023 Learned Image Compression With Gaussian-Laplacian-Logistic Mixture Model and Concatenated Residual Modules
abstract
Recently deep learning-based image compression methods have achieved significant achievements and gradually outperformed traditional approaches including the latest standard Versatile Video Coding (VVC) in both PSNR and MS-SSIM metrics. Two key components of learned image compression are the entropy model of the latent representations and the encoding/decoding network architectures. Various models have been proposed, such as autoregressive, softmax, logistic mixture, Gaussian mixture, and Laplacian. Existing schemes only use one of these models. However, due to the vast diversity of images, it is not optimal to use one model for all images, even different regions within one image. In this paper, we propose a more flexible discretized Gaussian-Laplacian-Logistic mixture model (GLLMM) for the latent representations, which can adapt to different contents in different images and different regions of one image more accurately and efficiently, given the same complexity. Besides, in the encoding/decoding network design part, we propose a concatenated residual blocks (CRB), where multiple residual blocks are serially connected with additional shortcut connections. The CRB can improve the learning ability of the network, which can further improve the compression performance. Experimental results using the Kodak, Tecnick-100 and Tecnick-40 datasets show that the proposed scheme outperforms all the leading learning-based methods and existing compression standards including VVC intra coding (4:4:4 and 4:2:0) in terms of the PSNR and MS-SSIM. The source code is available at https://github.com/fengyurenpingsheng.
Haisheng Fu, Feng Liang 0001, Bing Li 0022, Jie Liang 0001, Guohe Zhang, Dong Liu 0002, Chengjie Tu, Jingning Han
IEEE Trans. Image Process.7
2022 Learned Image Compression with Inception Residual Blocks and Multi-Scale Attention Module
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jiangning Han
PCS5
2021 Efficient neural network using pointwise convolution kernels with linear phase constraint
Feng Liang 0001, Zhichao Tian, Ming Dong 0002, Shuting Cheng, Hai Li 0001, Yiran Chen 0001, Guohe Zhang
Neurocomputing8
2021 An extended context-based entropy hybrid modeling for image compression
Haisheng Fu, Feng Liang 0001, Qian Zhang 0082, Jie Liang 0001, Chengjie Tu, Guohe Zhang
Signal Process. Image Commun.7
2020 Variable-Rate Multi-Frequency Image Compression using Modulated Generalized Octave Convolution
abstract
In this proposal, we design a learned multi-frequency image compression approach that uses generalized octave convolutions to factorize the latent representations into high-frequency (HF) and low-frequency (LF) components, and the LF components have lower resolution than HF components, which can improve the rate-distortion performance, similar to wavelet transform. Moreover, compared to the original octave convolution, the proposed generalized octave convolution (GoConv) and octave transposed-convolution (GoTConv) with internal activation layers preserve more spatial structure of the information, and enable more effective filtering between the HF and LF components, which further improve the performance. In addition, we develop a variable-rate scheme using the Lagrangian parameter to modulate all the internal feature maps in the autoencoder, which allows the scheme to achieve the large bitrate range of the JPEG AI with only three models. Experiments show that the proposed scheme achieves much better Y MS-SSIM than VVC. In terms of YUV PSNR, our scheme is very similar to HEVC.
Haisheng Fu, Qian Zhang 0082, Shang Wang 0006, Jie Liang 0001, Dong Liu 0002, Feng Liang 0001, Guohe Zhang, Chengjie Tu
MMSP9
2020 INOR - An Intelligent noise reduction method to defend against adversarial audio examples
Qingli Guo, Jing Ye 0001, Yiran Chen 0001, Yu Hu 0001, Yazhu Lan, Guohe Zhang, Xiaowei Li 0001
Neurocomputing6
2020 A low-cost and high-speed hardware implementation of spiking neural network
Guohe Zhang, Bing Li 0022, Jianxing Wu, Ran Wang 0007, Yazhu Lan, Shaochong Lei, Hai Li 0001, Yiran Chen 0001
Neurocomputing1
2020 FCDM: A Methodology Based on Sensor Pattern Noise Fingerprinting for Fast Confidence Detection to Adversarial Attacks
abstract
Deep neural networks (DNNs) have shown phenomenal success in many real-world applications. However, a concerning weakness of DNNs is their vulnerability to adversarial attacks. Although there exist some methods to detect adversarial attacks, they often suffer from high computational cost and constraints on certain types of attacks, and ignore external features that could aid during attack detection. In this article, we propose fast confidence detection method (FCDM), an innovative method for fast confidence detection of adversarial attacks based on measuring the integrity of sensor pattern noise fingerprinting embedded in input examples. We note that the existing adversarial detectors are often designed as a binary classifier to differentiate clean or adversarial examples. However, the detection of adversarial examples can be much more complicated than such a scenario. Our key insight is that the confidence level of detecting an input sample as an adversarial example is a more useful info for the system to properly take an action to resist potential attacks. The experimental results show that FCDM is capable to give a confidence distribution model of the most popular adversarial attacks. And, using the confidence distribution model, FCDM can quickly determine the confidence level of the input sample. Based on different properties of the confidence distribution models associated with these adversarial attacks, FCDM can provide early attack warning including even the possible attack types of the adversarial attack examples. FCDM also has the following advantages: 1) it is effective for both a white-box attack and black-box attack; 2) it do not depend on the class of adversarial attacks and can be used as both known attack defense and unknown attack defense; and 3) it does not need to know the details of the DNN model and does not affect the functionality of the DNN. Since fast confidence detection method (FCDM) is a computationally heavy task, we propose an FPGA-based accelerator based on a series of optimization techniques, such as the quantization, data reuse and operation replacement, etc. We implement our method on an FPGA platform and achieve a system clock frequency of 279 MHz with a power consumption of the only 0.7626 W. Moreover, in the real system performance test, we obtain a high efficiency of 29.740 IPS/W and a low latency of just 44.1 ms with very marginal accuracy loss.
Yazhu Lan, Kent W. Nixon, Qingli Guo, Guohe Zhang, Yuanchao Xu 0002, Hai Li 0001, Yiran Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Fast Confidence Detection: One Hot Way to Detect Adversarial Attacks via Sensor Pattern Noise Fingerprinting
abstract
Deep Neural Networks (DNNs) have shown phenomenal success in a wide range of real-world applications. However, a concerning weakness of DNNs is that they are vulnerable to adversarial attacks. Although there exist methods to detect adversarial attacks, they often suffer constraints on specific attack types and provide limited information to downstream systems. We specifically note that existing adversarial detectors are often binary classifiers, which differentiate clean or adversarial examples. However, detection of adversarial examples is much more complicated than such a scenario. Our key insight is that the confidence probability of detecting an input sample as an adversarial example will be more useful for the system to properly take action to resist potential attacks. In this work, we propose an innovative method for fast confidence detection of adversarial attacks based on integrity of sensor pattern noise embedded in input examples. Experimental results show that our proposed method is capable of providing a confidence distribution model of most of popular adversarial attacks. Furthermore, our presented method can provide early attack warning with even the attack types based on different properties of the confidence distribution models. Since fast confidence detection is a computationally heavy task, we propose an FPGA-Based hardware architecture based on a series of optimization techniques, such as incremental multi-level quantization and etc. We realize our proposed method on an FPGA platform and achieve a high efficiency of 29.740 IPS/W with a power consumption of only 0.7626W.
Yazhu Lan, Qingli Guo, Guohe Zhang, Yuanchao Xu 0002, Kent W. Nixon, Hai Li 0001, Yiran Chen 0001
FPGA3
2019 PUFPass: A password management mechanism based on software/hardware codesign
Qingli Guo, Jing Ye 0001, Bing Li 0017, Yu Hu 0001, Xiaowei Li 0001, Yazhu Lan, Guohe Zhang
Integr.7
2013 Test Patterns of Multiple SIC Vectors: Theory and Application in BIST Schemes
abstract
This paper proposes a novel test pattern generator (TPG) for built-in self-test. Our method generates multiple single-input change (MSIC) vectors in a pattern, i.e., each vector applied to a scan chain is an SIC vector. A reconfigurable Johnson counter and a scalable SIC counter are developed to generate a class of minimum transition sequences. The proposed TPG is flexible to both the test-per-clock and the test-per-scan schemes. A theory is also developed to represent and analyze the sequences and to extract a class of MSIC sequences. Analysis results show that the produced MSIC sequences have the favorable features of uniform distribution and low input transition density. The performances of the designed TPGs and the circuits under test with 45 nm are evaluated. Simulation results with ISCAS benchmarks demonstrate that MSIC can save test power and impose no more than 7.5% overhead for a scan design. It also achieves the target fault coverage without increasing the test length.
Feng Liang 0001, Shaochong Lei, Guohe Zhang, Kaile Gao
IEEE Trans. Very Large Scale Integr. Syst.4