Zhengning Wang

dblp:09/1064 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-4218-164XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Enhanced multimodal MRI classification of schizophrenia through cross-attention graph neural networks
Maomin Qian, Zhengning Wang, Weiyang Shi, Yunchun Chen, Huaning Wang, Wenming Liu, Yongfeng Yang, Ping Wan, Luxian Lv, Yuqing Song, Yuhui Du, Xiufeng Xu, Tianzai Jiang
Medical Image Anal.4
2026 Learning Efficient Meshflow and Optical Flow From Event Cameras
abstract
In this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively.
Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 LPUDC: A Laplacian pyramid neural network for restoring images from under-display cameras
Zhenyan Ding, Zhixiang Fang, Zhengning Wang, Lehan Ding, Zhenni Zeng, Xinglong Luo, Shaoqin Yuan, Binquan Leng
Neurocomputing3
2025 Fine-scale striatal parcellation using diffusion MRI tractography and graph neural networks
Mingqi Liu, Maomin Qian, Heping Tang, Zhengning Wang, Fengmei Lu
Medical Image Anal.9
2025 MTDA-STGCN: Modern Temporal and Dual-Attention-Based Spatiotemporal Graph Convolutional Network for 4D Trajectory Prediction
abstract
Four-dimensional (4D) trajectory prediction plays a critical role in modern air traffic management, enabling applications such as conflict detection, anomaly monitoring, and congestion mitigation. However, existing methods have limited information sources when modeling potential spatial correlations between aircraft in complex airspace scenarios, and their final trajectory inference ability is weak, resulting in lower prediction accuracy. Faced with these challenges, we propose Modern Temporal and Dual Attention based Spatiotemporal Graph Convolutional Network (MTDA-STGCN), which employs a self-attention mechanism to reconstruct the adjacency matrix to enhance the ability of capturing global node correlations. This adjacency matrix reconstructed with the self-attention mechanism is dynamically optimized throughout the training process of network, offering a more nuanced reflection of the inter-node relationships compared to traditional algorithms. Subsequently, our model uses graph attention to extract additional global features for modeling accuracy interactions between aircraft. Finally, the output is input into the Modern Temporal Prediction Network (MTPN) to obtain the predicted trajectory probability distribution. The experiments on real-world ADS-B datasets demonstrate that MTDA-STGCN outperforms existing 4D trajectory prediction algorithms on all datasets. The proposed dual-attention framework significantly enhances the capture of node spatial correlations, while the MTPN module effectively improves the accuracy of the predicted results.
Yuheng Kuang, Shuxuan Yuan, Yuding Zhang, Fanman Meng, Zhengning Wang
IEEE Trans. Intell. Transp. Syst.8
2025 An Energy-Efficient Block-Based Nonmaximum Suppression Engine for High-Parallel Postprocessing of Visual Object Detection
abstract
Nowadays, visual object detection (VOD) is widely used in many AI applications, such as autonomous driving, intelligent robotics, and smart surveillance. As an essential postprocessing step in VOD, nonmaximum suppression (NMS) is employed to generate bounding boxes as detection results. However, NMS is difficult to parallelize and computationally intensive, resulting in high processing latency and energy consumption. To address this issue, this brief proposes an energy-efficient block-based NMS engine that incorporates both algorithm- and hardware-level design techniques to improve processing speed and energy efficiency. These techniques include a block-NMS scheme, an adaptive hybrid sorting architecture (AHSA), and a reconfigurable pipeline-based block-NMS computation architecture. The proposed engine is implemented in 28-nm CMOS technology. Compared with the state-of-the-art designs, it achieves the highest performance (237.97 GOPS) and energy efficiency (8.71 TOPS/W), while delivering results fully equivalent to those of the original NMS.
Yuchuan Gong, Haojie Wei, Hongtao Guo, Jiahao Zheng 0004, Qingyuan Hou, Zherong Liu, Jingxiao Zheng, Ye Liu 0011, Zhengning Wang, Jun Zhou 0017
IEEE Trans. Very Large Scale Integr. Syst.11
2024 Efficient Meshflow and Optical Flow Estimation from Event Cameras
abstract
In this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280×720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (39× faster) of our EEMFlow model compared to recent state-of-the-art flow methods. Our code is available at https://github.com/boomluo02/EEMFlow.
Xinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin, Bing Zeng 0001, Shuaicheng Liu
CVPR3
2024 Reconstruction flow recurrent network for compressed video quality enhancement
Zhengning Wang, Xuhang Liu, Chuan Wang 0001, Ting Jiang 0005, Tianjiao Zeng, Zhenni Zeng, Guoqing Wang 0001, Shuaicheng Liu
Pattern Recognit.1
2023 Learning Optical Flow from Event Camera with Rendered Dataset
abstract
We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted foreground objects. The former case can produce real event values but with calculated flow labels, which are sparse and inaccurate. The latter case can generate dense flow labels but the interpolated events are prone to errors. In this work, we propose to render a physically correct event-flow dataset using computer graphics models. In particular, we first create indoor and outdoor 3D scenes by Blender with rich scene content variations. Second, diverse camera motions are included for the virtual capturing, producing images and accurate flow labels. Third, we render high-framerate videos between images for accurate events. The rendered dataset can adjust the density of events, based on which we further introduce an adaptive density module (ADM). Experiments show that our proposed dataset can facilitate event-flow learning, whereas previous approaches when trained on our dataset can improve their performances constantly by a relatively large margin. In addition, event-flow pipelines when equipped with our ADM can further improve performances. Our code is available at https://github.com/boomluo02/ADMFlow.
Xinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang, Ping Tan 0002, Shuaicheng Liu
ICCV4
2023 Stereo RGB and Deeper LIDAR-Based Network for 3D Object Detection in Autonomous Driving
abstract
3D object detection has become an emerging task in autonomous driving scenarios. Most of previous works process 3D point clouds using either projection-based or voxel-based models. However, both approaches contain some drawbacks. The voxel-based methods lack semantic information, while the projection-based methods suffer from numerous spatial information loss when projected to different views. In this paper, we propose the Stereo RGB and Deeper LIDAR (SRDL) framework which can utilize semantic and spatial information simultaneously such that the performance of network for 3D object detection can be improved naturally. Specifically, the network generates candidate boxes from stereo pairs and combines different region-wise features using a deep fusion scheme. The stereo strategy offers more information for prediction compared with prior works. Then, several local and global feature extractors are stacked in the segmentation module to capture richer deep semantic geometric features from point clouds. After aligning the interior points with fused features, the proposed network refines the prediction in a more accurate manner and encodes the whole box in a novel compact method. The decent experimental results on the challenging KITTI detection benchmark demonstrate the effectiveness of utilizing both stereo images and point clouds for 3D object detection.
Qingdong He, Zhengning Wang, Yijun Liu 0012, Shuaicheng Liu, Bing Zeng 0001
IEEE Trans. Intell. Transp. Syst.2
2022 SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds
abstract
Accurate 3D object detection from point clouds has become a crucial component in autonomous driving. However, the volumetric representations and the projection methods in previous works fail to establish the relationships between the local point sets. In this paper, we propose Sparse Voxel-Graph Attention Network (SVGA-Net), a novel end-to-end trainable network which mainly contains voxel-graph module and sparse-to-dense regression module to achieve comparable 3D detection tasks from raw LIDAR data. Specifically, SVGA-Net constructs the local complete graph within each divided 3D spherical voxel and global KNN graph through all voxels. The local and global graphs serve as the attention mechanism to enhance the extracted features. In addition, the novel sparse-to-dense regression module enhances the 3D box estimation accuracy through feature maps aggregation at different levels. Experiments on KITTI detection benchmark and Waymo Open dataset demonstrate the efficiency of extending the graph representation to 3D object detection and the proposed SVGA-Net can achieve decent detection accuracy.
Qingdong He, Zhengning Wang, Yijun Liu 0012
AAAI2
2022 DeepOIS: Gyroscope-Guided Deep Optical Image Stabilizer Compensation
abstract
Mobile captured images can be aligned using their gyroscope sensors. Optical image stabilizer (OIS) terminates this possibility by adjusting the images during the capturing. In this work, we propose a deep network that compensates for the motions caused by the OIS, such that the gyroscopes can be used for image alignment on the OIS cameras. To achieve this, we first record both videos and gyroscope readings with an OIS camera as training data. Then, we convert gyroscope readings into motion fields. Second, we propose an Essential Mixtures motion model for rolling shutter cameras, where an array of rotations within a frame are extracted as the ground-truth guidance. Third, we train a convolutional neural network with gyroscope motions as input to compensate for the OIS motion. Once finished, the compensation network can be applied for other scenes, where the image alignment is purely based on gyroscopes with no need for images contents, delivering strong robustness. Experiments show that our results are comparable with that of non-OIS cameras, and outperform image-based alignment results with a relatively large margin. Code and dataset is available at:https://github.com/lhaippp/DeepOIS.
Shuaicheng Liu, Haipeng Li 0001, Zhengning Wang, Jue Wang 0001, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 PD-GAN: Perceptual-Details GAN for Extremely Noisy Low Light Image Enhancement
abstract
Extremely noisy low light enhancement suffers from high-level noise, loss of texture detail, and color degradation. When recovering color or illumination for images taken in a dark environment, the challenge for networks is how to balance the enhancement for noise and texture details for a good visual effect. A single network is not suitable for solving the ill-posed problem of mapping the input image's noise to the clear target in the ground truth. To solve the problems, we pro-pose perceptual-details GAN (PD-GAN) utilizing Zero-DCE to initially recover illumination and combine residual dense-block Encoder-Decoder structure to suppress noise while finely adjusting the illumination. Besides, fractional differential gradient masks are integrated into the discriminator to enhance details. Experiment results demonstrate that PD-GAN outperforms other methods on the extremely low-light image dataset.
Yijun Liu 0012, Zhengning Wang, Deming Zhao
ICASSP2
2021 A Detection and Tracking Combined Network for Long-Term Tracking
Zhengning Wang, Deming Zhao, Yijun Liu 0012
ICIG (3)2
2021 OAENet: Oriented attention ensemble for accurate facial expression recognition
Zhengning Wang, Fanwei Zeng, Shuaicheng Liu, Bing Zeng 0001
Pattern Recognit.1
2020 Structure-preserving extremely low light image enhancement with fractional order differential mask guidance
abstract
Low visibility and high-level noise are two challenges for low-light image enhancement. In this paper, by introducing fractional order differential, we propose an end-to-end conditional generative adversarial network(GAN) to solve those two problems. For the problem of low visibility, we set up a global discriminator to improve the overall reconstruction quality and restore brightness information. For the high-level noise problem, we introduce fractional order differentiation into both the generator and the discriminator. Compared with conventional end-to-end methods, fractional order can better distinguish noise and high-frequency details, thereby achieving superior noise reduction effects while maintaining details. Finally, experimental results show that the proposed model obtains superior visual effects in low-light image enhancement. By introducing fractional order differential, we anticipate that our framework will enable high quality and detailed image recovery not only in the field of low-light enhancement but also in other fields that require details.
Yijun Liu 0012, Zhengning Wang, Ruixu Geng
MMAsia2
2020 Robust Heart Rate Monitoring for Quasi-Periodic Motions by Wrist-Type PPG Signals
abstract
Heart rate (HR) monitoring using photoplethysmography (PPG) is a promising feature in modern wearable devices. PPG is easily contaminated by motion artifacts (MA), hindering estimation of HR. For quasi-periodic motions, previous works generally focused on a few specific motions, such as walking and fast running. However, they may not work well for many different quasi-periodic motions where MA are very complex. In this paper, a robust HR monitoring scheme for different quasi-periodic motions using wrist-type PPG is proposed, which consists of dictionary learning for signal characteristics learning, human motion recognition for the current motion recognition and dictionary selection, sparse representation-based MA elimination for denoising, and spectral peak tracking for HR-related spectral peak tracking. The proposed scheme is robust to MA caused by different motions and has high accuracy. Experiments on six common quasi-periodic motions showed that the average absolute error of heart rate estimation was 2.40 beat per minute, and also showed that the proposed method is more robust than some state-of-the-art approaches for different motions.
Wenwen He, Yalan Ye, Li Lu 0001, Yunfei Cheng, Yunxia Li, Zhengning Wang
IEEE J. Biomed. Health Informatics6
2019 A Light-Weighted Network for Facial Landmark Detection via Combined Heatmap and Coordinate Regression
abstract
3D facial landmark, which offers more expressive and occlusive information than its 2D counterpart, receives more and more attention from researchers in recent years. The top performing algorithms for 3D facial landmark detection are mainly divided into two categories: the two-step approach and the volume representation method. However, the former lacks the relation of depth with plat and the latter leads to large computation. In this paper, we propose the Combined Heatmap and Coordinate Regression (CHCR), which is an end-to-end method for 3D facial landmark detection from a single 2D image. To achieve that, we innovatively present the combined heatmap of three channels, and each channel of heatmap records the likelihoods of landmarks location of any two different axes. Such representation maintains the relation of various view while decreases either the channels for encoding or dimension for decoding. Then we retrieve the 3D coordinate vectors from corresponding combined heatmap by coordinate regression. Hence, an encoder-decoder network with a simple CNN attached to an hourglass module is designed to cope with the whole process. Experiments show our model is extremely light-weighted and runs faster than any other methods while on high performance of accuracy.
Zhengning Wang, Longfei Feng, Fanwei Zeng, Xia Lv, Fengjun Zhang
ICME1
2018 A Fractional-Order Variational Framework for Retinex: Fractional-Order Partial Differential Equation-Based Formulation for Multi-Scale Nonlocal Contrast Enhancement with Texture Preserving
abstract
This paper discusses a novel conceptual formulation of the fractional-order variational framework for retinex, which is a fractional-order partial differential equation (FPDE) formulation of retinex for the multi-scale nonlocal contrast enhancement with texture preserving. The well-known shortcomings of traditional integer-order computation-based contrast-enhancement algorithms, such as ringing artefacts and staircase effects, are still in great need of special research attention. Fractional calculus has potentially received prominence in applications in the domain of signal processing and image processing mainly because of its strengths like long-term memory, nonlocality, and weak singularity, and because of the ability of a fractional differential to enhance the complex textural details of an image in a nonlinear manner. Therefore, in an attempt to address the aforementioned problems associated with traditional integer-order computation-based contrast-enhancement algorithms, we have studied here, as an interesting theoretical problem, whether it will be possible to hybridize the capabilities of preserving the edges and the textural details of fractional calculus with texture image multi-scale nonlocal contrast enhancement. Motivated by this need, in this paper, we introduce a novel conceptual formulation of the fractional-order variational framework for retinex. First, we implement the FPDE by means of the fractional-order steepest descent method. Second, we discuss the implementation of the restrictive fractional-order optimization algorithm and the fractional-order Courant-Friedrichs-Lewy condition. Third, we perform experiments to analyze the capability of the FPDE to preserve edges and textural details, while enhancing the contrast. The capability of the FPDE to preserve edges and textural details is a fundamental important advantage, which makes our proposed algorithm superior to the traditional integer-order computation-based contrast enhancement algorithms, especially for images rich in textural details.
Yi-Fei Pu, Patrick Siarry, Amitava Chatterjee, Zhengning Wang, Zhang Yi 0001, Yiguang Liu, Jiliu Zhou, Yan Wang 0015
IEEE Trans. Image Process.4
2017 Long-Distance/Environment Face Image Enhancement Method for Recognition
Zhengning Wang, Shanshan Ma, Mingyan Han, Shuaicheng Liu
ICIG (1)1
2017 Uncertain Region Identification for Stereoscopic Foreground Cutout
Taotao Yang, Shuaicheng Liu, Zhengning Wang, Bing Zeng 0001
ICIG (3)4
2016 Automatic Reflection Removal using Gradient Intensity and Motion Cues
abstract
We present a method to separate the background image and reflection from two photos that are taken in front of a transparent glass under slightly different viewpoints. In our method, the SIFT-flow between two images is first calculated and a motion hierarchy is constructed from the SIFT-flow at multiple levels of spatial smoothness. To distinguish background edges and reflection edges, we calculate a motion score for each edge pixel by its variance along the motion hierarchy. Alternatively, we make use of the so-called superpixels to group edge pixels into edge segments and calculate the motion scores by averaging over each segment. In the meantime, we also calculate an intensity score for each edge pixel by its gradient magnitude. We combine both motion and intensity scores to get a combination score. A binary labelling (for separation) can be obtained by thresholding the combination scores. The background image is finally reconstructed from the separated gradients. Compared to the existing approaches that require a sequence of images or a video clip for the separation, we only need two images, which largely improves its feasibility. Various challenging examples are tested to validate the effectiveness of our method.
Shuaicheng Liu, Taotao Yang, Bing Zeng 0001, Zhengning Wang, Guanghui Liu 0001
ACM Multimedia5
2016 Visible-light and near-infrared face recognition at a distance
Chun-Ting Huang, Zhengning Wang, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2013 Saliency detection using a central stimuli sensitivity based model
abstract
In this paper, a novel method is proposed to predict attention in image scenes by using a central stimuli sensitivity based saliency model. The proposed method is based on the general “center-surround” visual attention mechanism and the spatial frequency response of the human visual system (HVS). Following three biologically inspired principles, the saliency value is computed by two “scatter matrices” which are used to measure the similarity and distinctness within and between two classes, i.e., the center and surrounding regions, respectively. In order to detect salient objects with different size, the saliency of a pixel is estimated via the saliency support region of the pixel, which is the most salient region centered at the pixel with respect to the surrounding region. The proposed method which is compliant with human perceptual characteristics enables the prediction of human fixations. Experimental results on three eye tracking datasets verify the effectiveness of the method and show that the proposed method outperforms the state-of-the-art methods on the visual saliency detection task.
Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, Zhengning Wang, Guanghui Liu 0001
ISCAS4
2012 Saliency detection from joint embedding of spatial and color cues
abstract
Visual saliency detection provides an important methodology for many computer vision applications. In this paper, we propose a novel method to detect salient regions from an image. To detect pixel-level saliency, this method uses joint embedding of spatial and color cues, i.e., spatial constraint based saliency, color double-opponent saliency, and similarity distribution based saliency. Finally, a multi-layer structure is adopted to merge the three terms into a saliency map. In order to make the saliency map consistent, we perform a region-based saliency detection by incorporating a multi-scale segmentation technique. The proposed method was evaluated on the MSRA benchmark images. Experimental results show that our method outperforms the state-of-the-art methods on visual saliency detection by achieving both higher precision and better recall.
Linfeng Xu 0001, Hongliang Li 0001, Zhengning Wang
ISCAS3
2007 A Fast Transform Domain Based Algorithm for H.264/AVC Intra Prediction
abstract
Directional intra prediction is one of the new features adopted in H.264/AVC standard. In contrast to some previous coding standards, the prediction is performed in the spatial domain. Two types of intra prediction for luma - Intra16times16 and Intra4times4 are included in each profile. For high profile, there is an additional type - Intra8times8. Intra 16x16 and Intra4times4 support four and nine modes respectively. As a result, the encoder complexity is increased dramatically. In this paper we proposed a fast intra prediction algorithm using transform domain features of the target block to filter out the majority of candidate modes. Extensive simulations verify that the proposed method speeds up the intra prediction process by 67.5% on average without sacrificing the picture quality and compression ratio. A comparison between the proposed algorithm and others is also provided.
Zhengning Wang, Jun Yang 0005, Qiang Peng, Changqian Zhu
ICME1
2005 Residual Texture Based Fast Block-Size Selection for Inter-Frame Coding in H.264/AVC
abstract
One of the new features adopted in H.264/AVC is the utilization of flexible block size ranging from 16x16 to 4x4 in inter-frame coding. The aim is to reduce the error due to fixed block size prediction within a macroblock. However, this feature requires extremely computational complexity. In this paper, we proposed a residual texture based fast block size selection algorithm for inter-frame coding. Firstly, we perform a motion estimation (ME) for a macroblock and get the residual; then we predict the block size by analyzing the residual texture. Extensive simulations verify that the proposed method speeds up the block-size selection procedure by 50% without sacrificing picture quality and compression ratio.
Zhengning Wang, Qiang Peng
PDCAT1