EDBT 2026 Demo / reviewers in the wild / expert
Haiyong Zheng
dblp:05/6770
· DBLP profile ↗
54ranked-venue papers
2as first author
38since 2021 · last 2027
0000-0002-8027-0734ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Hierarchical long-range relation-aware multi-head graph attention network for satellite earth observation requirement gap filling
Sutong Yang, Hongbing Ji, Haiyong Zheng |
Expert Syst. Appl. | 6 |
| 2026 | Structured Bitmap-to-Mesh triangulation for geometry-aware discretization of image-derived domainsabstractWe introduce a template-driven triangulation framework for embedding discrete boundaries into a regular triangular grid, enabling raster- or segmentation-derived domains to support structure-preserving, numerically stable PDE discretization. Unlike constrained Delaunay triangulation (CDT), which requires global connectivity updates, our method retriangulates only boundary-intersecting triangles, preserving the base mesh and enabling synchronization-free parallel execution. To ensure determinism and scalability, all local intersection patterns are classified under discrete equivalence and triangle symmetry, forming a finite symbolic lookup table mapping each case to a conflict-free retriangulation template. The resulting mesh is provably closed, angle-bounded, and compatible with cotangent-based discretizations and finite element methods. Numerical experiments – including elliptic and parabolic PDEs, signal interpolation, and structural evaluation – demonstrate fewer slivers, more equilateral elements, and greater geometric fidelity near complex boundaries. These properties make the framework well suited for real-time geometric analysis and physically grounded simulation over image-derived domains. Wei Feng 0001, Haiyong Zheng |
Graph. Model. | 2 |
| 2026 | A quantum neural network with built-in self-attention mechanism
Shangshang Shi, Ruimin Shang, Haiyong Zheng, Guoqiang Zhong 0001, Yongjian Gu |
Neurocomputing | 6 |
| 2026 | Covert Backscatter Communication With Multitags
Jiahao Liu 0008, Jihong Yu, Bohan Li 0005, Qian Li 0010, Haiyong Zheng |
IEEE Internet Things J. | 5 |
| 2026 | HCSMamba: A Hierarchical Causal Scanning State Space model guided by physical priors for underwater image enhancement
Wenshuo Jia, Na Tian, Youjia Shao, Haiyong Zheng, Wencang Zhao |
Knowl. Based Syst. | 4 |
| 2026 | Revisiting Face Forgery Detection: From Facial Representation to Forgery DetectionabstractFace Forgery Detection (FFD), or Deepfake detection, aims to determine whether a digital face is real or fake. Due to different face synthesis algorithms with diverse forgery patterns, FFD models often overfit specific patterns in training datasets, resulting in poor generalization to other unseen forgeries. Existing FFD methods primarily leverage pre-trained backbones with general image representation capabilities and fine-tune them to identify facial forgery cues. However, these backbones lack domain-specific facial knowledge and insufficiently capture complex facial features, thus hindering effective implicit forgery cue identification and limiting generalization. Therefore, it is essential to revisit FFD workflow across the pre-training and fine-tuning stages, achieving an elaborate integration from facial representation to forgery detection to improve generalization. Specifically, we develop an FFD-specific pre-trained backbone with superior facial representation capabilities through self-supervised pre-training on real faces. We then propose a competitive fine-tuning framework that stimulates the backbone to identify implicit forgery cues through a competitive learning mechanism. Moreover, we devise a threshold optimization mechanism that utilizes prediction confidence to improve the inference reliability. Comprehensive experiments demonstrate that our method achieves excellent performance in FFD and extra face-related tasks, i.e., presentation attack detection. Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | G2LFormer: Global-to-local token mixing transformer for blind image inpainting and beyond
Haoru Zhao, Zonghui Guo, Shishi Qiao, Zhaorui Gu, Junyu Dong, Haiyong Zheng |
Pattern Recognit. | 7 |
| 2026 | Structured Analytic Mappings for Point Set RegistrationabstractAbstract. We present an analytic approximation model for nonrigid point set registration, grounded in the multivariate Taylor expansion of vector-valued functions. By exploiting the algebraic structure of Taylor expansions, we construct a structured function space spanned by truncated basis terms, allowing smooth deformations to be represented with low complexity and explicit form. To estimate mappings within this space, we develop a quasi-Newton optimization algorithm that progressively lifts the identity map into higher-order analytic forms. This structured framework unifies rigid, affine, and nonlinear deformations under a single closed-form formulation, without relying on kernel functions or high-dimensional parameterizations. The proposed model is embedded into a standard ICP loop—using (by default) nearest-neighbor correspondences—resulting in Analytic-ICP, an efficient registration algorithm with quasi-linear time complexity. Experiments on 2D and 3D datasets demonstrate that Analytic-ICP achieves higher accuracy and faster convergence than classical methods such as CPD and TPS-RPM, particularly for small and smooth deformations. Wei Feng 0001, Tengda Wei, Haiyong Zheng |
SIAM J. Imaging Sci. | 3 |
| 2026 | Energy-Efficient Covert Communications for Underwater Acoustic Backscatter SystemsabstractIn this work, we explore energy-efficient covert underwater acoustic backscatter communications (EC-UABCom) within the framework of the Internet of Underwater Things (IoUT). Specifically, a passive buoy node covertly transmits acoustic information passively to a maritime receiver by reflecting incident acoustic carrier signals from an autonomous underwater vehicle (AUV) transmitter. Simultaneously, the system exploits the uncertainty of underwater noise to mask the covert backscatter information, thereby evading detection by a submarine warden. To optimize the covert strategy, we first derive the warden’s optimal power-detection threshold that minimizes the detection error probability, accounting for underwater noise uncertainty. In response to this optimal detection strategy, we propose an energy-efficient covert policy by jointly optimizing the AUV’s transmit power and the buoy’s reflection coefficient to meet the covertness constraint. We then conduct a performance analysis, providing a closed-form expression for the expected detection error probability at the warden and outage probability at the receiver, thereby revealing their inherent trade-off. Numerical simulations validate the effectiveness of our approach compared to the state-of-the-art methods, demonstrating that higher transmit power and reflection coefficient degrade covertness while improving covert rate performance. Jiahao Liu 0008, Jihong Yu, Haiyong Zheng, Bohan Li 0005, Qian Li 0010, Jianping An |
IEEE Trans. Commun. | 3 |
| 2026 | Hierarchical Text-Guided Hashing for Open-World Image RetrievalabstractRecent advancements in deep learning have led to significant achievements in hashing for image retrieval. However, existing methods primarily operate under the assumption that training and testing data share the same distribution, meaning that the categories in the training and test sets are identical. This assumption may not hold in real-world scenarios, potentially limiting the effectiveness of these methods. In this work, we investigate the performance of existing deep hashing methods on unseen category data during retrieval tests and find a considerable performance decline. To address this issue, we propose a Hierarchical Text-guided Hashing (HTH) framework to mitigate the performance degradation in open-world image retrieval. Specifically, our method is trained in a self-supervised learning (SSL) framework using automatically synthesized coarse-to-fine textual descriptions. By combining the strengths of SSL in learning discriminative low- and mid-level features with the semantic richness of hierarchically structured text, our approach aims to enhance the model’s ability to generalize across unseen categories and complex open-world settings. Technically, we elaborately design a local attention pooling module to fuse the local patch information. Furthermore, we propose both hierarchical and fine-grained alignment modules, respectively applied to the global and local vision-language representations at different semantic levels, guiding the hash encoding to fully understand the visual primitives and extract discriminative and generalizable semantic information from images. Under the newly established large-scale ImageNet-CoG open evaluation protocol, our method demonstrates significant improvements in generalization compared to state-of-the-art and also possesses enhanced performance across various other open-world retrieval datasets and scenarios. Shishi Qiao, Miaonan Chen, Haiyong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Fine-Grained Visual Classification via Adaptive Attention Quantization TransformerabstractVision transformer (ViT) has recently demonstrated remarkable performance in fine-grained visual classification (FGVC). However, most existing ViT-based methods often overlook the varied focus of different attention heads, in which heads that attend to nondiscriminative regions would dilute the discriminative signal crucial for FGVC. To address such issues, we propose a novel adaptive attention quantization transformer (A2QTrans) for FGVC to select the key discriminative features by analyzing the heads’ attention, which comprises three key modules: the adaptive quantization selection (AQS) module, the background elimination (BE) module, and the dynamic hybrid optimization (DHO) module. Specifically, the AQS module dynamically selects the most discriminative features in a data-driven manner by quantizing the attention scores across multiple attention heads with a global, learnable threshold. This process effectively filters out generally irrelevant information from nondiscriminative tokens, thus concentrating attention on important regions. To address the nondifferentiability inherent in updating this threshold during binarization, our AQS module employs a straight-through estimator (STE) for discrete optimization, enabling end-to-end gradient backpropagation. In addition, we utilize the prior that background regions usually do not contain meaningful information, and design the BE module to further calibrate the focus of the attention heads to the main objects in images. Finally, the DHO module adaptively optimizes and integrates the attentive results of the AQS and BE modules to achieve optimal classification performance. Extensive experiments conducted on four challenging FGVC benchmark datasets and three ViT variants demonstrate A2QTrans’s superior performance, achieving state-of-the-art (SOTA) results. The source code is available athttps://github.com/Lishixian0817/A2QTrans Shishi Qiao, Shixian Li, Haiyong Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Face Forgery Video Detection via Temporal Forgery Cue UnravelingabstractFace Forgery Video Detection (FFVD) is a critical yet challenging task in determining whether a digital facial video is authentic or forged. Existing FFVD methods typically focus on isolated spatial or coarsely fused spatiotemporal information, failing to leverage temporal forgery cues thus resulting in unsatisfactory performance. We strive to unravel these cues across three progressive levels: momentary anomaly, gradual inconsistency, and cumulative distortion. Accordingly, we design a consecutive correlate module to capture momentary anomaly cues by correlating interactions among consecutive frames. Then, we devise a future guide module to unravel inconsistency cues by iteratively aggregating historical anomaly cues and gradually propagating them into future frames. Finally, we introduce a historical review module that unravels distortion cues via momentum accumulation from future to historical frames. These three modules form our Temporal Forgery Cue Unraveling (TFCU) framework, sequentially highlighting spatial discriminative features by unraveling temporal forgery cues bidirectionally between historical and future frames. Extensive experiments and ablation studies demonstrate the effectiveness of our TFCU method, achieving state-of-the-art performance across diverse unseen datasets and manipulation methods. Code is available at https://github.com/zhenglab/TFCU. Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan |
CVPR | 4 |
| 2025 | LPAH: Language-Guided Patch Aggregation Hashing for Fine-Grained Image Retrieval
Shishi Qiao, Miaonan Chen, Haiyong Zheng |
PRCV (12) | 4 |
| 2025 | Context-aware mutual learning for blind image inpainting and beyond
Haoru Zhao, Zhaorui Gu, Haiyong Zheng |
Expert Syst. Appl. | 5 |
| 2025 | Image sterilization via overwriting
Mingyao Tan, Lei Huang 0010, Zhaorui Gu, Xiaodong Wang 0006, Haiyong Zheng |
Inf. Sci. | 6 |
| 2025 | Reference-then-supervision framework for infrared and visible image fusion
Guihui Li, Zhensheng Shi, Zhaorui Gu, Haiyong Zheng |
Pattern Recognit. | 5 |
| 2025 | Image inpainting via Multi-scale Adaptive Priors
Haoru Zhao, Haiyong Zheng |
Pattern Recognit. | 5 |
| 2025 | Token Aggregation and Selection Hashing for Efficient Underwater Image RetrievalabstractThe scarcity of large-scale annotated data poses significant challenges for underwater visual analysis. Deep hashing methods offer promising solutions for efficient large-scale image retrieval tasks due to their exceptional computational and storage efficiency. However, underwater images suffer from inherent degradation (e.g., low contrast, color distortion), complex background noise, and fine-grained semantic distinctions, severely hindering the discriminability of learned hash codes. To address these issues, we propose Token Aggregation and Selection Hashing (TASH), the first deep hashing framework specifically designed for underwater image retrieval. Built upon a teacher-student self-distillation Vision Transformer (ViT) architecture, TASH incorporates three key innovations: (1) An Underwater Image Augmentation (UIA) module that simulates realistic degradation patterns (e.g., color shifts) to augment the student branch's input, explicitly enhancing model robustness to the diverse distortions encountered underwater; (2) A Multi-layer Token Aggregation (MTA) module that fuses features across layers, capturing hierarchical contextual information crucial for overcoming low contrast and resolving ambiguities in degraded underwater scenes; and (3) An Attention-based Token Selection (ATS) module that dynamically identifies and emphasizes the most discriminative tokens, eliminating the effect of background noise and enabling extracting subtle yet critical visual cues for distinguishing fine-grained underwater species. The resulting discriminative real-valued features are compressed into compact binary codes via a dedicated hash layer. Extensive experiments on two underwater datasets demonstrate that TASH significantly outperforms state-of-the-art methods, establishing new benchmarks for efficient and accurate underwater image retrieval. Shishi Qiao, Benqian Lin, Guanren Bu, Haiyong Zheng |
IEEE Signal Process. Lett. | 5 |
| 2024 | Spherical Pseudo-Cylindrical Representation for Omnidirectional Image Super-resolutionabstractOmnidirectional images have attracted significant attention in recent years due to the rapid development of virtual reality technologies. Equirectangular projection (ERP), a naive form to store and transfer omnidirectional images, however, is challenging for existing two-dimensional (2D) image super-resolution (SR) methods due to its inhomogeneous distributed sampling density and distortion across latitude. In this paper, we make one of the first attempts to design a spherical pseudo-cylindrical representation, which not only allows pixels at different latitudes to adaptively adopt the best distinct sampling density but also is model-agnostic to most off-the-shelf SR methods, enhancing their performances. Specifically, we start by upsampling each latitude of the input ERP image and design a computationally tractable optimization algorithm to adaptively obtain a (sub)-optimal sampling density for each latitude of the ERP image. Addressing the distortion of ERP, we introduce a new viewport-based training loss based on the original 3D sphere format of the omnidirectional image, which inherently lacks distortion. Finally, we present a simple yet effective recursive progressive omnidirectional SR network to showcase the feasibility of our idea. The experimental results on public datasets demonstrate the effectiveness of the proposed method as well as the consistently superior performance of our method over most state-of-the-art methods both quantitatively and qualitatively. Dongwei Ren, Haiyong Zheng, Junyu Dong, Yee-Hong Yang |
AAAI | 5 |
| 2024 | Video Harmonization with Triplet Spatio-Temporal Variation PatternsabstractVideo harmonization is an important and challenging task that aims to obtain visually realistic composite videos by automatically adjusting the foreground's appearance to harmonize with the background. Inspired by the short-term and long-term gradual adjustment process of manual har-monization, we present a Video Triplet Transformer frame-work to model three spatio-temporal variation patterns within videos, i.e., short-term spatial as well as long-term global and dynamic, for video-to-video tasks like video har-monization. Specifically, for short-term harmonization, we adjust foreground appearance to consist with background in spatial dimension based on the neighbor frames; for long-term harmonization, we not only explore global ap-pearance variations to enhance temporal consistency but also alleviate motion offset constraints to align similar con-textual appearances dynamically. Extensive experiments and ablation studies demonstrate the effectiveness of our method, achieving state-of-the-art performance in video harmonization, video enhancement, and video demoireing tasks. We also propose a temporal consistency metric to better evaluate the harmonized videos. Code is available at https://github.com/zhenglablVideoTripletTransformer. Zonghui Guo, Jie Zhang 0071, Shiguang Shan, Haiyong Zheng |
CVPR | 5 |
| 2024 | Multidomain Transfer Ensemble Learning for Wireless Fingerprinting LocalizationabstractMultidomain localization has emerged as an important learning paradigm for wireless fingerprinting localization, which leverages data from multiple related domains, known as source domains, to enhance location prediction accuracy in a target domain. However, it faces challenges due to label sparsity, feature heterogeneity, and domain correlation issues. Specifically, 1) intradomain classification-based localization models can only make predictions based on existing discrete labels, causing large errors in sparsely labeled areas; 2) feature heterogeneity across domains impedes effective interdomain model communication, limiting multidomain knowledge utilization; and 3) low correlation between some source domains and the target domain can result in negative transfer. In this article, we present multidomain transfer ensemble localization (MDTELoc) to tackle these challenges. To manage label sparsity and feature heterogeneity, we present a classification-to-regression (C2R) ensemble localization model. This model estimates continuous position values, addressing label sparsity, and fosters interdomain communication by sharing ensemble model weights, enhancing multidomain knowledge use. To mitigate the domain correlation issue, we design a multidomain transfer method to model and learn domain relationships via a domain covariance matrix, which handles both positive and negative domain correlations and identifies outlier domains. MDTELoc combines intradomain model ensemble and interdomain knowledge transfer in a Bayesian framework, allowing simultaneous estimation of model parameters and domain correlation. Experimental results from the DeepMIMO simulation data set and a real-world library data set demonstrate our framework’s superior performance over other state-of-the-art methods, reducing errors by over 12%. Lin Li 0028, Haiyong Zheng |
IEEE Internet Things J. | 2 |
| 2024 | Spatiotemporal self-supervised predictive learning for atmospheric variable prediction via multi-group multi-attention
Zhensheng Shi, Haiyong Zheng, Junyu Dong |
Knowl. Based Syst. | 2 |
| 2024 | Short-Term Earthquake Forecasting via Self-Supervised LearningabstractDue to the infrequency of major earthquakes (magnitude 5 or above) and the highly nonlinear nature of seismic activities, short-term earthquake forecasting (hours to weeks) faces the challenge of lacking seismic samples and struggling with acquiring seismic-related knowledge. To address these issues, we propose a novel self-supervised learning for earthquake forecasting (SSL-EF) approach, namely, self-supervised learning (SSL) for earthquake forecasting, which has two distinctive characteristics: 1) an efficient pretext task leveraging temporal signal patterns for knowledge acquisition and 2) the knowledge transfer mitigating the constraints posed by limited seismic samples. Specifically, we first design the prediction task as a pretext task, leveraging the past week’s observational data to predict the coming week’s data. Subsequently, we set the classification task as a downstream task, focusing on whether a major earthquake occurs in the coming week. Moreover, transferring knowledge from the prediction task facilitates training the classification model on a small-scale and balanced dataset obtained through undersampling. Compared with ten competing methods, our SSL-EF yields state-of-the-art area under the curve (AUC), with a 21.98% relative improvement on the real-world dataset. Zining Yu, Haiyong Zheng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Multi-object reconstruction of plankton digital holograms
Nan Wang 0013, Yanni Cui, Jia Yu 0016, Haiyong Zheng |
Multim. Tools Appl. | 7 |
| 2024 | Study on Geomagnetic Observations Associated With Three Major Earthquakes in Southwest China Through a Novel Deep Learning FrameworkabstractA novel method for detecting geomagnetic anomalies is proposed using a deep learning (DL) framework based on transformer architecture, which benefits from the self-attention mechanism therein that can assign attention weights to important features from the entire sequence. Given the rarity and variability of seismo-geomagnetic anomalies, it is difficult to establish correlations between anomalies and most data for seismic quiet periods. Thereby, we propose an anomaly transformer for the geomagnetism model (ATGM) with a two-branch structure to extract the sequence features and distribution features of anomalies. The first branch, termed sequence offset, measures whether geomagnetic sequences deviate from normal temporal dependencies. The second branch is distribution offset, to further assess the difference between the geomagnetic high-frequency data distribution and a prior distribution through its learnable kernel function. Finally, the ATGM is applied to analyze the preearthquake geomagnetic observations of the 2008 Wenchuan Ms 8.0 earthquake, the 2013 Lushan Ms 7.0 earthquake, and the 2014 Kangding Ms 6.5 earthquake. Significant anomalies could be extracted, occurring about 30–70 days before the earthquakes. Further comparing the sigmoidal-shaped growth in the cumulative numbers of preearthquake anomalies with the linear growth exhibited during normal periods and at distant stations, we can clearly distinguish the uniqueness and significance of the extracted anomalies. In addition, during the anomaly periods, the seismic activities and the strain observations around the epicenters also confirm the existence of abnormal underground activity preceding the earthquakes. Our proposed ATGM achieves a combined representation of manually designed and data-driven learned features, as well as develops preearthquake anomaly detection technology. Zining Yu, Xilong Jing, Haiyong Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Quantum recurrent neural networks for sequential learning
Rongbing Han, Shangshang Shi, Ruimin Shang, Haiyong Zheng, Guoqiang Zhong 0001, Yongjian Gu |
Neural Networks | 7 |
| 2023 | Transformer for Image Harmonization and BeyondabstractImage harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it should look like in the foreground referring to surrounding background. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our two vision Transformer frameworks and corresponding methods, as well as comprehensive experiments and empirical study, demonstrating the power of Transformer and investigating the Transformer for vision. Our methods achieve state-of-the-art performance on the image harmonization as well as four additional vision and graphics tasks, i.e., image enhancement, image inpainting, white-balance editing, and portrait relighting, indicating the superiority of our work. Code, models, more results and details can be found at the project website http://ouc.ai/project/HarmonyTransformer. Zonghui Guo, Zhaorui Gu, Junyu Dong, Haiyong Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | TransCNN-HAE: Transformer-CNN Hybrid AutoEncoder for Blind Image InpaintingabstractBlind image inpainting is extremely challenging due to the unknown and multi-property complexity of contamination in different contaminated images. Current mainstream work decomposes blind image inpainting into two stages: mask estimating from the contaminated image and image inpainting based on the estimated mask, and this two-stage solution involves two CNN-based encoder-decoder architectures for estimating and inpainting separately. In this work, we propose a novel one-stage Transformer-CNN Hybrid AutoEncoder (TransCNN-HAE) for blind image inpainting, which intuitively follows the inpainting-then-reconstructing pipeline by leveraging global long-range contextual modeling of Transformer to repair contaminated regions and local short-range contextual modeling of CNN to reconstruct the repaired image. Moreover, a Cross-layer Dissimilarity Prompt (CDP) is devised to accelerate the identifying and inpainting of contaminated regions. Ablation studies validate the efficacy of both TransCNN-HAE and CDP, and extensive experiments on various datasets with multi-property contaminations show that our method achieves state-of-the-art performance with much lower computational cost on blind image inpainting. Our code is available at https://github.com/zhenglab/TransCNN-HAE. Haoru Zhao, Zhaorui Gu, Haiyong Zheng |
ACM Multimedia | 4 |
| 2022 | Temporal Moment Localization via Natural Language by Utilizing Video Question Answers as a Special Variant and Bypassing NLP for CorporaabstractTemporal moment localization using natural language (TMLNL) is an emerging issue in computer vision for localizing a specific moment inside a long, untrimmed video. The goal of TMLNL is to obtain the video’s output moment, which is related to the input query in a substantial way. Previous research focused on the visual portion of TMLNL, such as objects, backdrops, and other visual attributes, but natural language processing (NLP) techniques were largely used for the textual portion. A long query requires sufficient context to properly localize moments within a long untrimmed video. Thus, as a consequence of not completely understanding how to handle queries, performances deteriorated, especially when the query was longer. In this paper, we treat the TMLNL challenge as a unique variation of VQA, which equally considers the visual elements by using our proposed VQA joint visual-textual framework (JVTF). However, we also manage complex and long input queries without employing natural language processing (NLP) by improving poorly graded to finely graded distinct granularity representations. Our suggested BCPN searches for insufficient context for long input queries using an approach called query handler (QH) and helps the JVTF find the most relevant moment. Previously, a recurrence of words was caused by increasing the number of encoding layers in transformers, LSTMs, and other NLP techniques; however, our QH ensured that repetition of word locations was reduced. The output of BCPN is combined with JVTF’s guided attention to further improve the end outcome. Therefore, we propose a novel bidirectional context predictor network (BCPN), in addition to a VQA joint visual-textual framework (JVTF), to address the equal importance of videos and queries. Through extensive experiments on three benchmark datasets, we show that the proposed BCPN outperforms the state-of-the-art methods by$IoU = 0.3 (2.65 \%) $,$IoU = 0.5 (2.49 \%)$, and$IoU = 0.7 (2.06 \%) $. Hafiza Sadia Nawaz, Zhensheng Shi, Yanhai Gan, Amanuel Hirpa Madessa, Junyu Dong, Haiyong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement With Multi-Scale PerceptionabstractDue to the wavelength-dependent light attenuation, refraction and scattering, underwater images usually suffer from color distortion and blurred details. However, due to the limited number of paired underwater images with undistorted images as reference, training deep enhancement models for diverse degradation types is quite difficult. To boost the performance of data-driven approaches, it is essential to establish more effective learning mechanisms that mine richer supervised information from limited training sample resources. In this paper, we propose a novel underwater image enhancement network, called SGUIE-Net, in which we introduce semantic information as high-level guidance via region-wise enhancement feature learning. Accordingly, we propose semantic region-wise enhancement module to better learn local enhancement features for semantic regions with multi-scale perception. After using them as complementary features and feeding them to the main branch, which extracts the global enhancement features on the original image scale, the fused features bring semantically consistent and visually superior enhancements. Extensive experiments on the publicly available datasets and our proposed dataset demonstrate the impressive performance of SGUIE-Net. The code and proposed dataset are available at https://trentqq.github.io/SGUIE-Net.html. Qi Qi 0008, Kunqian Li, Haiyong Zheng, Xiang Gao 0009, Guojia Hou, Kun Sun 0002 |
IEEE Trans. Image Process. | 3 |
| 2022 | Multi-Attention DenseNet: A Scattering Medium Imaging Optimization Framework for Visual Data Pre-Processing of Autonomous Driving SystemsabstractThe vision system is important for almost all kinds of autonomous driving systems. However, visual data interfered by scattering media, such as smoke, haze, water, and other non-uniform media will be degraded seriously, showing the characteristics of detail loss, poor contrast, low visibility, or color distortion. These characteristics can significantly interfere with the reliability of autonomous driving systems. In real environments the image degradation mechanism is complex, and the estimation of degradation parameters is difficult. This issue remains to be solved. In this study, we employed dense blocks as the framework and introduced the attention mechanism to our model from four dimensions: Multi-scale Attention, Channel Attention, Structure Attention, and ROI (region of interest) Attention. With the help of the training data provided by the weakly supervised model, the proposed method achieved excellent performance in the task of scattering medium imaging optimization in different scenes. Comparative experiments show that the proposed method is robust, and is superior to other state-of-the-art methods in image dehazing, and underwater image enhancement tasks. It is of great significance to improve the reliability of autonomous driving systems in underwater and severe weather environments. Peng Liu 0036, Chufeng Zhang, Haiyong Zheng |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | One-Shot Image-to-Image Translation via Part-Global Learning With a Multi-Adversarial FrameworkabstractIt is well known that humans can learn and recognize objects effectively from several limited image samples. However, learning from just a few images is still a tremendous challenge for existing main-stream deep neural networks. Inspired by analogical reasoning in the human mind, a feasible strategy is to “translate” the abundant images of a rich source domain to enrich the relevant yet different target domain with insufficient image data. To achieve this goal, we propose a novel, effective multi-adversarial framework (MA) based on part-global learning, which accomplishes the one-shot cross-domain image-to-image translation. In specific, we first devise a part-global adversarial training scheme to provide an efficient way for feature extraction and prevent discriminators from being overfitted. Then, a multi-adversarial mechanism is employed to enhance the image-to-image translation ability to unearth the high-level semantic representation. Moreover, a balanced adversarial loss function is presented, which aims to balance the training data and stabilize the training process. Extensive experiments demonstrate that the proposed approach can obtain impressive results on various datasets between two extremely imbalanced image domains and outperform state-of-the-art methods on one-shot image-to-image translation. Our code will be released with this paper athttps://github.com/zhengziqiang/OST. Ziqiang Zheng, Zhibin Yu 0002, Haiyong Zheng, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2021 | Intrinsic Image HarmonizationabstractCompositing an image usually inevitably suffers from inharmony problem that is mainly caused by incompatibility of foreground and background from two different images with distinct surfaces and lights, corresponding to material-dependent and light-dependent characteristics, namely, reflectance and illumination intrinsic images, respectively. Therefore, we seek to solve image harmonization via separable harmonization of reflectance and illumination, i.e., intrinsic image harmonization. Our method is based on an autoencoder that disentangles composite image into reflectance and illumination for further separate harmonization. Specifically, we harmonize reflectance through material-consistency penalty, while harmonize illumination by learning and transferring light from background to foreground, moreover, we model patch relations between foreground and background of composite images in an inharmony-free learning way, to adaptively guide our intrinsic image harmonization. Both extensive experiments and ablation studies demonstrate the power of our method as well as the efficacy of each component. We also contribute a new challenging dataset for benchmarking illumination harmonization. Code and dataset are at https://github.com/zhenglab/IntrinsicHarmony. Zonghui Guo, Haiyong Zheng, Zhaorui Gu |
CVPR | 2 |
| 2021 | Image Harmonization with TransformerabstractImage harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it looks like in the surrounding background near the foreground. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our harmonization Transformer frameworks without and with disentanglement, as well as comprehensive experiments and ablation study, demonstrating the power of Transformer and investigating the Transformer for vision. Our method achieves state-of-the-art performance on both image harmonization and image inpainting/enhancement, indicating its superiority. Our code and models are available at https://github.com/zhenglab/HarmonyTransformer. Zonghui Guo, Haiyong Zheng, Zhaorui Gu, Junyu Dong |
ICCV | 3 |
| 2021 | Painting from PartabstractThis paper studies the problem of painting the whole image from part of it, namely painting from part or part-painting for short, involving both inpainting and outpainting. To address the challenge of taking full advantage of both information from local domain (part) and knowledge from global domain (dataset), we propose a novel part-painting method according to the observations of relationship between part and whole, which consists of three stages: part-noise restarting, part-feature repainting, and part-patch refining, to paint the whole image by leveraging both feature-level and patch-level part as well as powerful representation ability of generative adversarial network. Extensive ablation studies show efficacy of each stage, and our method achieves state-of-the-art performance on both inpainting and outpainting benchmarks with free-form parts, including our new mask dataset for irregular outpainting. Our code and dataset are available at https://github.com/zhenglab/partpainting. Haoru Zhao, Yunhao Cheng, Haiyong Zheng, Zhaorui Gu |
ICCV | 4 |
| 2021 | Multi-Modal Multi-Action Video RecognitionabstractMulti-action video recognition is much more challenging due to the requirement to recognize multiple actions co-occurring simultaneously or sequentially. Modeling multi-action relations is beneficial and crucial to understand videos with multiple actions, and actions in a video are usually presented in multiple modalities. In this paper, we propose a novel multi-action relation model for videos, by leveraging both relational graph convolutional networks (GCNs) and video multi-modality. We first build multi-modal GCNs to explore modality-aware multi-action relations, fed by modality-specific action representation as node features, i.e., spatiotemporal features learned by 3D convolutional neural network (CNN), audio and textual embeddings queried from respective feature lexicons. We then joint both multi-modal CNN-GCN models and multi-modal feature representations for learning better relational action predictions. Ablation study, multi-action relation visualization, and boosts analysis, all show efficacy of our multi-modal multi-action relation modeling. Also our method achieves state-of-the-art performance on large-scale multi-action M-MiT benchmark. Our code is made publicly available at https://github.com/zhenglab/multi-action-video. Zhensheng Shi, Ju Liang, Haiyong Zheng, Zhaorui Gu, Junyu Dong |
ICCV | 4 |
| 2021 | Learning spectral normalized adversarial systems with stacked structure for high-quality 3D object generationabstractSummary This paper proposes a new method for generating 3D objects based on generative adversarial networks (GANs). Recently, GANs have been used in 3D object generation, but it is still very challenging to generate high‐quality 3D objects because of the complex data distribution over 3D objects. In this paper, we propose a system based on GAN that makes the generated objects more realistic. We use multiple generators and discriminators to enhance the ability of the model for learning complex distributions. Such a stacked structure can be considered as a coarse‐to‐fine or low‐to‐high–resolution mechanism. We employ the spectral normalization technology to control the Lipschitz constant of the discriminators by literally constraining the spectral norm of each layer to get a more stable training process. In this way, the proposed model can generate realistic and high‐quality 3D objects. Moreover, our system can also recover incomplete 3D objects into complete 3D objects. Experiments demonstrate that our model performs better in the quality of the generated objects than the baselines. Haoxu Zhang, Chenchen Qiu, Chao Wang 0022, Zhibin Yu 0002, Haiyong Zheng |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | Generative Adversarial Network with Multi-branch Discriminator for imbalanced cross-species image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yang Wu 0001, Haiyong Zheng, Minho Lee 0001 |
Neural Networks | 4 |
| 2020 | Spiral Generative Network for Image Extrapolation
Hongzhi Liu 0002, Haoru Zhao, Yunhao Cheng, Qingwei Song, Zhaorui Gu, Haiyong Zheng |
ECCV (19) | 7 |
| 2020 | CoTeRe-Net: Discovering Collaborative Ternary Relations in Videos
Zhensheng Shi, Cheng Guan, Liangjie Cao, Ju Liang, Zhaorui Gu, Haiyong Zheng |
ECCV (6) | 7 |
| 2020 | Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationabstractLearning spatiotemporal features is very effective but challenging for video understanding especially action recognition. In this paper, we propose Multi-Group Multi-Attention, dubbed MGMA, paying more attention to "where and when" the action happens, for learning discriminative spatiotemporal representation in videos. The contribution of MGMA is three-fold: First, by devising a new spatiotemporal separable attention mechanism, it can learn temporal attention and spatial attention separately for fine-grained spatiotemporal representation. Second, through designing a novel multi-group structure, it can capture multi-attention rendered spatiotemporal features better. Finally, our MGMA module is lightweight and flexible yet effective, so that can be easily embedded into any 3D Convolutional Neural Network (3D-CNN) architecture. We embed multiple MGMA modules into 3D-CNN to train an end-to-end, RGB-only model and evaluate on four popular benchmarks: UCF101 and HMDB51, Something-Something V1 and V2. Ablation study and experimental comparison demonstrate the strength of our MGMA, which achieves superior performance compared to state-of-the-arts. Our code is available at https://github.com/zhenglab/mgma. Zhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang, Zhaorui Gu, Haiyong Zheng |
ACM Multimedia | 7 |
| 2020 | Discriminative Region Proposal Adversarial Network for High-Quality Image-to-Image Translation
Chao Wang 0022, Wenjie Niu, Haiyong Zheng, Zhibin Yu 0002, Zhaorui Gu |
Int. J. Comput. Vis. | 4 |
| 2020 | KA-Ensemble: towards imbalanced image classification ensembling under-sampling and over-sampling
Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng |
Multim. Tools Appl. | 5 |
| 2020 | Depth map prediction from a single image with generative adversarial nets
Shaoyong Zhang, Chenchen Qiu, Zhibin Yu 0002, Haiyong Zheng |
Multim. Tools Appl. | 5 |
| 2020 | Fine-grained facial image-to-image translation with an attention based pipeline generative adversarial framework
Ziqiang Zheng, Chao Wang 0022, Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013 |
Multim. Tools Appl. | 7 |
| 2019 | Unpaired photo-to-caricature translation on faces in the wild
Ziqiang Zheng, Chao Wang 0022, Zhibin Yu 0002, Nan Wang 0013, Haiyong Zheng |
Neurocomputing | 5 |
| 2018 | Discriminative Region Proposal Adversarial Networks for High-Quality Image-to-Image Translation
Chao Wang 0022, Haiyong Zheng, Zhibin Yu 0002, Ziqiang Zheng, Zhaorui Gu |
ECCV (1) | 2 |
| 2018 | Teaching Squeeze-and-Excitation PyramidNet for Imbalanced Image Classification with GAN-based Curriculum LearningabstractImage classification with datasets that suffer from great imbalanced class distribution is a challenging task in computer vision field. In many real-world problems, the datasets are typically imbalanced and have a serious impact on the performance of classifiers. Although deep convolutional neural networks (DCNNs) have shown remarkable performance on image classification tasks in recent years, there are still few effective deep learning algorithms specifically for imbalanced image classification problems. To solve imbalanced image classification problem, in this paper, we explore a new deep learning algorithm called Squeeze-and-Excitation Deep Pyramidal Residual Network (SE-PyramidNet) combining with Generative Adversarial Network (GAN)-based curriculum learning. Firstly, we construct the refined Deep Pyramidal Residual Network by embedding the “Squeeze-and-Excitation” (SE) blocks. Secondly, towards the class imbalance problem, we adopt GAN to generate samples of minority classes. Finally, we draw lessons from the curriculum learning strategy by teaching our classifier training from original easy samples to generated complex samples, which improves the classification ability. Experimental results show that our method achieves around 0.5% gains for accuracy and 0.02 gains for F1 score respectively outperforming the state-of-the-art DCNNs. Angang Du, Chao Wang 0022, Haiyong Zheng, Nan Wang 0013 |
ICPR | 4 |
| 2018 | Unsupervised pixel-wise classification for Chaetoceros image segmentation
Fei Zhou 0007, Zhaorui Gu, Haiyong Zheng, Zhibin Yu 0002 |
Neurocomputing | 4 |
| 2017 | CGAN-plankton: Towards large-scale imbalanced class generation and fine-grained classificationabstractPlankton classification is becoming critically important as people concentrate more on oceans and global environment changing. Data of plankton species naturally exhibit imbalance in their class distribution. Meanwhile, it arouses fine-grained classification challenge. Although Convolutional Neural Networks (CNNs) have human-level performance on image classification task, they tend to be biased to large classes without considering the imbalance issue. In this paper, we introduce Generative Adversarial Network (GAN) based generative model to overcome these challenges. Our proposed model consists of fully convolutional layers, and includes three parts: a generative model G, a discriminative model D and a classification model C. We train generative and discriminative models on small classes data to learn a discriminative features through D model and reduce mode missing problem to some extent. We implement classification task using shared CNN layers of D model on whole data. Experimental results show that our model significantly improved the F1 score on an imbalanced plankton dataset with well-generated plankton images. Chao Wang 0022, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013 |
ICIP | 3 |
| 2017 | Automatic plankton image classification combining multiple view features via multiple kernel learningabstractBACKGROUND: Plankton, including phytoplankton and zooplankton, are the main source of food for organisms in the ocean and form the base of marine food chain. As the fundamental components of marine ecosystems, plankton is very sensitive to environment changes, and the study of plankton abundance and distribution is crucial, in order to understand environment changes and protect marine ecosystems. This study was carried out to develop an extensive applicable plankton classification system with high accuracy for the increasing number of various imaging devices. Literature shows that most plankton image classification systems were limited to only one specific imaging device and a relatively narrow taxonomic scope. The real practical system for automatic plankton classification is even non-existent and this study is partly to fill this gap. RESULTS: Inspired by the analysis of literature and development of technology, we focused on the requirements of practical application and proposed an automatic system for plankton image classification combining multiple view features via multiple kernel learning (MKL). For one thing, in order to describe the biomorphic characteristics of plankton more completely and comprehensively, we combined general features with robust features, especially by adding features like Inner-Distance Shape Context for morphological representation. For another, we divided all the features into different types from multiple views and feed them to multiple classifiers instead of only one by combining different kernel matrices computed from different types of features optimally via multiple kernel learning. Moreover, we also applied feature selection method to choose the optimal feature subsets from redundant features for satisfying different datasets from different imaging devices. We implemented our proposed classification system on three different datasets across more than 20 categories from phytoplankton to zooplankton. The experimental results validated that our system outperforms state-of-the-art plankton image classification systems in terms of accuracy and robustness. CONCLUSIONS: This study demonstrated automatic plankton image classification system combining multiple view features using multiple kernel learning. The results indicated that multiple view features combined by NLMKL using three kernel functions (linear, polynomial and Gaussian kernel functions) can describe and use information of features better so that achieve a higher classification accuracy. Haiyong Zheng, Ruchen Wang, Zhibin Yu 0002, Nan Wang 0013, Zhaorui Gu |
BMC Bioinform. | 1 |
| 2017 | Robust and automatic cell detection and segmentation from microscopic images of non-setae phytoplankton speciesabstractSaliency‐based marker‐controlled watershed method was proposed to detect and segment phytoplankton cells from microscopic images of non‐setae species. This method first improved IG saliency detection method by combining saturation feature with colour and luminance feature to detect cells from microscopic images uniformly and then produced effective internal and external markers by removing various specific noises in microscopic images for efficient performance of watershed segmentation automatically. The authors built the first benchmark dataset for cell detection and segmentation, including 240 microscopic images across multiple phytoplankton species with pixel‐wise cell regions labelled by a taxonomist, to evaluate their method. They compared their cell detection method with seven popular saliency detection methods and their cell segmentation method with six commonly used segmentation methods. The quantitative comparison validates that their method performs better on cell detection in terms of robustness and uniformity and cell segmentation in terms of accuracy and completeness. The qualitative results show that their improved saliency detection method can detect and highlight all cells, and the following marker selection scheme can remove the corner noise caused by illumination, the small noise caused by specks, and debris, as well as deal with blurred edges. Haiyong Zheng, Nan Wang 0013, Zhibin Yu 0002, Zhaorui Gu |
IET Image Process. | 1 |
| 2007 | A New Method of Steganalysis Based on Image Entropy
Xiaoyan Qiao, Guangrong Ji, Haiyong Zheng |
ICIC (3) | 3 |
| 2007 | An Approach for Image Compression of Algae Cell Using Multiwavelets
Guangrong Ji, Nengqiang Wang, Haiyong Zheng |
ICIC (3) | 4 |