EDBT 2026 Demo / reviewers in the wild / expert
Ziqiang Zheng
dblp:210/2302
· DBLP profile ↗
34ranked-venue papers
14as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 18 · 10 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent CuesabstractGlass is a prevalent material among solid objects in every-day life, yet segmentation methods struggle to distinguish it from opaque materials due to its transparency and reflection. While it is known that human perception relies on boundary and reflective-object features to distinguish glass objects, the existing literature has not yet sufficiently captured both properties when handling transparent objects. Hence, we propose incorporating both of these powerful visual cues via the Boundary Feature Enhancement and Reflection Feature Enhancement modules in a mutually beneficial way. Our proposed framework, TransCues, is a pyramidal transformer encoder-decoder architecture to segment transparent objects. We empirically show that these two modules can be used together effectively, improving overall performance across various benchmark datasets, including glass object semantic segmentation, mirror object semantic segmentation, and generic segmentation datasets. Our method outperforms the state-of-the-art by a large margin, achieving +4.2% mIoU on Trans10K-v2, +5.6% mIoU on MSD, +10.1% mIoU on RGBD-Mirror, +13.1% mIoU on TROSD, and +8.3% mIoU on Stanford2D3D, showing the effectiveness of our method against glass objects. Tuan-Anh Vu, Hai Nguyen-Truong, Ziqiang Zheng, Binh-Son Hua, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
WACV | 3 |
| 2026 | ORCA: Object Recognition and Comprehension for Archiving Marine SpeciesabstractMarine visual understanding is essential for monitoring and protecting marine ecosystems, enabling automatic and scalable biological surveys. However, progress is hindered by limited training data and the lack of a systematic task formulation that aligns domain-specific marine challenges with well-defined computer vision tasks, thereby limiting effective model application. To address this gap, we present ORCA, a multi-modal benchmark for marine research comprising 14,647 images from 478 species, with 42,217 bounding box annotations and 22,321 expert-verified instance captions. The dataset provides fine-grained visual and textual annotations that capture morphology-oriented attributes across diverse marine species. To catalyze methodological advances, we evaluate 18 state-of-the-art models on three tasks: object detection (closed-set and open-vocabulary), instance captioning, and visual grounding. Results highlight key challenges, including species diversity, morphological overlap, and specialized domain demands, underscoring the difficulty of marine understanding. ORCA thus establishes a comprehensive benchmark to advance research in marine domain. Yuk-Kwan Wong, Haixin Liang, Zeyu Ma 0002, Yiwei Chen 0003, Ziqiang Zheng, Rinaldi Gotama, Pascal Sebastian, Lauren D. Sparks, Sai-Kit Yeung |
WACV | 5 |
| 2026 | MarineEval: Assessing the Marine Intelligence of Vision-Language ModelsabstractWe have witnessed promising progress led by large language models (LLMs) and further vision language models (VLMs) in handling various queries as a general-purpose assistant. VLMs, as a bridge to connect the visual world and language corpus, receive both visual content and various text-only user instructions to generate corresponding responses. Though great success has been achieved by VLMs in various fields, in this work, we ask whether the existing VLMs can act as domain experts, accurately answering marine questions, which require significant domain expertise and address special domain challenges/requirements. To comprehensively evaluate the effectiveness and explore the boundary of existing VLMs, we construct the first large-scale marine VLM dataset and benchmark called MarineEval, with 2,000 image-based question-answering pairs. During our dataset construction, we ensure the diversity and coverage of the constructed data: 7 task dimensions and 20 capacity dimensions. The domain requirements are specially integrated into the data construction and further verified by the corresponding marine domain experts. We comprehensively benchmark 17 existing VLMs on our MarineEval and also investigate the limitations of existing models in answering marine research questions. The experimental results reveal that existing VLMs cannot effectively answer the domain-specific questions, and there is still a large room for further performance improvements. We hope our new benchmark and observations will facilitate future research. Yuk-Kwan Wong, Tuan-An To, Ziqiang Zheng, Sai-Kit Yeung |
WACV | 4 |
| 2026 | CamoVid60K: A Large-Scale Video Dataset for Moving Camouflaged Animals UnderstandingabstractAbstract We have been witnessing remarkable success led by the power of neural networks driven by a significant scale of training data in handling various computer vision tasks. However, less attention has been paid to monitoring the camouflaged animals, the masters of hiding themselves in the background. Robust and precise segmentation of camouflaged animals is challenging even for domain experts due to their similarity to the environment. Although several efforts have been made in camouflaged animal image segmentation, to the best of our knowledge, limited work exists on camouflaged animal video understanding (CAVU). Biologists often prefer videos for monitoring and understanding animal behaviors, as videos provide redundant information and temporal consistency. However, the scarcity of labeled video data significantly hinders progress in this area. To address these challenges, we present CamoVid60K , a diverse, large-scale, and accurately annotated video dataset of camouflaged animals. This dataset comprises 218 videos with 62,774 finely annotated frames, covering 70 animal categories, which surpasses all previous datasets in terms of the number of videos/frames and species included. CamoVid60K also offers more diverse downstream tasks in computer vision, such as camouflaged animal classification, detection, and task-specific segmentation (semantic, referring, motion), etc. We have benchmarked several state-of-the-art algorithms on the proposed CamoVid60K dataset, and the experimental results provide valuable insights for future research directions. Our dataset serves as a novel and challenging benchmark to stimulate the development of more powerful camouflaged animal video segmentation algorithms, with substantial room for further improvement. Tuan-Anh Vu, Ziqiang Zheng, Chengyang Song, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
Int. J. Comput. Vis. | 2 |
| 2025 | VideoDPO: Omni-Preference Alignment for Video Diffusion GenerationabstractRecent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user preferences, highlighting the need for preference alignment on pre-trained models. Although Direct Preference Optimization (DPO) [42] has demonstrated significant improvements in language and image generation [52], we pioneer its adaptation to video diffusion models and propose a VideoDPO pipeline by making several key adjustments. Unlike previous image alignment methods that focus solely on either (i) visual quality or (ii) semantic alignment between text and videos, we comprehensively consider both dimensions and construct a preference score accordingly, which we term the OmniScore. We design a pipeline to automatically collect preference pair data based on the proposed OmniScore and discover that re-weighting these pairs based on the score significantly impacts overall preference alignment. Our experiments demonstrate substantial improvements in both visual quality and semantic alignment, ensuring that no preference aspect is neglected. Code and data are available at https://videodpo.github.io/. Runtao Liu, Ziqiang Zheng, Yingqing He, Renjie Pi, Qifeng Chen 0001 |
CVPR | 3 |
| 2025 | CoraLSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-Supervised Guidance
Ziqiang Zheng, Yuk-Kwan Wong, Binh-Son Hua, Jianbo Shi, Sai-Kit Yeung |
ICCV | 1 |
| 2025 | All-in-One Image Compression and RestorationabstractVisual images corrupted by various types and levels of degradations are commonly encountered in practical image compression. However, most existing image compression methods are tailored for clean images, therefore struggling to achieve satisfying results on these images. Joint compression and restoration methods typically focus on a single type of degradation and fail to address a variety of degradations in practice. To this end, we propose a unified framework for all-in-one image compression and restoration, which incorporates the image restoration capability against various degradations into the process of image compression. The key challenges involve distinguishing authentic image content from degradations, and flexibly eliminating various degradations without prior knowledge. Specifically, the proposed framework approaches these challenges from two perspectives: i.e., content information aggregation, and degradation representation aggregation. Extensive experiments demonstrate the following merits of our model: 1) superior rate-distortion (RD) performance on various degraded inputs while preserving the performance on clean data; 2) strong generalization ability to real-world and unseen scenarios; 3) higher computing efficiency over compared methods. Our code is available at https://github.com/ZeldaM1/All-in-one. Jiacheng Li 0004, Ziqiang Zheng, Zhiwei Xiong |
WACV | 3 |
| 2025 | Removing Rebar Clutter Through Iterative F-k Migration in GPR DataabstractIn the ground-penetrating radar (GPR) detection of concrete structures, the reflection of rebar layers often obscures the useful signals below. In this letter, an effective and practical method for removing rebar clutter is proposed. It is based on iterative F-k migration and demigration, combined with real-time mask window and some classic GPR data processing steps. First, we calculate the wave velocity through travel time and layer thickness and migrate the B-scan data. Then, we create a mask window to extract the focused rebar reflection. Finally, the rebar clutter is restored through F-k demigration and removed from the original data. Meanwhile, multiple iterations are performed to ensure the complete removal of rebar clutter. The proposed method is not limited by data size and observation scale. The effectiveness of the proposed method is demonstrated by both numerical simulations and model field experiments. Junkai Ge, Huaifeng Sun, Ziqiang Zheng |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Deep Learning-Based GPR Imaging for Permafrost: A Source Wavelet-Independent Inversion ApproachabstractThe distribution and hydrological characteristics of permafrost play a crucial role in environmental stability. Ground-penetrating radar (GPR), as a non-invasive geophysical method, can effectively describe the layer structure of permafrost and obtain permittivity distribution through inversion. However, gradient-based full waveform inversion (FWI) poses significant computational challenges for large-scale GPR inversion, particularly in three-dimensional (3D) permafrost models. While deep learning offers a computationally efficient alternative, existing methods struggle to account for variations in field source wavelets. To address these limitations, we propose a deep learning-based, source-independent FWI approach for rapid GPR inversion. Our method leverages a cross-convolution strategy, where reference traces containing field source wavelets are convolved with simulated responses during training. During prediction, simulated standard wavelets are convolved with field data, ensuring data consistency between the training and inference stages. Meanwhile, we specifically designed a source-independent inversion network that integrates a forward-cycle head mechanism and a near-field interference loss. This approach enhances robustness against non-standard wavelets, making it particularly suitable for field applications. In order to ensure the consistency with real-world conditions, we construct a simulation dataset under 3D scene for training, and further discuss the differences between 2D and 3D GPR simulations. Finally, we apply the proposed method to GPR data from three sites in the Tanggula Mountains on the Tibetan Plateau. The results provide quantitative assessments of ice content in permafrost and water content in thawed zones, which align well with borehole data. Junkai Ge, Shirong Zhang, Huaifeng Sun, Ziqiang Zheng |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | UVEB: A Large-scale Benchmark and Baseline Towards Real-World Underwater Video EnhancementabstractLearning-based underwater image enhancement (UIE) methods have made great progress. However, the lack of large-scale and high-quality paired training samples has become the main bottleneck hindering the development of UIE. The inter-frame information in underwater videos can accelerate or optimize the UIE process. Thus, we constructed the first large-scale high-resolution underwater video enhancement benchmark (UVEB) to promote the development of underwater vision. It contains 1,308 pairs of video sequences and more than 453,000 high-resolution with 38% Ultra-High-Definition (UHD) 4K frame pairs. UVEB comes from multiple countries, containing various scenes and video degradation types to adapt to diverse and complex underwater environments. We also propose the first supervised underwater video enhancement method, UVE-Net. UVE-Net converts the current frame information into convolutional kernels and passes them to adjacent frames for efficient inter-frame information exchange. By fully utilizing the redundant degraded information of underwater videos, UVE-Net completes video enhancement better. experiments show the effective network design and good performance of UVE-Net. Yaofeng Xie, Lingwei Kong, Ziqiang Zheng, Zhibin Yu 0002 |
CVPR | 4 |
| 2024 | CoralSCOP: Segment any COral Image on this PlanetabstractUnderwater visual understanding has recently gained increasing attention within the computer vision community for studying and monitoring underwater ecosystems. Among these, coral reefs play an important and intricate role, often referred to as the rainforests of the sea, due to their rich bio-diversity and crucial environmental impact. Existing coral analysis, due to its technical complexity, requires significant manual work from coral biologists, therefore hindering scalable and comprehensive studies. In this paper, we introduce CoralSCop, the first foundation model designed for the automatic dense segmentation of coral reefs. CoralSCOP is developed to accurately assign labels to different coral entities, addressing the challenges in the semantic analysis of coral imagery. Its main objective is to identify and delineate the irregular boundaries between various coral individuals across different granularities, such as coral/non-coral, growth form, and genus. This task is challenging due to the semantic agnostic nature or fixed limited semantic categories of previous generic segmentation methods, which fail to adequately capture the complex characteristics of coral structures. By introducing a novel parallel semantic branch, CoralSCOP can produce high-quality coral masks with semantics that enable a wide range of downstream coral reef analysis tasks. We demonstrate that CoralSCOP exhibits a strong zero-shot ability to segment unseen coral images. To effectively train our foundation model, we propose CoralMask, a new dataset with 41,297 densely labeled coral images and 330,144 coral masks. We have conducted comprehensive and extensive experiments to demonstrate the advantages of CoralSCOP over existing generalist segmentation algorithms and coral reef analytical approaches. Ziqiang Zheng, Haixin Liang, Binh-Son Hua, Yue Him Wong, Put Ang, Apple Pui Yi Chui, Sai-Kit Yeung |
CVPR | 1 |
| 2024 | MarineInst: A Foundation Model for Marine Image Analysis with Instance Visual Description
Ziqiang Zheng, Yiwei Chen 0003, Tuan-Anh Vu, Binh-Son Hua, Sai-Kit Yeung |
ECCV (2) | 1 |
| 2024 | Instance-Dictionary Learning for Open-World Object Detection in Autonomous Driving ScenariosabstractThis paper addresses an important and valuable open-world object detection (OWOD) in autonomous driving scenarios, which aims to detect objects under bothdomain-agnosticandcategory-agnosticsettings simultaneously. Existing OWOD algorithms mainly focus on the detection of pre-defined object categories under various conditions (domain-agnostic) or instead perform zero-shot object detection (category-agnostic), separately. The knowledge gap between seen and unseen object categories poses challenges for models optimized with supervision from the only seen object categories. The domain difference across different scenarios also causes further challenges in aligning observations with different appearances. To address these two challenges simultaneously, we propose our Instance Dictionary Learning (IDL for short) for more robust and accurate OWOD performance. We first design a pre-training procedure to build up the mappings between region features and category semantic embeddings by introducing instance contrastive learning. The joint vision-semantic space is formulated through the more detailed instance-level “Dictionary”, which expresses the region-category correspondences and helps link the seen and unseen object categories. The domain discrimination is further designed for extracting the domain invariance feature representations in the further training procedure seamlessly. The proposed IDL could detect the unseen categories from unseen domains without any bounding box annotations while there is no obvious performance drop on detecting seen categories meanwhile. Comprehensive experiments have been conducted and our method could achieve a new state-of-the-art OWOD performance over previous algorithms. Zeyu Ma 0002, Ziqiang Zheng, Jiwei Wei, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Fully Unsupervised Domain-Agnostic Image RetrievalabstractRecent research in cross-domain image retrieval has focused on addressing two challenging issues: handling domain variations in the data and dealing with the lack of sufficient training labels. However, these problems have often been studied separately, limiting the practicality and significance of the research outcomes. The existing cross-domain setting is also restricted to cases where domain labels are known during training, and all samples have semantic category information or instance correspondences. In this paper, we propose a novel approach to address a more general and practical problem:fully unsupervised domain-agnostic image retrievalunder the domain-unknown setting, where no annotations are provided. Our approach tackles both thedomain variationandmissing labelschallenges simultaneously. We introduce a new fully unsupervised One-Shot Synthesis-based Contrastive learning method (termed OSSCo) to project images from different data distributions into a shared feature space for similarity measurement. To handle the domain-unknown setting, we propose One-Shot unpaired image-to-image Translation (OST) between a randomly selected one-shot image and the rest of the training images. By minimizing the global distance between the original images and the generated images from OST, the model learns domain-agnostic representations. To address the label-unknown setting, we employ contrastive learning with a synthesis-based transform module from the OST training. This allows for effective representation learning without any annotations or external constraints. We evaluate our proposed method on diverse datasets, and the results demonstrate its effectiveness. Notably, our approach achieves comparable performance to current state-of-the-art supervised methods. Ziqiang Zheng, Hao Ren 0002, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | A Fast and Efficient Method for 3-D Transient Electromagnetic Modeling Considering IP EffectabstractLate-time negative responses in central-loop transient electromagnetic (TEM) data are often linked to the induced polarization (IP) effect. Early methods for modeling the IP effect in TEM data try to avoid calculating the fractional derivative arising from considering the Cole-Cole model by either using the Fourier transform to convert frequency-domain responses to the time domain or approximating the fractional derivative in the time domain directly. The frequency-to-time conversion method suffer from accuracy issues if the number of frequencies calculated is small. The time-domain approximation method also has accuracy issues because of simplified Cole-Cole models. The Caputo series can approximate fractional derivatives accurately if historic electromagnetic (EM) fields are saved. However, the storage of historic EM fields leads to a significant memory consumption. We introduce the sum-of-exponentials (SOE) method to discretize fractional derivatives, which does not need to store field values from previous times except for the first two time-steps. We discretize the resulting partial differential equations from the SOE discretization using a finite-difference time-domain (FDTD) approach. Additionally, we improve computational efficiency by employing the direct-splitting strategy to transform large sparse matrices into smaller diagonally dominant tridiagonal matrices. We validate the accuracy and efficiency of our algorithm by comparing it with the Caputo approximation method using a chargeable half-space model. Furthermore, we compare our results with existing literature data for a chargeable anomaly in a nonchargeable half space. Finally, we analyze the response characteristics of the IP effect using a block-in-half space model. Qi Zhao 0023, Huaifeng Sun, Shangbin Liu, Xushan Lu, Xixian Bai, Ziqiang Zheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Pixel Bleach Network for Detecting Face Forgery Under CompressionabstractThe existing face forgery algorithms have achieved remarkable progress in how to generate reasonable facial images and can even successfully deceive human beings. Considering public security, face forgery detection is of vital importance, making it essential to design face forgery detection algorithms to detect forgery images over the Internet. Despite the great success achieved by the existing Deepfake detection algorithms, they usually failed to achieve satisfactory Deepfake detection performance when deployed to handle the forgery videos in practice. One significant reason is compression. The videos over the Internet are inevitably compressed considering the transmission efficiency. To address this issue, in this paper, we propose a generic, simple yet effective “bleaching” pre-processing module based on the generative model and the high-level feature representations to produce ableached image, which shares a similar appearance with the compressed images. The bleached images with recovered information can be identified accurately by the optimized Deepfake detection models without retraining. The proposed method has utilized a redesigned feature representation, which serves as a navigator to effectively and sufficiently alter the feature distribution in the high-dimensional space to remedy the difference between real facial images and forgery counterparts. Thus, the proposed method can successfully avoid misclassification. Comprehensive and extensive experiments are carried out on four low-quality Faceforensics++ datasets, demonstrating the effectiveness of our method in recovering the information loss caused by the compression artifacts across various backbones and compression. Congrui Li, Ziqiang Zheng, Yi Bin, Guoqing Wang 0001, Yang Yang 0002, Xuesheng Li, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2023 | Cross-Domain Autonomous Driving Perception Using Contrastive Appearance AdaptationabstractAddressing domain shifts for complex perception tasks in autonomous driving has long been a challenging problem. In this paper, we show that existing domain adaptation methods pay little attention to the content mismatch issue between source and target domains, thus weakening the domain adaptation per-formance and the decoupling of domain-invariant and domain-specific representations. To solve the aforementioned problems, we propose an image-level domain adaptation framework that aims at adapting source-domain images to the target domain with content-aligned source-target image pairs. Our framework consists of three mutually beneficial modules in a cycle: a cross-domain content alignment module to generate source-target pairs with consistent content representations in a self-supervised manner, a reference-guided image synthesis based on the generated content-aligned source-target image pairs, and a contrastive learning module to self-supervise domain-invariant feature extractor. Our contrastive appearance adaptation is task-agnostic and robust to complex perception tasks in autonomous driving. Our proposed method demonstrates state-of-the-art results in cross-domain object detection, semantic segmentation, and depth estimation as well as better image synthesis ability qualitatively and quantitatively. Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Yang Wu 0001, Sai-Kit Yeung |
IROS | 1 |
| 2023 | CompUDA: Compositional Unsupervised Domain Adaptation for Semantic Segmentation Under Adverse ConditionsabstractIn autonomous driving, performing robust semantic segmentation under adverse weather conditions is a long-standing challenge. Imperfect camera observations under adverse conditions result in images with reduced visibility, which hinders label annotation and semantic scene understanding based on these images. A common solution is to adopt semantic segmentation models trained in a source domain with ground truth labels and perform unsupervised domain adaptation (UDA) from the source domain to an unlabeled target domain that has adverse conditions. Due to imperfect visual observations in the target domain, such adaptation needs special treatment to achieve good performance. In this paper, we propose a new compositional unsupervised domain adaptation (CompUDA) method that disentangles the domain gap based on multiple factors including style, visibility, and image quality. The domain gaps caused by these individual factors can then be addressed separately by introducing the intermediate domains. Specifically, 1) to address the style gap, we perform source-to-intermediate domain adaptation and generate pseudo-labels for self-training in the target domain; 2) to address the visibility gap, we perform a geometry-aligned normal-to-adverse image translation and introduce a synthetic domain; 3) finally, to address the image quality gap between the synthetic and target domain, we perform a synthetic-to-real adaptation based on the generated pseudo-labels. Our compositional unsupervised domain adaptation can be used in conjunction with a wide variety of semantic segmentation methods and result in significant performance improvement across datasets. The codes are available at https://github.com/zhengziqiang/CompUDA. Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Sai-Kit Yeung |
IROS | 1 |
| 2023 | Open-Scenario Domain Adaptive Object Detection in Autonomous DrivingabstractExisting domain adaptive object detection algorithms (DAOD) have demonstrated their effectiveness in discriminating and localizing objects across scenarios. However, these algorithms typically assume a single source and target domain for adaptation, which is not representative of the more complex data distributions in practice. To address this issue, we propose a novel Open-Scenario Domain Adaptive Object Detection (OSDA), which leverages multiple source and target domains for more practical and effective domain adaptation. We are the first to increase the granularity of the background category by building the foundation model using contrastive vision-language pre-training in an open-scenario setting for better distinguishing foreground and background, which is under-explored in previous studies. The performance gains by introducing the pre-training have been observed and have validated the model's ability to detect objects across domains. To further fine-tune the model for domain-specific object detection, we propose a hierarchical feature alignment strategy to obtain a better common feature space among the various source and target domains. In the case of multi-source domains, the cross-reconstruction framework is introduced for learning more domain invariances. The proposed method is able to alleviate knowledge forgetting without any additional computational costs. Extensive experiments across different scenarios demonstrate the effectiveness of the proposed model. Zeyu Ma 0002, Ziqiang Zheng, Jiwei Wei, Xiaoyong Wei, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 2 |
| 2023 | DaCo: domain-agnostic contrastive learning for visual place recognition
Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001 |
Appl. Intell. | 2 |
| 2023 | ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image RetrievalabstractThe huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset). Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Robust Perception Under Adverse Conditions for Autonomous Driving Based on Data AugmentationabstractMany existing advanced deep learning-based autonomous systems have recently been used for autonomous vehicles. In general, a deep learning-based visual perception system heavily relies on visual perception to recognize and localize dynamic interest objects (e.g., pedestrians and cars) and indicative traffic signs and lights to assist autonomous vehicles in maneuvering safely. However, the performance of existing object recognition algorithms could degrade significantly under some adverse and challenging scenarios including rainy, foggy, and rainy night conditions. The raindrops, light reflection, and low illumination pose a great challenge to robust object recognition. Thus, A robust and accurate autonomous driving system has attracted growing attention from the computer vision community. To achieve robust and accurate visual perception, we target to build effective and efficient augmentation and fusion techniques based on visual perception under various adverse conditions. The unpaired image-to-image (I2I) synthesis is integrated for visual perception enhancement and effective synthesis-based augmentation. Besides, we design a two-branch architecture to utilize the information from both the original image and the enhanced image synthesized by I2I. We comprehensively and hierarchically investigate the performance improvement and limitation of the proposed system based on visual recognition tasks and network backbones. An extensive experimental analysis of various adverse weather conditions is also included. The experimental results have demonstrated the proposed system could promote the ability of autonomous vehicles for robust and accurate perception under adverse weather conditions. Ziqiang Zheng, Yujie Cheng, Zhichao Xin, Zhibin Yu 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Asynchronous Generative Adversarial Network for Asymmetric Unpaired Image-to-Image TranslationabstractThe unpaired image-to-image translation aims to translate input images from one source domain to some desired outputs in a target domain by learning from unpaired training data. Cycle-consistency constraint provides a general principle to estimate and measure forward and backward mapping functions between two domains. In many cases, the information entropy of images from the two domains is not equal, resulting in an information-rich domain and an information-poor domain. However, existing solutions based on cycle-consistency either completely discard the information asymmetry between the two domains (a common choice), which leads to inferior translation performance for the asymmetric unpaired image-to-image translation, or have to rely on special task-specific designs and introduce extra loss components. These elaborative designs especially for the relatively harder translation direction from the information-poor domain to the information-rich domain (poor-to-rich translation) require extra labor and are limited to some specific tasks. In this paper, we propose a novel asynchronous generative adversarial network named Async-GAN, which provides a model-agnostic framework for easily turning symmetrical models into powerful asymmetric counterparts that can handle asymmetric unpaired image-to-image translation much better. The key innovation is to iteratively build gradually improving intermediate domains for generating pseudo paired training samples, which provide stronger full supervision for assisting the poor-to-rich translation. Extensive experiments on various asymmetric unpaired translation tasks demonstrate the superiority of the proposal. Furthermore, the proposed training framework could be extended to various Cycle-GAN solutions and achieve a performance gain. Ziqiang Zheng, Yi Bin, Xiaoou Lv, Yang Wu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2023 | Composition-Aware Image Steganography Through Adversarial Self-Generated SupervisionabstractSteganography is an important and prevailing information hiding tool to perform secret message transmission in an open environment. Existing steganography methods can mainly fall into two categories: predefined rule-based and data-driven methods. The former is susceptible to the statistical attack, while the latter adopts the deep convolution neural networks to promote security. However, deep learning-based methods suffer from perceptible artificial artifacts or deep steganalysis. In this article, we introduce a novel composition-aware image steganography (CAIS) to guarantee both visual security and resistance to deep steganalysis through the self-generated supervision. The key innovation is an adversarial composition estimation module, which has integrated the rule-based composition method and generative adversarial network to help synthesize steganographic images with more naturalness. We first perform a rule-based image blending method to obtain infinite synthetically data-label pairs. Then, we utilize an adversarial composition estimation branch to recognize the message feature pattern from the composite image based on these self-generated data-label pairs. Through the adversarial training, we force the steganography function to synthesize steganographic images, which can fool the composition estimation network. Thus, the proposed CAIS can achieve better information hiding and higher security to resist deep steganalysis. Furthermore, an effective global-and-part checking is designed to alleviate visual artifacts caused by hiding secret information. We conduct a comprehensive analysis of CAIS from various aspects (e.g., security and robustness) to verify the superior performance of the proposed method. Comprehensive experimental results on three large-scale widely used datasets have demonstrated the superior performance of our CAIS compared with several state-of-the-art approaches. Ziqiang Zheng, Yuanmeng Hu, Yi Bin, Xing Xu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Not every sample is efficient: Analogical generative adversarial network for unpaired image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yubo Wang 0001, Zhijian Sun |
Neural Networks | 1 |
| 2022 | Energy-Guided Feature Fusion for Zero-Shot Sketch-Based Image Retrieval
Hao Ren 0002, Ziqiang Zheng, Hong Lu 0001 |
Neural Process. Lett. | 2 |
| 2022 | One-Shot Image-to-Image Translation via Part-Global Learning With a Multi-Adversarial FrameworkabstractIt is well known that humans can learn and recognize objects effectively from several limited image samples. However, learning from just a few images is still a tremendous challenge for existing main-stream deep neural networks. Inspired by analogical reasoning in the human mind, a feasible strategy is to “translate” the abundant images of a rich source domain to enrich the relevant yet different target domain with insufficient image data. To achieve this goal, we propose a novel, effective multi-adversarial framework (MA) based on part-global learning, which accomplishes the one-shot cross-domain image-to-image translation. In specific, we first devise a part-global adversarial training scheme to provide an efficient way for feature extraction and prevent discriminators from being overfitted. Then, a multi-adversarial mechanism is employed to enhance the image-to-image translation ability to unearth the high-level semantic representation. Moreover, a balanced adversarial loss function is presented, which aims to balance the training data and stabilize the training process. Extensive experiments demonstrate that the proposed approach can obtain impressive results on various datasets between two extremely imbalanced image domains and outperform state-of-the-art methods on one-shot image-to-image translation. Our code will be released with this paper athttps://github.com/zhengziqiang/OST. Ziqiang Zheng, Zhibin Yu 0002, Haiyong Zheng, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2021 | Direct Sparse Stereo Visual-Inertial Global OdometryabstractRobust and accurate localization plays a key role in autonomous driving and robot applications. To utilize the complementary properties of different sensors, we present a novel tightly-coupled approach to combine the local (stereo cameras, IMU) and global sensors (magnetometer, GNSS). We jointly optimize all the model parameters through one active window. The visual part integrates constraints from static stereo into the photometric bundle adjustment pipeline of dynamic multiview stereo. Accumulating IMU information between keyframes, magnetometer and GNSS measurements are all inserted into the active window as additional constrains among all the keyframes. Through these, our method can realize globally drift-free and locally accurate state estimation. We evaluate the effectiveness of our system on public datasets under with real-world experiments. Dingkun Zhou, Ziqiang Zheng |
ICRA | 4 |
| 2021 | Generative Adversarial Network with Multi-branch Discriminator for imbalanced cross-species image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yang Wu 0001, Haiyong Zheng, Minho Lee 0001 |
Neural Networks | 1 |
| 2020 | ForkGAN: Seeing into the Rainy Night
Ziqiang Zheng, Yang Wu 0001, Xinran Han, Jianbo Shi |
ECCV (3) | 1 |
| 2020 | Fine-grained facial image-to-image translation with an attention based pipeline generative adversarial framework
Ziqiang Zheng, Chao Wang 0022, Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013 |
Multim. Tools Appl. | 2 |
| 2019 | Learning Chinese Word Embeddings from Stroke, Structure and Pinyin of CharactersabstractChinese word embeddings have recently attracted much attention in natural language processing (NLP). Existing researches learn Chinese word embeddings based on characters, radicals, components and stroke n-gram. Besides abovementioned features, Chinese characters also own structure and pinyin features. In this paper, we design feature substring, a super set of radicals, components and stroke n-gram with structure and pinyin information, to integrate stroke, structure and pinyin features of Chinese characters and capture the semantics of Chinese words. Based on the feature substring, we propose a novel method ssp2vec to predict the contextual words based on the feature substrings of the target words for learning Chinese word embeddings. It is based on our observation that exploiting the morphological information (stroke and structure) and the phonetic information (pinyin) is crucial for capturing the meanings of Chinese words. Meanwhile, the phonetic information (pinyin) can assist the model to distinguish Chinese words. Experimental results on word analogy, word similarity, text classification and named entity recognition tasks show that the proposed method obtains better results than state-of-the-art approaches. Yun Zhang 0019, Yongguo Liu, Jiajing Zhu, Ziqiang Zheng, Shuangqing Zhai |
CIKM | 4 |
| 2019 | Unpaired photo-to-caricature translation on faces in the wild
Ziqiang Zheng, Chao Wang 0022, Zhibin Yu 0002, Nan Wang 0013, Haiyong Zheng |
Neurocomputing | 1 |
| 2018 | Discriminative Region Proposal Adversarial Networks for High-Quality Image-to-Image Translation
Chao Wang 0022, Haiyong Zheng, Zhibin Yu 0002, Ziqiang Zheng, Zhaorui Gu |
ECCV (1) | 4 |