Shuai Zheng 0005

dblp:165/9510-5 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-8560-8135ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MRCNet: Motion Reasoning Chain for Cross Modal Video Camouflaged Object Detection
abstract
Video camouflaged object detection (VCOD) aims to identify objects that seamlessly blend into their surroundings in video sequences. Traditional methods merely rely on visual cues to capture inter-frame motion that reveals camouflaged objects. However, the high similarity between camouflaged objects and their environments often renders pure reliance on visual cues unreliable. Additionally, random motions including camera shaking and abrupt scene transitions also inevitably bring noise into the identification process. To overcome these challenges, we propose a Motion Reasoning Chain Network (MRCNet), a novel cross-modal VCOD framework that emulates the human thought process when observing camouflaged objects, i.e., motion reasoning. Specifically, we introduce a generative sampling strategy grounded in multimodal large language models (MLLMs) to bridge the implicit knowledge space of MLLMs and the explicit representation space regarding the attributes of camouflaged objects, thereby enabling the effective establishment of the motion reasoning chain tailored for VCOD. This process provides semantic guidance for visual comprehension of camouflaged objects through motion and concept attribute reasoning. To improve the identification capability of camouflaged objects, we develop motion representation learning driven by the motion reasoning chain. It introduces hierarchical de-biased motion prototype learning to mitigate hallucinations of MLLMs, boosting the motion perception. To learn precise prompts for the visual foundation model, cross-modal prompt learning further incorporates the de-biased concept prototype into visual representations to enhance the visual comprehension of camouflaged objects. Extensive experiments across three datasets demonstrate that MRCNet achieves state-of-the-art results on both general metrics and spatiotemporal consistency metrics.
Wenjun Hui, Zhenfeng Zhu, Shuai Zheng 0005, Ming-Ming Cheng, Huchuan Lu, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 BiKT: Unleashing the Potential of GNNs via Bi-Directional Knowledge Transfer
abstract
Based on the message-passing paradigm, there has been an amount of research proposing diverse and impressive feature propagation mechanisms to improve the performance of GNNs. However, less focus has been put on feature transformation, another major operation of the message-passing framework. In this paper, we first empirically investigate the performance of the feature transformation operation in several typical GNNs. Unexpectedly, we notice that GNNs do not completely free up the power of the inherent feature transformation operation. By this observation, we propose the Bi-directional Knowledge Transfer (BiKT), a plug-and-play approach to unleash the potential of the feature transformation operations without modifying the original architecture. Taking the feature transformation operation as a derived representation learning model that shares parameters with the original GNN, the direct prediction by this model provides a topological-agnostic knowledge feedback that can further instruct the learning of GNN and the feature transformations therein. On this basis, BiKT not only allows us to acquire knowledge from both the GNN and its derived model but also promotes each other by injecting the knowledge into the other. In addition, a theoretical analysis is further provided to demonstrate that BiKT improves the generalization bound of the GNNs from the perspective of domain adaptation. An extensive group of experiments on up to 7 datasets with 5 typical GNNs demonstrates that BiKT brings up to 0.5% - 4% performance gain over the original GNN, which means a boosted GNN is obtained. Meanwhile, the derived model also shows a powerful performance to compete with or even surpass the original GNN, enabling us to flexibly apply it independently to some other specific downstream tasks.
Shuai Zheng 0005, Zhizhe Liu, Zhenfeng Zhu, Xingxing Zhang 0001, Jianxin Li 0002, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Evolution-consistent dynamic graph condensation
Dong Chen 0043, Shuai Zheng 0005, Yeyu Yan, Muhao Xu, Zhenfeng Zhu, Yao Zhao 0001
Pattern Recognit.2
2026 Multi-level decoupled trend learning for GNN-based multivariate time series prediction
Shaohan Li, Zhenfeng Zhu, Youru Li, Yeyu Yan, Shuai Zheng 0005, Pengyuan Li 0013, Yao Zhao 0001
Pattern Recognit.5
2026 Evolution Rather Than Degradation: Structure-Guided Elastic Consensus Learning for Multimodal Knowledge Graph Completion
Yameng Liu, Shuai Zheng 0005, Zhenfeng Zhu, Yunhui Xu, Yao Zhao 0001, Kunlun He
IEEE Trans. Knowl. Data Eng.2
2026 HarmoFGL: Harmonizing GNN Latent Factors for Federated Graph Learning
abstract
Federated graph learning (FGL), as a privacy-preserving paradigm for distributed graph data training, aims to resolve graph data isolation issues under the framework of federated learning (FL). Despite the significant efforts made by existing FGL methods, two key challenges are still not well addressed: 1) how to mitigate graph heterogeneity in clients arising from feature deviation and structural deviation and 2) how to devise a favorable aggregation mechanism to maximize the client's benefit from collaborative training with privacy preserving. To tackle these issues, we take a perspective of latent factor and propose a HarmoFGL framework by Harmonizing graph neural network (GNN) latent factors for Federated Graph Learning, achieving cross-client federated training by coordinating personalized aggregation and client-level representation in a symbiotic space. To alleviate feature deviation, an implicit feature crossing (IFC) approach is proposed through the disentanglement of higher order feature dependency into client-universal and client-specific interactions. As for the graph heterogeneity induced by structural deviation, we establish a cross-client symbiotic parameter space spanned by GNN latent factors, on which a client-level representation is derived to characterize the inherent properties of clients. On the server side, on the basis of client relevance-driven personalized parameter aggregation, graph Laplacian regularization on client-level representations is implemented for collaborative training. Experimental results on five public graph datasets and two medical datasets demonstrate the effectiveness of HarmoFGL.
Yeyu Yan, Zhenfeng Zhu, Shuai Zheng 0005, Kunlun He, Yao Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2026 Towards effective and efficient graph alignment without supervision
Songyang Chen, Youfang Lin, Yu Liu 0070, Shuai Zheng 0005, Lei Zou 0001
World Wide Web (WWW)4
2025 Towards Pre-trained Graph Condensation via Optimal Transport
abstract
Graph condensation (GC) aims to distill the original graph into a small-scale graph, mitigating redundancy and accelerating GNN training. However, conventional GC approaches heavily rely on rigid GNNs and task-specific supervision. Such a dependency severely restricts their reusability and generalization across various tasks and architectures. In this work, we revisit the goal of ideal GC from the perspective of GNN optimization consistency, and then a generalized GC optimization objective is derived, by which those traditional GC methods can be viewed nicely as special cases of this optimization paradigm. Based on this, \textbf{Pre}-trained \textbf{G}raph \textbf{C}ondensation (\textbf{PreGC}) via optimal transport is proposed to transcend the limitations of task- and architecture-dependent GC methods. Specifically, a hybrid-interval graph diffusion augmentation is presented to suppress the weak generalization ability of the condensed graph on particular architectures by enhancing the uncertainty of node states. Meanwhile, the matching between optimal graph transport plan and representation transport plan is tactfully established to maintain semantic consistencies across source graph and condensed graph spaces, thereby freeing graph condensation from task dependencies. To further facilitate the adaptation of condensed graphs to various downstream tasks, a traceable semantic harmonizer from source nodes to condensed nodes is proposed to bridge semantic associations through the optimized representation transport plan in pre-training. Extensive experiments verify the superiority and versatility of PreGC, demonstrating its task-independent nature and seamless compatibility with arbitrary GNNs.
Yeyu Yan, Shuai Zheng 0005, Wenjun Hui, Xiangkai Zhu, Dong Chen 0043, Zhenfeng Zhu, Yao Zhao 0001, Kunlun He
NeurIPS2
2025 SiGNN: A spike-induced graph neural network for dynamic graph representation learning
Dong Chen 0043, Shuai Zheng 0005, Muhao Xu, Zhenfeng Zhu, Yao Zhao 0001
Pattern Recognit.2
2024 Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object Detection
abstract
The Segment Anything Model (SAM), a prompt-driven foundational model, has demonstrated remarkable performance in natural image segmentation. However, its application in video camouflaged object detection (VCOD) en-counters challenges, chiefly stemming from the overlooked temporal-spatial associations and the unreliability of user-provided prompts for camouflaged objects that are difficult to discern with the naked eye. To tackle the above issues, we endow SAM with keen eyes and propose the Temporal-spatial Prompt SAM (TSP-SAM), a novel approach tailored for VCOD via an ingenious prompted learning scheme. Firstly, motion-driven self-prompt learning is employed to capture the camouflaged object, thereby bypassing the need for user-provided prompts. With the detected subtle motion cues across consecutive video frames, the overall movement of the camouflaged object is captured for more precise spa-tial localization. Subsequently, to eliminate the prompt bias resulting from inter-frame discontinuities, the long-range consistency within the video sequences is taken into account to promote the robustness of the self-prompts. It is also injected into the encoder of SAM to enhance the representational capabilities. Extensive experimental results on two benchmarks demonstrate that the proposed TSP-SAM achieves a significant improvement over the state-of-the-art methods. With the mIoU metric increasing by 7.8% and 9.6%, TSP-SAM emerges as a groundbreaking step forward in the field of VCOD.
Wenjun Hui, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001
CVPR3
2024 FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction
abstract
Multimodal electronic health record (EHR) data can offer a holistic assessment of a patient's health status, supporting various predictive healthcare tasks. Recently, several studies have embraced the multitask learning approach in the healthcare domain, exploiting the inherent correlations among clinical tasks to predict multiple outcomes simultaneously. However, existing methods necessitate samples to possess complete labels for all tasks, which places heavy demands on the data and restricts the flexibility of the model. Meanwhile, within a multitask framework with multimodal inputs, how to comprehensively consider the information disparity among modalities and among tasks still remains a challenging problem. To tackle these issues, a unified healthcare prediction model, also named by \textbf{FlexCare}, is proposed to flexibly accommodate incomplete multimodal inputs, promoting the adaption to multiple healthcare tasks. The proposed model breaks the conventional paradigm of parallel multitask prediction by decomposing it into a series of asynchronous single-task prediction. Specifically, a task-agnostic multimodal information extraction module is presented to capture decorrelated representations of diverse intra- and inter-modality patterns. Taking full account of the information disparities between different modalities and different tasks, we present a task-guided hierarchical multimodal fusion module that integrates the refined modality-level representations into an individual patient-level representation. Experimental results on multiple tasks from MIMIC-IV/MIMIC-CXR/MIMIC-NOTE datasets demonstrate the effectiveness of the proposed method. Additionally, further analysis underscores the feasibility and potential of employing such a multitask strategy in the healthcare domain. The source code is available at https://github.com/mhxu1998/FlexCare.
Muhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng 0005, Kunlun He, Yao Zhao 0001
KDD4
2024 Node-Oriented Spectral Filtering for Graph Neural Networks
abstract
Graph neural networks (GNNs) have shown remarkable performance on homophilic graph data while being far less impressive when handling non-homophilic graph data due to the inherent low-pass filtering property of GNNs. In general, since real-world graphs are often complex mixtures of diverse subgraph patterns, learning a universal spectral filter on the graph from the global perspective as in most current works may still suffer from great difficulty in adapting to the variation of local patterns. On the basis of the theoretical analysis of local patterns, we rethink the existing spectral filtering methods and propose theNode-oriented spectralFiltering forGraphNeuralNetwork (namely NFGNN). By estimating the node-oriented spectral filter for each node, NFGNN is provided with the capability of precise local node positioning via the generalized translated operator, thus discriminating the variations of local homophily patterns adaptively. Meanwhile, the utilization of re-parameterization brings a good trade-off between global consistency and local sensibility for learning the node-oriented spectral filters. Furthermore, we theoretically analyze the localization property of NFGNN, demonstrating that the signal after adaptive filtering is still positioned around the corresponding node. Extensive experimental results demonstrate that the proposed NFGNN achieves more favorable performance.
Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Youru Li, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Graph meets probabilistic generation model: A new perspective for graph disentanglement
Zouzhang Peng, Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Jian Cheng 0001, Honghui Dong, Yao Zhao 0001
Pattern Recognit.2
2024 The Devil Is in the Boundary: Boundary-Enhanced Polyp Segmentation
abstract
Due to the various appearance of the polyps and the tiny contrast between the polyp area and its surrounding background, accurate polyp segmentation has become a challenging task. To tackle this issue, we introduce a boundary-enhanced framework for polyp segmentation, called the Focused on Boundary Segmentation (FoBS) framework, that leverages multi-level collaboration among sample, feature, and optimization. It places greater emphasis on the polyp boundary to improve the accuracy of segmentation. Firstly, a boundary-aware mixup method is designed to improve the model’s awareness of the boundary. More importantly, we propose deformable laplacian-based feature refining to explicitly strengthen the representation ability of the boundary features. It employs a deformable Laplacian refinement function to capture discriminative information from a deformable perceptual field, thereby improving its ability to adapt to boundary variations. In addition, we introduce the self-adjusting refinement coefficient learning that enables adaptive control over the refinement strength at each location. Furthermore, we develop a location-sensitive compensation criterion that assigns more importance to the degraded feature after feature refinement during optimization. Extensive quantitative and qualitative experiments on four polyp benchmarks demonstrate the effectiveness of our method for automatic polyp segmentation. Our code is available at https://github.com/TFboys-lzz/ FoBS.
Zhizhe Liu, Shuai Zheng 0005, Xiaoyi Sun, Zhenfeng Zhu, Xuebing Yang, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 From Observation to Concept: A Flexible Multi-View Paradigm for Medical Report Generation
abstract
Automated radiology report generation aims to generate accurate and radiologist-like descriptions for the patient's images, which can greatly relieve the workload of radiologists. However, due to the data bias and long report problems, medical report generation has been a challenging task. In this article, we propose aFlexibleMulti-viewParadigm (FMVP) for medical report generation in a novel observation-to-concept manner. It first makes some medical observations automatically or with the help of a radiologist on the patient's image to obtain patient-related priori knowledge, just as radiologists do in practice. Furthermore, to bridge the gap betweenpretrainandgenerationphases, the hierarchical alignment is proposed to jointly conduct the implicit alignment between region-tag and the explicit global alignment of the image-report pair. Finally, a compatible decoder towards decoding the fused multi-view knowledge is proposed to capture more complementary information for the report generation, which breaks the traditional entrenched decoding mechanism guided by visual information. Extensive quantitative and qualitative experiments on the public MIMIC-CXR and IU-Xray datasets show that our model achieves competitive performance compared to state-of-the-art methods.
Zhizhe Liu, Zhenfeng Zhu, Shuai Zheng 0005, Kunlun He, Yao Zhao 0001
IEEE Trans. Multim.3
2023 Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation with Cross-Scale Distortion Awareness
abstract
Based on the Manhattan World assumption, most existing indoor layout estimation schemes focus on recovering layouts from vertically compressed 1D sequences. However, the compression procedure confuses the semantics of different planes, yielding inferior performance with ambiguous interpretability. To address this issue, we propose to disentangle this 1D representation by pre-segmenting orthogonal (vertical and horizontal) planes from a complex scene, explicitly capturing the geometric cues for indoor layout estimation. Considering the symmetry between the floor boundary and ceiling boundary, we also design a soft-flipping fusion strategy to assist the pre-segmentation. Besides, we present a feature assembling mechanism to effectively integrate shallow and deep features with distortion distribution awareness. To compensate for the potential errors in pre-segmentation, we further leverage triple attention to reconstruct the disentangled sequences for better performance. Experiments on four popular benchmarks demonstrate our superiority over existing SoTA solutions, especially on the 3DIoU metric. The code is available at https://github.com/zhijieshen-bjtu/DOPNet.
Zhijie Shen, Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Shuai Zheng 0005, Yao Zhao 0001
CVPR6
2023 Sylvester Equation Induced Collaborative Representation Learning for Recommendation
abstract
For an actual recommendation system, it generally involves a variety of heterogeneous interactive relationships, such as the typical user-user (U2U), item-item (I2I), and user-item (U2I) interaction relationships. With the application of graph neural networks (GNNs) in embedding various interactive relations, recommendation technology has made gratifying progress in recent years, which benefits lot from its powerful ability in relation modeling. However, most of the existing GNN-based methods fail to collaboratively explore the above heterogeneous multiple interactive relationships, including the internal correlations among multiple relationships and the intrinsic association behind different relationships. As a consequence, the user's personalized preference for the items to be recommended will not be well captured. In this paper, we propose aSylvester equation inducedCollaborativeRepresentationLearning framework (S-CRL) for recommendation system by utilizing the heterogeneous multiple interactive relationships. In particular, we ingeniously define a novel Sylvester equation to associate tactfully the multiple heterogeneous relations together. From the perspective of rating propagation, such Sylvester equation is shown theoretically to be the optimal solution of a local structure sensitive rating propagation function. Additionally, to seek more expressive embeddings about user and item, a layer-wise attention is introduced to aggregate the multi-hop information from U2U and I2I graphs, respectively, so as to promote the aggregation with the corresponding embeddings from the U2I interaction graph. Extensive experiments on three real-world datasets verify that our model achieves more favorable performance over currently representative methods.
Xingyuan Li 0002, Zhenfeng Zhu, Shuai Zheng 0005, Zhizhe Liu, Youru Li, Deqiang Kong, Yao Zhao 0001
IEEE Trans. Knowl. Data Eng.3
2022 SGT: Scene Graph-Guided Transformer for Surgical Report Generation
Shuai Zheng 0005, Zhizhe Liu, Youru Li, Zhenfeng Zhu, Yao Zhao 0001
MICCAI (8)2
2022 Attention-Enhanced Disentangled Representation Learning for Unsupervised Domain Adaptation in Cardiac Segmentation
Xiaoyi Sun, Zhizhe Liu, Shuai Zheng 0005, Zhenfeng Zhu, Yao Zhao 0001
MICCAI (8)3
2022 MFHI: Taking Modality-Free Human Identification as Zero-Shot Learning
abstract
Human identification is an important topic in event detection, person tracking, and public security. There have been numerous methods proposed for human identification, such as face identification, person re-identification, and gait identification. Typically, existing methods predominantly classify a queried image to a specific identity in an image gallery set (I2I). This is seriously limited for the scenario where only a textual description of the query or an attribute gallery set is available in a wide range of video surveillance applications (A2IorI2A). However, very few efforts have been devoted towards modality-free identification, i.e., identifying a query in a gallery set in a scalable way. In this work, we take an initial attempt, and formulate such a novelModality-FreeHumanIdentification (named MFHI) task as a generic zero-shot learning model in a scalable way. Meanwhile, it is capable of bridging the visual and semantic modalities by learning a discriminative prototype of each identity. In addition, the semantics-guided spatial attention is enforced on visual modality to obtain interpretable representations with both high global category-level and local attribute-level discrimination. Finally, we design and conduct an extensive group of experiments on two common challenging identification tasks, including face identification and person re-identification, demonstrating that our method outperforms a wide variety of state-of-the-art methods on modality-free human identification.
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Margin Preserving Self-Paced Contrastive Learning Towards Domain Adaptation for Medical Image Segmentation
abstract
To bridge the gap between the source and target domains in unsupervised domain adaptation (UDA), the most common strategy puts focus on matching the marginal distributions in the feature space through adversarial learning. However, such category-agnostic global alignment lacks of exploiting the class-level joint distributions, causing the aligned distribution less discriminative. To address this issue, we propose in this paper a novel margin preserving self-paced contrastive Learning (MPSCL) model for cross-modal medical image segmentation. Unlike the conventional construction of contrastive pairs in contrastive learning, the domain-adaptive category prototypes are utilized to constitute the positive and negative sample pairs. With the guidance of progressively refined semantic prototypes, a novel margin preserving contrastive loss is proposed to boost the discriminability of embedded representation space. To enhance the supervision for contrastive learning, more informative pseudo-labels are generated in target domain in a self-paced way, thus benefiting the category-aware distribution alignment for UDA. Furthermore, the domain-invariant representations are learned through joint contrastive learning between the two domains. Extensive experiments on cross-modal cardiac segmentation tasks demonstrate that MPSCL significantly improves semantic segmentation performance, and outperforms a wide variety of state-of-the-art methods by a large margin.
Zhizhe Liu, Zhenfeng Zhu, Shuai Zheng 0005, Yang Liu 0235, Yao Zhao 0001
IEEE J. Biomed. Health Informatics3
2022 Multi-Modal Graph Learning for Disease Prediction
abstract
Benefiting from the powerful expressive capability of graphs, graph-based approaches have been popularly applied to handle multi-modal medical data and achieved impressive performance in various biomedical applications. For disease prediction tasks, most existing graph-based methods tend to define the graph manually based on specified modality (e.g., demographic information), and then integrated other modalities to obtain the patient representation by Graph Representation Learning (GRL). However, constructing an appropriate graph in advance is not a simple matter for these methods. Meanwhile, the complex correlation between modalities is ignored. These factors inevitably yield the inadequacy of providing sufficient information about the patient's condition for a reliable diagnosis. To this end, we propose an end-to-end Multi-modal Graph Learning framework (MMGL) for disease prediction with multi-modality. To effectively exploit the rich information across multi-modality associated with the disease, modality-aware representation learning is proposed to aggregate the features of each modality by leveraging the correlation and complementarity between the modalities. Furthermore, instead of defining the graph manually, the latent graph structure is captured through an effective way of adaptive graph learning. It could be jointly optimized with the prediction model, thus revealing the intrinsic connections among samples. Our model is also applicable to the scenario of inductive learning for those unseen data. An extensive group of experiments on two disease prediction tasks demonstrates that the proposed MMGL achieves more favorable performance. The code of MMGL is available at https://github.com/SsGood/MMGL.
Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Yang Liu 0235, Yao Zhao 0001
IEEE Trans. Medical Imaging1
2021 CETransformer: Casual Effect Estimation via Transformer Based Representation Learning
Shuai Zheng 0005, Zhizhe Liu, Zhenfeng Zhu
PRCV (4)2
2020 Distribution-Induced Bidirectional Generative Adversarial Network for Graph Representation Learning
abstract
Graph representation learning aims to encode all nodes of a graph into low-dimensional vectors that will serve as input of many computer vision tasks. However, most existing algorithms ignore the existence of inherent data distribution and even noises. This may significantly increase the phenomenon of over-fitting and deteriorate the testing accuracy. In this paper, we propose a Distribution-induced Bidirectional Generative Adversarial Network (named DBGAN) for graph representation learning. Instead of the widely used Gaussian assumption, the prior distribution of latent representation in our DBGAN is estimated in a structure-aware way, which implicitly bridges the graph and content spaces by prototype learning. Thus discriminative and robust representations are generated for all nodes. Furthermore, to improve their generalization ability while preserving representation ability, the sample-level and distribution-level consistency are well balanced via a bidirectional adversarial learning framework. An extensive group of experiments is then carefully designed and presented, demonstrating that our DBGAN obtains remarkably more favorable trade-off between representation and robustness, and meanwhile is dimension-efficient, over currently available alternatives in various tasks.
Shuai Zheng 0005, Zhenfeng Zhu, Xingxing Zhang 0001, Zhizhe Liu, Jian Cheng 0001, Yao Zhao 0001
CVPR1
2020 Convolutional prototype learning for zero-shot recognition
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
Image Vis. Comput.4
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.22
2019 Edge Heuristic GAN for Non-Uniform Blind Deblurring
abstract
Non-uniform blur, mainly caused by camera shake and motions of multiple objects, is one of the most common causes of image quality degradation. However, the traditional blind deblurring methods based on blur kernel estimation do not perform well on complicated non-uniform motion blurs. However, recent studies show that GAN-based approaches achieve impressive performance on deblurring tasks. In this letter, to further improve the performance of GAN-based methods on deblurring tasks, we propose an edge heuristic multi-scale generative adversarial network (GAN), which uses the coarse-to-fine scheme to restore clear images in an end-to-end manner. In particular, an edge-generated network is designed to generate sharp edges as auxiliary information to guide the deblurring process. Furthermore, We propose a hierarchical content loss function for deblurring tasks. Extensive experiments on different datasets show that our method achieves state-of-the-art performance in dynamic scene deblurring.
Shuai Zheng 0005, Zhenfeng Zhu, Jian Cheng 0001, Yandong Guo, Yao Zhao 0001
IEEE Signal Process. Lett.1