Shiyin Zhang

dblp:272/0602 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-2128-0190ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Transformer for Heterogeneous Graphs
Shiyin Zhang, Zheng Fang 0007, Guojie Song
IEEE Big Data1
2025 Attribute-Enhanced Fine Tuning for Subject-Driven Generation
Shiyin Zhang, Guojie Song, Wenjun Ma, Rundong Cao
IEEE Big Data1
2025 IHE-Net:Hidden feature discrepancy fusion and triple consistency training for semi-supervised medical image segmentation
Mengyi Ju, Bing Wang 0020, Zutong Zhao, Shiyin Zhang, Zhihong Wei
Artif. Intell. Medicine4
2025 The Generated-bbox Guided Interactive Image Segmentation With Vision Transformers
abstract
ABSTRACT Existing click‐based interactive image segmentation methods typically initiate object extraction with the first click and iteratively refine the coarse segmentation through subsequent interactions. Unlike box‐based methods, click‐based approaches mitigate ambiguity when multiple targets are present within a single bounding box, but suffer from a lack of precise location and outline information. Inspired by instance segmentation, the authors propose a Generated‐bbox Guided method that provides location and outline information using an automatically generated bounding box, rather than a manually labelled one, minimising the need for extensive user interaction. Building on the success of vision transformers, the authors adopt them as the network architecture to enhance model's performance. A click‐based interactive image segmentation network named the Generated‐bbox Guided Coarse‐to‐Fine Network (GCFN) was proposed. GCFN is a two‐stage cascade network comprising two sub‐networks: Coarsenet and Finenet. A transformer‐based Box Detector was introduced to generate an initial bounding box from a inside click, that can provide location and outline information. Additionally, two feature enhancement modules guided by foreground and background information: the Foreground‐Background Feature Enhancement Module (FFEM) and the Pixel Enhancement Module (PEM) were designed. The authors evaluate the GCFN method on five popular benchmark datasets and demonstrate the generalisation capability on three medical image datasets.
Shiyin Zhang, Yafei Dong
IET Comput. Vis.1
2025 High-Precision Edge Detection Guided byFlow Fields
abstract
Edge detection is frequently employed to support downstream visual tasks. However, current edge detection methods still encounter two significant challenges: extracting complex textured targets and capturing valuable information from complex backgrounds. We propose FFED, a flow field-guided edge detection model. FFED integrates the three components of our design. FFED incorporates three designed components: the Feature Broadcast Module (FBM), the Antagonistic Bio-inspired Spatial Attention Module (ABSAM), a novel pixel difference convolution named ALS. The FBM serves as an implementation mode of the flow field, with its input pair selection strategy inspired by video processing. The FBM broadcasts high-level semantic features to high-resolution ones, preserving more meaningful texture details. Inspired by biological studies, we propose the ABSAM. ABSAM extracts valuable information from complex backgrounds by optimizing spatial modeling of data. The ALS exhibits enhanced capability in extracting gradient information and capturing subtle texture details that are easily overlooked. Experimental results demonstrate that FFED achieved competitive detection results on NYUD, BSDS500, and BIPED datasets, as well as good performance on industrial datasets. Additionally, the experiment verified the auxiliary effect of FFED on downstream visual tasks. The code is available at https://github.com/hanyuchen2022/Flow-field-guided-edge-detection-FFED-.
Shiyin Zhang, Zhenbing Zhao, Yongjie Zhai
IEEE Trans. Image Process.3
2025 Efficient hierarchical multiscale convolutional attention for accurate medical image segmentation
Bing Wang 0020, Zhihong Wei, Mengyi Ju, Zutong Zhao, Shiyin Zhang
Vis. Comput.5
2024 Exploring the Applicability of Spectral Recovery in Semantic Segmentation of RGB Images
abstract
Compared with RGB images, hyperspectral images (HSIs) offer a distinct advantage in that they can record continuous spectral bands of light reflectance in each pixel, reflecting the physical and chemical characteristics of materials. This capability enables differentiation between objects that may have similar textures but different spectral characteristics. It is desirable to recover spectral information from RGB images to improve semantic segmentation accuracy. Additionally, semantic information can serve as a guide for spectral information recovery, thereby ensuring the quality of the recovered spectral information. The two tasks are mutually beneficial in this regard. In light of these considerations, we propose a multi-task framework that exploits the complementary relationship between spectral recovery and semantic segmentation tasks, comprising a complementary spectral-semantic attentive fusion model (CSSF) that enables the two tasks to mutually facilitate each other by fusing information from both branches. Specifically, the proposed CSSF incorporates a window-based spectral-semantic attentive fusion (WSSAF) module to incorporate recovered spectral information into the segmentation process effectively, and a pixel-shuffle-based fusion (PSF) module to provide semantic guidance for spectral recovery. To evaluate the effectiveness of our approach, we built the first flower hyperspectral image dataset (FHRS) with corresponding segmentation annotations and RGB images. By doing so, we have made the first attempt to explore the complementary relationship between semantic segmentation and spectral recovery. Experimental results on both the FHRS dataset and the publicly available LIB-HSI dataset demonstrate that our proposed method has the ability to enhance both tasks by utilizing their complementary relationship, indicating the generalization ability of our method.
Zhuoran Du, Shikui Wei, Ting Liu 0012, Shunli Zhang 0005, Shiyin Zhang, Yao Zhao 0001
IEEE Trans. Multim.6
2023 Rethinking Parking Slot Detection with Rotated Bounding Box
abstract
Parking slot detection is an essential yet challenging task in the field of self-driving perception. During parking, vehicles often block part of the parking slots which makes the corners occluded. In addition, due to the impact of the external environment, the corners of the parking slot may be blurred. Existing parking slot detection algorithms based on parking slot markings are sensitive to the corners of the parking slots, which makes it difficult to cope with the above scenario. To address this problem, we propose a parking slot entrance line detection algorithm called RPSED, which is the first to apply rotating object detection to the parking slot entrance line. RPSED takes a different route from traditional corner detection methods by focusing on the entrance lines of parking slots to grasp the intricate geometric details inherent to parking slots, which solves the problem that existing parking slot detection algorithms cannot detect parking slots with blurred corners. To further improve the precision and recall of the model and make the model more generalizable, we propose a model ensemble strategy to match and select the results of multiple models. Moreover, we propose two manually optimized parking slot dataset named RPS2.0 and RPSV, which adds more annotations with obstructed corners or obscured configurations to the datasets ps2.0 and psv, making the model evaluation more reasonable and realistic. Experimental results on the RPS2.0 and RPSV benchmarks demonstrate the superiority of our approach compared to existing state-of-the-art methods.
Shikui Wei, Shiyin Zhang, Weiyan Xu, Yao Zhao 0001
MMAsia3
2023 Interactive Object Segmentation With Inside-Outside Guidance
abstract
This article explores how to harvest precise object segmentation masks while minimizing the human interaction cost. To achieve this, we propose a simple yet effective interaction scheme, named Inside-Outside Guidance (IOG). Concretely, we leverage an inside point that is clicked near the object center and two outside points at the symmetrical corner locations (top-left and bottom-right or top-right and bottom-left) of an almost-tight bounding box that encloses the target object. The interaction results in a total of one foreground click and four background clicks for segmentation. The advantages of our IOG are four-fold: 1) the two outside points can help remove distractions from other objects or background; 2) the inside point can help eliminate the unrelated regions inside the bounding box; 3) the inside and outside points are easily identified, reducing the confusion raised by the state-of-the-art DEXTR Maninis et al. 2018, in labeling some extreme samples; 4) it naturally supports additional click annotations for further correction. Despite its simplicity, our IOG not only achieves state-of-the-art performance on several popular benchmarks such as GrabCut Rother et al. 2004, PASCAL Everingham et al. 2010 and MS COCO Russakovsky et al. 2015, but also demonstrates strong generalization capability across different domains such as street scenes (Cityscapes Cordts et al. 2016), aerial imagery (Rooftop Sun et al. 2014 and Agriculture-Vision Chiu et al. 2020) and medical images (ssTEM Gerhard et al. 2013). Code is available at https://github.com/shiyinzhang/Inside-Outside-Guidancehttps://github.com/shiyinzhang/Inside-Outside-Guidance.
Shiyin Zhang, Shikui Wei, Jun Hao Liew, Kunyang Han, Yao Zhao 0001, Yunchao Wei
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Heterogeneous Feature Alignment and Fusion in Cross-Modal Augmented Space for Composed Image Retrieval
abstract
Composed image retrieval (CIR) aims at fusing a reference image and text feedback to search for the desired images. Compared to general image retrieval, it can model the users' search intent more comprehensively and search the target images more accurately, which has significant impacts in various real-world applications, such as E-commerce and Internet search. However, because of the existing heterogeneous semantic gap, the synthetic understanding and fusion of both image and text are difficult to implement. In this work, to tackle this difficult problem, we propose an end-to-end framework MCR, which uses text and images as retrieval queries. The framework mainly includes four pivotal modules. Specifically, we introduce the Relative Caption-aware Consistency (RCC) constraint to align text pieces and images in the database, which can effectually bridge the heterogeneous gap. The Multi-modal Complementary Fusion (MCF) and Cross-modal Guided Pooling (CGP) are constructed to mine multiple interactions between image local features and text word features and learn the complementary representation of the composed query. Furthermore, we develop a plug-and-play Weak-text Semantic Augment (WSA) module for datasets with short or incomplete query texts, which can supplement the weak-text features and is conducive to modeling an augmented semantic space. Extensive experiments demonstrate the practical superior performance over the existing state-of-the-art empirical algorithms on several benchmarks.
Huaxin Pang, Shikui Wei, Gangjian Zhang, Shiyin Zhang, Yao Zhao 0001
IEEE Trans. Multim.4
2022 Taxonomy-Enhanced Graph Neural Networks
abstract
Despite the recent success of Graph Neural Networks (GNNs), their learning pipeline is guided only by the input graph and the desired output of certain tasks, failing to capture useful patterns when not enough data are presented. Existing attempts incorporate auxiliary knowledge to mitigate this issue, most of which are not in a unified structure or hard to obtain. Noticing that nodes in graphs usually form implicit hierarchical structures, we proposed to integrate category taxonomies into the learning process of GNNs. A category taxonomy is a form of domain knowledge with a hierarchical tree structure, which is widely adopted in real-world scenarios. In this paper, we introduce Taxonomy-Enhanced Graph Neural Networks (Taxo-GNN). Specifically, we jointly optimize the taxonomy representation and node representation tasks, where categories in taxonomy are mapped to Gaussian distributions and nodes are embedded with the GNN framework. To characterize the bidirectional interaction between the taxonomy and the graph, the model is comprised of two modules, namely information distillation for taxonomy and knowledge fusion to graph. Information is first distilled from the graph and aligned with the hierarchical structure of the taxonomy in a bottom-to-top mechanism.After that, knowledge brought by the taxonomy is in turn fused to the graph convolution process, in the form of taxonomy-aware aggregation weights and taxonomy-augmented contexts. Extensive experiments on real-world datasets in multiple downstream tasks verify the effectiveness of our model.
Lingjun Xu, Shiyin Zhang, Guojie Song, Junshan Wang, Tianshu Wu
CIKM2
2022 Multi-Source Aggregation Transformer for Concealed Object Detection in Millimeter-Wave Images
abstract
The active millimeter wave scanner has been widely used for detecting objects concealed underneath a person’s clothing in the field of security inspection and anti-terrorism. However, the active millimeter wave (AMMW) images always suffer from low signal-noise ratio, motion blur, and small size objects, making it challenging to detect concealed objects efficiently and accurately. The scanner usually captures a sequence of images in different views around a human body at once, while the existing algorithms only utilize the single image without considering the relationships among images. In this paper, we design a multi-source aggregation transformer (MATR) with two different attention mechanisms to model spatial correlations within an image and contextual interactions across images. Specifically, a self-attention module is introduced to encode local relationships between the region proposals in each image, while a cross-attention mechanism is built to focus on modeling the cross-correlations between different images. Besides, to handle the problem of small objects in size and suppress the noise in AMMW images, we present a selective context module (SCM). It designs a dynamic selection mechanism to enhance the high-resolution feature with spatial details and make it more distinguishable from the noisy background. Experiments on two AMMW image datasets demonstrate that the proposed methods lead to a remarkable improvement compared to previous state-of-the-art and will benefit the concealed object detection in practice.
Ting Liu 0012, Shiyin Zhang, Yao Zhao 0001, Shikui Wei
IEEE Trans. Circuits Syst. Video Technol.4
2020 Interactive Object Segmentation With Inside-Outside Guidance
abstract
This paper explores how to harvest precise object segmentation masks while minimizing the human interaction cost. To achieve this, we propose an Inside-Outside Guidance (IOG) approach in this work. Concretely, we leverage an inside point that is clicked near the object center and two outside points at the symmetrical corner locations (top-left and bottom-right or top-right and bottom-left) of a tight bounding box that encloses the target object. This results in a total of one foreground click and four background clicks for segmentation. The advantages of our IOG is four-fold: 1) the two outside points can help to remove distractions from other objects or background; 2) the inside point can help to eliminate the unrelated regions inside the bounding box; 3) the inside and outside points are easily identified, reducing the confusion raised by the state-of-the-art DEXTR in labeling some extreme samples; 4) our approach naturally supports additional clicks annotations for further correction. Despite its simplicity, our IOG not only achieves state-of-the-art performance on several popular benchmarks, but also demonstrates strong generalization capability across different domains such as street scenes, aerial imagery and medical images, without fine-tuning. In addition, we also propose a simple two-stage solution that enables our IOG to produce high quality instance segmentation masks from existing datasets with off-the-shelf bounding boxes such as ImageNet and Open Images, demonstrating the superiority of our IOG as an annotation tool.
Shiyin Zhang, Jun Hao Liew, Yunchao Wei, Shikui Wei, Yao Zhao 0001
CVPR1