Ziyi Shen

dblp:167/9166 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-6867-1909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 ProtoMix-SAM: Prototype-Vicinal Sharpness Minimization for Low-Shot Covariate-Shift Generalization
Yuchuan Pu, Yaoran Yang, Ziyi Shen
ICIC (15)5
2026 RIGS: Risk-Adaptive Invariant Gating for Few-Shot Out-of-Distribution Generalization
Yaoran Yang, Yuchuan Pu, Ziyi Shen
ICIC (26)5
2026 Temporal-Consensus Sharpness Minimization for Data-Efficient Out-of-Distribution Generalization
Yaoran Yang, Yuchuan Pu, Ziyi Shen
ICIC (26)5
2026 PromptReg: Interactive Registration by "Corresponding Prompts" for Segment Anything Model (SAM)
abstract
Effectively establishing correspondence between two images is at the centre of image registration methods. Spatially omnipresent representations, including dense displacement fields (DDFs) and spatial (non-)rigid transformations, have been used to parameterise such correspondence. Alternatively, region-based representation uses paired regions of interest (ROIs) to represent region-level correspondence, while retaining its local and dense representation capability at pixel/voxel level if required. Thus, registration can be re-envisioned as a problem of segmenting corresponding paired ROIs in the to-be-registered images. In this work, we utilize models such as SAM, which are pre-trained on substantive datasets, to segment ROIs of the same class from two images, for a new training-free, non-iterative registration algorithm. First, a "corresponding prompt problem" is posed to find a corresponding Prompt Y on Image Y, given any vision Prompt X on Image X, such that the two respectively prompt-conditioned segmentations are a pair of corresponding ROIs from the two images. Second, we propose an "inverse prompt" solution to the corresponding prompt problem, by inverting Prompt X to the Image Y prompt space, where the Jacobian of prototypical features is used. Third, we propose a new registration algorithm that identifies multiple paired corresponding ROIs, by marginalizing the inverted Prompt X over both prompt and spatial spaces, random sampling Prompt X and spatial warping Image X. Comprehensive experiments were conducted on five applications of registering 3D prostate MR, 3D abdomen CT, 3D lung CT, 2D histopathology and, as a non-medical example, 2D aerial images. Based on metrics including Dice and target registration errors on anatomical structures, the proposed registration outperforms both intensity-based iterative algorithms and learning-based networks, even yielding competitive performance with weakly-supervised registration which requires fully-segmented training data.
Shiqi Huang 0001, Tingfa Xu, Jianan Li 0001, Shaheer U. Saeed, Ziyi Shen, Dean C. Barratt, Yipeng Hu
IEEE Trans. Image Process.5
2024 One Registration is Worth Two Segmentations
Shiqi Huang 0001, Tingfa Xu, Ziyi Shen, Shaheer U. Saeed, Wen Yan 0005, Dean C. Barratt, Yipeng Hu
MICCAI (12)3
2024 Nonrigid Reconstruction of Freehand Ultrasound Without a Tracker
abstract
Reconstructing 2D freehand Ultrasound (US) frames into 3D space without using a tracker has recently seen advances with deep learning. Predicting good frame-to-frame rigid transformations is often accepted as the learning objective, especially when the ground-truth labels from spatial tracking devices are inherently rigid transformations. Motivated by a) the observed nonrigid deformation due to soft tissue motion during scanning, and b) the highly sensitive prediction of rigid transformation, this study investigates the methods and their benefits in predicting nonrigid transformations for reconstructing 3D US. We propose a novel co-optimisation algorithm for simultaneously estimating rigid transformations among US frames, supervised by ground-truth from a tracker, and a nonrigid deformation, optimised by a regularised registration network. We show that these two objectives can be either optimised using meta-learning or combined by weighting. A fast scattered data interpolation is also developed for enabling frequent reconstruction and registration of non-parallel US frames, during training. With a new data set containing over 357,000 frames in 720 scans, acquired from 60 subjects, the experiments demonstrate that, due to an expanded thus easier-to-optimise solution space, the generalisation is improved with the added deformation estimation, with respect to the rigid ground-truth. The global pixel reconstruction error (assessing accumulative prediction) is lowered from 18.48 to 16.51 mm, compared with baseline rigid-transformation-predicting methods. Using manually identified landmarks, the proposed co-optimisation also shows potentials in compensating nonrigid tissue motion at inference, which is not measurable by tracker-provided ground-truth. The code and data used in this paper are made publicly available at https://github.com/QiLi111/NR-Rec-FUS .
Qi Li 0030, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, Yipeng Hu
MICCAI (4)2
2024 Poisson Ordinal Network for Gleason Group Estimation Using Bi-Parametric MRI
Yinsong Xu 0002, Ziyi Shen, Iani J. M. B. Gayo, Natasha Thorley, Shonit Punwani, Aidong Men, Dean C. Barratt, Qingchao Chen, Yipeng Hu
MICCAI (5)3
2024 Active learning using adaptable task-based prioritisation
abstract
Supervised machine learning-based medical image computing applications necessitate expert label curation, while unlabelled image data might be relatively abundant. Active learning methods aim to prioritise a subset of available image data for expert annotation, for label-efficient model training. We develop a controller neural network that measures priority of images in a sequence of batches, as in batch-mode active learning, for multi-class segmentation tasks. The controller is optimised by rewarding positive task-specific performance gain, within a Markov decision process (MDP) environment that also optimises the task predictor. In this work, the task predictor is a segmentation network. A meta-reinforcement learning algorithm is proposed with multiple MDPs, such that the pre-trained controller can be adapted to a new MDP that contains data from different institutes and/or requires segmentation of different organs or structures within the abdomen. We present experimental results using multiple CT datasets from more than one thousand patients, with segmentation tasks of nine different abdominal organs, to demonstrate the efficacy of the learnt prioritisation controller function and its cross-institute and cross-organ adaptability. We show that the proposed adaptable prioritisation metric yields converging segmentation accuracy for a new kidney segmentation task, unseen in training, using between approximately 40% to 60% of labels otherwise required with other heuristic or random prioritisation metrics. For clinical datasets of limited size, the proposed adaptable prioritisation offers a performance improvement of 22.6% and 10.2% in Dice score, for tasks of kidney and liver vessel segmentation, respectively, compared to random prioritisation and alternative active sampling strategies.
Shaheer U. Saeed, João Ramalhinho, Mark A. Pinnock, Ziyi Shen, Yunguan Fu, Nina Montaña Brown, Ester Bonmati, Dean C. Barratt, Stephen P. Pereira, Brian R. Davidson, Matthew J. Clarkson, Yipeng Hu
Medical Image Anal.4
2024 Combiner and HyperCombiner networks: Rules to combine multimodality MR images for prostate cancer localisation
Wen Yan 0005, Bernard Chiu, Ziyi Shen, Qianye Yang, Tom Syer, Zhe Min, Shonit Punwani, Mark Emberton, David Atkinson, Dean C. Barratt, Yipeng Hu
Medical Image Anal.3
2022 Collaborative Quantization Embeddings for Intra-subject Prostate MR Image Registration
Ziyi Shen, Qianye Yang, Yuming Shen, Francesco Giganti, Vasilis Stavrinides, Richard E. Fan, Caroline M. Moore, Mirabela Rusu, Geoffrey A. Sonn, Philip Torr 0001, Dean C. Barratt, Yipeng Hu
MICCAI (6)1
2022 SiamATL: Online Update of Siamese Tracking Network via Attentional Transfer Learning
abstract
Visual object tracking with semantic deep features has recently attracted much attention in computer vision. Especially, Siamese trackers, which aim to learn a decision making-based similarity evaluation, are widely utilized in the tracking community. However, the online updating of the Siamese fashion is still a tricky issue due to the limitation, which is a tradeoff between model adaption and degradation. To address such an issue, in this article, we propose a novel attentional transfer learning-based Siamese network (SiamATL), which fully exploits the previous knowledge to inspire the current tracker learning in the decision-making module. First, we explicitly model the template and surroundings by using an attentional online update strategy to avoid template pollution. Then, we introduce an instance-transfer discriminative correlation filter (ITDCF) to enhance the distinguishing ability of the tracker. Finally, we suggest a mutual compensation mechanism that integrates cross-correlation matching and ITDCF detection into the decision-making subnetwork to achieve online tracking. Comprehensive experiments demonstrate that our approach outperforms state-of-the-art tracking algorithms on multiple large-scale tracking datasets.
Bo Huang 0012, Tingfa Xu, Ziyi Shen, Shenwang Jiang, Bingqing Zhao, Ziyang Bian
IEEE Trans. Cybern.3
2022 Confidence-and-Refinement Adaptation Model for Cross-Domain Semantic Segmentation
abstract
With the rapid development of convolutional neural networks (CNNs), significant progress has been achieved in semantic segmentation. Despite the great success, such deep learning approaches require large scale real-world datasets with pixel-level annotations. However, considering that pixel-level labeling of semantics is extremely laborious, many researchers turn to utilize synthetic data with free annotations. But due to the clear domain gap, the segmentation model trained with the synthetic images tends to perform poorly on the real-world datasets. Unsupervised domain adaptation (UDA) for semantic segmentation recently gains an increasing research attention, which aims at alleviating the domain discrepancy. Existing methods in this scope either simply align features or the outputs across the source and target domains or have to deal with the complex image processing and post-processing problems. In this work, we propose a novel multi-level UDA model named Confidence-and-Refinement Adaptation Model (CRAM), which contains a confidence-aware entropy alignment (CEA) module and a style feature alignment (SFA) module. Through CEA, the adaptation is done locally via adversarial learning in the output space, making the segmentation model pay attention to the high-confident predictions. Furthermore, to enhance the model transfer in the shallow feature space, the SFA module is applied to minimize the appearance gap across domains. Experiments on two challenging UDA benchmarks “GTA5-to-Cityscapes” and “SYNTHIA-to-Cityscapes” demonstrate the effectiveness of CRAM. We achieve comparable performance with the existing state-of-the-art works with advantages in simplicity and convergence speed.
Xiaohong Zhang 0009, Yi Chen 0023, Ziyi Shen, Yuming Shen, Haofeng Zhang 0001, Yudong Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2021 You Never Cluster Alone
abstract
Recent advances in self-supervised learning with instance-level contrastive objectives facilitate unsupervised clustering. However, a standalone datum is not perceiving the context of the holistic cluster, and may undergo sub-optimal assignment. In this paper, we extend the mainstream contrastive learning paradigm to a cluster-level scheme, where all the data subjected to the same cluster contribute to a unified representation that encodes the context of each data group. Contrastive learning with this representation then rewards the assignment of each datum. To implement this vision, we propose twin-contrast clustering (TCC). We define a set of categorical variables as clustering assignment confidence, which links the instance-level learning track with the cluster-level one. On one hand, with the corresponding assignment variables being the weight, a weighted aggregation along the data points implements the set representation of a cluster. We further propose heuristic cluster augmentation equivalents to enable cluster-level contrastive learning. On the other hand, we derive the evidence lower-bound of the instance-level contrastive objective with the assignments. By reparametrizing the assignment variables, TCC is trained end-to-end, requiring no alternating steps. Extensive experiments show that TCC outperforms the state-of-the-art on benchmarked datasets.
Yuming Shen, Ziyi Shen, Jie Qin 0004, Philip Torr 0001, Ling Shao 0001
NeurIPS2
2021 BSCF: Learning background suppressed correlation filter tracker for wireless multimedia sensor networks
Bo Huang 0012, Tingfa Xu, Ziyi Shen, Shenwang Jiang, Jianan Li 0001
Ad Hoc Networks3
2021 Modeling and Enhancing Low-Quality Retinal Fundus Images
abstract
Retinal fundus images are widely used for the clinical screening and diagnosis of eye diseases. However, fundus images captured by operators with various levels of experience have a large variation in quality. Low-quality fundus images increase uncertainty in clinical observation and lead to the risk of misdiagnosis. However, due to the special optical beam of fundus imaging and structure of the retina, natural image enhancement methods cannot be utilized directly to address this. In this article, we first analyze the ophthalmoscope imaging system and simulate a reliable degradation of major inferior-quality factors, including uneven illumination, image blurring, and artifacts. Then, based on the degradation model, a clinically oriented fundus enhancement network (cofe-Net) is proposed to suppress global degradation factors, while simultaneously preserving anatomical retinal structures and pathological characteristics for clinical observation and analysis. Experiments on both synthetic and real images demonstrate that our algorithm effectively corrects low-quality fundus images without losing retinal details. Moreover, we also show that the fundus correction method can benefit medical image analysis applications, e.g., retinal vessel segmentation and optic disc/cup detection.
Ziyi Shen, Huazhu Fu, Jianbing Shen, Ling Shao 0001
IEEE Trans. Medical Imaging1
2020 Exploiting Semantics for Face Image Deblurring
Ziyi Shen, Wei-Sheng Lai, Tingfa Xu, Jan Kautz, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.1
2020 Transfer learning-based discriminative correlation filter for visual tracking
abstract
Most Correlation Filter (CF)-based tracking methods can hardly handle occlusion or severe deformation, due to the lack of effective utilization of previous target information. To overcome this, we propose a novel Transfer Learning-based Discriminative Correlation Filter (TLDCF), which extracts knowledge from multiple previous tracking tasks and applies the knowledge for a new tracking task through Instance-Transfer Learning (ITL) and Probability-Transfer Learning (PTL). ITL applies knowledge of Gaussian Mixture Modelling (GMM) target representations and multi-channel filters learned in previous frames to directly train a new correlation filter. This improves the robustness of tracker for heavy occlusion and large appearance variations. Meanwhile, PTL encodes the spatio-temporal relationship predicted by Kalman Filter (KF) into a shared Gaussian prior to suppress huge location drift caused by similar targets. For optimization, we develop an efficient Alternating Direction Method of Multipliers (ADMM) based algorithm to calculate CFs on each independent channel in real time. Extensive experiments on OTB-2013 and OTB-2015 datasets well demonstrate the effectiveness of the proposed method. In particular, our method improves AUC score of the two datasets by 5.5% and 3.9% respectively compared to baseline, and achieves competitive performance against recent state-of-the-art deep trackers.
Bo Huang 0012, Tingfa Xu, Jianan Li 0001, Ziyi Shen, Yiwen Chen 0002
Pattern Recognit.4
2020 Motion-Aware Rapid Video Saliency Detection
abstract
In this paper, we propose a computationally efficient and consistently accurate spatiotemporal salient object detection method to identify the most noticeable object in a video sequence. Intuitively, the underlying motion in a video is a more stable saliency indicator than the apparent color cues that often contain significant variations and complex structures. Based on this observation, we build an efficient and accurate spatiotemporal saliency detection method that uses motion information as a leverage to locate the most dynamic regions in a video sequence. We first analyze the optical flow field to obtain foreground priors, and then incorporate spatial saliency features such as appearance contrasts and compactness measures, into a multi-cue integration framework to combine various saliency cues and achieve temporal consistency. Rigorous experiments on the challenging SegTrackV1, SegTrackV2, and FBMS datasets demonstrate that our method generates comparable or superior performance to state-of-the-art methods while running almost 100× faster at only 0.08 sec/frame. Promising performance and rapid speed imply that the proposed spatiotemporal saliency method can be easily involved in various vision applications.
Wenguan Wang, Ziyi Shen, Jianbing Shen, Ling Shao 0001, Dacheng Tao
IEEE Trans. Circuits Syst. Video Technol.3
2019 Human-Aware Motion Deblurring
abstract
This paper proposes a human-aware deblurring model that disentangles the motion blur between foreground (FG) humans and background (BG). The proposed model is based on a triple-branch encoder-decoder architecture. The first two branches are learned for sharpening FG humans and BG details, respectively; while the third one produces global, harmonious results by comprehensively fusing multi-scale deblurring information from the two domains. The proposed model is further endowed with a supervised, human-aware attention mechanism in an end-to-end fashion. It learns a soft mask that encodes FG human information and explicitly drives the FG/BG decoder-branches to focus on their specific domains. Above designs lead to a fully differentiable motion deblurring network, which can be trained end-to-end. To further benefit the research towards Human-aware Image Deblurring, we introduce a large-scale dataset, named HIDE, which consists of 8,422 blurry and sharp image pairs with 65,784 densely annotated FG human bounding boxes. HIDE is specifically built to span a broad range of scenes, human object sizes, motion patterns, and background complexities. Extensive experiments on public benchmarks and our dataset demonstrate that our model performs favorably against the state-of-the-art motion deblurring methods, especially in capturing semantic details.
Ziyi Shen, Wenguan Wang, Xiankai Lu, Jianbing Shen, Haibin Ling, Tingfa Xu, Ling Shao 0001
ICCV1
2018 Deep Semantic Face Deblurring
abstract
In this paper, we present an effective and efficient face deblurring algorithm by exploiting semantic cues via deep convolutional neural networks (CNNs). As face images are highly structured and share several key semantic components (e.g., eyes and mouths), the semantic information of a face provides a strong prior for restoration. As such, we propose to incorporate global semantic priors as input and impose local structure losses to regularize the output within a multi-scale deep CNN. We train the network with perceptual and adversarial losses to generate photo-realistic results and develop an incremental training strategy to handle random blur kernels in the wild. Quantitative and qualitative evaluations demonstrate that the proposed face deblurring algorithm restores sharp images with more facial details and performs favorably against state-of-the-art methods in terms of restoration quality, face recognition and execution speed.
Ziyi Shen, Wei-Sheng Lai, Tingfa Xu, Jan Kautz, Ming-Hsuan Yang 0001
CVPR1
2018 Generating Reliable Online Adaptive Templates for Visual Tracking
abstract
Online adaption of visual tracking is a significant strategy to achieve good tracking performance since the appearance of the object target varies all along with the sequence. However, directly using the tracking results of previous frames to update the model will cause drifting, resulting in tracking failure. We propose a task-guided generative adversarial network (GAN), named TGGAN, to learn the general appearance distribution that a target may undergo through a sequence. Then the online adaption is simply to select templates from the images that are generated from the ground truth template in the first frame and a set of random vectors by the generator. This strategy helps the model alleviate drifting while still obtaining adaptivity. Tracking is treated as a template matching problem under a proposed Siamese matching network structure. Experiments show the effectiveness of the proposed online adaption strategy and the Siamese matching network.
Jie Guo 0004, Tingfa Xu, Shenwang Jiang, Ziyi Shen
ICIP4
2018 Non-uniform motion deblurring with Kernel grid regularization
Ziyi Shen, Tingfa Xu, Jinshan Pan, Jie Guo 0004
Signal Process. Image Commun.1
2017 Visual Tracking Via Sparse Representation With Reliable Structure Constraint
abstract
In this letter, we present a novel visual tracking algorithm based on sparse representation. In contrast to just use the target templates and the trivial templates to sparsely represent the target, we propose to further constrain the model with a set of discriminative weight maps. These weight maps contain the reliable structures of the target object. They help the model penalize the trivial template coefficients depending on the reliable structures of the target object. Then, the target object can be well represented by a sparse set of target templates together with a sparse set of target weight maps. We propose a unified objective function to integrate these two sparse representation problems together. This optimization problem can be well solved by the proposed iteration manner and a customized accelerated proximal gradient method. Furthermore, a novel weight map constructing method is proposed based on consistent motion property and forward-backward errors. Plenty of qualitative and quantitative evaluations demonstrate that our method performs favorably against the state-of-the-art methods in a wide range of tracking scenarios.
Jie Guo 0004, Tingfa Xu, Ziyi Shen, Guokai Shi
IEEE Signal Process. Lett.3
2015 AppTrace: Dynamic trace on Android devices
abstract
Mass vulnerabilities involved in the Android alternative applications could threaten the security of the launched device or users data. To analyze the alternative applications, generally, researchers would like to observe applications' runtime features first. Then they need to decompile the target application and read the complicated code to figure out what the application really does. Traditional dynamic analysis methodology, for instance, the TaintDroid, uses dynamic taint tracking technique to mark information at source APIs. However, TaintDroid is limited to constraint on requiring target application to run in custom sandbox that might be not compatible with all the Android versions. For solving this problem and helping analysts to have insight into the runtime behavior, this paper presents AppTrace, a novel dynamic analysis system that uses dynamic instrumentation technique to trace member methods of target application that could be deployed in any version above Android 4.0. The paper presents an evaluation of AppTrace with 8 apps from Google Play as well as 50 open source apps from F-Droid. The results show that AppTrace could trace methods of target applications successfully and notify users effectively when some sensitive APIs are invoked.
Lingzhi Qiu, Zixiong Zhang, Ziyi Shen, Guozi Sun
ICC3