Zijian Zhou 0002

dblp:73/1606-2 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-3315-3962ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 DC-RRG: Diagnosis-centered cascaded radiology report generation
Zinan Hong, Zijian Zhou 0002, Miaojing Shi, Jin-Gang Yu, Jingping Yun, Shuangping Huang
Expert Syst. Appl.3
2025 Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge Transfer
abstract
Generalized few-shot semantic segmentation (GFSS) aims to segment objects of both base and novel classes, using sufficient samples of base classes and few samples of novel classes. Representative GFSS approaches typically employ a two-phase training scheme, involving base class pre-training followed by novel class fine-tuning, to learn the classifiers for base and novel classes respectively. Nevertheless, distribution gap exists between base and novel classes in this process. To narrow this gap, we exploit effective knowledge transfer from base to novel classes. First, a novel prototype modulation module is designed to modulate novel class prototypes by exploiting the correlations between base and novel classes. Second, a novel classifier calibration module is proposed to calibrate the weight distribution of the novel classifier according to that of the base classifier. Furthermore, existing GFSS approaches suffer from a lack of contextual information for novel classes due to their limited samples, we thereby introduce a context consistency learning scheme to transfer the contextual knowledge from base to novel classes. Extensive experiments on PASCAL-5i and COCO-20i demonstrate that our approach significantly enhances the state of the art in the GFSS setting.
Xinyue Chen 0009, Miaojing Shi, Zijian Zhou 0002, Lianghua He, Sophia Tsoka
AAAI3
2025 MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
abstract
Autonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying on the model’s spatial understanding capabilities. Prior works typically express spatial information through textual representations of coordinates, resulting in semantic gaps between visual coordinate representations and textual descriptions. This oversight hinders the accurate transmission of spatial information and increases the expressive burden. To address this, we propose a novel Marker-based Prompt learning framework (MPDrive), which represents spatial coordinates by concise visual markers, ensuring linguistic expressive consistency and enhancing the accuracy of both visual perception and spatial expression in AD-VQA. Specifically, we create marker images by employing a detection expert to overlay object regions with numerical labels, converting complex textual coordinate generation into straightforward text-based visual marker predictions. Moreover, we fuse original and marker images as scene-level features and integrate them with detection priors to derive instance-level features. By combining these features, we construct dual-granularity visual prompts that stimulate the LLM’s spatial perception capabilities. Extensive experiments on the DriveLM and CODA-LM datasets show that MPDrive achieves state-of-the-art performance, particularly in cases requiring sophisticated spatial understanding.
Wenjie Peng, Zijian Zhou 0002, Miaojing Shi, Shuangping Huang
CVPR5
2025 Learning Flow Fields in Attention for Controllable Person Image Generation
abstract
Controllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person’s appearance or pose. However, prior methods often distort fine-grained details from the reference image, despite achieving high overall image quality. We attribute these distortions to inadequate attention to corresponding regions in the reference image. To address this, we thereby propose learning flow fields in attention (Leffa), which explicitly guides the target query to attend to the correct reference key in the attention layer during training. Specifically, it is realized via a regularization loss on top of the attention map within a diffusionbased baseline. Our extensive experiments show that Leffa achieves state-of-the-art performance in controlling appearance and pose, significantly reducing fine-grained detail distortion while maintaining high image quality. Additionally, we show that our loss is model-agnostic and can be used to improve the performance of other diffusion models.
Zijian Zhou 0002, Shikun Liu, Kam Woh Ng, Tian Xie 0003, Yuren Cong, Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang 0002, Miaojing Shi, Sen He 0001
CVPR1
2025 VLPrompt-PSG: Vision-Language Prompting for Panoptic Scene Graph Generation
Zijian Zhou 0002, Holger Caesar, Miaojing Shi
Int. J. Comput. Vis.1
2024 OpenPSG: Open-Set Panoptic Scene Graph Generation via Large Multimodal Models
Zijian Zhou 0002, Holger Caesar, Miaojing Shi
ECCV (10)1
2023 HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph Generation
abstract
Panoptic Scene Graph generation (PSG) is a recently proposed task in image scene understanding that aims to segment the image and extract triplets of subjects, objects and their relations to build a scene graph. This task is particularly challenging for two reasons. First, it suffers from a long-tail problem in its relation categories, making naive biased methods more inclined to high-frequency relations. Existing unbiased methods tackle the long-tail problem by data/loss rebalancing to favor low-frequency relations. Second, a subject-object pair can have two or more semantically overlapping relations. While existing methods favor one over the other, our proposed HiLo framework lets different network branches specialize on low and high frequency relations, enforce their consistency and fuse the results. To the best of our knowledge we are the first to propose an explicitly unbiased PSG method. In extensive experiments we show that our HiLo framework achieves state-of-the-art results on the PSG task. We also apply our method to the Scene Graph Generation task that predicts boxes instead of masks and see improvements over all baseline methods. Code is available at https://github.com/franciszzj/HiLo.
Zijian Zhou 0002, Miaojing Shi, Holger Caesar
ICCV1
2023 Text Promptable Surgical Instrument Segmentation with Vision-Language Models
abstract
In this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehension of surgical instruments and adaptability to new instrument types. Inspired by recent advancements in vision-language models, we leverage pretrained image and text encoders as our model backbone and design a text promptable mask decoder consisting of attention- and convolution-based prompting schemes for surgical instrument segmentation prediction. Our model leverages multiple text prompts for each surgical instrument through a new mixture of prompts mechanism, resulting in enhanced segmentation performance. Additionally, we introduce a hard instrument area reinforcement module to improve image feature comprehension and segmentation precision. Extensive experiments on several surgical instrument segmentation datasets demonstrate our model's superior performance and promising generalization capability. To our knowledge, this is the first implementation of a promptable approach to surgical instrument segmentation, offering significant potential for practical application in the field of robotic-assisted surgery. Code is available at https://github.com/franciszzj/TP-SIS.
Zijian Zhou 0002, Oluwatosin Alabi, Tom Vercauteren, Miaojing Shi
NeurIPS1
2020 Gaussian Vector: An Efficient Solution for Facial Landmark Detection
Yilin Xiong, Zijian Zhou 0002, Yuhao Dou, Zhizhong Su
ACCV (5)2
2019 A New Parallel Detection-Recognition Approach for End-to-End Scene Text Extraction
abstract
In this work, we present a new conceptually simple and flexible network for accurate scene text extraction, which handle text detection and text recognition concurrently. Apart from solving feature sharing by designing a unified network, we implement end-to-end training of the whole system by a set of new optimization strategies. More importantly, this method highlights its novel parallel detection-recognition structure, which constructs a loose connection between both detection and recognition. This loose connection is embodied in the definition of overall loss and derivative back propagation for the model parameter updating, which automatically balances the contribution of the two branches to the system performance. It is different from the existing end-to-end methods where two subtasks are connected serially and thus yielding heavy dependence of the predecessor text detection task on the follow-up text recognition task and sensitivity of recognition to detection noise. In addition, a simple Mask-Rectifier mechanism is applied to easily adapt our system to incidental text recognition with arbitrary orientation and shape. Experiment results on Incidental Scene Text ICDAR2015 dataset surpass the current state-of-the-art FOTS method, as suggests the effectiveness of the proposed approach.
Zijian Zhou 0002, Zhizhong Su, Shuangping Huang
ICDAR2