Zhao Zhang 0018

dblp:87/6853-18 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-1521-8163ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2025 RelationLMM: Large Multimodal Model as Open and Versatile Visual Relationship Generalist
abstract
Visual relationships are crucial for visual perception and reasoning, and cover tasks like Scene Graph Generation, Human-Object Interaction, and object affordance. Despite significant efforts, this field still suffers from the following limitations: specialists for a specific task without considering similar ones, strict and complex task formulations with limited flexibility, and underexploited reasoning with language and knowledge. To solve these limitations, we seek to build a new framework, one model for all tasks, over Large Multimodal Models (LMMs). LMMs offer the potential of unifying tasks, flexible forms, and reasoning with language. However, they fail to handle visual relationship tasks well. We find the obstacles include the conflicts between different tasks and insufficient instance-level information. We solve these problems by reforming the data for LMMs, rather than architectures, considering their strong language-in language-out capability. We propose to disassemble tasks into simple and common sub-tasks, verbally estimate instance confidence, and augment instance diversity, all without additional modules. These strategies help us build a visual relationship generalist, RelationLMM, with a simple architecture. Exhaustive experiments demonstrate RelationLMM is strong, generalizable and flexible to different tasks, with one model and one suite of weight.
Chi Xie 0001, Shuang Liang 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Link-Context Learning for Multimodal LLMs
abstract
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLLMs) and Large Language Models (LLMs) being trained on mega-scale datasets, recognizing unseen images or understanding novel concepts in a training-free manner remains a challenge. In-Context Learning (ICL) explores training-free few-shot learning, where models are encouraged to “learn to learn” from limited tasks and generalize to unseen tasks. In this work, we propose link-context learning (LCL), which emphasizes “reasoning from cause and effect” to augment the learning capabilities of MLLMs. LCL goes beyond traditional ICL by explicitly strengthening the causal relationship between the support set and the query set. By providing demonstrations with causal links, LCL guides the model to discern not only the analogy but also the underlying causal associations between data points, which empowers MLLMs to recognize unseen images and understand novel concepts more effectively. To facilitate the evaluation of this novel approach, we introduce the ISEKAI dataset, comprising exclusively of unseen generated image-label pairs designed for link-context learning. Extensive experiments show that our LCL-MLLM exhibits strong link-context learning capabilities to novel concepts over vanilla MLLMs. Code, demo, and dataset have been released.
Yan Tai, Weichen Fan, Zhao Zhang 0018, Ziwei Liu 0002
CVPR3
2023 Advancing Referring Expression Segmentation Beyond Single Image
abstract
Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expression. However, in broader real-world scenarios, it is not always possible to determine if the described object exists in a specific image. Generally, a collection of images is available, some of which potentially contain the target objects. To this end, we propose a more realistic setting, named Group-wise Referring Expression Segmentation (GRES), which expands RES to a group of related images, allowing the described objects to exist in a subset of the input image group. To support this new setting, we introduce an elaborately compiled dataset named Grouped Referring Dataset (GRD), containing complete group-wise annotations of the target objects described by given expressions. Moreover, we also present a baseline method named Grouped Referring Segmenter (GRSer), which explicitly captures the language-vision and intra-group vision-vision interactions to achieve state-of-the-art results on the proposed GRES setting and related tasks, such as Co-Salient Object Detection and traditional RES. Our dataset and codes are publicly released in https://github.com/shikras/d-cube.
Zhao Zhang 0018, Chi Xie 0001, Feng Zhu 0006, Rui Zhao 0001
ICCV2
2023 Described Object Detection: Liberating Object Detection with Flexible Expressions
abstract
Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to flexible language expressions for OVD and overcoming the limitation of REC only grounding the pre-existing object. We establish the research foundation for DOD by constructing a *Description Detection Dataset* ($D^3$). This dataset features flexible language expressions, whether short category names or long descriptions, and annotating all described objects on all images without omission. By evaluating previous SOTA methods on $D^3$, we find some troublemakers that fail current REC, OVD, and bi-functional methods. REC methods struggle with confidence scores, rejecting negative instances, and multi-target scenarios, while OVD methods face constraints with long and complex descriptions. Recent bi-functional methods also do not work well on DOD due to their separated training procedures and inference strategies for REC and OVD tasks. Building upon the aforementioned findings, we propose a baseline that largely improves REC methods by reconstructing the training data and introducing a binary classification sub-task, outperforming existing methods. Data and code are available at https://github.com/shikras/d-cube and related works are tracked in https://github.com/Charles-Xie/awesome-described-object-detection.
Chi Xie 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001, Shuang Liang 0001
NeurIPS2
2023 Sequential interactive image segmentation
abstract
Interactive image segmentation (IIS) is an important technique for obtaining pixel-level annotations. In many cases, target objects share similar semantics. However, IIS methods neglect this connection and in particular the cues provided by representations of previously segmented objects, previous user interaction, and previous prediction masks, which can all provide suitable priors for the current annotation. In this paper, we formulate a sequential interactive image segmentation (SIIS) task for minimizing user interaction when segmenting sequences of related images, and we provide a practical approach to this task using two pertinent designs. The first is a novel interaction mode. When annotating a new sample, our method can automatically propose an initial click proposal based on previous annotation. This dramatically helps to reduce the interaction burden on the user. The second is an online optimization strategy, with the goal of providing semantic information when annotating specific targets, optimizing the model with dense supervision from previously labeled samples. Experiments demonstrate the effectiveness of regarding SIIS as a particular task, and our methods for addressing it.
Zheng Lin 0005, Zhao Zhang 0018, Ziyue Zhu, Deng-Ping Fan, Xialei Liu
Comput. Vis. Media2
2023 Co-Salient Object Detection With Co-Representation Purification
abstract
Co-salient object detection (Co-SOD) aims at discovering the common objects in a group of relevant images. Mining a co-representation is essential for locating co-salient objects. Unfortunately, the current Co-SOD method does not pay enough attention that the information not related to the co-salient object is included in the co-representation. Such irrelevant information in the co-representation interferes with its locating of co-salient objects. In this paper, we propose a Co-Representation Purification (CoRP) method aiming at searching noise-free co-representation. We search a few pixel-wise embeddings probably belonging to co-salient regions. These embeddings constitute our co-representation and guide our prediction. For obtaining purer co-representation, we use the prediction to iteratively reduce irrelevant embeddings in our co-representation. Experiments on three datasets demonstrate that our CoRP achieves state-of-the-art performances on the benchmark datasets. Our source code is available at https://github.com/ZZY816/CoRP.
Ziyue Zhu, Zhao Zhang 0018, Zheng Lin 0005, Xing Sun 0001, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 FocusCut: Diving into a Focus View in Interactive Segmentation
abstract
Interactive image segmentation is an essential tool in pixel-level annotation and image editing. To obtain a high-precision binary segmentation mask, users tend to add interaction clicks around the object details, such as edges and holes, for efficient refinement. Current methods regard these repair clicks as the guidance to jointly determine the global prediction. However, the global view makes the model lose focus from later clicks, and is not in line with user intentions. In this paper, we dive into the view of clicks' eyes to endow them with the decisive role in object details again. To verify the necessity of focus view, we design a simple yet effective pipeline, named FocusCut, which integrates the functions of object segmentation and local refinement. After obtaining the global prediction, it crops click-centered patches from the original image with adaptive scopes to refine the local predictions progressively. Without user perception and parameters increase, our method has achieved state-of-the-art results. Extensive experiments and visualized results demonstrate that FocusCut makes hyper-fine segmentation possible for interactive image segmentation.
Zheng Lin 0005, Zheng-Peng Duan, Zhao Zhang 0018, Chunle Guo, Ming-Ming Cheng
CVPR3
2022 PAC-Net: Highlight Your Video via History Preference Modeling
Penghao Zhou, Chong Zhou, Zhao Zhang 0018, Xing Sun 0001
ECCV (34)4
2022 Multi-Mode Interactive Image Segmentation
abstract
Large-scale pixel-level annotations are scarce for current data-hungry medical image analysis models. For the fast acquisition of annotations, an economical and efficient interactive medical image segmentation method is urgently needed. However, current techniques usually fail in many cases, as their interaction styles cannot work on various inherent ambiguities of medical images, such as irregular shapes and fuzzy boundaries. To address this problem, we propose a multi-mode interactive segmentation framework for medical images, where diverse interaction modes can be chosen and allowed to cooperate with each other. In our framework, users can encircle the target regions with various initial interaction modes according to the structural complexity. Then, based on the initial segmentation, users can jointly utilize the region and boundary interactions to refine the mislabeled regions caused by different ambiguities. We evaluate our framework on extensive medical images, including X-ray, CT, MRI, ultrasound, endoscopy, and photo. Sufficient experimental results and user study show that our framework is a reliable choice for image annotation in various real scenes.
Zheng Lin 0005, Zhao Zhang 0018, Linghao Han, Shao-Ping Lu
ACM Multimedia2
2022 KnifeCut: Refining Thin Part Segmentation with Cutting Lines
abstract
Objects with thin structures remain challenging for current image segmentation techniques. Their outputs often do well in the main body but with thin parts unsatisfactory. In practical use, they inevitably need post-processing. However, repairing them is time-consuming and laborious, either in professional editing applications (e.g. PhotoShop) or by current interactive image segmentation methods (e.g. by click, scribble, and polygon). To refine the thin parts for unsatisfactory pre-segmentation, we propose an efficient interaction mode, where users only need to draw a line across the mislabeled thin part like cutting with a knife. This low-stress and intuitive action does not require the user to aim deliberately, and is friendly when using the mouse, touchpad, and mobile devices. Additionally, the line segment provides a contrasting prior because it passes through both the foreground and background regions and there must be thin part pixels on it. Based on the interaction idea, we propose KnifeCut, which offers the users two results, where one only focuses on the target thin part and the other provides the refinements for all thin parts that share similar features with the target one. To our best knowledge, KnifeCut is the first method to solve interactive thin structure refinement pertinently. Extensive experiments and visualized results further demonstrate its friendliness, convenience, and effectiveness. The project page is available on http://mmcheng.net/knifecut/.
Zheng Lin 0005, Zheng-Peng Duan, Zhao Zhang 0018, Chunle Guo, Ming-Ming Cheng
ACM Multimedia3
2021 Bilateral Attention Network for RGB-D Salient Object Detection
abstract
RGB-D salient object detection (SOD) aims to segment the most attractive objects in a pair of cross-modal RGB and depth images. Currently, most existing RGB-D SOD methods focus on the foreground region when utilizing the depth images. However, the background also provides important information in traditional SOD methods for promising performance. To better explore salient information in both foreground and background regions, this paper proposes a Bilateral Attention Network (BiANet) for the RGB-D SOD task. Specifically, we introduce a Bilateral Attention Module (BAM) with a complementary attention mechanism: foreground-first (FF) attention and background-first (BF) attention. The FF attention focuses on the foreground region with a gradual refinement style, while the BF one recovers potentially useful salient information in the background region. Benefited from the proposed BAM module, our BiANet can capture more meaningful foreground and background cues, and shift more attention to refining the uncertain details between foreground and background regions. Additionally, we extend our BAM by leveraging the multi-scale techniques for better SOD performance. Extensive experiments on six benchmark datasets demonstrate that our BiANet outperforms other state-of-the-art RGB-D SOD methods in terms of objective metrics and subjective visual comparison. Our BiANet can run up to 80 fps on 224×224 RGB-D images, with an NVIDIA GeForce RTX 2080Ti GPU. Comprehensive ablation studies also validate our contributions.
Zhao Zhang 0018, Zheng Lin 0005, Jun Xu 0019, Wenda Jin, Shao-Ping Lu, Deng-Ping Fan
IEEE Trans. Image Process.1
2021 Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
abstract
The use of RGB-D information for salient object detection (SOD) has been extensively explored in recent years. However, relatively few efforts have been put toward modeling SOD in real-world human activity scenes with RGB-D. In this article, we fill the gap by making the following contributions to RGB-D SOD: 1) we carefully collect a newSalientPerson (SIP) data set that consists of ~1 K high-resolution images that cover diverse real-world scenes from various viewpoints, poses, occlusions, illuminations, and background s; 2) we conduct a large-scale (and, so far, the most comprehensive) benchmark comparing contemporary methods, which has long been missing in the field and can serve as a baseline for future research, and we systematically summarize 32 popular models and evaluate 18 parts of 32 models on seven data sets containing a total of about 97k images; and 3) we propose a simple general architecture, called deep depth-depurator network (D3Net). It consists of a depth depurator unit (DDU) and a three-stream feature learning module (FLM), which performs low-quality depth map filtering and cross-modal feature learning, respectively. These components form a nested structure and are elaborately designed to be learned jointly. D3Net exceeds the performance of any prior contenders across all five metrics under consideration, thus serving as a strong model to advance research in this field. We also demonstrate that D3Net can be used to efficiently extract salient object masks from real scenes, enabling effective background-changing application with a speed of 65 frames/s on a single GPU. All the saliency maps, our new SIP data set, the D3Net model, and the evaluation tools are publicly available athttps://github.com/DengPingFan/D3NetBenchmark.
Deng-Ping Fan, Zheng Lin 0005, Zhao Zhang 0018, Menglong Zhu, Ming-Ming Cheng
IEEE Trans. Neural Networks Learn. Syst.3
2020 Interactive Image Segmentation With First Click Attention
abstract
In the task of interactive image segmentation, users initially click one point to segment the main body of the target object and then provide more points on mislabeled regions iteratively for a precise segmentation. Existing methods treat all interaction points indiscriminately, ignoring the difference between the first click and the remaining ones. In this paper, we demonstrate the critical role of the first click about providing the location and main body information of the target object. A deep framework, named First Click Attention Network (FCA-Net), is proposed to make better use of the first click. In this network, the interactive segmentation result can be much improved with the following benefits: focus invariance, location guidance, and error-tolerant ability. We then put forward a click-based loss function and a structural integrity strategy for better segmentation effect. The visualized segmentation results and sufficient experiments on five datasets demonstrate the importance of the first click and the superiority of our FCA-Net.
Zheng Lin 0005, Zhao Zhang 0018, Lin-Zhuo Chen, Ming-Ming Cheng, Shao-Ping Lu
CVPR2
2020 Gradient-Induced Co-Saliency Detection
Zhao Zhang 0018, Wenda Jin, Jun Xu 0019, Ming-Ming Cheng
ECCV (12)1
2018 Low Resolution Face Recognition and Reconstruction Via Deep Canonical Correlation Analysis
abstract
Low-resolution (LR) face identification is always a challenge in computer vision. In this paper, we propose a new LR face recognition and reconstruction method using deep canonical correlation analysis (DCCA). Unlike linear CCA-based methods, our proposed method can learn flexible nonlinear representations by passing LR and high-resolution (HR) image principal component features through multiple stacked layers of nonlinear transformation. As the nonlinear transformation in deep neural networks is implicit, we apply radial basis function based neural network to learn an explicit mapping between principal components and correlational features. In addition, we also design two residual compensation methods for identification and vision enhancement, respectively. The proposed approach is compared with existing LR face recognition and reconstruction algorithms. A number of experimental results on benchmark datasets have demonstrated the effectiveness and robustness of our method.
Zhao Zhang 0018, Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010
ICASSP1
2018 Learning Parallel Canonical Correlations for Scale-Adaptive Low Resolution Face Recognition
abstract
Low resolution is one of the main obstacles in the application of face recognition. Although many methods have been proposed to improve the problem, they assume that low-resolution (LR) face images have a uniform scale. In real scenarios, this prerequisite is very harsh. In this paper, we propose a scale-adaptive LR face recognition approach based on two-dimensional multi-set canonical correlation analysis (2DM-CCA), where face image matrix does not need to be previously transformed into a vector. In the proposed method, training sets with different resolutions are treated as different views, and then projected in parallel into a latent coherent space where the consistency of multi-view face data is maximally enhanced. When a new LR face image with an arbitrary scale is input, we first transform it by using the left and right projection matrices of an appropriate training view, and then reconstruct its high resolution facial feature by neighborhood reconstruction. Experimental results show that our proposed method is more effective and efficient than several existing methods.
Yun-Hao Yuan 0001, Zhao Zhang 0018, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Xiaobo Shen 0001
ICPR2
2017 Supervised Deep Canonical Correlation Analysis for Multiview Feature Learning
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Min Ruan, Zhao Zhang 0018
ICONIP (6)6
2017 Face Hallucination and Recognition Using Kernel Canonical Correlation Analysis
Zhao Zhang 0018, Yun-Hao Yuan 0001, Yun Li 0010, Bin Li 0006, Jipeng Qiang
ICONIP (6)1