Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qian Yu 0015

dblp:16/3790-15 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 78% Generative modeling · 10% Image recognition and object detection · 9%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
dichotomous image segmentation
1.622025
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity · ICLR 2025
Multi-View Aggregation Network for Dichotomous Image Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding
image segmentation
1.622025
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity · ICLR 2025
Multi-View Aggregation Network for Dichotomous Image Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
diffusion-based segmentation
0.912025
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity · ICLR 2025
Computer vision › Image recognition and object detection
object detection
0.812024
Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline · CVPR 2024
Computer vision › Segmentation and scene understanding
object segmentation
0.812024
Multi-View Aggregation Network for Dichotomous Image Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding › interactive segmentation
point-based segmentation
0.812024
Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline · CVPR 2024
Image and video processing
edge detection
0.312025
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity · ICLR 2025
Computer vision › 3D vision › 3d scene understanding › multi-view understanding › multi-view fusion
multi-view feature aggregation
0.212024
Multi-View Aggregation Network for Dichotomous Image Segmentation · CVPR 2024

Methods — techniques the papers use, named apart from their topics

one-step denoising · 1.7edge generation · 1.7diffusion model · 1.7multi-view aggregation · 0.8multi-dimensional collaborative network · 0.8encoder-decoder · 0.8distance-adaptive mask generation · 0.8attention fusion · 0.8
YearPublicationVenuePosition
2025 High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
abstract
In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets comprising billions of image-text pairs, such as SD V2.1, have revolutionized text-to-image synthesis by delivering exceptional quality, fine detail resolution, and strong contextual awareness, making them an attractive solution for high-resolution image segmentation. To this end, we propose DiffDIS, a diffusion-driven segmentation model that taps into the potential of the pre-trained U-Net within diffusion models, specifically designed for high-resolution, fine-grained object segmentation. By leveraging the robust generalization capabilities and rich, versatile image representation prior of the SD models, coupled with a task-specific stable one-step denoising approach, we significantly reduce the inference time while preserving high-fidelity, detailed generation. Additionally, we introduce an auxiliary edge generation task to not only enhance the preservation of fine details of the object boundaries, but reconcile the probabilistic nature of diffusion with the deterministic demands of segmentation. With these refined strategies in place, DiffDIS serves as a rapid object mask generation model, specifically optimized for generating detailed binary maps at high resolutions, while demonstrating impressive accuracy and swift processing. Experiments on the DIS5K dataset demonstrate the superiority of DiffDIS, achieving state-of-the-art results through a streamlined inference process. The source code will be publicly available at \href{https://github.com/qianyu-dlut/DiffDIS}{DiffDIS}.
Qian Yu 0015, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115, Lihe Zhang, Huchuan Lu
ICLR1
2024 Multi-View Aggregation Network for Dichotomous Image Segmentation
abstract
Dichotomous Image Segmentation (DIS) has recently emerged towards high-precision object segmentation from high-resolution natural images. When designing an effective DIS model, the main challenge is how to balance the semantic dispersion of high-resolution targets in the small receptive field and the loss of high-precision details in the large receptive field. Existing methods rely on tedious multiple encoder-decoder streams and stages to gradually complete the global localization and local refinement. Human visual system captures regions of interest by observing them from multiple views. Inspired by it, we model DIS as a multi-view object perception problem and provide a parsi-monious multi-view aggregation network (MVANet), which unifies the feature fusion of the distant view and close-up view into a single stream with one encoder-decoder structure. With the help of the proposed multi-view complementary localization and refinement modules, our approach established long-range, profound visual interactions across multiple views, allowing the features of the detailed close-up view to focus on highly slender structures. Experiments on the popular DIS-5K dataset show that our MVANet significantly outperforms state-of-the-art methods in both accuracy and speed. The source code and datasets will be publicly available at MVANet.
Qian Yu 0015, Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
CVPR1
2024 Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline
abstract
We conduct a comprehensive study on a new task named power battery detection (PBD), which aims to localize the dense cathode and anode plates endpoints from X-ray images to evaluate the quality of power batteries. Existing manufacturers usually rely on human eye observation to complete PBD, which makes it difficult to balance the accuracy and efficiency of detection. To address this issue and drive more attention into this meaningful task, we first elaborately collect a dataset, called X-ray PBD, which has 1,500 diverse X-ray images selected from thousands of power batteries of 5 manufacturers, with 7 different visual interference. Then, we propose a novel segmentation-based solution for PBD, termed multi-dimensional collaborative network (MDCNet). With the help of line and counting predictors, the representation of the point segmentation branch can be improved at both semantic and detail aspects. Besides, we design an effective distance-adaptive mask generation strategy, which can alleviate the visual challenge caused by the inconsistent distribution density of plates to provide MDCNet with stable supervision. Without any bells and whistles, our segmentation-based MDCNet consistently outperforms various other corner detection, crowd counting and general/tiny object detection-based so-lutions, making it a strong baseline that can help facilitate future research in PBD. Finally, we share some potential difficulties and works for future researches. The source code and datasets will be publicly available at X-ray PBD.
Xiaoqi Zhao 0003, Youwei Pang, Zhenyu Chen 0001, Qian Yu 0015, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu
CVPR4