Yutong Gao 0001

dblp:174/3887-1 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-6766-0703ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SinColor: Uncertainty-Guided Single-Step Diffusion for Image Colorization
abstract
Image colorization is a fundamental yet challenging task in computer vision, aiming to recover plausible and spatially coherent colors from grayscale images. Recent advancements in diffusion models have enabled significant progress in this field, yet existing methods predominantly rely on multi-step diffusion processes. While effective for generating high-frequency details, these approaches are suboptimal for colorization, as color information is inherently low-frequency, spatially smooth, and globally consistent. This mismatch leads to two critical limitations: 1) color artifacts and inconsistency due to excessive noise in the color space, and 2) high computational cost that hinders practical application. In this work, we propose a novel single-step diffusion framework for efficient and high-quality image colorization. We introduce a color uncertainty estimation (CUE) module to identify reliable and uncertain regions in the image, allowing the model to prioritize local certainty while reasoning about confused regions. To focus the model on low-frequency color generation, we directly encode the grayscale image into a latent representation, remove structural components in the output, and reconstruct the final image via efficient decoding. Extensive experiments on ImageNet, COCO-Stuff, and Extended COCO-Stuff demonstrate that our approach achieves state-of-the-art performance while reducing inference time by 98% and trainable parameters by 97% compared to leading multi-step diffusion methods. Our contributions include a systematic analysis of diffusion-based colorization, a lightweight yet effective uncertainty-aware framework, and comprehensive validation of its efficiency and effectiveness.
Yutong Gao 0001, Congyan Lang, Yidian Liu, Fayao Liu, Guoshun Nan, Yunchao Wei
IEEE Trans. Image Process.1
2026 Investigate Interactive Semantic Segmentation via an Uncertainty Mining View
abstract
With the rapid development of intelligence media, traditional semantic segmentation has shown excellent potential in application scenarios like autonomous driving. However, due to limited performance, traditional segmentation models usually lead to poor user experiences in applications that require high segmentation precision. Therefore, interactive semantic segmentation (ISS) is gaining the attention is gaining attention due to its capability to generate high-precision semantic segmentation results through a few user-provided clicks for experience improvement, which thus has a promising development prospect in fine-grained application scenarios,e.g., virtual reality, smart medical, data annotation,etc.. For good interaction efficiency, most existing interactive methods make efforts to conduct suitable click simulation strategies and reasonable click encoding methods, aiming at the robust understanding of diverse user clicks and translating comprehensible user intent,i.e., assign the correct category to the clicked area, for the neural network. Though proved effective, their designs ignore the uncertainty hiding in the extracted interaction features, which reflects the interaction difficulty and the user clicking intents. This can lead to inappropriate click simulation and click encoding, limiting the interaction efficiency. Hence we focus on exploring a reasonable ISS scheme via an uncertainty mining view. Specifically, we propose an uncertainty-based class-balanced click sampling (UCCS) simulation strategy by considering both the uncertainty of the click simulation region and its semantic imbalance, to form a reasonable click distribution. Furthermore, we propose a semantic uncertainty residual encoding (SURE) method to better embed the user's intention into the localization maps, by mining semantic confusion between the click and misprediction classes. We prove the effectiveness of our design through extensive experiments and initially analyze the importance of uncertainty mining for the ISS. Our model can achieve state-of-the-art performance on three semantic segmentation benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Xun Xu 0002, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Multim.1
2025 VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-Based Group Relative Policy Optimization
abstract
Understanding hour-long videos with multi-modal large language models (MM-LLMs) enriches the landscape of human-centered AI applications. However, for end-to-end video understanding with LLMs, uniformly sampling video frames results in LLMs being overwhelmed by a vast amount of irrelevant information as video length increases. Existing hierarchical key frame extraction methods improve the accuracy of video understanding but still face two critical challenges. 1) How can the interference of extensive redundant information in long videos be mitigated? 2) How can a model dynamically adapt to complex hierarchical structures while accurately identifying key frames? To address these issues, we propose VideoMiner, which iteratively segments, captions, and clusters long videos, forming a hierarchical tree structure. The proposed VideoMiner progresses from long videos to events to frames while preserving temporal coherence, effectively addressing the first challenge. To precisely locate key frames, we introduce T-GRPO, a tree-based group relative policy optimization in reinforcement learning method that guides the exploration of the VideoMiner. The proposed T-GRPO is specifically designed for tree structures, integrating spatiotemporal information at the event level while being guided by the question, thus solving the second challenge. We achieve superior performance in all long-video understanding tasks and uncover several interesting insights. Our proposed T-GRPO surprisingly incentivizes the model to spontaneously generate a reasoning chain. Additionally, the designed tree growth auxin dynamically adjusts the expansion depth, obtaining accuracy and efficiency gains. The code is publicly available at https://github.com/caoxinye/VideoMiner.
Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Yuqi Pan, Tianhao Hou, Yutong Gao 0001
ICCV9
2025 SkinMamba: Segmentation and Classification of Skin Cancer with Multi-level Context Understanding
abstract
Skin cancer accounts for nearly 40% of all cancer cases. Segmentation and classification of the lesions can help medical professionals delineate the boundaries of skin lesions and then categorize the type, ensuring timely and efficient intervention. However, the nuance of lesions, hairs over the skin, and blurred boundaries make such a task quite challenging. This motivates us to propose SkinMamba, a novel method that explores Mamba to learn the multi-level context of skin lesions, thereby enabling more accurate segmentation of affected areas and determination of lesion type. Specifically, we introduce a novel encoder termed SkinBlock, by integrating convolutional layers with the Mamba approach, and such an encoder can effectively capture the global and local context of the lesions. We feed the output segmentation clues and three features with different focal areas to the classifier. The classifier consists of the proposed SkinBlock and ResNet50. By doing so, our SkinMamba can properly tackle the challenges mentioned above. Experiments on two public benchmarks show the effectiveness of the proposed SkinMamba.
Guoshun Nan, Chengyao Jia, Yutong Gao 0001, Zuye Xiao
IJCNN5
2025 Mining Semantic Correlations Between Mispredictions and Corrections for Interactive Semantic Segmentation
abstract
Interactive semantic segmentation pursues high-quality segmentation results at the cost of a small number of user clicks. It is attracting more and more research attention for its convenience in labeling semantic pixel-level data. Existing interactive segmentation methods often pursue higher interaction efficiency by mining the latent information of user clicks or exploring efficient interaction manners. However, these works neglect to explicitly exploit the semantic correlations between user corrections and model mispredictions, thus suffering from two flaws. First, similar prediction errors frequently occur in actual use, causing users to repeatedly correct them. Second, the interaction difficulty of different semantic classes varies across images, but existing models use monotonic parameters for all images which lack semantic pertinence. Therefore, in this article, we explore the semantic correlations existing in corrections and mispredictions by proposing a simple yet effective online learning solution to the above problems, named correction-misprediction correlation mining (CM2). Specifically, we leverage the correction-misprediction similarities to design a confusion memory module (CMM) for automatic correction when similar prediction errors reappear. Furthermore, we measure the semantic interaction difficulty by counting the correction-misprediction pairs and design a challenge adaptive convolutional layer (CACL), which can adaptively switch different parameters according to interaction difficulties to better segment the challenging classes. Our method requires no extra training besides the online learning process and can effectively improve interaction efficiency. Our proposed CM2 achieves state-of-the-art results on three public semantic segmentation benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Chuan-Sheng Foo, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Multi-holder Role and Strange Attractor-Based Data Possession Proof in Medical Clouds
Jingchen Wu, Yutong Gao 0001
ICA3PP (2)5
2024 Enhancing Text-Image Person Re-identification via Intra-Class Relevance Learning
Yutong Gao 0001, Chaomurilige Wang
ICA3PP (2)2
2024 Towards Information Sharing Beetle Antennae Search Optimization
Xuan Liu 0008, Chenyan Wang, Lefeng Zhang, Xianggan Liu, Yutong Gao 0001
ICA3PP (2)6
2024 Language-Based Colorization with Sparse Attention and Multi-scale Cross-Modal Semantic Alignment
Yutong Gao 0001, Xuan Liu 0008, Lefeng Zhang, Xianggan Liu, Shan Jiang 0012
ICA3PP (5)2
2024 Data Poisoning Attack Against Reinforcement Learning from Human Feedback in Robot Control Tasks
Zihui Zhou, Yutong Gao 0001, Minfeng Qi
ICA3PP (1)2
2024 Poster: Four Eyes See More Than Two: Collective Intelligence Schemes for Virtual Cluster Allocation
abstract
With the benefit of virtualization technology and an ascending trend towards more cloud applications, virtual cluster (VC) allocation concerning different requirements of application deployment is vital to cloud datacenters and edge networks. Most of the state-of-the-art approaches for VC allocation problems are probability-based heuristics, and none of them absolutely outperforms any other approach in any case. Thus, it is necessary to collect and select effective and efficient approaches to jointly solve diverse and customizable VC allocation problems. This paper presents VCA-Solver, a VC allocation solver that can draw on collectively feasible schemes from multiple effective approaches. We describe the architecture of VCA-Solver and present a preliminary evaluation with a prototype implementation.
Xuan Liu 0008, Chenyan Wang, Xiangyu Qu, Chang Xu 0016, Yutong Gao 0001
MobiSys5
2024 Dynamic Interaction Dilation for Interactive Human Parsing
abstract
Interactive segmentation pursues generating high-quality pixel-level predictions with a few user-provided clicks, which is gaining attention for its convenience in segmentation data annotation. Users are allowed to iteratively refine the prediction by adding clicks until the result is satisfactory. Existing interactive methods usually transform the clicks into a set of localization maps by Euclidian distance computation or RGB texture extraction to guide the segmentation, which makes the click transformation a core module in interactive segmentation networks. However, when adopted in human images where large poses, occlusions, and bad illuminations are prevailing, prior transformation methods tend to cause uncorrectable overlapping across localization maps which are difficult to form a good match among human parts. Furthermore, the inappropriately transformed information is hard to be refined with the static transformation manner which is out of tune with the dynamically refined interaction process. Hence, we design a dynamic transformation scheme for interactive human parsing (IHP) named Dynamic Interaction Dilation Net (DID-Net), which serves as an initial attempt to break the limitations of static transformation while capturing long-range dependencies of clicks within each human part. Specifically, we construct a Dynamic Dilation Module (DD-Module) to dilate clicks radially in several directions assisted by human body edge detection to refine the dilation quality in each interaction iteration. Furthermore, we propose an Adaptive Interaction Excitation Block (AIE-Block) to exploit potential semantic clues buried in the dilated clicks. Our DID-Net achieves state-of-the-art performance on 3 public human parsing benchmarks.
Yutong Gao 0001, Congyan Lang, Fayao Liu, Yuanzhouhan Cao, Yunchao Wei
IEEE Trans. Multim.1
2023 Integrating topology beyond descriptions for zero-shot learning
Yutong Gao 0001, Congyan Lang, Yidong Li, Hongzhe Liu 0001, Fayao Liu
Pattern Recognit.2
2023 Clicking Matters: Towards Interactive Human Parsing
abstract
In this work, we focus on Interactive Human Parsing (IHP), which aims to segment a human image into multiple human body parts with guidance from users’ interactions. This new task inherits the class-aware property of human parsing, which cannot be well solved by traditional interactive image segmentation approaches that are generally class-agnostic. To tackle this new task, we first exploit user clicks to identify different human parts in the given image. These clicks are subsequently transformed into semantic-aware localization maps, which are concatenated with the RGB image to form the input of the segmentation network and generate the initial parsing result. To enable the network to better perceive user's purpose during the correction process, we investigate several principal ways for the refinement, and reveal that random-sampling-based click augmentation is the best way for promoting the correction effectiveness. Furthermore, we also propose a semantic-perceiving loss (SP-loss) to augment the training, which can effectively exploit the semantic relationships of clicks for better optimization. To the best knowledge, this work is the first attempt to tackle the human parsing task under the interactive setting. Our IHP solution achieves 85% mIoU on the benchmark LIP, 80% mIoU on PASCAL-Person-Part and CIHP, 75% mIoU on Helen with only 1.95, 3.02, 2.84 and 1.09 clicks per class respectively. These results demonstrate that we can simply acquire high-quality human parsing masks with only a few human effort. We hope this work can motivate more researchers to develop data-efficient solutions to IHP in the future.
Yutong Gao 0001, Liqian Liang, Congyan Lang, Songhe Feng, Yidong Li, Yunchao Wei
IEEE Trans. Multim.1
2021 Fine-Grained Semantic Image Synthesis with Object-Attention Generative Adversarial Network
abstract
Semantic image synthesis is a new rising and challenging vision problem accompanied by the recent promising advances in generative adversarial networks. The existing semantic image synthesis methods only consider the global information provided by the semantic segmentation mask, such as class label, global layout, and location, so the generative models cannot capture the rich local fine-grained information of the images (e.g., object structure, contour, and texture). To address this issue, we adopt a multi-scale feature fusion algorithm to refine the generated images by learning the fine-grained information of the local objects. We propose OA-GAN, a novel object-attention generative adversarial network that allows attention-driven, multi-fusion refinement for fine-grained semantic image synthesis. Specifically, the proposed model first generates multi-scale global image features and local object features, respectively, then the local object features are fused into the global image features to improve the correlation between the local and the global. In the process of feature fusion, the global image features and the local object features are fused through the channel-spatial-wise fusion block to learn ‘what’ and ‘where’ to attend in the channel and spatial axes, respectively. The fused features are used to construct correlation filters to obtain feature response maps to determine the locations, contours, and textures of the objects. Extensive quantitative and qualitative experiments on COCO-Stuff, ADE20K and Cityscapes datasets demonstrate that our OA-GAN significantly outperforms the state-of-the-art methods.
Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001
ACM Trans. Intell. Syst. Technol.6
2020 End-to-End Text-to-Image Synthesis with Spatial Constrains
abstract
Although the performance of automatically generating high-resolution realistic images from text descriptions has been significantly boosted, many challenging issues in image synthesis have not been fully investigated, due to shapes variations, viewpoint changes, pose changes, and the relations of multiple objects. In this article, we propose a novel end-to-end approach for text-to-image synthesis with spatial constraints by mining object spatial location and shape information. Instead of learning a hierarchical mapping from text to image, our algorithm directly generates multi-object fine-grained images through the guidance of the generated semantic layouts. By fusing text semantic and spatial information into a synthesis module and jointly fine-tuning them with multi-scale semantic layouts generated, the proposed networks show impressive performance in text-to-image synthesis for complex scenes. We evaluate our method both on single-object CUB dataset and multi-object MS-COCO dataset. Comprehensive experimental results demonstrate that our method significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.
Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001
ACM Trans. Intell. Syst. Technol.6
2018 Rapid Human Finding with Motion Segmentation for Mobile Robot
abstract
A computer vision method is presented for the mobile robot to find humans in scene. Face detection is used for confirming humans. In order to reduce regions of search, optical flow algorithm is used to segment the image in advance. Asymmetric problems in face detection are explained, and relative solutions are put forward by bootstrapping strategy and asymmetric adaboost algorithm. In addition, fisher discriminant analysis further improves the performance of face detection. Multi-view face models are trained to accommodate practical face detection application. At last, experiments demonstrate that our multi-view face detector achieves high detection accuracy and fast detection speed on both standard testing datasets and real-life images.
Yutong Gao 0001, Weimin Lei, Xie Xie
Int. J. Pattern Recognit. Artif. Intell.1