Chaoyi Hong

dblp:291/8504 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-1306-6634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 SPNet: Leveraging Sketch Proxies for Robust Cross-View Geo-Localization
abstract
Cross-view geo-localization (CVGL) determines the geographic location of a query image by matching it with the most similar GPS-tagged satellite images. Existing methods have improved cross-view feature consistency through complex network architectures but still face challenges under significant illumination variations. Inspired by cognitive psychology, we observe that humans can selectively ignore color distractions and focus on structural information under certain circumstances. Therefore, we propose an SPNet, leveraging sketches to enhance structural feature consistency and achieve illumination-robust CVGL. SPNet employs Sketch Proxy Training, which introduces both RGB and sketch images during training. By leveraging the structural consistency provided by sketches under color variations, it mitigates the network’s over-reliance on color cues and achieves robustness to illumination changes. To further integrate multi-source information, we design a Recurrent Progressive Channel Attention (RPCA) module that progressively selects and reweights channel features, effectively combining the semantic cues of RGB images with the structural information of sketches. In addition, we introduce a High-Pass Cross-Layer Connection (HPCL) to transmit high-frequency information across feature layers, emphasizing edges and details to reinforce structural modeling and cross-layer feature consistency. Extensive experiments demonstrate that our SPNet achieves superior performance on several CVGL datasets, including University-1652, SUES-200, CVUSA, and CVACT, and ranks first on the public leaderboard of the University-160k-WX dataset released by UAVM 2024, further validating its robustness and generalization capability.
Shuaiyuan Du, Chaoyi Hong, Zhiguo Cao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion
abstract
Radar-Camera depth estimation aims to predict dense and accurate metric depth by fusing input images and Radar data. Model efficiency is crucial for this task in pursuit of real-time processing on autonomous vehicles and robotic platforms. However, due to the sparsity of Radar returns, the prevailing methods adopt multi-stage frameworks with intermediate quasi-dense depth, which are time-consuming and not robust. To address these challenges, we propose TacoDepth, an efficient and accurate Radar-Camera depth estimation model with one-stage fusion. Specifically, the graph-based Radar structure extractor and the pyramid-based Radar fusion module are designed to capture and integrate the graph structures of Radar point clouds, delivering superior model efficiency and robustness without relying on the intermediate depth results. Moreover, TacoDepth can be flexible for different inference modes, providing a better balance of speed and accuracy. Extensive experiments are conducted to demonstrate the efficacy of our method. Compared with the previous state-of-the-art approach, TacoDepth improves depth accuracy and processing speed by 12.8% and 91.8%. Our work provides a new perspective on efficient Radar-Camera depth estimation.
Yiran Wang 0005, Jiaqi Li 0007, Chaoyi Hong, Ruibo Li, Liusheng Sun, Xiao Song 0002, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
CVPR3
2025 Dynamic Beauty is Easy to Find: A Large-Scale Composition-Aware Dataset and an End-to-End Framework for Video Reframing
abstract
Video reframing, which converts landscape-oriented (LO) to portrait-oriented (PO) video for some PO devices such as smartphones and tablets, faces challenges. Existing approaches mainly follow a multi-step pipeline to preserve video content that ignore composition quality due to lack of large-scale datasets. To address these challenges, we propose a fully automated composition-aware dataset using vision-language models and image composition assessment models, pairing LO videos with high-quality PO versions. We then propose an end-to-end model with an attention-aware backbone and a time-aware consistency module. Experiments show our approach outperforms others in efficiency and effectiveness, proving that composition awareness and end-to-end modeling are critical for video reframing.
Sitian Gu, Chaoyi Hong, Zhiguo Cao 0001
ACM Multimedia3
2025 NVDS$^{\mathbf{+}}$+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation
abstract
Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient and lacks robustness. An alternative is learning to enforce consistency from data, which requires well-designed models and sufficient video depth data. To address both challenges, we introduce NVDS that stabilizes inconsistent depth estimated by various single-image models in a plug-and-play manner. We also elaborate a large-scale Video Depth in the Wild (VDW) dataset, which contains 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset. Additionally, a bidirectional inference strategy is designed to improve consistency by adaptively fusing forward and backward predictions. We instantiate a model family ranging from small to large scales for different applications. The method is evaluated on VDW dataset and three public benchmarks. To further prove the versatility, we extend NVDS to video semantic segmentation and several downstream applications like bokeh rendering, novel view synthesis, and 3D reconstruction. Experimental results show that our method achieves significant improvements in consistency, accuracy, and efficiency. Our work serves as a solid baseline and data foundation for learning-based video depth estimation.
Yiran Wang 0005, Min Shi 0004, Jiaqi Li 0007, Chaoyi Hong, Zihao Huang 0001, Juewen Peng, Zhiguo Cao 0001, Jianming Zhang 0001, Ke Xian, Guosheng Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Unifying Automatic and Interactive Matting with Pretrained ViTs
abstract
Automatic and interactive matting largely improve image matting by respectively alleviating the need for auxil-iary input and enabling object selection. Due to different settings on whether prompts exist, they either suffer from weakness in instance completeness or region details. Also, when dealing with different scenarios, directly switching between the two matting models introduces inconvenience and higher workload. Therefore, we wonder whether we can al-leviate the limitations of both settings while achieving unification to facilitate more convenient use. Our key idea is to offer saliency guidance for automatic mode to enable its attention to detailed regions, and also refine the instance completeness in interactive mode by replacing the binary mask guidance with a more probabilistic form. With different guidance for each mode, we can achieve unification through adaptable guidance, defined as saliency information in automatic mode and user cue for interactive one. It is instantiated as candidate feature in our method, an automatic switch for class token in pretrained ViTs and average feature of user prompts, controlled by the existence of user prompts. Then we use the candidate feature to generate a probabilistic similarity map as the guidance to alleviate the over-reliance on binary mask. Extensive experiments show that our method can adapt well to both automatic and inter-active scenarios with more light-weight framework. Code available at github.com/coconut/SMat.
Zixuan Ye, Wenze Liu, He Guo 0005, Yujia Liang, Chaoyi Hong, Hao Lu 0003, Zhiguo Cao 0001
CVPR5
2023 Infusing Definiteness into Randomness: Rethinking Composition Styles for Deep Image Matting
abstract
We study the composition style in deep image matting, a notion that characterizes a data generation flow on how to exploit limited foregrounds and random backgrounds to form a training dataset. Prior art executes this flow in a completely random manner by simply going through the foreground pool or by optionally combining two foregrounds before foreground-background composition. In this work, we first show that naive foreground combination can be problematic and therefore derive an alternative formulation to reasonably combine foregrounds. Our second contribution is an observation that matting performance can benefit from a certain occurrence frequency of combined foregrounds and their associated source foregrounds during training. Inspired by this, we introduce a novel composition style that binds the source and combined foregrounds in a definite triplet. In addition, we also find that different orders of foreground combination lead to different foreground patterns, which further inspires a quadruplet-based composition style. Results under controlled experiments on four matting baselines show that our composition styles outperform existing ones and invite consistent performance improvement on both composited and real-world datasets. Code is available at: https://github.com/coconuthust/composition_styles
Zixuan Ye, Yutong Dai 0001, Chaoyi Hong, Zhiguo Cao 0001, Hao Lu 0003
AAAI3
2022 SymmNeRF: Learning to Explore Symmetry Prior for Single-View View Synthesis
Xingyi Li 0005, Chaoyi Hong, Yiran Wang 0005, Zhiguo Cao 0001, Ke Xian, Guosheng Lin
ACCV (1)2
2022 Class-attribute inconsistency learning for novelty detection
Shuaiyuan Du, Chaoyi Hong, Yinpeng Chen, Zhiguo Cao 0001
Pattern Recognit.2
2021 Composing Photos Like a Photographer
abstract
We show that explicit modeling of composition rules benefits image cropping. Image cropping is considered a promising way to automate aesthetic composition in professional photography. Existing efforts, however, only model such professional knowledge implicitly, e.g., by ranking from comparative candidates. Inspired by the observation that natural composition traits always follow a specific rule, we propose to learn such rules in a discriminative manner, and more importantly, to incorporate learned composition clues explicitly in the model. To this end, we introduce the concept of the key composition map (KCM) to encode the composition rules. The KCM can reveal the common laws hidden behind different composition rules and can inform the cropping model of what is important in composition. With the KCM, we present a novel cropping-by-composition paradigm and instantiate a network to implement composition-aware image cropping. Extensive experiments on two benchmarks justify that our approach enables effective, interpretable, and fast image cropping.
Chaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong
CVPR1
2020 Parallel Network to Learn Novelty from the Known
Shuaiyuan Du, Chaoyi Hong, Chen Feng 0002, Zhiguo Cao 0001
ICPR2