Yunfan Ye

dblp:175/2728 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-5723-9627ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Video plot segmentation
Xichen Tan, Yuanjing Luo, Yunfan Ye, Chengyu Wang 0008, Fang Liu 0002, Zhiping Cai
Expert Syst. Appl.3
2026 StarBurst: Aiding Design Ideation Through AI-Generated Remote Associations
abstract
Remote associations play a crucial role in enhancing creativity during design ideation, yet designers face challenges in effectively creating and integrating them. Our formative study (N = 5) shows the potential of AI-generated remote associations to facilitate this process, but there are still challenges to understand and apply them. To address these, we propose StarBurst, which supports design ideation through AI-generated remote associations. It consists of three core components: (1) generating diverse remote associations from images to expand creative possibilities; (2) constructing an attribute map to explore connections between associative elements; and (3) providing suggestions to integrate these associations into the final design idea. Through two forms of user studies (N = 32, N = 16), we found that StarBurst outperformed designers in generating remote associations and provided more effective support for diverse and creative idea development compared to the baseline system. Additionally, we discussed how users’ usage patterns and perceptions influence StarBurst’s effectiveness.
Runqi Fang, Fang Liu 0002, Yunfan Ye, Shenglan Cui, Ming Yin 0001
Int. J. Hum. Comput. Interact.4
2026 Non-Local Guided Neural Fields for 4D CT Reconstruction
abstract
Dynamic CT reconstruction plays a crucial role in both medical and industrial applications. However, existing 4D CT reconstruction methods typically rely on complex regularization techniques or external large-scale training datasets, posing challenges for reconstruction quality and generalization when handling complex object motion and varied imaging modes. Neural Radiance Fields (NeRF) offer a promising approach to dynamic CT reconstruction, but existing NeRF-based methods often assume that the scene is low-rank, limiting their representation capabilities. To address these issues, we propose NG-NeRF. First, we combine 3D and 4D hash grids for scene representation, effectively reducing temporal redundancy in static regions of dynamic scenes while improving the model’s representation capabilities and efficiency. Next, we design a non-local hash attention module to establish non-local dependencies between the features of different hash grids. This guides the model to adaptively select features based on hash table load information, significantly alleviating hash collisions and achieving the decoupling of dynamic and static regions. Besides, we introduce global continuity by employing mask positional encoding, which helps reduce the noise often introduced by grid features. Our experimental results on medical and industrial datasets demonstrate that the proposed method outperforms existing state-of-the-art methods by 5.84 dB and 3.4 dB, respectively, and exhibits excellent generalization ability across different 4D CT scenarios.
Qingyang Zhou, Yunfan Ye, Zhihuang Liu, Zhiping Cai
IEEE Trans. Circuits Syst. Video Technol.2
2025 ALLVB: All-in-One Long Video Understanding Benchmark
abstract
From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effectively evaluating the long-sequence modeling capabilities of MLLMs. This highlights the urgent need for a comprehensive and integrated long video understanding benchmark to assess the ability of MLLMs thoroughly. To this end, we propose ALLVB (ALL-in-One Long Video Understanding Benchmark). ALLVB's main contributions include: 1) It integrates 9 major video understanding tasks. These tasks are converted into video QA formats, allowing a single benchmark to evaluate 9 different video understanding capabilities of MLLMs, highlighting the versatility, comprehensiveness, and challenging nature of ALLVB. 2) A fully automated annotation pipeline using GPT-4o is designed, requiring only human quality control, which facilitates the maintenance and expansion of the benchmark. 3) It contains 1,376 videos across 16 categories, averaging nearly 2 hours each, with a total of 252k QAs. To the best of our knowledge, it is the largest long video understanding benchmark in terms of the number of videos, average duration, and number of QAs. We have tested various mainstream MLLMs on ALLVB, and the results indicate that even the most advanced commercial models have significant room for improvement. This reflects the benchmark's challenging nature and demonstrates the substantial potential for development in long video understanding.
Xichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu 0002, Zhiping Cai
AAAI3
2025 Spatiotemporal-Aware Neural Fields for Dynamic CT Reconstruction
abstract
We propose a dynamic Computed Tomography (CT) reconstruction framework called STNF4D (SpatioTemporal-aware Neural Fields). First, we represent the 4D scene using four orthogonal volumes and compress these volumes into more compact hash grids. Compared to the plane decomposition method, this method enhances the model's capacity while keeping the representation compact and efficient. However, in densely predicted high-resolution dynamic CT scenes, the lack of constraints and hash conflicts in the hash grid features lead to obvious dot-like artifact and blurring in the reconstructed images. To address these issues, we propose the Spatiotemporal Transformer (ST-Former) that guides the model in selecting and optimizing features by sensing the spatiotemporal information in different hash grids, significantly improving the quality of reconstructed images. We conducted experiments on medical and industrial datasets covering various motion types, sampling modes, and reconstruction resolutions. Experimental results show that our method outperforms the second-best by 5.99 dB and 4.11 dB in medical and industrial scenes, respectively.
Qingyang Zhou, Yunfan Ye, Zhiping Cai
AAAI2
2025 ROICtrl: Boosting Instance Control for Visual Generation
abstract
Natural language often struggles to accurately associate positional and attribute information with multiple instances, which limits current text-based visual generation models to simpler compositions featuring only a few dominant instances. To address this limitation, this work enhances diffusion models by introducing regional instance control, where each instance is governed by a bounding box paired with a free-form caption. Previous methods in this area typically rely on implicit position encoding or explicit attention masks to separate regions of interest (ROIs), resulting in either inaccurate coordinate injection or large computational overhead. Inspired by ROI-Align in object detection, we introduce a complementary operation called ROI-Unpool. Together, ROI-Align and ROI- Unpool enable explicit, efficient, and accurate ROI manipulation on high-resolution feature maps for visual generation. Building on ROI-Unpool, we propose ROICtrl, an adapter for pretrained diffusion models that enables precise regional instance control. ROICtrl is compatible with community-finetuned diffusion models, as well as with existing spatial-based add-ons (e.g., ControlNet, T2I- Adapter) and embedding-based add-ons (e.g., IP-Adapter, ED-LoRA), extending their applications to multi-instance generation. Experiments show that ROICtrl achieves superior performance in regional instance control while significantly reducing computational costs.
Yuchao Gu, Yipin Zhou, Yunfan Ye, Yixin Nie, Licheng Yu, Pingchuan Ma 0002, Qinghong Lin, Zheng Shou 0001
CVPR3
2025 HumanSAM: Classifying Human-Centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
Yunfan Ye, Fan Zhang 0144, Qingyang Zhou, Yuchuan Luo, Zhiping Cai
ICCV2
2025 Where Watermark Meets Beauty: Expert-Guided Aesthetic Visible Watermarking for Digital Artworks
abstract
In the era of widespread digital art dissemination, visible watermarks provide immediate copyright identification by overlaying visible markers, addressing the lag issue of invisible watermarks that are difficult to prevent in advance due to post hoc evidence collection, thus meeting artists' needs for preemptive prevention and explicit protection. However, existing approaches struggle to balance aesthetics and functionality. To address this challenge, we conducted an exploratory study with watermarking experts, identifying key principles, six common design patterns, and a systematic watermarking workflow. Based on these insights, we developed an end-to-end, perceptual-aware framework for aesthetic-preserving watermark embedding, modeled after expert workflows in 5 phases. Using the Chain-of-Thought strategy, we optimized prompt instructions to guide the Vision-Language Model in emulating experts' decision-making, generating effective watermarking schemes and conducting objective visual evaluations. Iterative feedback optimization ensures watermarked images adhere to aesthetic principles. Quantitative and qualitative experiments demonstrate the system's superiority over baseline methods in preserving aesthetics and ensuring effective copyright protection.
Changjuan Ran, Fang Liu 0002, Runqi Fang, Shenglan Cui, Yunfan Ye
ACM Multimedia6
2025 Integrating conceptual and visual representations with domain expertise for scalable visual plagiarism detection
Shenglan Cui, Fang Liu 0002, Yunfan Ye, Mohan Zhang
Expert Syst. Appl.4
2024 DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection
abstract
Limited by the encoder-decoder architecture, learning-based edge detectors usually have difficulty predicting edge maps that satisfy both correctness and crispness. With the recent success of the diffusion probabilistic model (DPM), we found it is especially suitable for accurate and crisp edge detection since the denoising process is directly applied to the original image size. Therefore, we propose the first diffusion model for the task of general edge detection, which we call DiffusionEdge. To avoid expensive computational resources while retaining the final performance, we apply DPM in the latent space and enable the classic cross-entropy loss which is uncertainty-aware in pixel level to directly optimize the parameters in latent space in a distillation manner. We also adopt a decoupled architecture to speed up the denoising process and propose a corresponding adaptive Fourier filter to adjust the latent features of specific frequencies. With all the technical designs, DiffusionEdge can be stably trained with limited resources, predicting crisp and accurate edge maps with much fewer augmentation strategies. Extensive experiments on four edge detection benchmarks demonstrate the superiority of DiffusionEdge both in correctness and crispness. On the NYUDv2 dataset, compared to the second best, we increase the ODS, OIS (without post-processing) and AC by 30.2%, 28.1% and 65.1%, respectively. Code: https://github.com/GuHuangAI/DiffusionEdge.
Yunfan Ye, Kai Xu 0004, Yuhang Huang 0006, Renjiao Yi, Zhiping Cai
AAAI1
2024 Learning Cross-Hand Policies of High-DOF Reaching and Grasping
Qijin She, Shishun Zhang, Yunfan Ye, Ruizhen Hu, Kai Xu 0004
ECCV (30)3
2024 Large Scale Self-Supervised Pretraining for Active Speaker Detection
abstract
In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show that with pretraining we not only have a more stable supervised training due to better audio-visual features used for initialization, but also improve the ASD mean average precision by 23% on a challenging dataset collected with Google Nest Hub Max devices capturing real user interactions.
Otavio Braga, Keith Johnson, Alice Chuang, Yunfan Ye, Olivier Siohan
ICASSP5
2024 FedStyle: Style-Based Federated Learning Crowdsourcing Framework for Art Commissions
abstract
The unique artistic style is crucial to artists’ occupational competitiveness, yet prevailing Art Commission Platforms rarely support style-based retrieval. Meanwhile, the fast-growing generative AI techniques aggravate artists’ concerns about releasing personal artworks to public platforms. To achieve artistic style-based retrieval without exposing personal artworks, we propose FedStyle, a style-based federated learning crowdsourcing framework. It allows artists to train local style models and share model parameters rather than artworks for collaboration. However, most artists possess a unique artistic style, resulting in severe model drift among them. FedStyle addresses such extreme data heterogeneity by having artists learn their abstract style representations and align with the server, rather than merely aggregating model parameters lacking semantics. Besides, we introduce contrastive learning to meticulously construct the style representation space, pulling artworks with similar styles closer and keeping different ones apart in the embedding space. Extensive experiments on the proposed datasets demonstrate the superiority of FedStyle.
Changjuan Ran, Yeting Guo, Fang Liu 0002, Shenglan Cui, Yunfan Ye
ICME5
2024 Learning accurate template matching with differentiable coarse-to-fine correspondence refinement
abstract
Template matching is a fundamental task in computer vision and has been studied for decades. It plays an essential role in manufacturing industry for estimating the poses of different parts, facilitating downstream tasks such as robotic grasping. Existing methods fail when the template and source images have different modalities, cluttered backgrounds, or weak textures. They also rarely consider geometric transformations via homographies, which commonly exist even for planar industrial parts. To tackle the challenges, we propose an accurate template matching method based on differentiable coarse-to-fine correspondence refinement. We use an edge-aware module to overcome the domain gap between the mask template and the grayscale image, allowing robust matching. An initial warp is estimated using coarse correspondences based on novel structure-aware information provided by transformers. This initial alignment is passed to a refinement network using references and aligned images to obtain sub-pixel level correspondences which are used to give the final geometric transformation. Extensive evaluation shows that our method to be significantly better than state-of-the-art methods and baselines, providing good generalization ability and visually plausible results even on unseen real data.
Zhirui Gao, Renjiao Yi, Zheng Qin 0002, Yunfan Ye, Chenyang Zhu 0002, Kai Xu 0004
Comput. Vis. Media4
2024 Fixing the Double Agent Vulnerability of Deep Watermarking: A Patch-Level Solution Against Artwork Plagiarism
abstract
Increasing artwork plagiarism incidents stresses the urgent need for proper copyright protection on behalf of the creators. The latest development in this context focuses on embedding watermarks via deep encoder-decoder networks. However, we find that deep watermarking has a serious vulnerability on its robustness when facing deliberate plagiarism. To manifest it, we construct an attack that misuses watermarking encoder as a plagiarism lookout for bypassing copyright detection. As a remedy, we propose a patch-level deep watermarking framework (DIPW) to retain copyright evidence in essential patches with plagiarism resistance, inspired by a user study observation that subject elements in artworks are the principal plagiarism entities. Technically, DIPW adaptively finds the embedding patches by identifying a subset of non-overlapping and feature-rich objects; and tailors the model with dual-distortion losses and adversarial plagiarism noise injection for robustness. Experimental results demonstrate the superiority of DIPW in facilitating better robustness, secrecy, and imperceptibility with acceptable time burden.
Yuanjing Luo, Tongqing Zhou, Shenglan Cui, Yunfan Ye, Fang Liu 0002, Zhiping Cai
IEEE Trans. Circuits Syst. Video Technol.4
2024 STEdge: Self-Training Edge Detection With Multilayer Teaching and Regularization
abstract
Learning-based edge detection has hereunto been strongly supervised with pixel-wise annotations which are tedious to obtain manually. We study the problem of self-training edge detection, leveraging the untapped wealth of large-scale unlabeled image datasets. We design a self-supervised framework with multilayer regularization and self-teaching. In particular, we impose a consistency regularization which enforces the outputs from each of the multiple layers to be consistent for the input image and its perturbed counterpart. We adopt L0-smoothing as the "perturbation" to encourage edge prediction lying on salient boundaries following the cluster assumption in self-supervised learning. Meanwhile, the network is trained with multilayer supervision by pseudo labels which are initialized with Canny edges and then iteratively refined by the network as the training proceeds. The regularization and self-teaching together attain a good balance of precision and recall, leading to a significant performance boost over supervised methods, with lightweight refinement on the target dataset. Through extensive experiments, our method demonstrates strong cross-dataset generality and can improve the original performance of edge detectors after self-training and fine-tuning.
Yunfan Ye, Renjiao Yi, Zhiping Cai, Kai Xu 0004
IEEE Trans. Neural Networks Learn. Syst.1
2023 NEF: Neural Edge Fields for 3D Parametric Curve Reconstruction from Multi-View Images
abstract
We study the problem of reconstructing 3D feature curves of an object from a set of calibrated multi-view images. To do so, we learn a neural implicit field representing the density distribution of 3D edges which we refer to as Neural Edge Field (NEF). Inspired by NeRF [20], NEF is optimized with a view-based rendering loss where a 2D edge map is rendered at a given view and is compared to the ground-truth edge map extracted from the image of that view. The rendering-based differentiable optimization of NEF fully exploits 2D edge detection, without needing a supervision of 3D edges, a 3D geometric operator or cross-view edge correspondence. Several technical designs are devised to ensure learning a range-limited and view-independent NEF for robust edge extraction. The final parametric 3D curves are extracted from NEF with an iterative optimization method. On our benchmark with synthetic data, we demonstrate that NEF outperforms existing state-of-the-art methods on all metrics. Project page: https://yunfan1202.github.io/NEF/.
Yunfan Ye, Renjiao Yi, Zhirui Gao, Chenyang Zhu 0002, Zhiping Cai, Kai Xu 0004
CVPR1
2023 Image captioning for cultural artworks: a case study on ceramics
Baoying Zheng, Fang Liu 0002, Mohan Zhang, Tongqing Zhou, Shenglan Cui, Yunfan Ye, Yeting Guo
Multim. Syst.6
2023 Delving Into Crispness: Guided Label Refinement for Crisp Edge Detection
abstract
Learning-based edge detection usually suffers from predicting thick edges. Through extensive quantitative study with a new edge crispness measure, we find that noisy human-labeled edges are the main cause of thick predictions. Based on this observation, we advocate that more attention should be paid on label quality than on model design to achieve crisp edge detection. To this end, we propose an effective Canny-guided refinement of human-labeled edges whose result can be used to train crisp edge detectors. Essentially, it seeks for a subset of over-detected Canny edges that best align human labels. We show that several existing edge detectors can be turned into a crisp edge detector through training on our refined edge maps. Experiments demonstrate that deep models trained with refined edges achieve significant performance boost of crispness from 17.4% to 30.6%. With the PiDiNet backbone, our method improves ODS and OIS by 12.2% and 12.6% on the Multicue dataset, respectively, without relying on non-maximal suppression. We further conduct experiments and show the superiority of our crisp edge detection for optical flow estimation and image segmentation.
Yunfan Ye, Renjiao Yi, Zhirui Gao, Zhiping Cai, Kai Xu 0004
IEEE Trans. Image Process.1
2015 Dynamic Min-Cut Clustering for Energy Savings in Ultra-Dense Networks
abstract
Femtocells are envisioned as a key solution to embrace the ever-increasing high data rate and thus are extensively deployed. Given that numerous femtocell access points (FAPs) deployed in ultra- dense networks (UDNs) lead to significant energy consumption, boosting their energy efficiency (EE) becomes an important issue to be addressed. However, most existing works either focus on homogeneous networks or assume that FAPs are equally distributed, which is not realistic in a dense network as random deployments cause severe interference. This paper explores the realistic scenario of randomly distributed FAPs in heterogeneous networks and proposes a clustering approach combined with an active FAP selection algorithm to boost both spectral and energy efficiency without manual configuration. Taking into account traffic load and interference, the paper reduces the complexity from the Bell Number to polynomial time by exploiting a graph-based Min-Cut strategy to cluster FAPs and allocate orthogonal resources in one cluster to mitigate interference that in turn improves EE. Simulation results confirm the effectiveness of the framework.
Yunfan Ye, Hongtao Zhang 0001, Yang Liu 0024
VTC Fall1