Yufan Deng

dblp:185/7084 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
abstract
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yu Bai 0018, Huashan Sun, Yiguan Lin, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Yang Gao 0016, Heyan Huang
ACL (1)10
2026 A Bidding-Collaborative Model for Weapon-Target Assignment against Integrated Defense Systems
Xiaozhan Li, Yuanhang Li, Yufan Deng, Guangquan Cheng
ICORES3
2026 DeepELIC: Deep encrypted lossy image compression network via compressive sensing unfolding
Fangyuan Gao, Yufan Deng, Xin Deng 0002, Zhenyu Guan 0002, Mai Xu
Pattern Recognit.2
2026 A Novel Visible-Infrared Image Compression Framework for High-Value Target Protection
abstract
The joint compression of visible-infrared images is crucial for military and surveillance applications. The challenge lies in the protection of high-value targets (HVT) while maintaining high compression efficiency. This paper proposes a novel dual-stream compression framework that effectively addresses this challenge. In our framework, the sensitive HVT infrared signatures are concealed within the visible image stream, while residual infrared image is encoded separately. This dual-stream compression framework introduces three key innovations. 1) HVT protection: The HVT information is physically isolated and hidden within public visible images through a dedicated concealment stream; 2) Key-conditioned reconstruction: A novel decoding mechanism enables active camouflage by replacing HVTs with plausible background content when unauthorized access is detected; 3) Unified optimization: The framework integrates compression efficiency and HVT protection within an endto- end trainable network. Extensive experiments demonstrate that our approach achieves state-of-the-art compression performance while providing superior HVT protection, significantly outperforming traditional encrypt-then-compress methods. The code and weights are open-source athttps://github.com/eecoder-dyf/rgbir-compress.
Yufan Deng, Xin Deng 0002, Shengxi Li, Xiaowan Hu, Mai Xu
IEEE Signal Process. Lett.1
2025 Video-Bench: Human-Aligned Video Generation Benchmark
abstract
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate generated video quality across multiple dimensions but often lack alignment with human judgments; and large language model (LLM)-based benchmarks, though capable of human-like reasoning, are constrained by a limited understanding of video quality metrics and cross-modal consistency. To address these challenges and establish a benchmark that better aligns with human preferences, this paper introduces Video-Bench, a comprehensive benchmark featuring a rich prompt suite and extensive evaluation dimensions. This benchmark represents the first attempt to systematically leverage MLLMs across all dimensions relevant to video generation assessment in generative models. By incorporating few-shot scoring and chain-of-query techniques, Video-Bench provides a structured, scalable approach to generated video evaluation. Experiments on advanced models including Sora demonstrate that Video-bench achieve superior alignment with human preferences across all dimensions. Moreover, in instances where our framework’s assessments diverge from human evaluations, it consistently offers more objective and accurate insights, suggesting an even greater potential advantage over traditional human judgment.
Yiwen Yuan, Yuling Wu, Yufan Deng, Chak Tou Leong, Hanwen Du, Junchen Fu, Youhua Li, Chi Zhang 0007, Li-jia Li, Yongxin Ni
CVPR6
2025 MTPNet: Multi-Grained Target Perception for Unified Activity Cliff Prediction
abstract
Activity cliff prediction is a critical task in drug discovery and material design. Existing computational methods are limited to handling single binding targets, which restricts the applicability of these prediction models. In this paper, we present the Multi-Grained Target Perception network (MTPNet) to incorporate the prior knowledge of interactions between the molecules and their target proteins. Specifically, MTPNet is a unified framework for activity cliff prediction, which consists of two components: Macro-level Target Semantic (MTS) guidance and Micro-level Pocket Semantic (MPS) guidance. By this way, MTPNet dynamically optimizes molecular representations through multi-grained protein semantic conditions. To our knowledge, it is the first time to employ the receptor proteins as guiding information to effectively capture critical interaction details. Extensive experiments on 30 representative activity cliff datasets demonstrate that MTPNet significantly outperforms previous approaches, achieving an average RMSE improvement of 18.95% on top of several mainstream GNN architectures. Overall, MTPNet internalizes interaction patterns through conditional deep learning to achieve unified predictions of activity cliffs, helping to accelerate compound optimization and design. Codes are available at: https://github.com/ZishanShu/MTPNet.
Zishan Shu, Yufan Deng, Hongyu Zhang 0002, Zhiwei Nie, Jie Chen 0001
IJCAI2
2025 OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
abstract
Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus, consisting of (i) OpenS2V‑Eval, a fine‑grained benchmark, and (ii) OpenS2V‑5M, a million‑scale dataset.In contrast to existing S2V benchmarks inherited from VBench that focus on global and coarse-grained assessment of generated videos, OpenS2V-Eval focuses on the model's ability to generate subject-consistent videos with natural subject appearance and identity fidelity. For these purposes, OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data. Furthermore, to accurately align human preferences with S2V benchmarks, we propose three automatic metrics, NexusScore, NaturalScore and GmeScore, to separately quantify subject consistency, naturalness, and text relevance in generated videos. Building on this, we conduct a comprehensive evaluation of 18 representative S2V models, highlighting their strengths and weaknesses across different content. Moreover, we create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triplets. Specifically, we ensure subject‐information diversity in our dataset by (1) segmenting subjects and building pairing information via cross‐video associations and (2) prompting GPT-4o on raw frames to synthesize multi-view representations. Through OpenS2V-Nexus, we deliver a robust infrastructure to accelerate future S2V generation research.
Shenghai Yuan 0002, Xianyi He, Yufan Deng, Jinfa Huang, Bin Lin 0014, Chongyang Ma, Jiebo Luo 0001, Li Yuan 0007
NeurIPS3
2024 DragVideo: Interactive Drag-Style Video Editing
Yufan Deng, Ruida Wang, Yu-Wing Tai, Chi-Keung Tang
ECCV (56)1
2024 VideoTetris: Towards Compositional Text-to-Video Generation
abstract
Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris
Ling Yang 0006, Yuan Gao 0015, Yufan Deng, Xintao Wang 0002, Zhaochen Yu, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Bin Cui 0001
NeurIPS5
2023 A Two-stage hybrid CNN-Transformer Network for RGB Guided Indoor Depth Completion
abstract
The indoor captured raw depth images usually contain large in-homogeneous missing regions. Most existing methods are designed for the outdoor sparse depth completion, which struggle in completing the indoor depth with large holes. In this paper, to solve this problem, we propose a hybrid CNN-Transformer network for RGB guided indoor depth completion. The proposed network is composed of two stages to achieve depth completion in a coarse-to-fine manner. In the first stage, we propose a CNN based self-completion module (SCM) with cross scale attention to restore a coarse depth image. In the second stage, we further refine the completed depth image with the guidance of RGB image by proposing a guided completion module (GCM). To fully explore the guidance from the RGB image, we design a cross-modal Transformer (CMT) block to fuse the features from the depth and RGB modalities at different scales. Extensive experiments on NYUv2 and SUN RGB-D datasets demonstrate the superior performance of the proposed method over other state-of-the-art methods both quantitatively and qualitatively. The code is available at https://github.com/eecoder-dyf/ICME-2023-depth-completion.
Yufan Deng, Xin Deng 0002, Mai Xu
ICME1
2023 MASIC: Deep Mask Stereo Image Compression
abstract
Stereo image compression (SIC) aims to simultaneously compress a pair of left and right stereoscopic images, which can achieve higher compression efficiency than single image compression. In this paper, to benefit the SIC tasks, we collect a large real-world stereo image dataset, namely Palace, which is composed of hundreds of stereo image pairs at high-resolution. More importantly, we propose a novel mask stereo image compression network, namely MASIC, which can jointly compress the stereo images with high compression efficiency. Specifically, we first estimate the homography matrix between the stereo images through a regression model. Then, the left image is spatially transformed by the homography matrix, so that only the residual information needs to be encoded for the right image. To avoid the wrong guidance between stereo image pair, we propose a mask prediction module (MPM) to generate a multi-channel guided mask to navigate both the encoding and decoding processes. Based on the guided mask, we introduce a new mask conditional stereo entropy (MCSE) model, to fully explore the correlation between the stereo images in entropy coding. In the decoder, we develop a stereo decoding module to simultaneously decode the stereo images and enhance their compression quality. Experimental results show that our MASIC significantly advances the performance of SIC both quantitatively and qualitatively on a variety of datasets, and is robust to the change of parallax level between stereo images. The software codes are available athttps://github.com/eecoder-dyf/MASIC.
Xin Deng 0002, Yufan Deng, Radu Timofte, Mai Xu
IEEE Trans. Circuits Syst. Video Technol.2
2017 Enhanced intra prediction for inter pictures
abstract
This paper presents a novel intra prediction method for inter pictures (i.e. P pictures and B pictures), denominated enhanced intra prediction (EIP). The traditional intra prediction only uses the reconstructed pixels to the left and above to derive intra prediction blocks. While the proposed method combines the left-above and right-below pixels to strengthen the prediction efficiency of the intra blocks in inter pictures. For accessing the pixels below and to the right of intra coding units (CUs), the encoding and decoding structures are adjusted to guarantee that all the inter CUs are reconstructed before all the intra CUs. With more available reference samples around the intra CUs, EIP achieves better prediction results. The proposed method is implemented on top of the H.265/HEVC reference software (HM-16.12), and the experimental results show that approximate 0.5% BD-rate reduction is achieved under Random Access (RA) and Low Delay P (LP) configurations.
Kui Fan, Ronggang Wang, Ge Li 0002, Wen Gao 0001, Yufan Deng, Shensian Syu, Ming-Jong Jou
ICME5
2017 Polar square projection for panoramic video
abstract
Panoramic video provides an immersive experience by presenting a 360° spherical video content. Due to the limitations of coding and storage technology, the spherical panoramic video needs to be projected onto the two-dimensional plane for storage and encoding. In this paper, we propose a polar square projection scheme. We project the area near the poles of the sphere into two square planes and a latitude circle on sphere is projected to a square circle on squares plane, in addition, the rest of area on sphere is projected into a rectangle by means of equal area projection. Experimental results show our proposed projection can obtain a gain of 11.63% BD-rate compared to the equirectangular projection.
Ronggang Wang, Zhenyu Wang 0002, Kui Fan, Yufan Deng, Shensian Syu, Ming-Jong Jou
VCIP5
2016 Performance evaluation of fractal dimension method based on box-covering algorithm in complex network
abstract
Complex network is becoming increasingly used in our daily life and society. The measurement of fractal dimension and self-similar trait in complex network have been a vital part in the research of complex system. In the recent years, more and more methods of measuring the fractal dimension have been presented. The different methods that based on box-covering algorithms was evaluated by the experiment with contrastive analysis. Moreover, we transform the unweighted network into the weighted network and combine with the traditional method. The results of the experiments show that the max-excluded mass burning(MEMB) algorithm is performed the best of all in the original complex network and the weighted network.
Yufan Deng
CSCWD1