Zixuan Ye

dblp:228/3343 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding
abstract
Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer questions about multiple instances or shots over the whole video. We attribute this challenge to two issues: 1) the lack of multi-shot multi-instance annotations of existing datasets, and 2) the negligence of instance-aware modeling of current VideoLLMs. Therefore, we first introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and question-answering pairs tailored for multi-shot and multi-instance scenarios. Moreover, since the existing VideoLLMs neglect the explicit modeling of instance-related features, we propose a novel Instance Prompt-guided Transformer, named IPFormer, to achieve instance-aware videounderstanding. In the IPFormer, we design a simple but effective instance-aware feature injection module, which encodes instance features as instance prompts via an attention-based connector. By this means, IPFormer can aggregate instance-specific information across multiple shots. Extensive experiments not only show that our dataset and model significantly improve multi-shot video understanding. but also show that our MultiClip-Bench can provide valuable training data and benchmarks for various video understanding tasks.
Yujia Liang, Jile Jiao, Xuetao Feng, Xinchen Liu, Zixuan Ye
AAAI7
2025 Training Matting Models Without Alpha Labels
abstract
The labeling difficulty has been a longstanding problem in deep image matting. To escape from fine labels, this work explores using rough annotations such as trimaps coarsely indicating the foreground/background as supervision. We present that the cooperation between learned semantics from indicated known regions and proper assumed matting rules can help infer alpha values at transition areas. Inspired by the nonlocal principle in traditional image matting, we build a directional distance consistency loss (DDC loss) at each pixel neighborhood to constrain the alpha values conditioned on the input image. DDC loss forces the distance of similar pairs on the alpha matte and on its corresponding image to be consistent. In this way, the alpha values can be propagated from learned known regions to unknown transition areas. With only images and trimaps, a matting model can be trained under the supervision of a known loss and the proposed DDC loss. Experiments on AM-2K and P3M-10K dataset show that our paradigm achieves comparable performance with the fine-label-supervised baseline, while sometimes offers even more satisfying results than human-labeled ground truth.
Wenze Liu, Zixuan Ye, Hao Lu 0003, Zhiguo Cao 0001, Xiangyu Yue 0001
AAAI2
2025 StyleMaster: Stylize Your Video with Artistic Generation and Translation
abstract
Style control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to transfer one video to the desired style. Our first observation is that the style extraction stage matters, whereas existing methods emphasize global style but ignore local textures. In order to bring texture features while preventing content leakage, we filter content-related patches while retaining style ones based on prompt-patch similarity; for global style extraction, we generate a paired style dataset through model illusion to facilitate contrastive learning, which greatly enhances the absolute style consistency. Moreover, to fill in the image-to-video gap, we train a lightweight motion adapter on still videos, which implicitly enhances stylization extent, and enables our image-trained model to be seamlessly applied to videos. Benefited from these efforts, our approach, StyleMaster, not only achieves significant improvement in both style resemblance and temporal coherence, but also can easily generalize to video style transfer with a gray tile ControlNet. Extensive experiments and visualizations demonstrate that StyleMaster significantly outperforms competitors, effectively generating high-quality stylized videos that align with textual content and closely resemble the style of reference images.
Zixuan Ye, Huijuan Huang 0001, Xintao Wang 0004, Pengfei Wan 0001, Di Zhang 0026, Wenhan Luo
CVPR1
2024 In-Context Matting
He Guo 0005, Zixuan Ye, Zhiguo Cao 0001, Hao Lu 0003
CVPR2
2024 Unifying Automatic and Interactive Matting with Pretrained ViTs
abstract
Automatic and interactive matting largely improve image matting by respectively alleviating the need for auxil-iary input and enabling object selection. Due to different settings on whether prompts exist, they either suffer from weakness in instance completeness or region details. Also, when dealing with different scenarios, directly switching between the two matting models introduces inconvenience and higher workload. Therefore, we wonder whether we can al-leviate the limitations of both settings while achieving unification to facilitate more convenient use. Our key idea is to offer saliency guidance for automatic mode to enable its attention to detailed regions, and also refine the instance completeness in interactive mode by replacing the binary mask guidance with a more probabilistic form. With different guidance for each mode, we can achieve unification through adaptable guidance, defined as saliency information in automatic mode and user cue for interactive one. It is instantiated as candidate feature in our method, an automatic switch for class token in pretrained ViTs and average feature of user prompts, controlled by the existence of user prompts. Then we use the candidate feature to generate a probabilistic similarity map as the guidance to alleviate the over-reliance on binary mask. Extensive experiments show that our method can adapt well to both automatic and inter-active scenarios with more light-weight framework. Code available at github.com/coconut/SMat.
Zixuan Ye, Wenze Liu, He Guo 0005, Yujia Liang, Chaoyi Hong, Hao Lu 0003, Zhiguo Cao 0001
CVPR1
2024 SCAPE: A Simple and Strong Category-Agnostic Pose Estimator
Yujia Liang, Zixuan Ye, Wenze Liu, Hao Lu 0003
ECCV (23)2
2024 Video Bokeh Rendering: Make Casual Videography Cinematic
abstract
Bokeh is a wide-aperture optical effect that creates aesthetic blurring in photography. However, achieving this effect typically demands expensive professional equipment and expertise. To make such cinematic techniques more accessible, bokeh rendering aims to generate the desired bokeh effects from all-in-focus inputs captured by smartphones. Previous efforts in bokeh rendering primarily focus on static images. However, when extended to video inputs, these methods exhibit flicker and artifacts due to a lack of temporal consistency modeling. Meanwhile, they cannot utilize information like occluded objects from adjacent frames, which are necessary for bokeh rendering. Moreover, the difficulties of capturing all-in-focus and bokeh video pairs result in a shortage of data for training video bokeh models. To tackle these challenges, we propose the Video Bokeh Renderer (VBR), the model designed specifically for video bokeh rendering.VBR leverages implicit feature space alignment and aggregation to model temporal consistency and exploit complementary information from adjacent frames. On the data front, we introduce the first Synthetic Video Bokeh (SVB) dataset, synthesizing authentic bokeh effects using ray-tracing techniques. Furthermore, to improve the robustness of the model to inaccurate disparity maps, we employ a set of augmentation strategies to simulate corrupted disparity inputs during training. Experimental results on both synthetic and real-world data demonstrate the effectiveness of our method.
Yawen Luo, Min Shi 0004, Liao Shen, Yachuan Huang, Zixuan Ye, Juewen Peng, Zhiguo Cao 0001
ACM Multimedia5
2023 Infusing Definiteness into Randomness: Rethinking Composition Styles for Deep Image Matting
abstract
We study the composition style in deep image matting, a notion that characterizes a data generation flow on how to exploit limited foregrounds and random backgrounds to form a training dataset. Prior art executes this flow in a completely random manner by simply going through the foreground pool or by optionally combining two foregrounds before foreground-background composition. In this work, we first show that naive foreground combination can be problematic and therefore derive an alternative formulation to reasonably combine foregrounds. Our second contribution is an observation that matting performance can benefit from a certain occurrence frequency of combined foregrounds and their associated source foregrounds during training. Inspired by this, we introduce a novel composition style that binds the source and combined foregrounds in a definite triplet. In addition, we also find that different orders of foreground combination lead to different foreground patterns, which further inspires a quadruplet-based composition style. Results under controlled experiments on four matting baselines show that our composition styles outperform existing ones and invite consistent performance improvement on both composited and real-world datasets. Code is available at: https://github.com/coconuthust/composition_styles
Zixuan Ye, Yutong Dai 0001, Chaoyi Hong, Zhiguo Cao 0001, Hao Lu 0003
AAAI1
2023 Robust Beamforming for Intelligent Reflecting Surface Aided Dual-Functional Radar-Communication System
abstract
Intelligent reflecting surface (IRS) has recently gained significant academic interest as a prospective contender for improving wireless communication system coverage and spectral efficiency. This paper investigates a robust beamforming design of an IRS-aided dual-functional radar-communication (DFRC) system in the presence of channel uncertainty, as opposed to the idealistic assumption of perfect channel state information (CSI) in the existing literature. The optimization is carried out by minimizing the transmit power while ensuring the detection performance and the achievable rate of the user meets the quality of service (QoS) requirement, which turns out to be a non-convex and intractable problem. To circumvent this issue, we alternatively update the transmit beamforming vector and the phase shifts at the IRS using the block coordinate descent (BCD) algorithm. Afterwards, the resulting two sub-problems can be efficiently solved with the help of approximation and transformation techniques. Simulation results have validated the convergence and effectiveness of the proposed algorithm.
Zixuan Ye, Dongqi Luo, Jihong Zhu 0001
WCNC1
2022 SAPA: Similarity-Aware Point Affiliation for Feature Upsampling
abstract
We introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling kernels. The kernels encourage not only semantic smoothness but also boundary sharpness in the upsampled feature maps. Such properties are particularly useful for some dense prediction tasks such as semantic segmentation. The key idea of our formulation is to generate similarity-aware kernels by comparing the similarity between each encoder feature point and the spatially associated local region of decoder features. In this way, the encoder feature point can function as a cue to inform the semantic cluster of upsampled feature points. To embody the formulation, we further instantiate a lightweight upsampling operator, termed Similarity-Aware Point Affiliation (SAPA), and investigate its variants. SAPA invites consistent performance improvements on a number of dense prediction tasks, including semantic segmentation, object detection, depth estimation, and image matting. Code is available at: https://github.com/poppinace/sapa
Hao Lu 0003, Wenze Liu, Zixuan Ye, Hongtao Fu, Zhiguo Cao 0001
NeurIPS3
2018 CORDIC Framework for Quaternion-based Joint Angle Computation to Classify Arm Movements
abstract
We present a novel architecture for arm movement classification based on kinematic properties (joint angle and position), computed from MARG sensors, using a quaternion-based gradient-descent method and a 2-link model of the upper limb. The design based on Coordinate Rotation Digital Computer framework was validated on stroke survivors and healthy subjects performing three elementary arm movements (reach and retrieve, lift arm, rotate arm), involved in `making-a-cup-of-tea' an archetypal daily activity, achieved an overall accuracy of 78% and 85% respectively. The design coded in System Verilog, was synthesized using STMicroelectronics 130 nm technology, occupies 340K NAND2 equivalent area and consumes 292 nW @ 150 Hz, besides being functionally verified up to 25 MHz making it suitable for real-time high speed operations. The orientation, arm position and the joint angle, are computed on-the-fly, with the classification performed at the end of movement duration.
Dwaipayan Biswas, Zixuan Ye, Evangelos B. Mazomenos, Michael Jöbges, Koushik Maharatna
ISCAS2