VLDB 2026 Research / reviewers in the wild / expert
Zhifei Yang 0004
dblp:56/6850-4
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-0054-8896ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SepPrune: Structured Pruning for Efficient Deep Speech SeparationabstractAlthough deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose SepPrune, the first structured pruning framework specifically designed to compress deep speech separation models and reduce their computational cost. SepPrune begins by analyzing the computational structure of a given model to identify layers with the highest computational burden. It then introduces a differentiable masking strategy to enable gradient-driven channel selection. Based on the learned masks, SepPrune prunes redundant channels and fine-tunes the remaining parameters to recover performance. Extensive experiments demonstrate that this learnable pruning paradigm yields substantial advantages for channel pruning in speech separation models, outperforming existing methods. Notably, a model pruned with SepPrune can recover 85% of the performance of a pre-trained model (trained over hundreds of epochs) with only one epoch of fine-tuning, and achieves convergence 36x faster than training from scratch. Zhifei Yang 0004, Zeyu Dong, Zhengtao Yao, Haoyan Xu, Yingli Tian, Yao Lu 0041 |
AAAI | 4 |
| 2026 | Gradient as Conditions: Rethinking HOG for All-in-one Image RestorationabstractAll-in-one image restoration (AIR) aims to address diverse degradations within a unified model by leveraging informative degradation conditions to guide the restoration process. However, existing methods often rely on implicitly learned priors, which may entangle feature representations and hinder performance in complex or unseen scenarios. Histogram of Oriented Gradients (HOG) as a classical gradient representation, we observe that it has strong discriminative capability across diverse degradations, making it a powerful and interpretable prior for AIR. Based on this insight, we propose HOGformer, a Transformer-based model that integrates learnable HOG features for degradation-aware restoration. The core of HOGformer is a Dynamic HOG-aware Self-Attention (DHOGSA) mechanism, which adaptively models long-range spatial dependencies conditioned on degradation-specific cues encoded by HOG descriptors. To further adapt the heterogeneity of degradations in AIR, we propose a Dynamic Interaction Feed-Forward (DIFF) module that facilitates channel–spatial interactions, enabling robust feature transformation under diverse degradations. Besides, we propose a HOG loss to explicitly enhance structural fidelity and edge sharpness. Extensive experiments on a variety of benchmarks, including adverse weather and natural degradations, demonstrate that HOGformer achieves state-of-the-art performance and generalizes well to complex real-world scenarios. Zhifei Yang 0004 |
AAAI | 2 |
| 2025 | MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationabstractControllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. However, current graph-based methods for scene generation are constrained to text-based inputs and exhibit insufficient adaptability to flexible user inputs, hindering the ability to precisely control object geometry. To address this issue, we propose MMGDreamer, a dual-branch diffusion model for scene generation that incorporates a novel Mixed-Modality Graph, visual enhancement module, and relation predictor. The mixed-modality graph allows object nodes to integrate textual and visual modalities, with optional relationships between nodes. It enhances adaptability to flexible user inputs and enables meticulous control over the geometry of objects in the generated scenes. The visual enhancement module enriches the visual fidelity of text-only nodes by constructing visual representations using text embeddings. Furthermore, our relation predictor leverages node representations to infer absent relationships between nodes, resulting in more coherent scene layouts. Extensive experimental results demonstrate that MMGDreamer exhibits superior control of object geometry, achieving state-of-the-art scene generation performance. Zhifei Yang 0004, Keyang Lu, Jiaxing Qi, Hanqi Jiang, Ruifei Ma, Shenglin Yin, Yifan Xu 0028, Mingzhe Xing, Jieyi Long, Xiangde Liu, Guangyao Zhai |
AAAI | 1 |
| 2025 | DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language ModelsabstractRuizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Ziyang Wang, Tony Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruizhe Chen, Wenhao Chai, Zhifei Yang 0004, Tony Q. S. Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu |
ACL (1) | 3 |
| 2025 | UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsabstractThis paper introduces UoRA, a novel parameter-efficient fine-tuning (PEFT) approach for large language models (LLMs). UoRA achieves state-of-the-art efficiency by leveraging a low-rank approximation method that reduces the number of trainable parameters without compromising performance. Unlike existing methods such as LoRA and VeRA, UoRA employs a re-parametrization mechanism that eliminates the need to adapt frozen projection matrices while maintaining shared projection layers across the model. This results in halving the trainable parameters compared to LoRA and outperforming VeRA in computation and storage efficiency. Comprehensive experiments across various benchmarks demonstrate UoRA’s superiority in achieving competitive fine-tuning performance with minimal computational overhead. We demonstrate its performance on GLUE and E2E benchmarks and is effectiveness in instruction-tuning large language models and image classification models. Our contributions establish a new paradigm for scalable and resource-efficient fine-tuning of LLMs. Jinman Zhao, Zhifei Yang 0004, Yibo Zhong, Shuhao Guan, Linbo Cao |
ACL (1) | 3 |
| 2025 | MaRI: Material Retrieval Integration across DomainsabstractAccurate material retrieval is critical for creating realistic 3D assets. Existing methods rely on datasets that capture shape-invariant and lighting-varied representations of materials, which are scarce and face challenges due to limited diversity and inadequate real-world generalization. Most current approaches adopt traditional image search techniques. They fall short in capturing the unique properties of material spaces, leading to suboptimal performance in retrieval tasks. Addressing these challenges, we introduce MaRI, a framework designed to bridge the feature space gap between synthetic and real-world materials. MaRI constructs a shared embedding space that harmonizes visual and material attributes through a contrastive learning strategy by jointly training an image and a material encoder, bringing similar materials and images closer while separating dissimilar pairs within the feature space. To support this, we construct a comprehensive dataset comprising high-quality synthetic materials rendered with controlled shape variations and diverse lighting conditions, along with real-world materials processed and standardized using material transfer techniques. Extensive experiments demonstrate the superior performance, accuracy, and generalization capabilities of MaRI across diverse and complex material retrieval tasks, outperforming existing methods. Zhifei Yang 0004, Yangfan He, Huixiong Zhang |
CVPR | 2 |
| 2025 | Spatial 3D-LLM : Exploring Spatial Awareness in 3D Vision-Language ModelsabstractNew era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting independent objects to perform these tasks, which limits their spatial awareness due to insufficient representation of the richness inherent in 3D scenes. To overcome these limitations, we propose Spatial 3D-LLM, a 3D MLLM specifically designed to enhance spatial awareness for 3D vision-language tasks by enriching the spatial embeddings of 3D scenes. Spatial 3D-LLM integrates an LLM backbone with a progressive spatial awareness scheme that progressively captures spatial information as the perception field expands, generating location-enriched 3D scene embeddings to serve as visual prompts. Furthermore, we introduce two novel tasks: 3D object distance measurement and 3D layout editing, and construct a 3D instruction dataset, MODEL, to evaluate the model’s spatial awareness capabilities. Experimental results demonstrate that Spatial 3D-LLM achieves state-of-the-art performance across a wide range of 3D vision-language tasks, revealing the improvements stemmed from our progressive spatial awareness scheme of mining more profound spatial information. Our code is available at https://github.com/bjshuyuan/Spatial-3D-LLM. Zeju Li, Yifan Xu 0028, Jiaxing Qi, Zhifei Yang 0004, Ruifei Ma, Xiangde Liu |
ICME | 5 |
| 2025 | CTSG: Integrating Context and Way Topology Into Scene Graph for Zero-shot NavigationabstractA robust environment representation is critical for enabling robot systems to accomplish embodied navigation tasks. While offering efficient and sparse representations of environments compared to dense semantic maps, traditional 3D Scene Graphs often rely on multi-level semantic hierarchies that risk semantic discrepancies between high-level nodes and objects. Furthermore, separating semantic context from way topological relationships creates a disconnect between scene interpretation and actionable navigation strategies. To address these challenges, we propose CTSG, a hierarchical 3D scene graph mapping framework for zero-shot object navigation that supports both visual and textual queries. CTSG features a dual-layer structure: an object layer and a novel conway layer (contextual information + way topology). The conway layer integrates topological waypoints with rich multi-modal context, enhancing the continuity of environmental semantics. By aligning observations with navigation-centric perspectives, CTSG bridges the gap between scene understanding and task execution. We validate our method through simulation and real-world experiments across diverse environments, demonstrating robust performance in both visual target and language-guided navigation scenarios. Ruifei Ma, Yifan Xu 0028, Yanping Fang, Zhifei Yang 0004, Jiaxing Qi |
IROS | 5 |
| 2025 | BridgeGLM: Bridging Graph and Language Spaces for Domain Generalization
Jiaxing Qi, Yifan Xu 0028, Zhifei Yang 0004, Ruifei Ma, Kuifei Yu |
ACM Multimedia | 3 |
| 2023 | Mask to Reconstruct: Cooperative Semantics Completion for Video-text RetrievalabstractRecently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo. Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou |
ACM Multimedia | 2 |