VLDB 2026 Research / reviewers in the wild / expert
Zilu Guo
dblp:299/5046
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous DrivingabstractIn autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free image editing offers a promising solution for improving model robustness by training data enhancement without any modifications to pre-trained diffusion models. Nevertheless, inversion-based methods often suffer from limited effectiveness and inherent inaccuracies, while recent rectified-flow-based approaches struggle to preserve objects with accurate 3D geometry. In this paper, we propose DriveFlow, a Rectified Flow Adaptation method for training data enhancement in autonomous driving based on pre-trained Text-to-Image flow models. Based on frequency decomposition, DriveFlow introduces two strategies to adapt noise-free editing paths derived from text-conditioned velocities. 1) High-Frequency Foreground Preservation: DriveFlow incorporates a high-frequency alignment loss for foreground to maintain precise 3D object geometry. 2) Dual-Frequency Background Optimization: DriveFlow also conducts dual-frequency optimization for background, balancing editing flexibility and semantic consistency. Comprehensive experiments validate the effectiveness and efficiency of DriveFlow, demonstrating comprehensive performance improvements on all categories across OOD scenarios. Yiming Yang 0001, Chaoda Zheng, Yifan Zhang 0004, Shuaicheng Niu, Zilu Guo, Gui Gui, Shuguang Cui, Zhen Li 0026 |
AAAI | 6 |
| 2026 | PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Modelsabstract3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current approaches attempt to transfer knowledge from 2D MLLMs to expand 3D instruction data, but still face modality and domain gaps. To this end, we introduce PiSA-Engine (Point-Self-Augmented-Engine), a new framework for generating instruction point-language datasets enriched with 3D spatial semantics. We observe that existing 3D MLLMs offer a comprehensive understanding of point clouds for annotation, while 2D MLLMs excel at cross-validation by providing complementary information. By integrating holistic 2D and 3D insights from off-the-shelf MLLMs, PiSA-Engine enables a continuous cycle of high-quality data generation. We select PointLLM as the baseline and adopt this co-evolution training framework to develop an enhanced 3D MLLM, termed PointLLM-PiSA. Additionally, we identify limitations in previous 3D benchmarks, which often feature coarse language captions and insufficient category diversity, resulting in inaccurate evaluations. To address this gap, we further introduce PiSA-Bench, a comprehensive 3D benchmark covering six key aspects with detailed and diverse labels. Experimental results demonstrate PointLLM-PiSA’s state-of-the-art performance in zero-shot 3D object captioning and generative classification on our PiSA-Bench, achieving significant improvements of 46.45% (+8.33%) and 63.75% (+16.25%), respectively. Project page: CG-ops/PiSA. Zilu Guo, Zhihao Yuan, Chaoda Zheng, Pengshuo Qiu, Dongzhi Jiang, Renrui Zhang, Chun-Mei Feng 0001, Zhen Li 0026 |
WACV | 1 |
| 2025 | DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion GenerationabstractIn autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from performance degradation, known as Out-of-Distribution (OOD) problems. To address this, controllable Text-to-Image (T2I) diffusion offers a potential solution for training data enhancement, which is required to generate diverse OOD scenarios with precise 3D object geometry. Nevertheless, existing controllable T2I approaches are restricted by the limited scale of training data or struggle to preserve all annotated 3D objects. In this paper, we present DriveGEN, a method designed to improve the robustness of 3D detectors in Driving via Training-Free Controllable Text-to-Image Diffusion Generation. Without extra diffusion model training, DriveGEN consistently preserves objects with precise 3D geometry across diverse OOD generations, consisting of 2 stages: 1) Self-Prototype Extraction: We empirically find that self-attention features are semantic-aware but require accurate region selection for 3D objects. Thus, we extract precise object features via layouts to capture 3D object geometry, termed self-prototypes. 2) Prototype-Guided Diffusion: To preserve objects across various OOD scenarios, we perform semantic-aware feature alignment and shallow feature alignment during denoising. Extensive experiments demonstrate our effectiveness in improving 3D detection. The code is available at github.com/Hongbin98/DriveGEN. Zilu Guo, Yifan Zhang 0004, Shuaicheng Niu, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
CVPR | 2 |
| 2025 | EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionabstractDiffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and metric stability during long-time generation, yielding state-of-the-art performance compared to existing methods. The main page of our paper can be found in https://emotivetalk.github.io/. Yuzhe Weng, Zilu Guo, Jun Du 0002, Shutong Niu, Jiefeng Ma, Cong Liu 0006, Qingfeng Liu |
CVPR | 4 |
| 2025 | An Investigation on Audio-Prompt and Structure Guided Long-Duration Music Generation Based on Diffusion ModelsabstractExisting text-to-music generation models face two main limitations: they either only generate approximately 10 seconds of music, falling short of users’ needs for longer compositions, or they require longer-duration datasets and increased output feature dimensions for training, leading to higher data and computational costs. We systematically explored two training-free strategies for long-duration music generation based on diffusion models. Both methods divide the task into multiple stages, where each stage generates a segment of music. The outputs from each stage are then concatenated to form the complete composition. The first, called the audio-prompt-guided strategy, employs the tail of audio features from the previous stage to serve as a prefix of the current stage, using masking techniques to guide the generation of the remainder of the current stage. The second, the structure-guided strategy, leverages the structure of previous audio representations as intermediate states in the current stage generation, initiating the inference sampling process from the intermediate time step. Our experiments demonstrate that existing short-form music generation systems are capable of generating long-duration music. Our demo page is provided at https://longmusic.github.io/long-music/. Zilu Guo, Jun Du 0002 |
ICME | 2 |
| 2025 | QA-MDT: Quality-aware Masked Diffusion Transformer for Enhanced Music GenerationabstractText-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often scarce in available datasets. Most open-source datasets frequently suffer from issues like low-quality waveforms and low text-audio consistency, hindering the advancement of music generation models. To address these challenges, we propose a novel quality-aware training paradigm for generating high-quality, high-musicality music from large-scale, quality-imbalanced datasets. Additionally, by leveraging unique properties in the latent space of musical signals, we adapt and implement a masked diffusion transformer (MDT) model for the TTM task, showcasing its capacity for quality control and enhanced musicality. Furthermore, we introduce a three-stage caption refinement approach to address low-quality captions' issue. Experiments show state-of-the-art (SOTA) performance on benchmark datasets including MusicCaps and the Song-Describer Dataset with both objective and subjective metrics. Demo audio samples are available at https://qa-mdt.github.io/, code and pretrained checkpoints are open-sourced at https://github.com/ivcylc/OpenMusic. Ruoyu Wang 0029, Jun Du 0002, Yixuan Sun, Zilu Guo, Zhengrong Zhang, Jianqing Gao |
IJCAI | 6 |
| 2025 | Controllable Conformer for Speech Enhancement and RecognitionabstractWe propose a novel approach to speech enhancement, termed Controllable ConforMer for Speech Enhancement (CCMSE), which leverages a Conformer-based architecture integrated with a control factor embedding module. Our method is designed to optimize speech quality for both human auditory perception and automatic speech recognition (ASR). It is observed that while mild denoising typically preserves speech naturalness, stronger denoising can improve human auditory tasks but often at the cost of ASR accuracy due to increased distortion. To address this, we introduce an algorithm that balances these trade-offs. By utilizing differential equations to interpolate between outputs at varying levels of denoising intensity, our method effectively combines the robustness of mild denoising with the clarity of stronger denoising, resulting in enhanced speech that is well-suited for both human and machine listeners. Experimental results on the CHiME-4 dataset validate the effectiveness of our approach. Zilu Guo, Jun Du 0002, Sabato Marco Siniscalchi, Qingfeng Liu |
IEEE Signal Process. Lett. | 1 |
| 2025 | DSNet: A Novel Way to Use Atrous Convolutions in Semantic SegmentationabstractAtrous convolutions are employed as a method to increase the receptive field in semantic segmentation tasks. However, in previous works of semantic segmentation, it was rarely employed in the shallow layers of the model. We revisit the design of atrous convolutions in modern convolutional neural networks (CNNs), and demonstrate that the concept of using large kernels to apply atrous convolutions could be a more powerful paradigm. We propose three guidelines to apply atrous convolutions more efficiently: Do not only use atrous convolutions, Avoiding the “Atrous Disasters”, Appropriate fusion mechanisms make it perfect. Following these guidelines, we propose DSNet, a Dual-Branch CNN architecture, which incorporates atrous convolutions in the shallow layers of the model architecture, as well as pretraining the nearly entire encoder on ImageNet to achieve better performance. To demonstrate the effectiveness of our approach, our models achieve a new state-of-the-art trade-off between accuracy and speed on ADE20K, Cityscapes and BDD datasets. Specifically, DSNet achieves 40.0% mIOU with inference speed of 179.2 FPS on ADE20K, and 80.4% mIOU with speed of 81.9 FPS on Cityscapes. Additionally, we propose a novel multi-scale attention fusion module, MSAF. It demonstrates outstanding performance in classification as well as downstream tasks such as segmentation. Source code and models are available at Github:https://github.com/takaniwa/DSNet. Zilu Guo, Liuyang Bian, Hu Wei, Huasheng Ni |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | A Variance-Preserving Interpolation Approach for Diffusion Models With Applications to Single Channel Speech Enhancement and RecognitionabstractIn this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM, and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach. Code and audio examples are available onlinehttps://github.com/zelokuo/VPIDM. Zilu Guo, Qing Wang 0008, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Variance-Preserving-Based Interpolation Diffusion Models for Speech Enhancement
Zilu Guo, Jun Du 0002, Chin-Hui Lee 0001, Wenbin Zhang 0002 |
INTERSPEECH | 1 |
| 2022 | Joint Optimization of the Module and Sign of the Spectral Real Part Based on CRN for Speech Denoising
Zilu Guo, Xu Xu 0003, Zhongfu Ye |
INTERSPEECH | 1 |
| 2021 | Automatically Paraphrasing via Sentence Reconstruction and Round-trip TranslationabstractParaphrase generation plays key roles in NLP tasks such as question answering, machine translation, and information retrieval. In this paper, we propose a novel framework for paraphrase generation. It simultaneously decodes the output sentence using a pretrained wordset-to-sequence model and a round-trip translation model. We evaluate this framework on Quora, WikiAnswers, MSCOCO and Twitter, and show its advantage over previous state-of-the-art unsupervised methods and distantly-supervised methods by significant margins on all datasets. For Quora and WikiAnswers, our framework even performs better than some strongly supervised methods with domain adaptation. Further, we show that the generated paraphrases can be used to augment the training data for machine translation to achieve substantial improvements. Zilu Guo, Zhongqiang Huang, Kenny Q. Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, Fei Huang 0002 |
IJCAI | 1 |