VLDB 2026 Research / reviewers in the wild / expert
Lanxiang Zhou
dblp:291/8654
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0009-7003-287XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FASTER: Face Attribute Sliders with Semantic RewardsabstractLarge-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing multi-step gradients, leading to considerable time and memory consumption. In this study, we propose FASTER(Face Attribute Sliders wiTh sEmantic Rewards), an effective method that employs stable diffusion models for face attribute editing. The key idea is to identify a low-rank attribute editing direction by leveraging attribute reward and S-CLIP reward between the original face and the edited face. This process helps to establish the desired face attribute slider. To acquire the edited face, we introduce an efficient one-step reward technique by utilizing denoised results at random timesteps for learning. This technique reduces training time by 6x. FASTER achieves 98.67% editing accuracy while simultaneously improving attribute preservation by nearly 10% compared to other methods on the CelebA-HQ dataset, all without compromising identity information. Jingyan Chen, Lanxiang Zhou, Han Fang 0002, Zerun Feng, Chao Ban, Hao Sun 0038, Jiani Hu |
ICASSP | 2 |
| 2025 | ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training FrameworkabstractThe rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video utilization and catastrophic forgetting due to the task-specific nature of vision instruction tuning. Take case of multiple-choice Video Question Answering (VideoQA) as an example, requiring only selecting correct answers makes it easier for LLMs to shortcut rather than fully comprehend video content. Moreover, the fixed task format of ’’predicting the correct option" risks catastrophic forgetting, which can cause LLMs to lose their origin cognitive abilities. To address this, we propose ViCo, a training framework that enables LLMs to perform a specific auxiliary tasks during vision instruction tuning, which helps the model leverage video information more thoroughly while mitigating catastrophic forgetting, thereby improving modality alignment and enhancing performance. We evaluate our method on several strong VideoQA benchmarks, where it outperforms all state-of-the-art methods, demonstrating its effectiveness and generality. Additionally, we also provide extensive analysis on the video information utilization and catastrophic forgetting resulting. The code will be made available at https://github.com/dunknsabsw/ViCo. Zhenda Yu, Jiayu Shen, Lanxiang Zhou, Han Fang 0002, Xianghao Zang, Chao Ban, Jingfeng Chen, Zhongjiang He, Hao Sun 0038, Yanmei Kang |
ICASSP | 4 |
| 2024 | ProTA: Probabilistic Token Aggregation for Text-Video RetrievalabstractText-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%). Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038 |
ICME | 5 |
| 2024 | GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningabstractRecent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality. Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038 |
ACM Multimedia | 7 |
| 2023 | Mask to Reconstruct: Cooperative Semantics Completion for Video-text RetrievalabstractRecently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo. Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou |
ACM Multimedia | 7 |
| 2023 | Exploring Frequency Attention Learning and Contrastive Learning for Face Forgery Detection
Neng Fang, Lanxiang Zhou |
PRCV (5) | 5 |
| 2020 | DFH-GAN: A Deep Face Hashing with Generative Adversarial NetworkabstractFace Image retrieval is one of the key research directions in computer vision field. Thanks to the rapid development of deep neural network in recent years, deep hashing has achieved good performance in the field of image retrieval. But for large-scale face image retrieval, the performance needs to be further improved. In this paper, we propose Deep Face Hashing with GAN (DFH-GAN), a novel deep hashing method for face image retrieval, which mainly consists of three components: a generator network for generating synthesized images, a discriminator network with a shared CNN to learn multi-domain face feature, and a hash encoding network to generate compact binary hash codes. The generator network is used to perform data augmentation so that the model could learn from both real images and diverse synthesized images. We adopt a two-stage training strategy. In the first stage, the GAN is trained to generate fake images, while in the second stage, to make the network convergence faster. The model inherits the trained shared CNN of discriminator to train the DFH model by using many different supervised loss functions not only in the last layer but also in the middle layer of the network. Extensive experiments on two widely used datasets demonstrate that DFH-GAN can generate high-quality binary hash codes and exceed the performance of the state-of-the-art model greatly. Lanxiang Zhou, Bo Xiao 0006, Qianfang Xu |
ICPR | 1 |