VLDB 2026 Research / reviewers in the wild / expert
Zhixuan Liu
dblp:286/2761
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-8158-7127ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object MemoryabstractReasoning Video Object Segmentation (ReasonVOS) is a challenging task that requires stable object segmentation across video sequences using implicit and complex textual inputs. Previous methods fine-tune Multimodal Large Language Models (MLLMs) to produce segmentation outputs, which demand substantial resources. Additionally, some existing methods are coupled in the processing of spatio-temporal information, which affects the temporal stability of the model to some extent. To address these issues, we propose Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory (SDAM). We aim to design a training-free reasoning video segmentation framework that outperforms existing methods requiring fine-tuning, using only pre-trained models. Meanwhile, we propose an Adaptive Object Memory module that selects and memorizes key objects based on motion cues in different video sequences. Finally, we propose Spatio-temporal Decoupling for stable temporal propagation. In the spatial domain, we achieve precise localization and segmentation of target objects, while in the temporal domain, we leverage key object temporal information to drive stable cross-frame propagation. Our method achieves excellent results on five benchmark datasets, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS, ReasonVOS, and ReVOS. Zhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang Li |
AAAI | 3 |
| 2025 | MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments
Zhixuan Liu, Haokun Zhu, Jonathan Francis, Soonmin Hwang, Ji Zhang 0003, Jean Oh |
ICCV | 1 |
| 2025 | Emergent Response Planning in LLMsabstractIn this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations encode future outputs beyond the next token}$. Through simple probing, we demonstrate that LLM prompt representations encode global attributes of their entire responses, including $\textit{structure attributes}$ (e.g., response length, reasoning steps), $\textit{content attributes}$ (e.g., character choices in storywriting, multiple-choice answers at the end of response), and $\textit{behavior attributes}$ (e.g., answer confidence, factual consistency). In addition to identifying response planning, we explore how it scales with model size across tasks and how it evolves during generation. The findings that LLMs plan ahead for the future in their hidden representations suggest potential applications for improving transparency and generation control. Zhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang 0026, Chaochao Lu |
ICML | 3 |
| 2025 | A Vehicle-Infrastructure Collaborative Environment Perception Approach Based on Sparse BEV FeaturesabstractOvercoming the limitations of individual-vehicle line-of-sight (LOS) sensing holds significant importance for guaranteeing the safety of driving. With a wider perception field, vehicle-infrastructure (VI) collaborative perception can provide vehicles with more comprehensive perception assistance, which has received widespread attention in recent years. However, the perception data fusion between infrastructure and vehicles is still impeded by issues such as large data volume and complex processing procedures, constituting a threat to driving safety. To deal with these issues, this paper proposes a VI collaborative environment perception approach based on sparse bird's eye view (BEV) features. By leveraging the representation of sparse BEV, features can be fused within a unified perspective in a lightweight manner, thereby enhancing the efficiency of feature fusion and reducing redundancy. Additionally, we present a solution for processing the overlapping features between the EGO-vehicle and road side unit (RSU) by taking the union of the coordinate points. Finally, the applicable vehicle and RSU datasets are collected through Carla. The experimental results demonstrate that the proposed approach can effectively mitigate the limitations of individual-vehicle perception by compensating for occluded information and provide a more comprehensive perception field. Zhixuan Liu, Yuchuan Fu, Changle Li, Nan Cheng 0001, Ruijin Sun |
VTC2025-Spring | 1 |
| 2024 | SCoFT: Self-Contrastive Fine-Tuning for Equitable Image GenerationabstractAccurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with harmful stereotypes and misrepresentations of cultures. We improve inclusive representation in generated images by (1) engaging with communities to collect a culturally representative dataset that we call the Cross-Cultural Understanding Benchmark (CCUB) and (2) proposing a novel Self-Contrastive Fine-Tuning (SCoFT, pronounced /sô ft/) method that leverages the model's known biases to self-improve. SCoFT is designed to prevent overfitting on small datasets, encode only high-level information from the data, and shift the generated distribution away from misrepresentations encoded in a pretrained model. Our user study conducted on 51 participants from 5 different countries based on their self-selected national cultural affiliation shows that fine-tuning on CCUB consistently generates images with higher cultural relevance and fewer stereotypes when compared to the Stable Diffusion baseline, which is further improved with our SCoFT technique. Resources and code are at https://ariannaliu.github.io/SCoFT. Zhixuan Liu, Peter Schaldenbrand, Beverley-Claire Okogwu, Wenxuan Peng, Youngsik Yun, Andrew Hundt, Jihie Kim, Jean Oh |
CVPR | 1 |
| 2024 | Depth-Enhanced Alignment for Label-Free 3D Semantic Segmentation
Shangjin Xie, Zibo Chen, Zhixuan Liu, Wei-Shi Zheng 0001 |
ICPR (18) | 4 |
| 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language ModelsabstractLarge language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance.
Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned $\texttt{gpt2}$s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2.0, we show that reusing off-the-shelf small models (e.g., $\texttt{zephyr-7b-beta}$ and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against $\texttt{gpt-4-turbo}$ (e.g., $34.4\% \rightarrow 37.9\%$ for $\texttt{Llama-3-70B-Instruct}$ and $16.0\% \rightarrow 20.1\%$ for $\texttt{gpt-3.5-turbo-instruct}$), despite the small models' low win rates $\approx 10.0\%$. Zhanhui Zhou, Zhixuan Liu, Jie Liu 0047, Zhichen Dong, Chao Yang 0026, Yu Qiao 0001 |
NeurIPS | 2 |
| 2023 | Grasp Region Exploration for 7-DoF Robotic Grasping in Cluttered ScenesabstractRobotic grasping is a fundamental skill for robots, but it is quite challenging in cluttered scenes. In cluttered scenes, the precise prediction of high-quality grasp configurations such as rotation and grasping width while avoiding collisions is essential. To accomplish this, the grasp detection models require the capabilities of stronger fine-grained information extracted around the grasp points. However, due to the computational resource restriction, point clouds are usually downsampled in existing networks, which inevitably make some potentially important points discarded. To overcome this problem, we propose a Grasp Region Exploration module to explore the area covered by high-quality grasps. Based on the grasp region, we enhance the point density around the grasp points to mitigate the loss of information caused by downsampling. Furthermore, we devise the Grasp Region Attention module to dynamically aggregate features of various points within the grasp region, such as the grasp point and contact points. The proposed method achieves state-of-the-art performance on the large-scale GraspNet-1Billion dataset. We also conduct real-world experiments on a Franka Emika Panda robot and show that the robot can grasp objects in cluttered scenes with a high success rate. Zibo Chen, Zhixuan Liu, Shangjin Xie, Wei-Shi Zheng 0001 |
IROS | 2 |
| 2022 | TransGrasp: A Multi-Scale Hierarchical Point Transformer for 7-DoF Grasp DetectionabstractRobotic grasping pose detection that predicts the configuration of the robotic gripper for object grasping is fundamental in robot manipulation. Based on point clouds, most of the existing methods predict grasp pose with the hierarchical PointNet++ backbone, while the non-local geometric information is underexplored. In this work, we address the 7-DoF (6- DoF with the grasp width) grasp detection by introducing a one- stage Transformer-based hierarchical multi-scale model dubbed TransGrasp. Empowered by TransGrasp, the point features are enhanced via acquiring multi-scale shape awareness in the whole scene. By directly modeling the long-range relevance, our pipeline is aware of object contour to avoid collisions and able to apply analogy reasoning for long-distance geometric structures. The evaluation results on the large scale GraspNet- 1Billion dataset demonstrate the effectiveness of the proposed TransGrasp. The real robot experiments on an ABB YUMI robot with an Azure Kinect DK camera and an ABB Smart two-finger gripper show high success rates in both single object and cluttered scenes. Zhixuan Liu, Zibo Chen, Shangjin Xie, Wei-Shi Zheng 0001 |
ICRA | 1 |
| 2022 | StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing TranslationabstractGenerating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control of the style of image to be generated. We present an approach for generating styled drawings for a given text description where a user can specify a desired drawing style using a sample image. Inspired by a theory in art that style and content are generally inseparable during the creative process, we propose a coupled approach, known here as StyleCLIPDraw, whereby the drawing is generated by optimizing for style and content simultaneously throughout the process as opposed to applying style transfer after creating content in a sequence. Based on human evaluation, the styles of images generated by StyleCLIPDraw are strongly preferred to those by the sequential approach. Although the quality of content generation degrades for certain styles, overall considering both content and style, StyleCLIPDraw is found far more preferred, indicating the importance of style, look, and feel of machine generated images to people as well as indicating that style is coupled in the drawing process itself. Our code, a demonstration, and style evaluation data are publicly available. Peter Schaldenbrand, Zhixuan Liu, Jean Oh |
IJCAI | 2 |
| 2022 | Learning multimodal relationship interaction for visual relationship detection
Zhixuan Liu, Wei-Shi Zheng 0001 |
Pattern Recognit. | 1 |