VLDB 2026 Research / reviewers in the wild / expert
Xiwen Liang
dblp:226/6507
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Structured Preference Optimization for Vision-Language Long-Horizon Task PlanningabstractXiwen Liang, Min Lin, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Jiaqi Chen, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xiwen Liang, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang |
EMNLP | 1 |
| 2025 | PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block AssemblyabstractWhile vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressive benchmark designed to assess VLMs on physical understanding and planning through robotic 3D block assembly tasks. PhyBlock integrates a novel four-level cognitive hierarchy assembly task alongside targeted Visual Question Answering (VQA) samples, collectively aimed at evaluating progressive spatial reasoning and fundamental physical comprehension, including object properties, spatial relationships, and holistic scene understanding. PhyBlock includes 2600 block tasks (400 assembly tasks, 2200 VQA tasks) and evaluates models across three key dimensions: partial completion, failure diagnosis, and planning robustness. We benchmark 23 state-of-the-art VLMs, highlighting their strengths and limitations in physically grounded, multi-step planning. Our empirical findings indicate that the performance of VLMs exhibits pronounced limitations in high-level planning and reasoning capabilities, leading to a notable decline in performance for the growing complexity of the tasks.Error analysis reveals persistent difficulties in spatial orientation and dependency reasoning.We position PhyBlock as a unified testbed to advance embodied reasoning, bridging vision-language understanding and real-world physical problem-solving. Jiajun Wen 0003, Rongtao Xu, Xiwen Liang, Bingqian Lin, Ziming Wei 0001, Haokun Lin, Mingfei Han 0002, Meng Cao 0002, Bokui Chen, Ivan Laptev, Xiaodan Liang |
NeurIPS | 5 |
| 2025 | AATM: An Anonymous Authentication Protocol for Time Span of Membership With Self-Blindness and AccountabilityabstractInternet of Things (IoT) devices using subscription services (e.g. connected vehicles accessing entertainment programs) often purchase membership credentials from service providers with limited usage counts or validity periods, we call them pay-per-use or time span of membership services. However, users’ access records, usage preferences, and habits are collected by network adversarys or membership providers for creating users’ profiles, targeted advertising, and even for being sold maliciously. To deal with these problems, lots of anonymous authentication protocols are proposed to provide users with pseudonyms to conceal their real identities. Although these protocols effectively prevent network adversarys from compromising users’ privacy, membership service providers can still gather users’ behavioral privacy via their membership credentials. Therefore, several scholars proposed k-times anonymous authentication protocols and self-blind credentials to enhance users’ privacy protection, but the k-times anonymous authentication protocols are only for pay-per-use membership services and the schemes of self-blind credentials are lack of regulating malicious users. To address these issues, this article proposes an anonymous authentication protocol for time span of membership (AATM) with self-blindness and accountability. Specifically, we utilize Structure Preserving Signatures on Equivalence Classes (SPS-EQ) and Signatures with Flexible Public Key (SFPK) to build accountable, self-blinding credentials that ensure that every time a user visits a member, he or she can create a brand new identity on their own, which not only prevents users from being linked by service providers, but also supports conditional fair regulation. Security and performance analyses show that AATM is better than the state-of-the-art schemes in terms of security and privacy-preserving capabilities, and its computation cost also meets the practical application requirements. Qiuyun Lyu, Xiwen Liang, Shaopeng Cheng, Yizhi Ren, Chengli Xu, Weizhi Meng 0001, Duohe Ma |
IEEE Internet Things J. | 2 |
| 2023 | NLIP: Noise-Robust Language-Image Pre-trainingabstractLarge-scale cross-modal pre-training paradigms have recently shown ubiquitous success on a wide range of downstream tasks, e.g., zero-shot classification, retrieval and image captioning. However, their successes highly rely on the scale and quality of web-crawled data that naturally contain much incomplete and noisy information (e.g., wrong or irrelevant contents). Existing works either design manual rules to clean data or generate pseudo-targets as auxiliary signals for reducing noise impact, which do not explicitly tackle both the incorrect and incomplete challenges at the same time. In this paper, to automatically mitigate the impact of noise by solely mining over existing data, we propose a principled Noise-robust Language-Image Pre-training framework (NLIP) to stabilize pre-training via two schemes: noise-harmonization and noise-completion. First, in noise-harmonization scheme, NLIP estimates the noise probability of each pair according to the memorization effect of cross-modal transformers, then adopts noise-adaptive regularization to harmonize the cross-modal alignments with varying degrees. Second, in noise-completion scheme, to enrich the missing object information of text, NLIP injects a concept-conditioned cross-modal decoder to obtain semantic-consistent synthetic captions to complete noisy ones, which uses the retrieved visual concepts (i.e., objects’ names) for the corresponding image to guide captioning generation. By collaboratively optimizing noise-harmonization and noise-completion schemes, our NLIP can alleviate the common noise effects during image-text pre-training in a more efficient way. Extensive experiments show the significant performance improvements of our NLIP using only 26M data over existing pre-trained models (e.g., CLIP, FILIP and BLIP) on 12 zero-shot classification datasets (e.g., +8.6% over CLIP on average accuracy), MSCOCO image captioning (e.g., +1.9 over BLIP trained with 129M data on CIDEr) and zero-shot image-text retrieval tasks. Runhui Huang, Yanxin Long, Jianhua Han, Hang Xu 0004, Xiwen Liang, Chunjing Xu, Xiaodan Liang |
AAAI | 5 |
| 2023 | Visual Exemplar Driven Task-Prompting for Unified Perception in Autonomous DrivingabstractMulti-task learning has emerged as a powerful paradigm to solve a range of tasks simultaneously with good efficiency in both computation resources and inference time. However, these algorithms are designed for different tasks mostly not within the scope of autonomous driving, thus making it hard to compare multi-task methods in autonomous driving. Aiming to enable the comprehensive evaluation of present multi-task learning methods in autonomous driving, we extensively investigate the performance of popular multi-task methods on the large-scale driving dataset, which covers four common perception tasks, i.e., object detection, semantic segmentation, drivable area segmentation, and lane detection. We provide an in-depth analysis of current multi-task learning methods under different common settings and find out that the existing methods make progress but there is still a large performance gap compared with single-task baselines. To alleviate this dilemma in autonomous driving, we present an effective multi-task framework, VE-Prompt, which introduces visual exemplars via task-specific prompting to guide the model toward learning high-quality task-specific representations. Specifically, we generate visual exemplars based on bounding boxes and color-based markers, which provide accurate visual appearances of target categories and further mitigate the performance gap. Furthermore, we bridge transformer-based encoders and convolutional layers for efficient and accurate unified perception in autonomous driving. Comprehensive experimental results on the diverse self-driving dataset BDD100K show that the VE-Prompt improves the multi-task baseline and further surpasses single-task models. Xiwen Liang, Minzhe Niu, Jianhua Han, Hang Xu 0004, Chunjing Xu, Xiaodan Liang |
CVPR | 1 |
| 2022 | Contrastive Instruction-Trajectory Learning for Vision-Language NavigationabstractThe vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction. Previous works learn to navigate step-by-step following an instruction. However, these works may fail to discriminate the similarities and discrepancies across instruction-trajectory pairs and ignore the temporal continuity of sub-instructions. These problems hinder agents from learning distinctive vision-and-language representations, harming the robustness and generalizability of the navigation policy. In this paper, we propose a Contrastive Instruction-Trajectory Learning (CITL) framework that explores invariance across similar data samples and variance across different ones to learn distinctive representations for robust navigation. Specifically, we propose: (1) a coarse-grained contrastive learning objective to enhance vision-and-language representations by contrasting semantics of full trajectory observations and instructions, respectively; (2) a fine-grained contrastive learning objective to perceive instructions by leveraging the temporal information of the sub-instructions; (3) a pairwise sample-reweighting mechanism for contrastive learning to mine hard samples and hence mitigate the influence of data sampling bias in contrastive learning. Our CITL can be easily integrated with VLN backbones to form a new learning paradigm and achieve better generalizability in unseen environments. Extensive experiments show that the model with CITL surpasses the previous state-of-the-art methods on R2R, R4R, and RxR. Xiwen Liang, Fengda Zhu, Yi Zhu 0004, Bingqian Lin, Xiaodan Liang |
AAAI | 1 |
| 2022 | Visual-Language Navigation Pretraining via Prompt-based Environmental Self-explorationabstractVision-language navigation (VLN) is a challenging task due to its large searching space in the environment.To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale datasets.However, the conventional fine-tuning methods require extra human-labeled navigation data and lack self-exploration capabilities in environments, which hinders their generalization of unseen scenes.To improve the ability of fast cross-domain adaptation, we propose Prompt-based Environmental Self-exploration (ProbES), which can self-explore the environments by sampling trajectories and automatically generates structured instructions via a large-scale cross-modal pretrained model (CLIP).Our method fully utilizes the knowledge learned from CLIP to build an in-domain dataset by self-exploration without human labeling.Unlike the conventional approach of fine-tuning, we introduce prompt-based learning to achieve fast adaptation for language embeddings, which substantially improves the learning efficiency by leveraging prior knowledge.By automatically synthesizing trajectoryinstruction pairs in any environment without human supervision and efficient prompt-based learning, our model can adapt to diverse visionlanguage navigation tasks, including VLN and REVERIE.Both qualitative and quantitative results show that our ProbES significantly improves the generalization ability of the navigation model * . Xiwen Liang, Fengda Zhu, Hang Xu 0004, Xiaodan Liang |
ACL (1) | 1 |
| 2022 | ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsabstractVision-Language Navigation (VLN) is a challenging task that requires an embodied agent to perform action-level modality alignment, i.e., make instruction-asked actions sequentially in complex visual environments. Most existing VLN agents learn the instruction-path data directly and cannot sufficiently explore action-level alignment knowledge inside the multi-modal inputs. In this paper, we propose modAlity-aligneD Action PrompTs (ADAPT), which provides the VLN agent with action prompts to enable the explicit learning of action-level modality alignment to pursue successful navigation. Specifically, an action prompt is defined as a modality-aligned pair of an image sub-prompt and a text sub-prompt, where the former is a single-view observation and the latter is a phrase like “walk past the chair”. When starting navigation, the instruction-related action prompt set is retrieved from a prebuilt action prompt base and passed through a prompt encoder to obtain the prompt feature. Then the prompt feature is concatenated with the original instruction feature and fed to a multilayer transformer for action prediction. To collect high-quality action prompts into the prompt base, we use the Contrastive Language-Image Pretraining (CLIP) model which has powerful cross-modality alignment ability. A modality alignment loss and a sequential consistency loss are further introduced to enhance the alignment of the action prompt and enforce the agent to focus on the related prompt sequentially. Experimental results on both R2R and RxR show the superiority of ADAPT over state-of-the-art methods. Bingqian Lin, Yi Zhu 0004, Zicong Chen, Xiwen Liang, Jianzhuang Liu, Xiaodan Liang |
CVPR | 4 |
| 2022 | Effective Adaptation in Multi-Task Co-Training for Unified Autonomous DrivingabstractAiming towards a holistic understanding of multiple downstream tasks simultaneously, there is a need for extracting features with better transferability. Though many latest self-supervised pre-training methods have achieved impressive performance on various vision tasks under the prevailing pretrain-finetune paradigm, their generalization capacity to multi-task learning scenarios is yet to be explored. In this paper, we extensively investigate the transfer performance of various types of self-supervised methods, e.g., MoCo and SimCLR, on three downstream tasks, including semantic segmentation, drivable area segmentation, and traffic object detection, on the large-scale driving dataset BDD100K. We surprisingly find that their performances are sub-optimal or even lag far behind the single-task baseline, which may be due to the distinctions of training objectives and architectural design lied in the pretrain-finetune paradigm. To overcome this dilemma as well as avoid redesigning the resource-intensive pre-training stage, we propose a simple yet effective pretrain-adapt-finetune paradigm for general multi-task training, where the off-the-shelf pretrained models can be effectively adapted without increasing the training overhead. During the adapt stage, we utilize learnable multi-scale adapters to dynamically adjust the pretrained model weights supervised by multi-task objectives while leaving the pretrained knowledge untouched. Furthermore, we regard the vision-language pre-training model CLIP as a strong complement to the pretrain-adapt-finetune paradigm and propose a novel adapter named LV-Adapter, which incorporates language priors in the multi-task model via task-specific prompting and alignment between visual and textual features. Our experiments demonstrate that the adapt stage significantly improves the overall performance of those off-the-shelf pretrained models and the contextual features generated by LV-Adapter are of general benefits for downstream tasks. Xiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu 0004, Chunjing Xu, Xiaodan Liang |
NeurIPS | 1 |
| 2022 | Configurable Graph Reasoning for Visual Relationship DetectionabstractVisual commonsense knowledge has received growing attention in the reasoning of long-tailed visual relationships biased in terms of object and relation labels. Most current methods typically collect and utilize external knowledge for visual relationships by following the fixed reasoning path of {subject, object → predicate} to facilitate the recognition of infrequent relationships. However, the knowledge incorporation for such fixed multidependent path suffers from the data set biased and exponentially grown combinations of object and relation labels and ignores the semantic gap between commonsense knowledge and real scenes. To alleviate this, we propose configurable graph reasoning (CGR) to decompose the reasoning path of visual relationships and the incorporation of external knowledge, achieving configurable knowledge selection and personalized graph reasoning for each relation type in each image. Given a commonsense knowledge graph, CGR learns to match and retrieve knowledge for different subpaths and selectively compose the knowledge routed path. CGR adaptively configures the reasoning path based on the knowledge graph, bridges the semantic gap between the commonsense knowledge, and the real-world scenes and achieves better knowledge generalization. Extensive experiments show that CGR consistently outperforms previous state-of-the-art methods on several popular benchmarks and works well with different knowledge graphs. Detailed analyses demonstrated that CGR learned explainable and compelling configurations of reasoning paths. Yi Zhu 0004, Xiwen Liang, Bingqian Lin, Qixiang Ye, Jianbin Jiao, Liang Lin 0004, Xiaodan Liang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | SOON: Scenario Oriented Object Navigation With Graph-Based ExplorationabstractThe ability to navigate like a human towards a language-guided target from anywhere in a 3D embodied environment is one of the ‘holy grail’ goals of intelligent robots. Most visual navigation benchmarks, however, focus on navigating toward a target from a fixed starting point, guided by an elaborate set of instructions that depicts step-by-step. This approach deviates from real-world problems in which human-only describes what the object and its surrounding look like and asks the robot to start navigation from any-where. Accordingly, in this paper, we introduce a Scenario Oriented Object Navigation (SOON) task. In this task, an agent is required to navigate from an arbitrary position in a 3D embodied environment to localize a target following a scene description. To give a promising direction to solve this task, we propose a novel graph-based exploration (GBE) method, which models the navigation state as a graph and introduces a novel graph-based exploration approach to learn knowledge from the graph and stabilize training by learning sub-optimal trajectories. We also propose a new large-scale benchmark named From Anywhere to Object (FAO) dataset. To avoid target ambiguity, the descriptions in FAO provide rich semantic scene information includes: object attribute, object relationship, region description, and nearby region description. Our experiments reveal that the proposed GBE outperforms various state-of-the-arts on both FAO and R2R datasets. And the ablation studies on FAO validates the quality of the dataset. Fengda Zhu, Xiwen Liang, Yi Zhu 0004, Qizhi Yu, Xiaojun Chang, Xiaodan Liang |
CVPR | 2 |
| 2019 | Learning mean progressive scattering using binomial truncated loss for image dehazingabstractIn this study, the authors propose a novel progressive dehazing network to address the single image haze removal problem based on a new mean progressive scattering model. Different from methods that learn atmosphere light and transmission maps with different networks, these two variables are optimised in a unified network. Following the methodology of traditional prior‐based methods that estimate a coarse transmission map first, a progressive refinement branch in the decoder has been designed to restore the fine‐scale transmission map. To improve the prediction accuracy of the transmission map, a novel binomial truncated loss that assigns weights to error values according to the probabilities of error occurrences has been proposed. An ablation study is conducted to verify the effectiveness of the components in the proposed method. Experiments in the synthetic datasets and real images demonstrate that the proposed method outperforms other state‐of‐the‐art methods. Bin Qiu, Xiwen Liang, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001 |
IET Image Process. | 2 |