VLDB 2026 Research / reviewers in the wild / expert
Zhaohui Hou
dblp:202/8276
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modality and Task Adaptation for Enhanced Zero-shot Composed Image RetrievalabstractAs a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens that effectively represent the input semantics. However, the inversion-based methods suffer from two inherent issues: First, the task discrepancy exists because inversion training and CIR inference involve different objectives. Second, the modality discrepancy arises from the input feature distribution mismatch between training and inference. To this end, we propose a lightweight post-hoc framework, consisting of two components: (1) A new text-anchored triplet construction pipeline leverages a large language model (LLM) to transform a standard image-text dataset into a triplet dataset, where a textual description serves as the target of each triplet. (2) The MoTa-Adapter, a novel parameter-efficient fine-tuning method, adapts the dual encoder to the CIR task using our constructed triplet data. Specifically, on the text side, multiple sets of learnable task prompts are integrated via a Mixture-of-Experts (MoE) layer to capture task-specific priors and handle different types of modifications. On the image side, MoTa-Adapter modulates the inversion network's input to better match the downstream text encoder. In addition, an entropy-based optimization strategy is proposed to assign greater weight to challenging samples, thus improving adaptation efficiency. Experiments show that, with the incorporation of our proposed components, inversion-based methods achieve significant improvements, reaching state-of-the-art performance across four widely-used benchmarks. Haiwen Li, Delong Liu, Zhaohui Hou, Zeliang Ma, Zhicheng Zhao 0001 |
AAAI | 3 |
| 2026 | RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic DataabstractPrecise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while instruction-based approaches lack accurate spatial control and tend to unintentionally modify background regions. To address these issues, we propose two key contributions. First, we develop a fully automated and self-improving pipeline for synthetic data generation. This pipeline utilizes a Large Language Model (LLM) to generate diverse prompts, a Diffusion Transformer (DiT) fine-tuned evolutionarily to synthesize high-quality images, and a Multimodal LLM (MLLM) combined with open-set object detector for automated quality control and annotation. This process produces the Remove/Add Dataset (RAD), consisting of over 514,510 high-quality image pairs, each richly annotated with bounding boxes, segmentation masks, and a variety of editing instructions. Second, based on RAD, we introduce Remove/Add Anything (RAA), a novel editing framework with precise spatial control. Built upon a diffusion-based inpainting model, RAA achieves high editing accuracy by conditioning on both textual instructions and an explicitly defined region of interest (ROI), enabling efficient fine-tuning while maintaining global visual coherence. Extensive experiments demonstrate that RAA significantly outperforms existing open-source methods on both addition and removal tasks, and even slightly surpasses costly proprietary models. Delong Liu, Haotian Hou, Zhaohui Hou, Shihao Han, Mingjie Zhan, Zhicheng Zhao 0001 |
AAAI | 3 |
| 2025 | UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame OrganizerabstractRecently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks. Delong Liu, Zhaohui Hou, Mingjie Zhan, Shihao Han, Zhicheng Zhao 0001 |
AAAI | 2 |
| 2025 | SpiritSight Agent: Advanced GUI Agent with One LookabstractGraphical User Interface (GUI) agents demonstrate promising potential in assisting human-computer interaction, automating human user’s navigation on digital devices. An ideal GUI agent is expected to achieve high accuracy, low latency, and compatibility for different GUI platforms. Recent vision-based approaches have shown promise by leveraging advanced Vision Language Models (VLMs). While they generally meet the requirements of compatibility and low latency, these vision-based GUI agents tend to have low accuracy due to their limitations in element grounding. To address this issue, we propose SpiritSight, a vision-based, end-to-end GUI agent that excels in GUI navigation tasks across various GUI platforms. First, we create a multi-level, large-scale, high-quality GUI dataset called GUI-Lasagne using scalable methods, empowering SpiritSight with robust GUI understanding and grounding capabilities. Second, we introduce the Universal Block Parsing (UBP) method to resolve the ambiguity problem inherited from the dynamic resolution strategy, further enhancing SpiritSight’s ability to ground GUI objects. Through these efforts, SpiritSight agent outperforms other advanced methods on diverse GUI benchmarks, demonstrating its superior capability and compatibility in GUI navigation tasks. The models and code will be made available upon publication. Ziming Cheng, Junting Pan, Zhaohui Hou, Mingjie Zhan |
CVPR | 4 |
| 2025 | Think Twice: Empowering Action Recognition Models with Human-Like Deep ReasoningabstractWhen engaged in complex visual cognition, humans tend to rely on their experience and make decisions after thinking again and again. Inspired by this, we pour similar capability into action recognition and propose a new Think Twice framework, that is, think twice about similar categories that are easy to confuse, thus obtaining performance improvement. Firstly, based on visual similarity, a large language model is applied to cluster all categories of a given dataset into disjoint cliques. Accordingly, a textual prompt for each clique will be generated. Secondly, through the first inference, pseudo-labels are obtained, and then the prompt corresponding to its clique is assigned to each sample. Thirdly, a prompt learning method is integrated to enable the framework to simulate human-like iterative thinking, yielding a final decision. Our proposed framework requires minimal parameters while achieving state-of-the-art parameter-efficient fine-tuning(PEFT) performance across four datasets. Our code is available at https://github.com/KangRuan6/ThinkTwice. Xiangning Ruan, Baoxing Xie, Zhaohui Hou, Qixiang Yin, Zhicheng Zhao 0001 |
ICME | 3 |
| 2025 | Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person RetrievalabstractPerson retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet diverse application requirements. To address the above limitations, we propose a new Composed Person Retrieval (CPR) task, which combines visual and textual queries to identify individuals of interest from large-scale person image databases. Nevertheless, the foremost difficulty of the CPR task is the lack of available annotated datasets. Therefore, we first introduce a scalable automatic data synthesis pipeline, which decomposes complex multimodal data generation into the creation of textual quadruples followed by identity-consistent image synthesis using fine-tuned generative models. Meanwhile, a multimodal filtering method is designed to ensure the resulting SynCPR dataset retains 1.15 million high-quality and fully synthetic triplets. Additionally, to improve the representation of composed person queries, we propose a novel Fine-grained Adaptive Feature Alignment (FAFA) framework through fine-grained dynamic alignment and masked feature reasoning. Moreover, for objective evaluation, we manually annotate the Image-Text Composed Person Retrieval (ITCPR) test set. The extensive experiments demonstrate the effectiveness of the SynCPR dataset and the superiority of the proposed FAFA framework when compared with the state-of-the-art methods. All code and data will be provided at https://github.com/Delong-liu-bupt/Composed_Person_Retrieval. Delong Liu, Haiwen Li, Zhaohui Hou, Zhicheng Zhao 0001 |
NeurIPS | 3 |
| 2023 | Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting TranscriptsabstractMeetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content.To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and efficient meeting summarization.RbS first leverages a self-supervised paradigm to annotate essential contents by reconstructing the meeting transcripts.Secondly, we propose a relative positional bucketing (RPB) algorithm to equip (conventional) summarization models to generate the summary.Despite the additional reconstruction process, our proposed RPB significantly compressed the input, leading to faster processing and reduced memory consumption compared to traditional summarization methods.We validate the effectiveness and efficiency of our method through extensive evaluations and analysis.On two meeting summarization datasets, AMI and ICSI, our approach outperforms previous state-of-the-art approaches without relying on large-scale pretraining or expert-grade annotating tools. Haochen Tan, Han Wu 0004, Wei Shao 0009, Xinyun Zhang 0001, Mingjie Zhan, Zhaohui Hou, Ding Liang, Linqi Song |
EMNLP | 6 |
| 2021 | Rethinking Anchor-Object Matching and Encoding in Rotating Object DetectionabstractRotating object detection is more challenging than horizontal object detection because of the multi-orientation of the objects involved. In the recent anchor-based rotating object detector, the IoU-based matching mechanism has some mismatching and wrong-matching problems. Moreover, the encoding mechanism does not correctly reflect the location relationships between anchors and objects. In this paper, RBox-Diff-based matching (RDM) mechanism and angle-first encoding (AE) method are proposed to solve these problems. RDM optimizes the anchor-object matching by replacing IoU (Intersection-over-Union) with a new concept called RBox-Diff, while AE optimizes the encoding mechanism to make the encoding results consistent with the relative position between objects and anchors more. The proposed methods can be easily applied to most of the anchor-based rotating object detectors without introducing extra parameters. The extensive experiments on DOTA-v1.0 dataset show the effectiveness of the proposed methods over other advanced methods. Zhaohui Hou, Pingyu Wang, Zhicheng Zhao 0001 |
VCIP | 2 |
| 2019 | Low-Thrust Trajectory Design Optimization with Stochastic ConvolutionabstractIn this paper, a new stochastic based method for low-thrust trajectory design optimization is proposed. Indirect optimization methods based on optimal control theory (OCT) provide accurate solutions but convergence is difficult. In our new approach, an augmented Gaussian is used to relax the terminal conditions of the Two Point Boundary Value Problem (TPBVP), and model overall impacts of initial guess of costates. Objective function of the low-thrust trajectory design is transformed into a quasi-quadratic problem using quadratic convolution technique. The quasi-quadratic form objective function contains information matrix of the impacts due to varied costates and terminal conditions. A gradient based searching algorithm is then used to search the optimal initial guess. Based on the method, two algorithms are developed for resolving the TPBVP of low thrust trajectory design. Convergence performances of the algorithms are demonstrated through two low-thrust interplanetary trajectory design problems. Possible extension of the algorithm to multi-objective evolutionary optimization is also discussed. Liqiang Hou, Shufan Wu, Zhaohui Hou, Zhongcheng Mu |
CEC | 3 |
| 2018 | Optimal Multi-Gravity-Assist Trajectories Design with Likelihood AnalysisabstractA likelihood assisted optimization strategy for the complex mixed-integer nonlinear programming problem, Multiple Gravity Assist (MGA) trajectory design is proposed. In MGA design, both the total velocity ΔV and MGA sequence are considered, and possible transfers are incrementally built and explored at each planetary encounter. Dimension of searching subspace increases exponentially with the number of MGA. Traditional MGA design uses heuristic based branching and pruning techniques to reduce the computational cost due to the exponentially increased searching spaces. In this work, a new stochastic based searching strategy without branching-pruning operations is proposed instead. A stochastic type metric is introduced and used to measure similarity level of the spacecraft trajectory to the expected transfer orbit. Swing-by planets are sampled and new transfer arcs are generated with respect to the metric values. Log-likelihood of the orbital transfers is constructed, and put into the optimization as the constraint to be maximized. Based on the strategy, the MGA design problem is translated into a continuous non-linear optimization problem, and two algorithms for the MGA trajectory design are proposed. In the first algorithm, tuning parameters of the similarity function are put into the design space as extended parameters, while in the second algorithm, an adaptive update mechanism for the tuning parameters is designed. Tuning parameters are updated with population of the objective function values and prior data set of tuning parameters. Effectiveness of the algorithms are demonstrated through design optimization of MarcoPolo mission. Simulation results show that significant efficiency improvement of the searching process can be obtained. Liqiang Hou, Zhaohui Hou |
CEC | 2 |
| 2017 | Multi-objective optimization with Proper Orthogonal Decomposition and Gaussian predictive distributionabstractA numerical study of multi-objective optimization with Proper Orthogonal Decomposition (POD) and Gaussian predictive distribution on the well known ZDT and DLZT benchmark set, is presented and discussed. Based on the algorithm, design optimization of the Global Trajectory Optimization Problems (GTOP) database based on trajectory models of real-world interplanetary space mission, Cassini is presented. The trajectory models are formulated as nonlinear optimization problems and are known to be difficult to solve. In this contribution, for part of the standard ZDT and DTLZ series problems, the proposed algorithm is able to solve these benchmarks to their optimal solutions within 20 generations. Liqiang Hou, Zhaohui Hou |
CEC | 4 |