VLDB 2026 Research / reviewers in the wild / expert
Jihao Liu
dblp:167/0509
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DuaFed: A clustered federated learning framework via dual-domain feature alignment for tackling data heterogeneity
Shining Zhang, Xingwei Wang 0001, Rongfei Zeng, Jihao Liu, Yu Gu 0002, Min Huang 0001 |
Knowl. Based Syst. | 5 |
| 2025 | LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical EncodingabstractYuxuan Hu, Jihao Liu, Ke Wang, Jinliang Zheng, Weikang Shi, Manyuan Zhang, Qi Dou, Rui Liu, Aojun Zhou, Hongsheng Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jihao Liu, Ke Wang 0036, Jinliang Zheng, Weikang Shi, Manyuan Zhang, Qi Dou 0001, Rui Liu 0019, Aojun Zhou, Hongsheng Li 0001 |
EMNLP | 2 |
| 2025 | MM-instruct: Generated visual instructions for large multimodal model alignmentabstractThis paper presents MM-Instruct, an automated pipeline for generating diverse and high-quality visual instruction data to better align large multimodal models (LMMs) with real-world use cases. While previous works have focused on question-answering data, their generated instruction datasets pose challenges for broader application scenarios. Additionally, manually collecting diverse instruction data at scale from users is prohibitively costly. To mitigate these issues, MM-Instruct leverages ChatGPT to automatically generate diverse instructions from a limited set of seed instructions through augmentation and summarization. It then uses an open-sourced large language model (LLM) to construct instruction-following answers and builds a large-scale visual instruction dataset. Evaluating the LLaVA-Instruct models trained with the generated data shows significant improvements in instruction-following capabilities compared to LLaVA-1.5 models. MM-Instruct releases its synthetic dataset containing diverse instructions and high-quality instruction-answer pairs to support training LMMs for real-world applications. Xin Huang 0027, Jihao Liu, Jinliang Zheng, Boxiao Liu, Jia Wang 0025, Yu Liu 0015, Hongsheng Li 0001, Osamu Yoshie |
Neurocomputing | 2 |
| 2024 | EasyDrag: Efficient Point-Based Manipulation on Diffusion ModelsabstractGenerative models are gaining increasing popularity, and the demand for precisely generating images is on the rise. However, generating an image that perfectly aligns with users' expectations is extremely challenging. The shapes of objects, the poses of animals, the structures of landscapes, and more may not match the user's desires, and this applies to real images as well. This is where point-based image editing becomes essential. An excellent image editing method needs to meet the following criteria: user-friendly interaction, high performance, and good generalization capability. Due to the limitations of StyleGAN, DragGAN exhibits limited robustness across diverse scenarios, while DragDiffusion lacks user-friendliness due to the necessity of LoRA fine-tuning and masks. In this paper, we introduce a novel interactive point-based image editing framework, called EasyDrag, that leverages pretrained diffusion models to achieve high-quality editing outcomes and user-friendship. Extensive experimentation demonstrates that our approach surpasses DragDiffusion in terms of both image quality and editing precision for point-based image manipulation tasks. The code will be available on https://github.com/Ace-Pegasus/EasyDrag. Xingzhong Hou, Boxiao Liu, Yi Zhang 0108, Jihao Liu, Yu Liu 0015, Haihang You |
CVPR | 4 |
| 2024 | GLID: Pre-training a Generalist Encoder-Decoder Vision ModelabstractThis paper proposes a GeneraLIst encoder-Decoder (GLID) pre-training method for better handling various downstream computer vision tasks. While self-supervised pre-training approaches, e.g., Masked Autoencoder, have shown success in transfer learning, task-specific sub-architectures are still required to be appended for differ-ent downstream tasks, which cannot enjoy the benefits of large-scale pre-training. GLID overcomes this challenge by allowing the pre-trained generalist encoder-decoder to be fine-tuned on various vision tasks with minimal task-specific architecture modifications. In the GLID training scheme, pre-training pretext task and other downstream tasks are modeled as “query-to-answer” problems, including the pre-training pretext task and other downstream tasks. We pre-train a task-agnostic encoder-decoder with query-mask pairs. During fine-tuning, GLID maintains the pre-trained encoder-decoder and queries, only replacing the topmost linear transformation layer with task-specific linear heads. This minimizes the pretrain-finetune architecture inconsis-tency and enables the pre-trained model to better adapt to downstream tasks. GLID achieves competitive performance on various vision tasks, including object detection, image segmentation, pose estimation, and depth estimation, outper-forming or matching specialist models such as Mask2Former, DETR, ViTPose, and BinsFormer. Jihao Liu, Jinliang Zheng, Yu Liu 0015, Hongsheng Li 0001 |
CVPR | 1 |
| 2024 | DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningabstractMultimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: $1)$ extracting both local and global task progressions; $2)$ enforcing temporal consistency of visual representation; $3)$ capturing trajectory-level language grounding. Most existing methods approach these via separate objectives, which often reach sub-optimal solutions. In this paper, we propose a universal unified objective that can simultaneously extract meaningful task progression information from image sequences and seamlessly align them with language instructions. We discover that via implicit preferences, where a visual trajectory inherently aligns better with its corresponding language instruction than mismatched pairs, the popular Bradley-Terry model can transform into representation learning through proper reward reparameterizations. The resulted framework, DecisionNCE, mirrors an InfoNCE-style objective but is distinctively tailored for decision-making tasks, providing an embodied representation learning framework that elegantly extracts both local and global task progression features, with temporal consistency enforced through implicit time contrastive learning, while ensuring trajectory-level instruction grounding via multimodal joint encoding. Evaluation on both simulated and real robots demonstrates that DecisionNCE effectively facilitates diverse downstream policy learning tasks, offering a versatile solution for unified representation and reward learning. Project Page: https://2toinf.github.io/DecisionNCE/ Jinliang Zheng, Yinan Zheng, Liyuan Mao, Sijie Cheng, Jihao Liu, Yu Liu 0015, Ya-Qin Zhang, Xianyuan Zhan |
ICML | 8 |
| 2024 | Instruction-Guided Visual MaskingabstractInstruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targeted local region of an image. To achieve more accurate and nuanced multimodal instruction following, we introduce Instruction-guided Visual Masking (IVM), a new versatile visual grounding model that is compatible with diverse multimodal models, such as LMM and robot model. By constructing visual masks for instruction-irrelevant regions, IVM-enhanced multimodal models can effectively focus on task-relevant image regions to better align with complex instructions. Specifically, we design a visual masking data generation pipeline and create an IVM-Mix-1M dataset with 1 million image-instruction pairs. We further introduce a new learning technique, Discriminator Weighted Supervised Learning (DWSL) for preferential IVM training that prioritizes high-quality data samples. Experimental results on generic multimodal tasks such as VQA and embodied robotic control demonstrate the versatility of IVM, which as a plug-and-play tool, significantly boosts the performance of diverse multimodal models, yielding new state-of-the-art results across challenging multimodal benchmarks. Code, model and data are available at https://github.com/2toinf/IVM. Jinliang Zheng, Sijie Cheng, Yinan Zheng, Jihao Liu, Yu Liu 0015, Xianyuan Zhan |
NeurIPS | 6 |
| 2023 | MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision TransformersabstractIn this paper, we propose Mixed and Masked AutoEncoder (MixMAE), a simple but efficient pretraining method that is applicable to various hierarchical Vision Transformers. Existing masked image modeling (MIM) methods for hierarchical Vision Transformers replace a random subset of input tokens with a special [MASK] symbol and aim at reconstructing original image tokens from the corrupted image. However, we find that using the [MASK] symbol greatly slows down the training and causes pretraining-finetuning inconsistency, due to the large masking ratio (e.g., 60% in SimMIM). On the other hand, MAE does not introduce [MASK] tokens at its encoder at all but is not applicable for hierarchical Vision Transformers. To solve the issue and accelerate the pretraining of hierarchical models, we replace the masked tokens of one image with visible tokens of another image, i.e., creating a mixed image. We then conduct dual reconstruction to reconstruct the two original images from the mixed input, which significantly improves efficiency. While MixMAE can be applied to various hierarchical Transformers, this paper explores using Swin Transformer with a large window size and scales up to huge model size (to reach 600M parameters). Empirical results demonstrate that MixMAE can learn high-quality visual representations efficiently. Notably, MixMAE with Swin-B/W14 achieves 85.1% top-1 accuracy on ImageNet-1K by pretraining for 600 epochs. Besides, its transfer performances on the other 6 datasets show that MixMAE has better FLOPs / performance tradeoff than previous popular MIM methods. Jihao Liu, Xin Huang 0027, Jinliang Zheng, Yu Liu 0015, Hongsheng Li 0001 |
CVPR | 1 |
| 2023 | GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D UnderstandingabstractMulti-view camera-based 3D detection is a challenging problem in computer vision. Recent works leverage a pretrained LiDAR detection model to transfer knowledge to a camera-based student network. However, we argue that there is a major domain gap between the LiDAR BEV features and the camera-based BEV features, as they have different characteristics and are derived from different sources. In this paper, we propose Geometry Enhanced Masked Image Modeling (GeoMIM) to transfer the knowledge of the LiDAR model in a pretrain-finetune paradigm for improving the multi-view camera-based 3D detection. GeoMIM is a multi-camera vision transformer with Cross-View Attention (CVA) blocks that uses LiDAR BEV features encoded by the pretrained BEV model as learning targets. During pretraining, GeoMIM’s decoder has a semantic branch completing dense perspective-view features and the other geometry branch reconstructing dense perspective-view depth maps. The depth branch is designed to be camera-aware by inputting the camera’s parameters for better transfer capability. Extensive results demonstrate that GeoMIM outperforms existing methods on nuScenes benchmark, achieving state-of-the-art performance for camera-based 3D object detection and 3D segmentation. Jihao Liu, Boxiao Liu, Qihang Zhang, Yu Liu 0015, Hongsheng Li 0001 |
ICCV | 1 |
| 2022 | UniNet: Unified Architecture Search with Convolution, Transformer, and MLP
Jihao Liu, Xin Huang 0027, Guanglu Song, Hongsheng Li 0001, Yu Liu 0015 |
ECCV (21) | 1 |
| 2022 | TokenMix: Rethinking Image Mixing for Data Augmentation in Vision Transformers
Jihao Liu, Boxiao Liu, Hang Zhou 0009, Hongsheng Li 0001, Yu Liu 0015 |
ECCV (26) | 1 |
| 2020 | Rotate-and-Render: Unsupervised Photorealistic Face Rotation From Single-View ImagesabstractThough face rotation has achieved rapid progress in recent years, the lack of high-quality paired training data remains a great hurdle for existing methods. The current generative models heavily rely on datasets with multi-view images of the same person. Thus, their generated results are restricted by the scale and domain of the data source. To overcome these challenges, we propose a novel unsupervised framework that can synthesize photo-realistic rotated faces using only single-view image collections in the wild. Our key insight is that rotating faces in the 3D space back and forth, and re-rendering them to the 2D plane can serve as a strong self-supervision. We leverage the recent advances in 3D face modeling and high-resolution GAN to constitute our building blocks. Since the 3D rotation-and-render on faces can be applied to arbitrary angles without losing details, our approach is extremely suitable for in-the-wild scenarios (i.e. no paired data are available), where existing methods fall short. Extensive experiments demonstrate that our approach has superior synthesis quality as well as identity preservation over the state-of-the-art methods, across a wide range of poses and domains. Furthermore, we validate that our rotate-and-render framework naturally can act as an effective data augmentation engine for boosting modern face recognition systems even on strong baseline models. Hang Zhou 0009, Jihao Liu, Ziwei Liu 0002, Yu Liu 0015, Xiaogang Wang 0001 |
CVPR | 2 |
| 2020 | Learning Where to Focus for Efficient Video Object Detection
Zhengkai Jiang 0001, Yu Liu 0015, Ceyuan Yang, Jihao Liu, Peng Gao 0007, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
ECCV (16) | 4 |
| 2019 | Differentiable Kernel EvolutionabstractThis paper proposes a differentiable kernel evolution (DKE) algorithm to find a better layer-operator for the convolutional neural network. Unlike most of the other neural architecture searching (NAS) technologies, we consider the searching space in a fundamental scope: kernel space, which encodes the assembly of basic multiplyaccumulate (MAC) operations into a conv-kernel. We first deduce a strict form of the generalized convolutional operator by some necessary constraints and construct a continuous searching space for its extra freedom-of-degree, namely, the connection of each MAC. Then a novel unsupervised greedy evolution algorithm called gradient agreement guided searching (GAGS) is proposed to learn the optimal location for each MAC in the spatially continuous searching space. We leverage DKE on multiple kinds of tasks such as object classification, face/object detection, large-scale finegrained and recognition, with various kinds of backbone architecture. Not to mention the consistent performance gain, we found the proposed DKE can further act as an autodilated operator, which makes it easy to boost the performance of miniaturized neural networks in multiple tasks. Yu Liu 0015, Jihao Liu, Xiaogang Wang 0001, Ailing Zeng |
ICCV | 2 |