Chengfei Cai

dblp:239/3394 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NAHA: Towards Efficient Adaptation of Foundation Models via Hierarchical Adaptive Nyström Attention in Computational Pathology
Xiake Zhang, Jun Xu 0005, Jun Li 0011, Mingxia Liu 0001, Yiping Jiao, Chengfei Cai
ICIC (20)6
2026 IPMMG: Information propagation with multi-granularity morphology-guided for nuclear segmentation and classification
Dawei Fan, Jun Li 0004, Chengfei Cai, Lihui Lin, Riqing Chen, Lifang Wei
Expert Syst. Appl.3
2026 Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
Yue Ma 0016, Zexuan Yan, Hongfa Wang, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Harry Shum, Zhifeng Li 0001, Wei Liu 0005, Qifeng Chen 0001
Int. J. Comput. Vis.9
2026 Gestalt-Inspired Feature Integration Network With Entropy Uncertainty Modeling for Pathology Image Segmentation
abstract
The accuracy and stability of pathology image segmentation have become critical factors in clinical applications such as cancer screening and tumor grading. However, the presence of complex local structures, uncertain regions, and subtle morphological variations in pathological images continues to pose significant challenges. Most existing feature fusion approaches rely on the simplistic aggregation of extracted features, neglecting the unique characteristics and relative importance of distinct feature representations, which ultimately limits their potential to enhance model performance. To address these issues, we propose a Gestalt-Inspired Feature Integration Network (GeNet), a novel architecture inspired by Gestalt theory that mirrors the human visual system's ability to derive holistic understanding from partial information. Embracing the principle that 'the whole is greater than the sum of its parts,' GeNet introduces a mechanism to synergistically leverage multi-scale information, which assesses the similarity between features to achieve a more meaningful fusion of global context and local detail. Given the variability in target appearance within pathological images, we use information entropy to quantify feature uncertainty, allowing the model to prioritize uncertain regions and reduce the occurrence of ambiguous results. To explicitly eliminate multi-feature redundancy and misalignment, the refinement block utilizes parallel convolutional recalibration to fully leverage the advantages of various features. Extensive experiments on multiple pathological image segmentation datasets, including GlaS, GCaSeg, and EBHI-Seg, demonstrate that GeNet achieves high accuracy and strong robustness, offering a new perspective for joint modeling of global and local features in medical image analysis.
Dawei Fan, Jiamei Wen, Mingyue Han, Jun Li 0004, Chengfei Cai, Changcai Yang, Riqing Chen, Lifang Wei
IEEE J. Biomed. Health Informatics6
2025 Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts
abstract
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. Additionally, current I2V methods require users not only to describe the target motion but also to provide redundant detailed descriptions of frame contents.These two issues hinder the practical utilization of current I2V tools. In this paper, we propose a practical framework, named Follow-Your-Click, to achieve image animation with a simple user click (for specifying what to move) and a motion prompt (for specifying how to move). Technically, we propose the first-frame masking strategy, which significantly improves the video generation quality, and a motion-augmented module equipped with a motion prompt dataset to improve the motion prompt following abilities of our model. To further control the motion speed, we propose flow-based motion magnitude control to control the speed of target movement more precisely. Extensive experiments compared with 7 baselines, including both commercial tools and research methods on 8 metrics, suggest the superiority of our approach.
Yue Ma 0016, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Jixuan Ying, Chengfei Cai, Zhifeng Li 0001, Harry Shum, Wei Liu 0005, Qifeng Chen 0001
AAAI8
2025 A Hierarchical Geometry-Guided Transformer for Histological Subtyping of Primary Liver Cancer
abstract
Primary liver malignancies are widely recognized as the most heterogeneous and prognostically diverse cancers of the digestive system. Among these, hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma (ICC) emerge as the two principal histological subtypes, demonstrating significantly greater complexity in tissue morphology and cellular architecture than other common tumors. The intricate representation of features in Whole Slide Images (WSIs) encompasses abundant crucial information for liver cancer histological subtyping, regarding hierarchical pyramid structure, tumor microenvironment (TME), and geometric representation. However, recent approaches have not adequately exploited these indispensable effective descriptors, resulting in a limited understanding of histological representation and suboptimal subtyping performance. To mitigate these limitations, A hieRarchical Geometry-gUided tranSformer (ARGUS) is proposed to advance histological subtyping in liver cancer by capturing the macro-meso-micro hierarchical information within the TME. Extensive experiments on public and private cohorts demonstrate that our ARGUS achieves state-of-the-art (SOTA) performance in histological subtyping of liver cancer, which provide an effective diagnostic tool for primary liver malignancies in clinical practice. Related code will be available to public.
Anwen Lu, Yiping Jiao, Geyang Xu, Hongyi Gong, Chengfei Cai, Jun Chen 0005, Jun Xu 0005
BIBM6
2025 HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
abstract
We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at HunyuanPortrait.
Zunnan Xu, Zhentao Yu, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Chengfei Cai, Shiyu Tang, Qin Lin 0003, Xiu Li 0001, Qinglin Lu
CVPR9
2025 SC-AGR: Spatially-Constrained Attention for Context-Aware Graph Representation in Histopathology Whole Slide Image Analysis
Chengfei Cai, Jun Li 0011, Jun Xu 0005
ICIC (25)1
2025 MurreNet: Modeling Holistic Multimodal Interactions Between Histopathology and Genomic Profiles for Survival Prediction
Chengfei Cai, Jun Li 0011, Pengbo Xu, Jiquan Ma, Jun Xu 0005
MICCAI (15)2
2025 Predicting ustekinumab treatment response in Crohn's disease using pre-treatment biopsy images
abstract
MOTIVATION: Crohn's disease (CD) exhibits substantial variability in response to biological therapies such as ustekinumab (UST), a monoclonal antibody targeting interleukin-12/23. However, predicting individual treatment responses remains difficult due to the lack of reliable histopathological biomarkers and the morphological complexity of tissue. While recent deep learning methods have leveraged whole-slide images (WSIs), most lack effective mechanisms for selecting relevant regions and integrating patch-level evidence into robust patient-level predictions. Therefore, a framework that captures local histological cues and global tissue context is needed to improve prediction performance. RESULTS: We propose a novel clustering-enhanced weakly supervised learning framework to predict UST treatment response from pre-treatment WSIs of CD patients. First, patches from WSIs were encoded using a pre-trained vision foundation model, and k-means clustering was applied to identify representative morphological patterns. Discriminative patches associated with treatment outcomes were selected via a DenseNet-based classifier, with Grad-CAM used to enhance interpretability. To aggregate patch-level predictions, we adopted a multi-instance learning approach, from which whole-slide features were extracted using both patch likelihood histograms and bag-of-words representations. These features were subsequently used to train a classifier for final response prediction. Experimental results on an independent test set demonstrated that our WSI-level model achieved superior predictive performance with an AUC of 0.938 (95% CI: 0.879-0.996), sensitivity of 0.951, and specificity of 0.825, outperforming baseline patch-level models. These findings suggest that our method enables accurate, interpretable, and scalable prediction of biological therapy response in CD, potentially supporting personalized treatment strategies in clinical settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/caicai2526/USTAIM.
Chengfei Cai, Rui-dong Chen, Jieyu Chen, Jun Li 0011, Caiyun Lv, Yiping Jiao, Lanqing Wu, Qianyun Shi, Jun Xu 0005
Bioinform.1
2025 Prediction of molecular subtypes for endometrial cancer based on hierarchical foundation model
abstract
MOTIVATION: Endometrial cancer is a prevalent gynecological malignancy that requires accurate identification of its molecular subtypes for effective diagnosis and treatment. Four molecular subtypes with different clinical outcomes have been identified: POLE mutation, mismatch repair deficient, p53 abnormal, and no specific molecular profile. However, determining these subtypes typically relies on expensive gene sequencing. To overcome this limitation, we propose a novel method that utilizes hematoxylin and eosin-stained whole slide images to predict endometrial cancer molecular subtypes. RESULTS: Our approach leverages a hierarchical foundation model as a backbone, fine-tuned from the UNI computational pathology foundation model, to extract tissue embedding from different scales. We have achieved promising results through extensive experimentation on the Fudan University Shanghai Cancer Center cohort (N = 364). Our model demonstrates a macro-average AUROC of 0.879 (95% CI, 0.853-0.904) in a five-fold cross-validation. Compared to the current state-of-the-art molecular subtypes prediction for endometrial cancer, our method outperforms in terms of predictive accuracy and computational efficiency. Moreover, our method is highly reproducible, allowing for ease of implementation and widespread adoption. This study aims to address the cost and time constraints associated with traditional gene sequencing techniques. By providing a reliable and accessible alternative to gene sequencing, our method has the potential to revolutionize the field of endometrial cancer diagnosis and improve patient outcomes. AVAILABILITY AND IMPLEMENTATION: The codes and data used for generating results in this study are available at https://github.com/HaoyuCui/hi-UNI for GitHub and https://doi.org/10.5281/zenodo.14627478 for Zenodo.
Haoyu Cui, Qinhao Guo, Jun Xu 0005, Chengfei Cai, Yiping Jiao, Wenlong Ming, Xiangxue Wang
Bioinform.5
2025 Global and Local Semantic Completion Learning for Vision-Language Pre-Training
abstract
Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-local alignment, i.e., associations between image patches and text tokens. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations to local features of the other modality. Therefore, in this paper, we propose a novel Global and Local Semantic Completion Learning (GLSCL) task to facilitate global-local alignment and local-local alignment simultaneously. Specifically, the GLSCL task complements the missing semantics of masked data and recovers global and local features by cross-modal interactions. Our GLSCL consists of masked global semantic completion (MGSC) and masked local token completion (MLTC). MGSC promotes learning more representative global features, which have a great impact on the performance of downstream tasks, while MLTC reconstructs modal-fusion local tokens, further enhancing accurate comprehension of multimodal data. To evaluate the proposed approaches on cross-modal alignment, we develop a validation benchmark called ALIGN-BENCH. Moreover, we present a flexible vision encoder, enabling our model to simultaneously perform image-text and video-text multimodal tasks. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.
Rongcheng Tu, Yatai Ji, Jie Jiang 0015, Weijie Kong, Chengfei Cai, Hongfa Wang, Yujiu Yang 0001, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Robustness-Guided Image Synthesis for Data-Free Quantization
abstract
Quantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks.
Jianhong Bai, Huanpeng Chu, Hualiang Wang, Zuozhu Liu, Ruizhe Chen, Xiaoxuan He, Lianrui Mu, Chengfei Cai, Haoji Hu
AAAI9
2024 SeqFRT: Towards Effective Adaption of Foundation Model via Sequence Feature Reconstruction in Computational Pathology
abstract
Given the intricate situation of modelling gigapixel images, the usage of multiple instance learning (MIL) framework has recently increased to support clinical practice, encompassing cancer diagnosis, subtyping, survival prediction and other tasks. In current practice, most state-of-the-art MIL proposals typically apply a frozen pre-trained CNN or a pathological foundation model for feature extraction. While this paradigm lacks the capability for sequence feature fine-tuning within the downstream-specific tasks, which hinders the continuous performance promotion in Whole Slide Images (WSIs) Analysis. To address this issue, we propose a Sequence Feature Reconstruction Transformer (SeqFRT) for optimizing feature extraction of the foundation model, which can capture more discriminative features within pathological instance sequences. The proposed model comprises three main modules: 1) an offline foundation model as the pathological feature extractor; 2) a sequence position optimization architecture which aims at refining the correlations between instances in both sequential ordering and transpositional ordering; 3) a sequence sparsity enhancement strategy is designed to reconstruct the sequence feature and extract the latent representations instead of redundant information. Extensive experiments on six benchmark datasets for three computational pathology tasks demonstrated our model’s superiority over the state-of-the-art MIL methods. The source code is available at https://github.com/caicai2526/SeqFRT-MIL.
Chengfei Cai, Jun Li 0011, Yiping Jiao, Jun Xu 0005
BIBM1
2024 Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
abstract
We present Follow-Your-Emoji, a diffusion-based framework for portrait animation, which animates a reference portrait with target landmark sequences. The main challenge of portrait animation is to preserve the identity of the reference portrait and transfer the target expression to this portrait while maintaining temporal consistency and fidelity. To address these challenges, Follow-Your-Emoji equipped the powerful Stable Diffusion model with two well-designed technologies. Specifically, we first adopt a new explicit motion signal, namely expression-aware landmark, to guide the animation process. We discover this landmark can not only ensure the accurate motion alignment between the reference portrait and target motion during inference but also increase the ability to portray exaggerated expressions (i.e., large pupil movements) and avoid identity leakage. Then, we propose a facial fine-grained loss to improve the model’s ability of subtle expression perception and reference portrait appearance reconstruction by using both expression and facial masks. Accordingly, our method demonstrates significant performance in controlling the expression of freestyle portraits, including real humans, cartoons, sculptures, and even animals. By leveraging a simple and effective progressive generation strategy, we extend our model to stable long-term animation, thus increasing its potential application value. To address the lack of a benchmark for this field, we introduce EmojiBench, a comprehensive benchmark comprising diverse portrait images, driving videos, and landmarks. We show extensive evaluations on EmojiBench to verify the superiority of Follow-Your-Emoji. The code, training dataset and benchmark will be found in https://github.com/mayuelala/FollowYourEmoji.
Yue Ma 0016, Hongfa Wang, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Harry Shum, Wei Liu 0005, Qifeng Chen 0001
SIGGRAPH Asia8
2023 Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning
abstract
Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-to-local alignment. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations. Therefore, in this paper, we propose a novel Semantic Completion Learning (SCL) task, complementary to existing masked modeling tasks, to facilitate global-to-local alignment. Specifically, the SCL task complements the missing semantics of masked data by capturing the corresponding information from the other modality, promoting learning more representative global features which have a great impact on the performance of downstream tasks. Moreover, we present a flexible vision encoder, which enables our model to perform image-text and video-text multimodal tasks simultaneously. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.
Yatai Ji, Rongcheng Tu, Jie Jiang 0015, Weijie Kong, Chengfei Cai, Hongfa Wang, Yujiu Yang 0001, Wei Liu 0005
CVPR5
2023 Unsupervised Hashing with Semantic Concept Mining
abstract
Recently, to improve the unsupervised image retrieval performance, plenty of unsupervised hashing methods have been proposed by designing a semantic similarity matrix, which is based on the similarities between image features extracted by a pre-trained CNN model. However, most of these methods tend to ignore high-level abstract semantic concepts contained in images. Intuitively, concepts play an important role in calculating the similarity among images. In real-world scenarios, each image is associated with some concepts, and the similarity between two images will be larger if they share more identical concepts. Inspired by the above intuition, in this work, we propose a novel Unsupervised Hashing with Semantic Concept Mining, called UHSCM, which leverages a VLP model to construct a high-quality similarity matrix. Specifically, a set of randomly chosen concepts is first collected. Then, by employing a vision-language pretraining (VLP) model with the prompt engineering which has shown strong power in visual representation learning, the set of concepts is denoised according to the training images. Next, the proposed method UHSCM applies the VLP model with prompting again to mine the concept distribution of each image and construct a high-quality semantic similarity matrix based on the mined concept distributions. Finally, with the semantic similarity matrix as guiding information, a novel hashing loss with a modified contrastive loss based regularization item is proposed to optimize the hashing network. Extensive experiments on three benchmark datasets show that the proposed method outperforms the state-of-the-art baselines in the image retrieval task.
Rongcheng Tu, Xianling Mao, Qinghong Lin, Chengfei Cai, Weize Qin, Wei Wei 0002, Hongfa Wang, Heyan Huang
Proc. ACM Manag. Data4
2023 Unsupervised Cross-Modal Hashing With Modality-Interaction
abstract
Recently, numerous unsupervised cross-modal hashing methods have been proposed to deal the image-text retrieval tasks for the unlabeled cross-modal data. However, when these methods learn to generate hash codes, almost all of them lack modality-interaction in the following two aspects: 1) The instance similarity matrix used to guide the hashing networks training is constructed without image-text interaction, which fails to capture the fine-grained cross-modal cues to elaborately characterize the intrinsic semantic similarity among the datapoints. 2) The binary codes used for quantization loss are inferior because they are generated by directly quantizing a simple combination of continuous hash codes from different modalities without the interaction among these continuous hash codes. Such problems will cause the generated hash codes to be of poor quality and degrade the retrieval performance. Hence, in this paper, we propose a novel Unsupervised Cross-modal Hashing with Modality-interaction, termed UCHM. Specifically, by optimizing a novel hash-similarity-friendly loss, a modality-interaction-enabled (MIE) similarity generator is first trained to generate a superior MIE similarity matrix for the training set. Then, the generated MIE similarity matrix is utilized as guiding information to train the deep hashing networks. Furthermore, during the process of training the hashing networks, a novel bit-selection module is proposed to generate high-quality unified binary codes for the quantization loss with the interaction among continuous codes from different modalities, thereby further enhancing the retrieval performance. Extensive experiments on two widely used datasets show that the proposed UCHM outperforms state-of-the-art techniques on cross-modal retrieval tasks.
Rongcheng Tu, Jie Jiang 0015, Qinghong Lin, Chengfei Cai, Shangxuan Tian, Hongfa Wang, Wei Liu 0005
IEEE Trans. Circuits Syst. Video Technol.4
2023 Deep Cross-Modal Proxy Hashing
abstract
Due to the high retrieval efficiency and low storage cost for cross-modal search tasks, cross-modal hashing methods have attracted considerable attention from the researchers. For the supervised cross-modal hashing methods, how to make the learned hash codes sufficiently preserve semantic information contained in the label of datapoints is the key to further enhance the retrieval performance. Hence, almost all supervised cross-modal hashing methods usually depend on defining similarities between datapoints with the label information to guide the hashing model learning fully or partly. However, the defined similarity between datapoints can only capture the label information of datapoints partially and misses abundant semantic information, which then hinders the further improvement of retrieval performance. Thus, in this paper, different from previous works, we propose a novel cross-modal hashing method without defining the similarity between datapoints, called Deep Cross-modal Proxy Hashing (DCPH). Specifically, DCPH first trains a proxy hashing network to transform each category information of a dataset into a semantic discriminative hash code, called proxy hash code. Each proxy hash code can preserve the semantic information of its corresponding category well. Next, without defining the similarity between datapoints to supervise the training process of the modality-specific hashing networks, we propose a novelmargin-dynamic-softmax lossto directly utilize the proxy hashing codes as supervised information. Finally, by minimizing the novelmargin-dynamic-softmax loss, the modality-specific hashing networks can be trained to generate hash codes that can simultaneously preserve the cross-modal similarity and abundant semantic information well. Extensive experiments on three benchmark datasets show that the proposed method outperforms the state-of-the-art baselines in the cross-modal retrieval tasks.
Rongcheng Tu, Xianling Mao, Rongxin Tu, Bin-Bin Bian, Chengfei Cai, Hongfa Wang, Wei Wei 0002, Heyan Huang
IEEE Trans. Knowl. Data Eng.5
2022 Egocentric Video-Language Pretraining
abstract
Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.
Qinghong Lin, Jinpeng Wang 0001, Mattia Soldan, Michael Wray, Rui Yan 0001, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu 0005, Zheng Shou 0001
NeurIPS11