Jian Wang 0108

dblp:39/449-108 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0003-4144-1753ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks
abstract
The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrity of clinical decision-making. To address this challenge, we introduce MHB, a novel and comprehensive benchmark framework designed to evaluate LLM reliability in two complex, high-stakes clinical contexts: multi-turn medical dialogues and clinical case report analysis. The core of our contribution is a systematic methodology for generating adversarial test cases by injecting ``hallucination traps" into realistic medical data, guided by a fine-grained taxonomy of clinical errors. MHB, comprising 4,695 samples and 20,288 evaluation rubrics, underwent a rigorous, two-stage validation by a panel of 60 licensed physicians from top-tier hospitals, ensuring high clinical realism and consistency. This comprehensive assessment of leading LLMs revealed significant, clinically relevant shortcomings across the board. Even the best-performing model, Claude-4-Sonnet, exhibited a hallucination rate of 29.1%, with some open-source models exceeding 57.0%. All models struggled with specific traps, like fabricated medical data or non-existent guidelines, highlighting prevalent systemic weaknesses.
Jianrong Lu, Xingyun Zheng, Jian Wang 0108, Yechao Zhang
AAAI5
2026 PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge, where the LLM's ability to generate responses based on the combination of a given query and retrieved documents is crucial. However, most benchmarks focus on overall RAG system performance, rarely assessing LLM-specific capabilities. Current benchmarks emphasize broad aspects such as noise robustness, but lack a systematic and granular evaluation framework on document utilization. To this end, we introduce Placeholder-RAG-Benchmark, a multi-level fine-grained benchmark, emphasizing the following progressive dimensions: (1) multi-level filtering abilities, (2) combination abilities, and (3) reference reasoning. To provide a more nuanced understanding of LLMs' roles in RAG systems, we formulate an innovative placeholder-based approach to decouple the contributions of the LLM's parametric knowledge and the external knowledge. Experiments demonstrate the limitations of representative LLMs in the RAG system's generation capabilities, particularly in error resilience and context faithfulness. Our benchmark provides a reproducible framework for developing more reliable and efficient RAG systems.
Zhehao Tan, Yihan Jiao, Dan Yang 0004, Duolin Sun, Jian Wang 0108, Jinjie Gu
AAAI9
2026 PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
abstract
Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional comparisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks.
Jiangwei Lao, Qi Zhu 0010, Congyun Jin, Shinan Liu, Zhihong Lu 0002, Lihe Zhang, Jian Wang 0108
AAAI11
2026 Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
abstract
Lei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen, Zhixuan Chu, Jian Wang, Jinjie Gu, Kui Ren. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhixuan Chu, Jian Wang 0108, Jinjie Gu, Kui Ren 0001
ACL (1)6
2026 WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning
abstract
Junjie Wang, Zequn Xie, Dan Yang, Jie Feng, Yue Shen, Duolin Sun, Meixiu Long, Yihan Jiao, Zhehao Tan, Jian Wang, Peng Wei, Jinjie Gu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zequn Xie, Dan Yang 0004, Duolin Sun, Meixiu Long, Yihan Jiao, Zhehao Tan, Jian Wang 0108, Jinjie Gu
ACL (1)10
2026 MedNQS: Medical State-Aware Dual-Stage Next-Turn Question Suggestion for Online Medical Consultations
Dongsheng Bi, Jian Wang 0108, Jinjie Gu
SIGIR4
2026 TOOL-CURE: Tool Selection via Curriculum-Enhanced Reinforcement Learning with Sample Screening for LLMs
abstract
Large language models (LLMs) are increasingly deployed as intelligent agents capable of executing complex real-world tasks through external tool interactions, but effective tool selection remains challenging due to the inherent limitations of real-world training data. These datasets suffer from severe tool imbalance following long-tail distributions, data scarcity for specialized tools, logic conflicts between user queries and available tools, rapidly evolving toolsets, and the presence of subpar samples including partially correct and dirty examples. Existing supervised fine-tuning (SFT) approaches struggle with these multifaceted challenges as they require abundant high-quality data, treat all labeled examples as ground truth regardless of quality, and lack the flexibility to generalize beyond specific query-tool pairings seen during training. While reinforcement learning (RL) offers a promising alternative through outcome-based learning, vanilla approaches like Group Relative Policy Optimization (GRPO) suffer from training instability due to conflicting reward signals and inefficient learning from weak signals. To address these issues, we propose TOOL-CURE, a novel method with two key improvements to GRPO: Proficiency-Scaled Curriculum Learning (PSCL), which organizes training into a two-stage curriculum that builds foundational skills on easier samples before progressing to harder ones, and Online Policy Guarding via Sample Screening (OPGSS), which continuously assesses rollout quality and masks dirty samples to prevent noisy gradients from destabilizing policy updates. Our approach enables stable and efficient learning from heterogeneous real-world data, resulting in a robust tool-selection agent that demonstrates significant improvements in accuracy and generalization capability. Our code is available at: https://github.com/einnullnull/TOOL-CURE.git.
Jie Zhang 0115, Dongsheng Bi, Tao Sun 0018, Jian Wang 0108, Yiwei Wang 0002
WSDM5
2025 SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
abstract
Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O.
Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004
CVPR8
2025 ADMIRE: ADaptive method to enhance Multiple Image REsolutions in text-rich multi-image understanding
Qipeng Zhu, Zhihong Lu 0002, Jiangwei Lao, Congyun Jin, Yingzhe Peng, Qi Zhu 0010, Lianzhen Zhong, Jiajia Liu 0002, Jian Wang 0108
KDD (2)12
2024 Towards Better Vision-Inspired Vision-Language Models
abstract
Vision-language (VL) models have achieved unprece-dented success recently, in which the connection module is the key to bridge the modality gap. Nevertheless, the abun-dant visual clues are not sufficiently exploited in most existing methods. On the vision side, most existing approaches only use the last feature of the vision tower, without using the low-level features. On the language side, most existing meth-ods only introduce shallow vision-language interactions. In this paper, we present a vision-inspired vision-language con-nection module, dubbed as VIVL, which efficiently exploits the vision cue for VL models. To take advantage of the lower-level information from the vision tower, a feature pyramid extractor (FPE) is introduced to combine features from differ-ent intermediate layers, which enriches the visual cue with negligible parameters and computation overhead. To en-hance VL interactions, we propose deep vision-conditioned prompts (DVCP) that allows deep interactions of vision and language features efficiently. Our VIVL exceeds the previous state-of-the-art method by 18.1 CIDEr when training from scratch on the COCO caption task, which greatly improves the data efficiency. When used as a plug-in module, VIVL consistently improves the performance for various backbones and VL frameworks, delivering new state-of-the-art results on multiple benchmarks, e.g., NoCaps and VQAv2.
Yun-Hao Cao, Kaixiang Ji, Chuanyang Zheng, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen, Ming Yang 0007
CVPR6
2024 SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
abstract
Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications.
Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001
CVPR12
2024 EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yongjun Zhang 0002, Jian Wang 0108, Liheng Zhong, Jingdong Chen, Ming Yang 0007
ECCV (68)5
2024 POA: Pre-training Once for Models of All Sizes
Xin Guo 0010, Jiangwei Lao, Lei Yu 0005, Lixiang Ru, Jian Wang 0108, Guo Ye, Huimei He, Jingdong Chen, Ming Yang 0007
ECCV (3)6
2024 Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual Recognition
abstract
Long-tailed recognition (LTR) aims to learn balanced models from extremely unbalanced training data. Fine-tuning pretrained foundation models has recently emerged as a promising research direction for LTR. However, we observe that the fine-tuning process tends to degrade the intrinsic representation capability of pretrained models and lead to model bias towards certain classes, thereby hindering the overall recognition performance. To unleash the intrinsic representation capability of pretrained foundation models, in this work, we propose a new Parameter-Efficient Complementary Expert Learning (PECEL) for LTR. Specifically, PECEL consists of multiple experts, where individual experts are trained via Parameter-Efficient Fine-Tuning (PEFT) and encouraged to learn different expertise on complementary sub-categories via the proposed sample-aware logit adjustment loss. By aggregating the predictions of different experts, PECEL effectively achieves a balanced performance on long-tailed classes. Nevertheless, learning multiple experts generally introduces extra trainable parameters. To ensure parameter efficiency, we further propose a parameter sharing strategy which decomposes and shares the parameters in each expert. Extensive experiments on 4 LTR benchmarks show that the proposed PECEL can effectively learn multiple complementary experts without increasing the trainable parameters and achieve new state-of-the-art performance.
Lixiang Ru, Xin Guo 0010, Lei Yu 0005, Jiangwei Lao, Jian Wang 0108, Jingdong Chen, Yansheng Li 0001, Ming Yang 0007
ACM Multimedia6
2024 Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
abstract
This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This architecture not only leverages global and local visual contexts effectively, but also facilitates the flexible extension of visual tokens through a compound token scaling strategy, allowing up to a 16x increase in the token count post pre-training. Consequently, Chain-of-Sight requires significantly fewer visual tokens in the pre-training phase compared to the fine-tuning phase. This intentional reduction of visual tokens during pre-training notably accelerates the pre-training process, cutting down the wall-clock training time by $\sim$73\%. Empirical results on a series of vision-language benchmarks reveal that the pre-train acceleration through Chain-of-Sight is achieved without sacrificing performance, matching or surpassing the standard pipeline of utilizing all visual tokens throughout the entire training process. Further scaling up the number of visual tokens for pre-training leads to stronger performances, competitive to existing approaches in a series of benchmarks.
Kaixiang Ji, Biao Gong, Zhiwu Qing, Kecheng Zheng, Jian Wang 0108, Jingdong Chen, Ming Yang 0007
NeurIPS7
2024 Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer
abstract
Abstract Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performance of self-attention mechanism in the language field, transformers tailored for visual data have drawn significant attention and triumphed over CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train and fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier works in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to a vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights to help other researchers and practitioners, and inspire more interesting research in other fields, such as remote sensing, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on the MS COCO dataset demonstrate that vision transformer based detectors trained from scratch can also achieve similar performance to their counterparts with ImageNet pre-training.
Weixiang Hong 0001, Wang Ren, Jiangwei Lao, Lele Xie, Liheng Zhong, Jian Wang 0108, Jingdong Chen, Honghai Liu 0001
Int. J. Comput. Vis.6
2023 Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic Segmentation
abstract
In order to tackle video semantic segmentation task at a lower cost, e.g., only one frame annotated per video, lots of efforts have been devoted to investigate the utilization of those unlabeled frames by either assigning pseudo labels or performing feature enhancement. In this work, we propose a novel feature enhancement network to simultaneously model short- and long-term temporal correlation. Compared with existing work that only leverage short-term correspondence, the long-term temporal correlation obtained from distant frames can effectively expand the temporal perception field and provide richer contextual prior. More importantly, modeling adjacent and distant frames together can alleviate the risk of over-fitting, hence produce high-quality feature representation for the distant unlabeled frames in training set and unseen videos in testing set. To this end, we term our method SSLTM, short for Simultaneously Short- and Long-Term Temporal Modeling. In the setting of only one frame annotated per video, SSLTM significantly outperforms the state-of-the-art methods by 2% ∼ 3% mIoU on the challenging VSPW dataset. Furthermore, when working with a pseudo label based method such as MeanTeacher, our final model only exhibits 0.13% mIoU less than the ceiling performance (i.e., all frames are manually annotated).
Jiangwei Lao, Weixiang Hong 0001, Xin Guo 0010, Jian Wang 0108, Jingdong Chen
CVPR5
2023 Uncertainty-guided Learning for Improving Image Manipulation Detection
abstract
Image manipulation detection (IMD) is of vital importance as faking images and spreading misinformation can be malicious and harm our daily life. IMD is the core technique to solve these issues and poses challenges in two main aspects: (1) Data Uncertainty, i.e., the manipulated artifacts are often hard for humans to discern and lead to noisy labels, which may disturb model training; (2) Model Uncertainty, i.e., the same object may hold different categories (tampered or not) due to manipulation operations, which could potentially confuse the model training and result in unreliable outcomes. Previous works mainly focus on solving the model uncertainty issue by designing meticulous features and networks, however, the data uncertainty problem is rarely considered. In this paper, we address both problems by introducing an uncertainty-guided learning framework, which measures data and model uncertainties by a novel Uncertainty Estimation Network (UEN). UEN is trained under dynamic supervision, and outputs estimated uncertainty maps to refine manipulation detection results, which significantly alleviates the learning difficulties. To our knowledge, this is the first work to embed uncertainty modeling into IMD. Extensive experiments on various datasets demonstrate state-of-the-art performance, validating the effectiveness and generalizability of our method.
Kaixiang Ji, Feng Chen 0047, Xin Guo 0010, Yadong Xu, Jian Wang 0108, Jingdong Chen
ICCV5
2023 Wall-to-Wall Above-Ground Biomass Estimation with Alos-2 Palsar-2 L-Band SAR Data and GEDI
abstract
Under the impact of climate change, monitoring forest carbon stock becomes an important task to evaluate the changes in carbon sequestrated from the atmosphere. Forest carbon stock estimation is still a challenging task, due to limited data sources that have a high correlation with above-ground biomass. With the help of the NASA Global Ecosystem Dynamics Investigation (GEDI) mission, above-ground biomass (AGB) can be measured by using the LiDAR data provided. However, GEDI data is sparse since it only samples about 4% of the Earth’s land surface between 51.6° N&S. Previous studies demonstrated L-Band SAR’s promising ability in retrieving forest stem volumes and estimating above-ground biomass. In this work, we propose a Deep Learning based workflow which utilizes PALSAR-2 L-Band images and GEDI to generate wall-to-wall above-ground biomass maps of North America. The workflow uses Convolutional Neural Network as the DL model and leverages both PALSAR-2 L-Band images and GEDI Relative Heights data to estimate the dense above-ground biomass maps. The results show that, by fusing GEDI Level 2 Relative Heights data with PALSAR-2 L-Band SAR data, it is possible to achieve a significantly high correlation with GEDI level 4 AGB data, as the final R-squared score of our model is as high as 0.83.
Xin Guo 0010, Liheng Zhong, Jian Wang 0108, Jingdong Chen
IGARSS4
2023 Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NER
abstract
The challenge posed by multimodal named entity recognition (MNER) is mainly two-fold: (1) bridging the semantic gap between text and image and (2) matching the entity with its associated object in image. Existing methods fail to capture the implicit entity-object relations, due to the lack of corresponding annotation. In this paper, we propose a bidirectional generative alignment method named BGA-MNER to tackle these issues. Our BGA-MNER consists of image2text and text2image generation with respect to entity-salient content in two modalities. It jointly optimizes the bidirectional reconstruction objectives, leading to aligning the implicit entity-object relations under such direct and powerful constraints. Furthermore, image-text pairs usually contain unmatched components which are noisy for generation. A stage-refined context sampler is proposed to extract the matched cross-modal content for generation. Extensive experiments on two benchmarks demonstrate that our method achieves state-of-the-art performance without image input during inference.
Feng Chen 0047, Jiajia Liu 0002, Kaixiang Ji, Wang Ren, Jian Wang 0108, Jingdong Chen
ACM Multimedia5
2022 Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer
abstract
Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performances of self-attention mech-anism in the language field, transformers tailored for visual data have drawn numerous attention and triumphed CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train & fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier work in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights can help other re-searchers and practitioners, and inspire more interesting research in other fields, such as semantic segmentation, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on MS COCO datasets demonstrate that vision transformer based detectors trained from scratch can also achieve similar performances to their counterparts with ImageNet pre-training.
Weixiang Hong 0001, Jiangwei Lao, Wang Ren, Jian Wang 0108, Jingdong Chen
CVPR4
2022 Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001
ECCV (27)5
2022 CRET: Cross-Modal Retrieval Transformer for Efficient Text-Video Retrieval
abstract
Given a text query, the text-to-video retrieval task aims to find the relevant videos in the database. Recently, model-based (MDB) methods have demonstrated superior accuracy than embedding-based (EDB) methods due to their excellent capacity of modeling local video/text correspondences, especially when equipped with large-scale pre-training schemes like ClipBERT. Generally speaking, MDB methods take a text-video pair as input and harness deep models to predict the mutual similarity, while EDB methods first utilize modality-specific encoders to extract embeddings for text and video, then evaluate the distance based on the extracted embeddings. Notably, MDB methods cannot produce explicit representations for text and video, instead, they have to exhaustively pair the query with every database item to predict their mutual similarities in the inference stage, which results in significant inefficiency in practical applications.
Kaixiang Ji, Jiajia Liu 0002, Weixiang Hong 0001, Liheng Zhong, Jian Wang 0108, Jingdong Chen
SIGIR5
2021 GilBERT: Generative Vision-Language Pre-Training for Image-Text Retrieval
abstract
Given a text/image query, image-text retrieval aims to find the relevant items in the database. Recently, visual-linguistic pre-training (VLP) methods have demonstrated promising accuracy on image-text retrieval and other visual-linguistic tasks. These VLP methods are typically pre-trained on a large amount of image-text pairs, then fine-tuned on various downstream tasks. Nevertheless, due to the natural modality incompleteness in image-text retrieval, i.e., the query is either image or text rather than an image-text pair, the naive application of VLP to image-text retrieval results in significant inefficiency. Moreover, existing VLP methods cannot extract comparable representations for a single-modal query and multi-modal database items. In this work, we propose a generative visual-linguistic pre-training approach, termed as GilBERT, to simultaneously learn generic representations of image-text data and complete the missing modality for incomplete pairs. In testing phase, the proposed GilBERT facilitates efficient vector-based retrieval by providing unified feature embedding for query and database items. Moreover, the generative training not only makes GilBERT compatible with non-parallel text/image corpus, but also enables GilBERT to model the image-text relationships without suffering massive randomly-sampled negative samples, leading to superior experimental performances. Extensive experiments demonstrate the advantages of GilBERT in image-text retrieval, in terms of both efficiency and accuracy.
Weixiang Hong 0001, Kaixiang Ji, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen
SIGIR4
2020 Automatic Car Damage Assessment System: Reading and Understanding Videos as Professional Insurance Inspectors
abstract
We demonstrate a car damage assessment system in car insurance field based on artificial intelligence techniques, which can exempt insurance inspectors from checking cars on site and help people without professional knowledge to evaluate car damages when accidents happen. Unlike existing approaches, we utilize videos instead of photos to interact with users to make the whole procedure as simple as possible. We adopt object and video detection and segmentation techniques in computer vision, and take advantage of multiple frames extracted from videos to achieve high damage recognition accuracy. The system uploads video streams captured by mobile devices, recognizes car damage on the cloud asynchronously and then returns damaged components and repair costs to users. The system evaluates car damages and returns results automatically and effectively in seconds, which reduces laboratory costs and decreases insurance claim time significantly.
Xin Guo 0010, Qingpei Guo, Jian Wang 0108, Qing Wang 0068, Chen Jiang 0006, Furong Xu
AAAI5