Yumao Lu

dblp:44/292 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 50% Language models and text generation · 24% Representation and self-supervised learning · 10%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.812024
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks · CVPR 2024
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting
0.612022
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022
Computer vision › Vision and language
image captioning
0.612022
Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022
Natural language and speech › Language models and text generation
in-context learning
0.612022
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering
0.612022
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022
Natural language and speech › Language models and text generation › language modeling
scaling behavior
0.612022
Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.612022
SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022
Computer vision › Vision and language
video captioning
0.612022
SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022
Computer vision › Vision and language
vision-language pretraining
0.612022
Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022
Computer vision › Vision and language
visual question answering
0.612022
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022
Computer vision › Vision and language
video-language understanding
0.212022
SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212022
Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022
Information retrieval
web search
0.222009
Improving Web Search Relevance with Semantic Features · EMNLP 2009
Context sensitive stemming for web search · SIGIR 2007
Information retrieval › web search › search personalization
location-based personalization
0.112010
Personalize web search results with user's location · SIGIR 2010
Information retrieval
personalized search
0.112010
Personalize web search results with user's location · SIGIR 2010
Information retrieval › query understanding
query intent
0.112010
Personalize web search results with user's location · SIGIR 2010
Information retrieval › ranking
search relevance
0.112009
Improving Web Search Relevance with Semantic Features · EMNLP 2009
Information retrieval
indexing
0.112007
Context sensitive stemming for web search · SIGIR 2007
Information retrieval › indexing
stemming
0.112007
Context sensitive stemming for web search · SIGIR 2007

Methods — techniques the papers use, named apart from their topics

sequence-to-sequence · 0.8automated image annotation · 0.8web data collection · 0.6transformer scaling · 0.6transformer · 0.6sparse attention mask · 0.6noisy data training · 0.6image captioning · 0.6end-to-end training · 0.6GPT-3 · 0.6probabilistic model · 0.1semantic features · 0.1context-sensitive stemming · 0.1
YearPublicationVenuePosition
2024 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
abstract
We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for various computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform diverse tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with un-precedented zero-shot and fine-tuning capabilities.
Bin Xiao 0004, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng 0001, Ce Liu 0001, Lu Yuan 0001
CVPR6
2022 An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
abstract
Knowledge-based visual question answering (VQA) involves answering questions that require external knowledge not present in the image. Existing methods first retrieve knowledge from external resources, then reason over the selected knowledge, the input image, and question for answer prediction. However, this two-step approach could lead to mismatches that potentially limit the VQA performance. For example, the retrieved knowledge might be noisy and irrelevant to the question, and the re-embedded knowledge features during reasoning might deviate from their original meanings in the knowledge base (KB). To address this challenge, we propose PICa, a simple yet effective method that Prompts GPT3 via the use of Image Captions, for knowledge-based VQA. Inspired by GPT-3’s power in knowledge retrieval and question answering, instead of using structured KBs as in previous work, we treat GPT-3 as an implicit and unstructured KB that can jointly acquire and process relevant knowledge. Specifically, we first convert the image into captions (or tags) that GPT-3 can understand, then adapt GPT-3 to solve the VQA task in a few-shot manner by just providing a few in-context VQA examples. We further boost performance by carefully investigating: (i) what text formats best describe the image content, and (ii) how in-context examples can be better selected and used. PICa unlocks the first use of GPT-3 for multimodal tasks. By using only 16 examples, PICa surpasses the supervised state of the art by an absolute +8.6 points on the OK-VQA dataset. We also benchmark PICa on VQAv2, where PICa also shows a decent few-shot performance.
Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Yumao Lu, Zicheng Liu 0001
AAAI5
2022 Scaling Up Vision-Language Pretraining for Image Captioning
abstract
In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON O, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art Vin VL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million imagetext pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M11The dataset is released at https://github.com/xiaoweihu/ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail vi-sual concepts when used in a zero-shot manner.
Xiaowei Hu 0006, Zhe Gan, Zhengyuan Yang, Zicheng Liu 0001, Yumao Lu
CVPR6
2022 SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning
abstract
The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to video captioning data. In this work, we present SwinBERT, an end-to-end transformer-based model for video captioning, which takes video frame patches directly as inputs, and outputs a natural language description. Instead of leveraging multiple 2D/3D feature extractors, our method adopts a video transformer to encode spatial-temporal representations that can adapt to variable lengths of video input without dedicated design for different frame rates. Based on this model architecture, we show that video captioning can benefit significantly from more densely sampled video frames as opposed to previous successes with sparsely sampled video frames for video-and-language understanding tasks (e.g., video question answering). Moreover, to avoid the inherent redundancy in consecutive video frames, we propose adaptively learning a sparse attention mask and optimizing it for task-specific performance improvement through better long-range video sequence modeling. Through extensive experiments on 5 video captioning datasets, we show that Swinbert achieves across-the-board performance improvements over previous methods, often by a large margin. The learned sparse attention masks in addition push the limit to new state of the arts, and can be transferred between different video lengths and between different datasets. Code is available at https://github.com/microsoft/SwinBERT.
Chung-Ching Lin, Faisal Ahmed 0001, Zhe Gan, Zicheng Liu 0001, Yumao Lu
CVPR7
2022 UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Faisal Ahmed 0001, Zicheng Liu 0001, Yumao Lu
ECCV (36)7
2010 Personalize web search results with user's location
abstract
We build a probabilistic model to identify implicit local intent queries, and leverage user's physical location to improve Web search results for these queries. Evaluation on commercial search engine shows significant improvement on search relevance and user experience.
Yumao Lu, Fuchun Peng, Benoît Dumoulin
SIGIR1
2009 Context sensitive synonym discovery for web search queries
abstract
We propose a simple yet effective approach to context sensitive synonym discovery for Web search queries based on co-click analysis; i.e., analyzing queries leading to clicking same documents. In addition to deriving word based synonyms, we also derive concept based synonyms with the help of query segmentation. Evaluation results show that this approach dramatically outperforms the thesaurus based synonym replacement method in keeping search intent, from accuracy of 40% to above 80%.
Fuchun Peng, Huihsin Tseng, Yumao Lu, Benoît Dumoulin
CIKM4
2009 Improving Web Search Relevance with Semantic Features
Yumao Lu, Fuchun Peng, Gilad Mishne, Benoît Dumoulin
EMNLP1
2008 Parallel randomized sampling for support vector machine (SVM) and support vector regression (SVR)
Yumao Lu, Vwani P. Roychowdhury
Knowl. Inf. Syst.1
2008 Distributed Parallel Support Vector Machines in Strongly Connected Networks
abstract
In this paper, we propose a distributed parallel support vector machine (DPSVM) training mechanism in a configurable network environment for distributed data mining. The basic idea is to exchange support vectors among a strongly connected network (SCN) so that multiple servers may work concurrently on distributed data set with limited communication cost and fast training speed. The percentage of servers that can work in parallel and the communication overhead may be adjusted through network configuration. The proposed algorithm further speeds up through online implementation and synchronization. We prove that the global optimal classifier can be achieved iteratively over an SCN. Experiments on a real-world data set show that the computing time scales well with the size of the training data for most networks. Numerical results show that a randomly generated SCN may achieve better performance than the state of the art method, cascade SVM, in terms of total training time.
Yumao Lu, Vwani P. Roychowdhury, Lieven Vandenberghe
IEEE Trans. Neural Networks1
2007 Context sensitive stemming for web search
abstract
Traditionally, stemming has been applied to Information Retrieval tasks by transforming words in documents to the their root form before indexing, and applying a similar transformation to query terms. Although it increases recall, this naive strategy does not work well for Web Search since it lowers precision and requires a significant amount of additional computation.
Fuchun Peng, Nawaaz Ahmed, Yumao Lu
SIGIR4
2006 Coupling feature selection and machine learning methods for navigational query identification
abstract
It is important yet hard to identify navigational queries in Web search due to a lack of sufficient information in Web queries, which are typically very short. In this paper we study several machine learning methods, including naive Bayes model, maximum entropy model, support vector machine (SVM), and stochastic gradient boosting tree (SGBT), for navigational query identification in Web search. To boost the performance of these machine techniques, we exploit several feature selection methods and propose coupling feature selection with classification approaches to achieve the best performance. Different from most prior work that uses a small number of features, in this paper, we study the problem of identifying navigational queries with thousands of available features, extracted from major commercial search engine results, Web search user click data, query log, and the whole Web's relational content. A multi-level feature extraction system is constructed.Our results on real search data show that 1) Among all the features we tested, user click distribution features are the most important set of features for identifying navigational queries. 2) In order to achieve good performance, machine learning approaches have to be coupled with good feature selection methods. We find that gradient boosting tree, coupled with linear SVM feature selection is most effective. 3) With carefully coupled feature selection and classification approaches, navigational queries can be accurately identified with 88.1% F1 score, which is 33% error rate reduction compared to the best uncoupled system, and 40% error rate reduction compared to a well tuned system without feature selection.
Yumao Lu, Fuchun Peng, Nawaaz Ahmed
CIKM1
2006 Parallel Randomized Support Vector Machine
Yumao Lu, Vwani P. Roychowdhury
PAKDD1
2004 Dynamic traffic controls for Web-server networks
Yumao Lu
Comput. Networks2