VLDB 2026 Research / reviewers in the wild / expert
Yumao Lu
dblp:44/292
· DBLP profile ↗
14ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Vision and language · 50% Language models and text generation · 24% Representation and self-supervised learning · 10% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model |
0.8 | 1 | 2024 | Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks · CVPR 2024 |
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting |
0.6 | 1 | 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022 |
Computer vision › Vision and language
image captioning |
0.6 | 1 | 2022 | Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022 |
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering |
0.6 | 1 | 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022 |
Natural language and speech › Language models and text generation › language modeling
scaling behavior |
0.6 | 1 | 2022 | Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.6 | 1 | 2022 | SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022 |
Computer vision › Vision and language
video captioning |
0.6 | 1 | 2022 | SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022 |
Computer vision › Vision and language
vision-language pretraining |
0.6 | 1 | 2022 | Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022 |
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA · AAAI 2022 |
Computer vision › Vision and language
video-language understanding |
0.2 | 1 | 2022 | SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · CVPR 2022 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
0.2 | 1 | 2022 | Scaling Up Vision-Language Pretraining for Image Captioning · CVPR 2022 |
Information retrieval
web search |
0.2 | 2 | 2009 | Improving Web Search Relevance with Semantic Features · EMNLP 2009 Context sensitive stemming for web search · SIGIR 2007 |
Information retrieval › web search › search personalization
location-based personalization |
0.1 | 1 | 2010 | Personalize web search results with user's location · SIGIR 2010 |
Information retrieval
personalized search |
0.1 | 1 | 2010 | Personalize web search results with user's location · SIGIR 2010 |
Information retrieval › query understanding
query intent |
0.1 | 1 | 2010 | Personalize web search results with user's location · SIGIR 2010 |
Information retrieval › ranking
search relevance |
0.1 | 1 | 2009 | Improving Web Search Relevance with Semantic Features · EMNLP 2009 |
Information retrieval
indexing |
0.1 | 1 | 2007 | Context sensitive stemming for web search · SIGIR 2007 |
Information retrieval › indexing
stemming |
0.1 | 1 | 2007 | Context sensitive stemming for web search · SIGIR 2007 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence · 0.8automated image annotation · 0.8web data collection · 0.6transformer scaling · 0.6transformer · 0.6sparse attention mask · 0.6noisy data training · 0.6image captioning · 0.6end-to-end training · 0.6GPT-3 · 0.6probabilistic model · 0.1semantic features · 0.1context-sensitive stemming · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Florence-2: Advancing a Unified Representation for a Variety of Vision TasksabstractWe introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for various computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform diverse tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with un-precedented zero-shot and fine-tuning capabilities. Bin Xiao 0004, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng 0001, Ce Liu 0001, Lu Yuan 0001 |
CVPR | 6 |
| 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAabstractKnowledge-based visual question answering (VQA) involves answering questions that require external knowledge not present in the image. Existing methods first retrieve knowledge from external resources, then reason over the selected knowledge, the input image, and question for answer prediction. However, this two-step approach could lead to mismatches that potentially limit the VQA performance. For example, the retrieved knowledge might be noisy and irrelevant to the question, and the re-embedded knowledge features during reasoning might deviate from their original meanings in the knowledge base (KB). To address this challenge, we propose PICa, a simple yet effective method that Prompts GPT3 via the use of Image Captions, for knowledge-based VQA. Inspired by GPT-3’s power in knowledge retrieval and question answering, instead of using structured KBs as in previous work, we treat GPT-3 as an implicit and unstructured KB that can jointly acquire and process relevant knowledge. Specifically, we first convert the image into captions (or tags) that GPT-3 can understand, then adapt GPT-3 to solve the VQA task in a few-shot manner by just providing a few in-context VQA examples. We further boost performance by carefully investigating: (i) what text formats best describe the image content, and (ii) how in-context examples can be better selected and used. PICa unlocks the first use of GPT-3 for multimodal tasks. By using only 16 examples, PICa surpasses the supervised state of the art by an absolute +8.6 points on the OK-VQA dataset. We also benchmark PICa on VQAv2, where PICa also shows a decent few-shot performance. Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Yumao Lu, Zicheng Liu 0001 |
AAAI | 5 |
| 2022 | Scaling Up Vision-Language Pretraining for Image CaptioningabstractIn recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON O, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art Vin VL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million imagetext pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M11The dataset is released at https://github.com/xiaoweihu/ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail vi-sual concepts when used in a zero-shot manner. Xiaowei Hu 0006, Zhe Gan, Zhengyuan Yang, Zicheng Liu 0001, Yumao Lu |
CVPR | 6 |
| 2022 | SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningabstractThe canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to video captioning data. In this work, we present SwinBERT, an end-to-end transformer-based model for video captioning, which takes video frame patches directly as inputs, and outputs a natural language description. Instead of leveraging multiple 2D/3D feature extractors, our method adopts a video transformer to encode spatial-temporal representations that can adapt to variable lengths of video input without dedicated design for different frame rates. Based on this model architecture, we show that video captioning can benefit significantly from more densely sampled video frames as opposed to previous successes with sparsely sampled video frames for video-and-language understanding tasks (e.g., video question answering). Moreover, to avoid the inherent redundancy in consecutive video frames, we propose adaptively learning a sparse attention mask and optimizing it for task-specific performance improvement through better long-range video sequence modeling. Through extensive experiments on 5 video captioning datasets, we show that Swinbert achieves across-the-board performance improvements over previous methods, often by a large margin. The learned sparse attention masks in addition push the limit to new state of the arts, and can be transferred between different video lengths and between different datasets. Code is available at https://github.com/microsoft/SwinBERT. Chung-Ching Lin, Faisal Ahmed 0001, Zhe Gan, Zicheng Liu 0001, Yumao Lu |
CVPR | 7 |
| 2022 | UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Faisal Ahmed 0001, Zicheng Liu 0001, Yumao Lu |
ECCV (36) | 7 |
| 2010 | Personalize web search results with user's locationabstractWe build a probabilistic model to identify implicit local intent queries, and leverage user's physical location to improve Web search results for these queries. Evaluation on commercial search engine shows significant improvement on search relevance and user experience. Yumao Lu, Fuchun Peng, Benoît Dumoulin |
SIGIR | 1 |
| 2009 | Context sensitive synonym discovery for web search queriesabstractWe propose a simple yet effective approach to context sensitive synonym discovery for Web search queries based on co-click analysis; i.e., analyzing queries leading to clicking same documents. In addition to deriving word based synonyms, we also derive concept based synonyms with the help of query segmentation. Evaluation results show that this approach dramatically outperforms the thesaurus based synonym replacement method in keeping search intent, from accuracy of 40% to above 80%. Fuchun Peng, Huihsin Tseng, Yumao Lu, Benoît Dumoulin |
CIKM | 4 |
| 2009 | Improving Web Search Relevance with Semantic Features
Yumao Lu, Fuchun Peng, Gilad Mishne, Benoît Dumoulin |
EMNLP | 1 |
| 2008 | Parallel randomized sampling for support vector machine (SVM) and support vector regression (SVR)
Yumao Lu, Vwani P. Roychowdhury |
Knowl. Inf. Syst. | 1 |
| 2008 | Distributed Parallel Support Vector Machines in Strongly Connected NetworksabstractIn this paper, we propose a distributed parallel support vector machine (DPSVM) training mechanism in a configurable network environment for distributed data mining. The basic idea is to exchange support vectors among a strongly connected network (SCN) so that multiple servers may work concurrently on distributed data set with limited communication cost and fast training speed. The percentage of servers that can work in parallel and the communication overhead may be adjusted through network configuration. The proposed algorithm further speeds up through online implementation and synchronization. We prove that the global optimal classifier can be achieved iteratively over an SCN. Experiments on a real-world data set show that the computing time scales well with the size of the training data for most networks. Numerical results show that a randomly generated SCN may achieve better performance than the state of the art method, cascade SVM, in terms of total training time. Yumao Lu, Vwani P. Roychowdhury, Lieven Vandenberghe |
IEEE Trans. Neural Networks | 1 |
| 2007 | Context sensitive stemming for web searchabstractTraditionally, stemming has been applied to Information Retrieval tasks by transforming words in documents to the their root form before indexing, and applying a similar transformation to query terms. Although it increases recall, this naive strategy does not work well for Web Search since it lowers precision and requires a significant amount of additional computation. Fuchun Peng, Nawaaz Ahmed, Yumao Lu |
SIGIR | 4 |
| 2006 | Coupling feature selection and machine learning methods for navigational query identificationabstractIt is important yet hard to identify navigational queries in Web search due to a lack of sufficient information in Web queries, which are typically very short. In this paper we study several machine learning methods, including naive Bayes model, maximum entropy model, support vector machine (SVM), and stochastic gradient boosting tree (SGBT), for navigational query identification in Web search. To boost the performance of these machine techniques, we exploit several feature selection methods and propose coupling feature selection with classification approaches to achieve the best performance. Different from most prior work that uses a small number of features, in this paper, we study the problem of identifying navigational queries with thousands of available features, extracted from major commercial search engine results, Web search user click data, query log, and the whole Web's relational content. A multi-level feature extraction system is constructed.Our results on real search data show that 1) Among all the features we tested, user click distribution features are the most important set of features for identifying navigational queries. 2) In order to achieve good performance, machine learning approaches have to be coupled with good feature selection methods. We find that gradient boosting tree, coupled with linear SVM feature selection is most effective. 3) With carefully coupled feature selection and classification approaches, navigational queries can be accurately identified with 88.1% F1 score, which is 33% error rate reduction compared to the best uncoupled system, and 40% error rate reduction compared to a well tuned system without feature selection. Yumao Lu, Fuchun Peng, Nawaaz Ahmed |
CIKM | 1 |
| 2006 | Parallel Randomized Support Vector Machine
Yumao Lu, Vwani P. Roychowdhury |
PAKDD | 1 |
| 2004 | Dynamic traffic controls for Web-server networks
Yumao Lu |
Comput. Networks | 2 |