VLDB 2026 Research / reviewers in the wild / expert
Ke Gao 0012
dblp:81/2423-12
· DBLP profile ↗
39ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0005-2150-160XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel GenerationabstractDeveloping high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While large language models (LLMs) offer promise for automation, both general-purpose and finetuned LLMs suffer from two fundamental and conflicting limitations: correctness and efficiency. The key reason is that existing LLM-based approaches directly generate the entire optimized low-level programs, requiring exploration of an extremely vast space encompassing both optimization policies and implementation codes. To address the challenge of exploring an intractable space, we propose Macro Thinking Micro Coding (MTMC), a hierarchical framework inspired by the staged optimization strategy of human experts. It decouples optimization strategy from implementation details, ensuring efficiency through high-level strategy and correctness through low-level implementation. Specifically, Macro Thinking employs reinforcement learning to guide lightweight LLMs in efficiently exploring and learning semantic optimization strategies that maximize hardware utilization. Micro Coding leverages general-purpose LLMs to incrementally implement the stepwise optimization proposals from Macro Thinking, avoiding full-kernel generation errors. Together, they effectively navigate the vast optimization space and intricate implementation details, enabling LLMs for high-performance GPU kernel generation. Comprehensive results on widely adopted benchmarks demonstrate the superior performance of MTMC on GPU kernel generation in both accuracy and running time. On KernelBench, MTMC achieves near 100% and 70% accuracy at Levels 1-2 and 3, over 50% than SOTA general-purpose and domain-finetuned LLMs, with up to 7.3× speedup over LLMs, and 2.2× over expert-optimized PyTorch Eager kernels. On the more challenging TritonBench, MTMC attains up to 59.64% accuracy and 34× speedup. All models and datasets will be made publicly available. Xinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen, Qi Guo 0001, Yuanbo Wen 0001, Hang Qin, Ruizhi Chen, Qirui Zhou, Ke Gao 0012, Ling Li 0001 |
AAAI | 10 |
| 2025 | QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language ModelsabstractAs a crucial operator in numerous scientific and engineering computing applications, the automatic optimization of General Matrix Multiplication (GEMM) with full utilization of ever-evolving hardware architectures (e.g. GPUs and RISC-V) is of paramount importance. While Large Language Models (LLMs) can generate functionally correct code for simple tasks, they have yet to produce high-performance code. The key challenge resides in deeply understanding diverse hardware architectures and crafting prompts that effectively unleash the potential of LLMs to generate high-performance code. In this paper, we propose a novel prompt mechanism called QiMeng-GEMM which enables LLMs to comprehend the architectural characteristics of different hardware platforms and automatically search for the optimization combinations for GEMM. The key of QiMeng-GEMM is a set of informative, adaptive, and iterative meta-prompts. Based on this, a searching strategy for optimal combinations of meta-prompts is used to iteratively generate high-performance code. Extensive experiments conducted on 4 leading LLMs, various paradigmatic hardware platforms, and representative matrix dimensions unequivocally demonstrate QiMeng-GEMM’s superior performance in auto-generating optimized GEMM code. Compared to vanilla prompts, our method achieves a performance enhancement of up to 113×. Even when compared to human experts, our method can reach 115% of cuBLAS on NVIDIA GPUs and 211% of OpenBLAS on RISC-V CPUs. Notably, while human experts often take months to optimize GEMM, our approach reduces the development cost by over 240×. Qirui Zhou, Yuanbo Wen 0001, Ruizhi Chen, Ke Gao 0012, Weiqiang Xiong, Ling Li 0001, Qi Guo 0001, Yunji Chen |
AAAI | 4 |
| 2025 | QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware PrimitivesabstractComputation-intensive tensor operators constitute over 90% of the computations in Large Language Models (LLMs) and Deep Neural Networks. Automatically and efficiently generating high-performance tensor operators with hardware primitives is crucial for diverse and ever-evolving hardware architectures like RISC-V, ARM, and GPUs, as manually optimized implementation takes at least months and lacks portability. LLMs excel at generating high-level language codes, but they struggle to fully comprehend hardware characteristics and produce high-performance tensor operators. We introduce a tensor-operator auto-generation framework with a one-line user prompt (QiMeng-TensorOp), which enables LLMs to automatically exploit hardware characteristics to generate tensor operators with hardware primitives, and tune parameters for optimal performance across diverse hardware. Experimental results on various hardware platforms, SOTA LLMs, and typical tensor operators demonstrate that QiMeng-TensorOp effectively unleashes the computing capability of various hardware platforms, and automatically generates tensor operators of superior performance. Compared with vanilla LLMs, QiMeng-TensorOp achieves up to 1291× performance improvement. Even compared with human experts, QiMeng-TensorOp could reach 251% of OpenBLAS on RISC-V CPUs, and 124% of cuBLAS on NVIDIA GPUs. Additionally, QiMeng-TensorOp also significantly reduces development costs by 200× compared with human experts. Xuzhi Zhang, Shaohui Peng, Qirui Zhou, Yuanbo Wen 0001, Qi Guo 0001, Ruizhi Chen, Xinguo Zhu, Weiqiang Xiong, Haixin Chen, Congying Ma, Ke Gao 0012, Yunji Chen, Ling Li 0001 |
IJCAI | 11 |
| 2025 | Coft: Making Large Language Models Better Zero-Shot Learners for Code GenerationabstractThe Chain-of-Thought (CoT) prompting mechanism has effectively enhanced the performance of large language models (LLMs) across a variety of natural language processing (NLP) tasks, including complex zero-shot learning scenarios. Recent studies suggest that this effectiveness arises from CoT's capacity to direct LLMs' attention toward task-relevant keywords. However, traditional CoT methods yield only marginal improvements in the realm of code generation, particularly for models with fewer than 10 billion parameters. We posit that this limitation stems from the substantial disparity between the logical structure and representational form of code compared to natural language. Considering the training and deployment costs, enhancing the performance of small LLMs through advanced prompting and instruction-tuning is essential. In this paper, we introduce COFT (Chain of Functional Triggers), a novel prompting strategy specifically designed for code generation tasks. The design of COFT is based on the following important observation: An optimal CoT tailored for code generation should clearly indicate the core functionality of each critical step, while employing standard identifiers prevalent within the coding domain. Extensive experiments conducted on representative small LLMs ($<10 \mathrm{B}$) benchmarks demonstrate that our COFT substantially outperforms vanilla CoT methods. In challenging zeroshot scenarios and the Pass@1 metric, COFT can improve the performance of fundation LLMs by up to 35.3 %. These empirical findings support our hypothesis that an appropriate design for CoT alongside instruction tuning can fully activate even smallersized LLMs, making them better zero-shot learners for code generation. The source code of COFT and the constructed instruction-tuning dataset will be released. Yongjie Qian, Ke Gao 0012, Haixin Chen, Yuchen Tong, Ling Li 0001 |
ICPC | 3 |
| 2025 | EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationabstractSpeculative decoding is an effective and lossless method for Large Language Model (LLM) inference acceleration. It employs a smaller model to generate a draft token sequence, which is then verified by the original base model. In multi-GPU systems, inference latency can be further reduced through tensor parallelism (TP), while the optimal TP size of the draft model is typically smaller than that of the base model, leading to GPU idling during the drafting stage. We observe that such inefficiency stems from the sequential execution of layers, which is seemingly natural but actually unnecessary. Therefore, we propose EasySpec, a layer-parallel speculation strategy that optimizes the efficiency of multi-GPU utilization. EasySpec breaks the inter-layer data dependencies in the draft model, enabling multiple layers to run simultaneously across multiple devices as ``fuzzy'' speculation. After each drafting-and-verification iteration, the draft model’s key-value cache is calibrated in a single forward pass, preventing long-term fuzzy-error accumulation at minimal additional latency. EasySpec is a training-free and plug-in method. We evaluated EasySpec on several mainstream open-source LLMs, using smaller versions of models from the same series as drafters. The results demonstrate that EasySpec can achieve a peak speedup of 4.17x compared to vanilla decoding, while preserving the original distributions of the base LLMs. Specifically, the drafting stage can be accelerated by up to 1.62x with a maximum speculation accuracy drop of only 7\%. The code is available at https://github.com/Yize-Wu/EasySpec. Yize Wu, Ke Gao 0012, Ling Li 0001 |
NeurIPS | 2 |
| 2025 | A High-Performance and Memory-Efficient RISC-V Operating System Optimization for AIoTabstractThe openness and flexibility of the RISC-V instruction set architecture (ISA) have driven its widespread adoption in AIoT (Artificial Intelligence of Things) devices. However, existing operating systems (OSes) for real RISC-V hardware often suffer from poor application performance and large memory footprints. To address these issues, we propose an OS optimization scheme tailored for RISC-V in AIoT devices. First, we introduce an application-transparent performance enhancement mechanism that leverages both coarse- and fine-grained process management to improve the performance of applications, particularly in AI inference. Second, we design a low-memory-footprint software stack through theoretical analysis and careful trade-offs in the adoption of software components. Lastly, we develop a lightweight OS image construction strategy algorithm tailored for RISC-V in AIoT. Using our OS optimization scheme, we build PolyOS from scratch to reduce the OS image size, thereby further lowering memory footprint. Across four real RISC-V hardware platforms, PolyOS achieves up to a 142% overall system performance improvement and up to 5.90× speedup in AI inference applications compared to baseline OSes (Armbian, Nucleisys, etc.). It also significantly reduces the runtime memory footprint of the standard C library, OpenCV, QuickJS, and AI inference applications, while shrinking the OS image size to 1/3.14–1/23.48 of its baseline OS. Limin Cheng, Ke Gao 0012, Jiageng Yu, Ruizhi Chen, Ling Li 0001 |
SMC | 2 |
| 2024 | Privacy-preserving Compression for Efficient Collaborative InferenceabstractCollaborative inference accelerates DNN inference tasks of resource-limited devices (e.g., clients) by offloading model slices to resource-rich devices (e.g., servers). During the inference procedure, outputs of model slices are transmitted among devices, causing significant intermediate data transmission overhead and posing a risk of privacy leakage of the client’s input data. Quantization has been widely used in collaborative inference to enhance communication efficiency. However, traditional quantization cannot prevent data privacy leaks. Besides, perturbation-based privacy protection methods, such as adding Laplace noise to the intermediate data of collaborative inference, do not consider communication efficiency. In this paper, we introduce Layered Laplace Random Quantization to simultaneously achieve communication efficiency and data privacy protection in collaborative inference by compressing the intermediate data with Laplace quantization noise. We also propose stability training to recover the accuracy loss caused by our method. Evaluation results show that our method achieved an average inference latency speedup of 1.2x-1.3x for different DNN models compared with the baseline methods while achieving comparable data privacy protection and recoverable accuracy loss. Yuzhe Luo, Ji Qi 0002, Jiageng Yu, Ruizhi Chen, Ke Gao 0012, Ling Li 0001 |
ICPADS | 5 |
| 2023 | GANHead: Towards Generative Animatable Neural Head AvatarsabstractTo bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting. Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai |
CVPR | 6 |
| 2023 | NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and NavigationabstractVisual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acquiring a large number of posed images and dense 3D labels in the real world is challenging and costly. In this paper, we present a novel visual localization method that achieves accurate localization while using only a few posed images compared to other localization methods. To achieve this, we first use a few posed images with coarse pseudo-3D labels provided by NeRF to train a coordinate regression network. Then a coarse pose is estimated from the regression network with PNP. Finally, we use the image-based visual servo (IBVS) with the scene prior provided by NeRF for pose optimization. Furthermore, our method can provide effective navigation prior, which enable navigation based on IBVS without using custom markers and depth sensor. Extensive experiments on 7-Scenes and 12-Scenes datasets demonstrate that our method outperforms state-of-the-art methods under the same setting, with only 5\% to 25\% training data. Furthermore, our framework can be naturally extended to the visual navigation task based on IBVS, and its effectiveness is verified in simulation experiments. Yuanze Wang, Yichao Yan, Dian-xi Shi, Wenhan Zhu, Jianqiang Xia, Jeff Tan, Songchang Jin, Ke Gao 0012, Xiaokang Yang 0001 |
NeurIPS | 8 |
| 2022 | Unbiased Manifold Augmentation for Coarse Class Subdivision
Baoming Yan, Ke Gao 0012 |
ECCV (25) | 2 |
| 2021 | Progressive Domain Expansion Network for Single Domain GeneralizationabstractSingle domain generalization is a challenging case of model generalization, where the models are trained on a single domain and tested on other unseen domains. A promising solution is to learn cross-domain invariant representations by expanding the coverage of the training domain. These methods have limited generalization performance gains in practical applications due to the lack of appropriate safety and effectiveness constraints. In this paper, we propose a novel learning framework called progressive domain expansion network (PDEN) for single domain generalization. The domain expansion subnetwork and representation learning subnetwork in PDEN mutually benefit from each other by joint learning. For the domain expansion subnetwork, multiple domains are progressively generated in order to simulate various photometric and geometric transforms in unseen domains. A series of strategies are introduced to guarantee the safety and effectiveness of the expanded domains. For the domain invariant representation learning subnetwork, contrastive learning is introduced to learn the domain invariant representation in which each class is well clustered so that a better decision boundary can be learned to improve it’s generalization. Extensive experiments on classification and segmentation have shown that PDEN can achieve up to 15.28% improvement compared with the state-of-the-art single-domain generalization methods. Codes will be released soon at https://github.com/lileicv/PDEN Ke Gao 0012, Juan Cao 0001, Ziyao Huang 0002, Yepeng Weng, Xiaoyue Mi, Zhengze Yu, Boyang Xia |
CVPR | 2 |
| 2019 | Cost-free Transfer Learning Mechanism: Deep Digging Relationships of Action CategoriesabstractEnd-to-end deep learning for video action recognition is always a data-hungry task because training data with detailed annotations will never be sufficient for such a difficult problem. To fully explore the potential of existing labeled categories, we propose a new transfer learning mechanism deeply digging the local and global relationships between different action categories. As for local relation mining, a weakly-constrained hierarchy structure is built to prioritize the correlation for action pairs. And for global, the consistency of joint distribution of visual and semantic feature for all action categories is measured by the maximum mean discrepancy. As far as we know, this is the first time to explicitly explore the local and global relationships between video action categories. The experimental results show that the proposed relationships of action categories can improve the classification results on most representative models without any extra annotation. Surprisingly, the boost is even more obvious when the training samples are few. We achieve state-of-art experimental results on well-known datasets. Wanneng Wang, Ke Gao 0012, Juan Cao 0001 |
ACM Multimedia | 3 |
| 2019 | Scene-adaptive coded aperture imaging
Yike Ma, Ke Gao 0012, Yongdong Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Context-adaptive matching for optical flow
Yueran Zu, Wenzhong Tang, Xiuguo Bao, Ke Gao 0012 |
Multim. Tools Appl. | 5 |
| 2018 | Not All Words Are Equal: Video-specific Information Loss for Video Captioning
Jiarong Dong, Ke Gao 0012, Xiaokai Chen, Junbo Guo, Juan Cao 0001, Yongdong Zhang 0001 |
BMVC | 2 |
| 2017 | Task-Driven Dynamic Fusion: Reducing Ambiguity in Video DescriptionabstractIntegrating complementary features from multiple channels is expected to solve the description ambiguity problem in video captioning, whereas inappropriate fusion strategies often harm rather than help the performance. Existing static fusion methods in video captioning such as concatenation and summation cannot attend to appropriate feature channels, thus fail to adaptively support the recognition of various kinds of visual entities such as actions and objects. This paper contributes to: 1)The first in-depth study of the weakness inherent in data-driven static fusion methods for video captioning. 2) The establishment of a task-driven dynamic fusion (TDDF) method. It can adaptively choose different fusion patterns according to model status. 3) The improvement of video captioning. Extensive experiments conducted on two well-known benchmarks demonstrate that our dynamic fusion method outperforms the state-of-the-art results on MSVD with METEOR scores 0.333, and achieves superior METEOR scores 0.278 on MSR-VTT-10K. Compared to single features, the relative improvement derived from our fusion method are 10.0% and 5.7% respectively on two datasets. Xishan Zhang, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001 |
CVPR | 2 |
| 2017 | Matryoshka Peek: Toward Learning Fine-Grained, Robust, Discriminative Features for Product SearchabstractIn sharp contrast to the traditional category/subcategory level image retrieval, product image search aims to find the images containing the exact same product. This is a challenging problem because in addition to being robust under different imaging conditions such as varying viewpoints and illumination changes, the features should also be able to distinguish the specific product among many similar products. Consequently, it is important to utilize a large dataset, containing many product classes, to learn a strongly discriminative representation. Building such a dataset requires laborious manual annotation. Toward learning fine-grained, robust, discriminative features for product image search, we present a novel paradigm that can construct the required dataset without any human annotation. Unlike other fine-grained recognition works that rely on high-quality annotated datasets and are very narrowly focused on a specific object category, our method handles multiple object classes and requires minimum human effort. First, an ImageNet pretrained model is used to generate product clusters. As the original features from ImageNet are not discriminative, the clusters generated by this unsupervised procedure contain much noise. We alleviate noise by explicitly modeling noise distribution and automatically detecting errors during learning. The proposed paradigm is general, requires minimum human efforts, and is applicable to any deep learning task where fine-grained discriminative features are desired. Extensive experiments on the ALISC dataset have demonstrated that our approach is sound and effective, surpassing the baseline GoogleNet model by 15.09%. Zawlin Kyaw, Shuhan Qi, Ke Gao 0012, Hanwang Zhang, Jun Xiao 0001, Xuan Wang 0002, Tat-Seng Chua |
IEEE Trans. Multim. | 3 |
| 2017 | Trip Outfits Advisor: Location-Oriented Clothing RecommendationabstractWhen packing for a journey, have you ever asked “what clothes should I take with me?” Wearing appropriate and aesthetically pleasing clothing when traveling is a concern for many of us. Our data observation of photos from several popular travel websites reveals that people's choice of clothing items and their color combinations have strong correlations with the weather, the season, and the main type of attraction at the destination. This leads to an interesting and novel problem: can the correlation between clothing and locations be automatically learned from social photos and leveraged for location-oriented clothing recommendations? In this paper, we systematically study this problem and propose a hybrid multilabel convolutional neural network combined with the support vector machine (mCNN-SVM) approach to capture the intrinsic and complex correlations between clothing attributes and location attributes. Specifically, we adapt the CNN architecture to multilabel learning and fine-tune it using each fine-grained clothing item. Then, the recognized items are fed to the SVM to learn the correlations. Experiments on three fashion datasets and a benchmark journey outfit dataset show that our proposed approach outperforms several baselines by over 10.52-16.38% in terms of the mAP for clothing item recognition and outperforms several alternative methods by over 9.59-29.41% in terms of the mAP when ranking clothing by appropriateness for travel destinations. Finally, an interesting case study demonstrates the effectiveness of our method by answering what items to wear, how to match them, and how to dress in an aesthetically pleasing manner for a journey. Xishan Zhang, Jia Jia 0001, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Efficient Perceptual Region Detector Based on Object Boundary
Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
MMM (2) | 2 |
| 2015 | Maximally Visual-Homogeneous Region Detector for Large Scale Image RetrievalabstractConventional local detectors often extract numerous small repeated regions in textured areas, which easily results in false matching. In order to find representative and distinctive local invariant regions, this paper proposes a Maximally Visual-Homogeneous Region (MVHR) detector. The main contributions can be summarized as 2 parts: (1) Being different from original MSER which employs single pixel intensity as ranking unit, we propose a novel sorting method based on visual homogeneity analysis on a local patch. (2) Identifying the observation scale has a close relationship with visual homogeneity analysis, a heuristic scale selection algorithm is developed to choose a proper scale according to the changes of visual homogeneity evaluation over a range of scales. Experiments demonstrate our detector can find less but representative regions with high repeatability, while still perserving competitive precision compared to the state-of-art detectors for large scale image retrieval. Ke Gao 0012, Jintao Li 0001 |
ICMR | 2 |
| 2014 | Salient region detection : Integrate both global and local cuesabstractVisual saliency detection provides an alternative methodology to semantic image understanding in many applications such as region-based image retrieval and adaptive compression of images. In this paper, we propose an approach which utilizes both global and local cues to extract saliency information. Our method can achieve better performance than existing saliency detection methods in terms of precision and recall rates. The main contributions are threefold: 1) a new model which can better describe the color perception of human beings is proposed. Based on this model, a global color contrast cue is also presented. 2) as supplements, two other global cues and one local cues are also presented to capture as much saliency information as we can. 3) a CRF model is used to integrate these cues and generate the final saliency map. Experimental results indicate that our proposed approach is effective and practicable. Tiancai Ye, Dongming Zhang 0004, Ke Gao 0012, Guoqing Jin, Yongdong Zhang 0001, Qingsheng Yuan |
ICME | 3 |
| 2014 | A Representative Local Region Detector Based On Color-Contrast-MSERabstractIn order to extract representative local invariant regions in textured natural images, we propose a Color-Contrast-MSER (CCM) detector with color-contrast pixel ranking, which can reduce the number of meaningless regions extracted from backgrounds. The main contributions are threefold: (1) In contrast with the original MSER[3] which adopts intensity pixel ranking, we develop a new pixel ranking mechanism based on color contrast analysis. (2) In this paper, the pixel ranking value of each pixel is defined as the color contrast between a kernel-sized window and the background. Therefore we propose an adaptive background scale selection mechanism that simulates the background color distribution as the benchmark for color contrast. (3) The experimental results demonstrate that compared with the original MSER detector[3], our Color-Contrast-MSER (CCM) detector can extract more representative local regions with competitive repeatability score at only 50% computational time and 10% memory cost. Ke Gao 0012, Sheng Tang, Yongdong Zhang 0001 |
ICMR | 2 |
| 2014 | Monte Carlo Sampling based Salient Region DetectionabstractIn this paper, a simple and effective method is proposed for salient region detection. Based on the observation that salient regions tend to be compact, connected and surrounded, our original idea is to exploit these three kinds of prior knowledge. However, concepts of spatial structure (such as connectivity and surroundedness) only have definite meanings in binary images. Thus, a Monte Carlo Sampling based Saliency model is proposed. Our model has two main advantages over other methods. Firstly, the result of each sampling process is a binary map which can greatly simplify the combination with prior knowledge of spatial structure. Secondly, our method is naturally parallelized because every sampling process is independent with each other, which makes our method very efficient. Experimental results on two datasets show that, compared with eleven state-of-the-art methods, our approach has a competitive performance and also runs very fast. Tiancai Ye, Dongming Zhang 0004, Guoqing Jin, Ke Gao 0012, Xiaoguang Gu, Yongdong Zhang 0001 |
ICMR | 4 |
| 2014 | Efficient binary code indexing with pivot based locality sensitive clustering
Wei Zhang 0043, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
Multim. Tools Appl. | 2 |
| 2013 | Learning Affine Robust Binary Codes Based on Locality Preserving Hash
Wei Zhang 0043, Ke Gao 0012, Dongming Zhang 0004, Jintao Li 0001 |
MMM (1) | 2 |
| 2013 | Robust common visual pattern discovery using graph matching
Hongtao Xie 0001, Yongdong Zhang 0001, Ke Gao 0012, Sheng Tang, Kefu Xu, Li Guo 0001, Jintao Li 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Accurate off-line query expansion for large-scale mobile visual search
Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Shouxun Lin |
Signal Process. | 1 |
| 2012 | Visual stem mapping and Geometric Tense coding for Augmented Visual VocabularyabstractThis paper addresses the problem of affine distortions caused by viewpoint changes for the application of image retrieval. We study how to expand the visual words from a query image for better retrieval recall without the sacrifice of retrieval precision and efficiency. Our main contribution is the building of visual dictionaries that retain the mapping relationships between visual words extracted from different viewpoints of the same object. Additionally, in each mapping rule we record the affine transformation in which the two visual words are related, as a compact code of viewpoints relationships. By analogizing the concepts of verb stem and verb tense in text, we use Visual Stems to denote visual words extracted from robust local patches, and record the relationships between their affine variants as visual stem mapping rules, including the geometric relationships coded as Geometric Tenses. In this way, our method augments original visual vocabulary with sufficient and accurate expansion information. In query phase, only the objects corresponding to the same visual stems and coherent geometric tense codes will be regarded as similar ones. Moreover, the mapping rules can be learned offline with only one sample for each object. Experiments show that our method can support efficient object retrieval with high recall, requiring little extra time and space cost over traditional visual vocabularies. Ke Gao 0012, Yongdong Zhang 0001, Wei Zhang 0043, Junhai Xia, Shouxun Lin |
CVPR | 1 |
| 2012 | A method for detecting salient regions using integrated featuresabstractWe develop a novel algorithm for detecting salient regions. By analyzing the advantages and disadvantages of the existing methods, five principles for designing salient region detection algorithms are summarized. Based on these principles, we propose a novel method that generates saliency map with highlighted salient regions by utilizing two different features, namely visual saliency value and spatial weight. The visual saliency value is determined based on local contrast differences and low-level feature frequencies. The spatial weight is computed by analyzing the size and location of salient regions. Experimental results show that the proposed algorithm outperforms 7 state-of-the-art methods on the public image set. Zhendong Mao 0001, Yongdong Zhang 0001, Ke Gao 0012, Dongming Zhang 0004 |
ACM Multimedia | 3 |
| 2012 | Geometric context-preserving progressive transmission in mobile visual searchabstractProgressive transmission is very effective to reduce retrieval latency in mobile visual search. However, the acceleration effects of existing progressive transmission strategies are often limited because of the neglect of geometric information in the query image. This paper proposes an effective and efficient geometric context-preserving progressive transmission method, which is suitable for mobile visual search. Here a query image is divided into blocks and local features in the same block are used as query units rather than a single feature. Since clustered features with geometric information are more discriminative, only a few of them could support correct matching with high precision. Thus our method significantly decreases the number of features needed for transmission, and dramatically reduces the retrieval latency. Experiments on Stanford dataset for mobile visual search show that, with comparable precision, we uses 43% less retrieval time than existing progressive transmission method. Moreover, we establish and release a large-scale image dataset called MVSBench which is more difficult and suitable for mobile visual search. It contains 75500 images and considers many variations like view change, blur, scale, illumination and rotation. MVSBench is another major contribution of this paper, and our method also outperforms other strategies on this dataset. Junhai Xia, Ke Gao 0012, Dongming Zhang 0004, Zhendong Mao 0001 |
ACM Multimedia | 2 |
| 2011 | Local geometric consistency constraint for image retrievalabstractIn state-of-the-art image retrieval systems, an image is represented by bag-of-features (BOF). As BOF representation discards geometric relationships among local features, exploiting geometric constraints as post-processing procedure has been shown to greatly improve retrieval precision. However, full geometric constraints are computationally expensive and weak geometric constraints have limited range of applications. To efficiently handle common transformations and deformations, we present a novel local geometric consistency constraint (LGC) method. It utilizes the local similarity characteristic of deformations, and measures the pairwise geometric similarity of matches between two sets of local features. Besides, we propose a new method to accurately calculate the transformation matrix between two matched features, with the information provided by their local neighbors. Experiments performed on famous datasets show the excellent performance of our method. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICIP | 2 |
| 2011 | Pairwise weak geometric consistency for large scale image searchabstractState-of-the-art image search systems mostly build on bag-of-features (BOF) representation. As BOF ignores geometric relationships among local features, geometric consistency constraints have been proposed to improve search precision. However, exploiting full geometric constraints are too computational expensive. Weak geometric constraints have strong assumptions and can only deal with uniform transformations. To handle view point changes and nonrigid deformations, in this paper we present a novel pairwise weak geometric consistency constraint (P-WGC) method. It utilizes the local similarity characteristic of deformations, and measures the pairwise geometric similarity of matches between two sets of local features. Experiments performed on four famous datasets and a dataset of one million of images show a significant improvement due to P-WGC as well as its efficiency. Further improvement of search accuracy is obtained when it is combined with full geometric verification. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICMR | 2 |
| 2011 | Common visual pattern discovery via graph matchingabstractDiscovering common visual patterns (CVPs) between two images is a challenging problem, due to the significant photometric and geometric transformations, and the high computational cost. In this paper, we formulate CVPs discovery as a graph matching problem, depending on pairwise geometric compatibility between feature correspondences. To efficiently find all CVPs, we propose two algorithms--Preliminary Initialization Optimization (PIO) and Post Agglomerative Combining (PAC). PIO reduces the search space of CVPs discovery based on the internal homogeneity of CVPs, while PAC refines the discovery result in an agglomerative way. Experiments on object recognition and near-duplicate image re-trieval validate the effectiveness and efficiency of our method. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001, Huamin Ren |
ACM Multimedia | 2 |
| 2011 | Efficient approximate nearest neighbor search with integrated binary codesabstractNearest neighbor search in Euclidean space is a fundamental problem in multimedia retrieval. The difficulty of exact nearest neighbor search has led to approximate solutions that sacrifice precision for efficiency. Among such solutions, approaches that embed data into binary codes in Hamming space have gained significant success for their efficiency and practical memory requirements. However, binary code searching only finds a big and coarse set of similar neighbors in Hamming space, and hence expensive Euclidean distance based ranking of the coarse set is needed to find nearest neighbors. Therefore, to improve nearest neighbor search efficiency, we proposed a novel binary code method called Integrated Binary Code (IBC) to get a compact set of similar neighbors. Experiments on public datasets show that our method is more efficient and effective than state-of-the-art in approximate nearest neighbor search. Wei Zhang 0043, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ACM Multimedia | 2 |
| 2011 | Efficient Feature Detection and Effective Post-Verification for Large Scale Near-Duplicate Image SearchabstractState-of-the-art near-duplicate image search systems mostly build on the bag-of-local features (BOF) representation. While favorable for simplicity and scalability, these systems have three shortcomings: 1) high time complexity of the local feature detection; 2) discriminability reduction of local descriptors due to BOF quantization; and 3) neglect of the geometric relationships among local features after BOF representation. To overcome these shortcomings, we propose a novel framework by using graphics processing units (GPU). The main contributions of our method are: 1) a new fast local feature detector coined Harris-Hessian (H-H) is designed according to the characteristics of GPU to accelerate the local feature detection; 2) the spatial information around each local feature is incorporated to improve its discriminability, supplying semi-local spatial coherent verification (LSC); and 3) a new pairwise weak geometric consistency constraint (P-WGC) algorithm is proposed to refine the search result. Additionally, part of the system is implemented on GPU to improve efficiency. Experiments conducted on reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of H-H, LSC, and P-WGC. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Sheng Tang, Jintao Li 0001 |
IEEE Trans. Multim. | 2 |
| 2010 | GPU-based fast scale invariant interest point detectorabstractTo take full advantage of the powerful computing capability of graphics processing units (GPU) to speed up local feature detection, we present a novel GPU-based scale invariant interest point detector, coined Harris-Hessian(H-H). H-H detects Harris points in low scale and refines their location and scale in higher scale-space with the determinant of Hessian matrix. Compared to the existing methods, H-H significantly reduces the pixel-level computation complexity and has better parallelism. The experiment results show that with the assistance of GPU, H-H achieves up to a 10-20x speedup than CPU-based method. It only takes 6.3ms to detect a 640 × 480 image with high detection accuracy, meeting the need of real-time detection. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ICASSP | 2 |
| 2010 | Data-oriented locality sensitive hashingabstractLocality Sensitive Hashing (LSH) has been proposed as a scalable and high-dimensional index for approximate similarity search. Euclidean LSH is a variation of LSH and has been successfully used in many multimedia applications. However, hash functions of the basic Euclidean LSH project data points over randomly selected directions, which reduces accuracy when data are non-uniformly distributed. So more hash tables are needed to guarantee the accuracy, and thus more memory is consumed. Since heavy memory cost is a significant drawback of Euclidean LSH, we propose Data-Oriented LSH to reduce memory consumption when data are non-uniformly distributed. Most of existing methods are query-directed, such as multi-probe and query expansion methods. We focused on the hash table construction, and thus the query-directed methods can be applied to our index to improve further. The experiment shows that to achieve the same accuracy, our method uses less time and less memory compared with original Euclidean LSH. Wei Zhang 0043, Ke Gao 0012, Yongdong Zhang 0001, Jintao Li 0001 |
ACM Multimedia | 2 |
| 2009 | Logo detection based on spatial-spectral saliency and partial spatial contextabstractLogo detection is important for brand advertising and surveillance applications. The central issues of this technology are fast localization and accurate matching. Based on key traits analysis of common logos, this paper presents a two-stage detection scheme based on spatialspectral saliency (SSS) and partial spatial context (PSC). SSS speeds up logo location and avoid the impact of cluttered background. PSC filters false matching using spatial consistency of local invariant points. The integration of SSS and PSC result in faster localization and increased accuracy. Experiments on a dataset of nearly 10,000 web images containing several popular logo types are presented. The results indicate that our method is applicable and precise for different logo detection scenarios. Ke Gao 0012, Shouxun Lin, Yongdong Zhang 0001, Sheng Tang, Dongming Zhang 0004 |
ICME | 1 |
| 2008 | Object retrieval based on spatially frequent items with informative patchesabstractSpatial relation of local image patches plays an important role in object-based image retrieval. An approach called spatial frequent items is proposed as an extension of Bag-of-Words method by introducing spatial relations between patches. Spatial frequent items are defined as frequent pairs of adjacent local image patches in polar coordinates, and exploited using data mining. Based on these frequent configurations, we develop a method to encode patches and their spatial relations for image indexing and retrieval. Besides, to avoid the interference of background patches, informative patches are filtrated based on their local entropy and self-similarity in the preprocess stage. Experimental results demonstrate that our method can be 8.6% more effective than the state-of-art object retrieval methods. Ke Gao 0012, Shouxun Lin, Junbo Guo, Dongming Zhang 0004, Yongdong Zhang 0001 |
ICME | 1 |