Zhiwei Li 0006

dblp:47/3951-6 · DBLP profile ↗
← Back
50ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 22 · 4 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorSystems, architecture and hardware · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2024 Continuity Preserving Online CenterLine Graph Learning
Yunhui Han, Zhiwei Li 0006
ECCV (72)3
2022 A Resource-Efficient Pipelined Architecture for Real-Time Semi-Global Stereo Matching
abstract
It is still a grand challenge to implement a high-accuracy and high-performance stereo matching algorithm on a resource-limited hardware platform in stereo vision systems. This paper proposes a resource-efficient pipelined hardware architecture with four-cycle time-sharing for the semi-global matching (SGM) algorithm with weighted path cost aggregation. To save hardware resources, we also combined image down-sampling and disparity skipping in the SGM algorithm. The presented architecture is synthesized and implemented on a Zynq-7 FPGA board, which results in a throughput of${1280 \times 960/62.5}$fps with 75 disparity levels at the maximum frequency of 216 MHz. To improve the accuracy of the disparity map at close range, we also adapt the presented architecture with two-cycle time-sharing, and the disparity range is increased to 128, which attains the processing of${1280 \times 960/116}$fps at 200 MHz on VCU-118 FPGA board; the throughput reaches 18245 MDE/s. The result shows that the whole architecture only takes 50465 LUTs, 48046 Registers, 125.5 BRAMs with 128 disparity levels, which is much more efficient than the latest reference work.
Zhimin Lu, Zhiwei Li 0006, Song Chen 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 DT-Loc: Monocular Visual Localization on HD Vector Map Using Distance Transforms of 2D Semantic Detections
abstract
Localizing a vehicle on a prebuilt HD vector map is a prerequisite for many autonomous driving applications. Existing visual localization approaches usually require a separate local feature layer to function. The separate localization layer suffers from the robustness issue inherited from the local features. Also, it could be difficult to create a feature layer that aligns perfectly with an existing vector map. In this paper, we propose a monocular visual localization method that exploits the vector map directly as the localization layer. The method detects semantic traffic elements from the images and matches them with the vectors in the map. To deal with the harmful problem of false matches, we propose to align the vector map to the distance transforms of the semantic detections, which enables a non-explicit and differentiable data association process. The system is able to achieve centimeter and sub-meter accuracies in lateral and longitudinal directions, respectively.
Chi Zhang 0069, Hao Liu 0007, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS7
2021 Robust LiDAR Localization on an HD Vector Map without a Separate Localization Layer
abstract
Many autonomous driving applications nowadays come along with a prebuilt vector map for routing and planning purposes. In order to localize on this map, traditional LiDAR localization methods usually require a separate localization layer to function. On one hand, the separate layer occupies large storage and is not convenient to update. On the other hand, the potential of the vector map itself has not been fully exploited by existing methods. In this paper, we present a LiDAR localization system that leverages the vector map directly as the localization layer. A semantic extraction module is developed to match the heterogeneous data between LiDAR measurements and the 3D vector elements. A local map maintenance module is introduced to keep the system function robustly when there are not enough vector matches. The system adopts an optimization-based framework and infers 6-DOF poses. Experiments show that the proposed system is able to achieve centimeter accuracy robustly in both highway and urban environments, without a separate localization layer.
Chi Zhang 0069, Liwen Liu, Zhoupeng Xue, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS7
2021 AVP-Loc: Surround View Localization and Relocalization Based on HD Vector Map for Automated Valet Parking
abstract
Localization is a crucial prerequisite for automated valet parking, in which a vehicle is required to navigate itself in a GPS-denied parking lot. Traditional visual localization methods usually build a feature map and use it for future localizations. However, the feature map is not robust to changes in illumination, appearance, and viewing perspective. To deal with this issue, we need a more stable map. In this paper, we propose to use the parking lot’s HD vector map directly for localization. The vector representation is ultimately stable but brings challenges in data association as well. To this end, we present a novel data association method to match the surround-view images with the vector map. In addition, we also propose a closed-form relocalization strategy by exploiting distinctive road mark combinations in the vector map. Experiments show that the proposed method is able to achieve centimeter-level localization accuracy in a multi-floor parking lot.
Chi Zhang 0069, Hao Liu 0007, Zhijun Xie, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS7
2020 Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
abstract
State-of-the-art deep learning based stereo matching approaches treat disparity estimation as a regression problem, where loss function is directly defined on true disparities and their estimated ones. However, disparity is just a byproduct of a matching process modeled by cost volume, while indirectly learning cost volume driven by disparity regression is prone to overfitting since the cost volume is under constrained. In this paper, we propose to directly add constraints to the cost volume by filtering cost volume with unimodal distribution peaked at true disparities. In addition, variances of the unimodal distributions for each pixel are estimated to explicitly model matching uncertainty under different contexts. The proposed architecture achieves state-of-the-art performance on Scene Flow and two KITTI stereo benchmarks. In particular, our method ranked the 1st place of KITTI 2012 evaluation and the 4th place of KITTI 2015 evaluation (recorded on 2019.8.20). The codes of AcfNet are available at: https://github.com/youmi-zym/AcfNet.
Youmin Zhang 0005, Xiao Bai 0001, Suihanjin Yu, Zhiwei Li 0006, Kuiyuan Yang
AAAI6
2019 Low-Resource Hardware Architecture for Semi-Global Stereo Matching
abstract
The semi-global matching algorithm is usually used for generating high-quality and real-time disparity maps in stereo vision systems. To reduce the hardware-resource consumption, we present a multi-stage pipeline hardware architecture with timesharing reuse for semi-global stereo matching. Combined with image down-sampling, jumping disparity, and a post processing, the presented architecture is used in a practical advanced driver-assistance system (ADAS), which is implemented on a Zynq-7 FPGA chip. The whole stereo matching architecture consumes 19,603 LUTs and 61.5 BRAM (36 KB), and the throughput is 2857 Million Disparity Estimation per second (MDE/S), which corresponds to a throughput of 31 fps when processing images with 1280∗960 resolution and 75 disparity levels.
Zhiwei Li 0006, Lan Yao, Song Chen 0001, Feng Wu 0001
ISCAS2
2019 High throughput hardware architecture for accurate semi-global matching
Yan Li 0068, Zhiwei Li 0006, Song Chen 0001
Integr.2
2018 DenseASPP for Semantic Segmentation in Street Scenes
abstract
Semantic image segmentation is a basic street scene understanding task in autonomous driving, where each pixel in a high resolution image is categorized into a set of semantic labels. Unlike other scenarios, objects in autonomous driving scene exhibit very large scale changes, which poses great challenges for high-level feature representation in a sense that multi-scale information must be correctly encoded. To remedy this problem, atrous convolution[14]was introduced to generate features with larger receptive fields without sacrificing spatial resolution. Built upon atrous convolution, Atrous Spatial Pyramid Pooling (ASPP)[2] was proposed to concatenate multiple atrous-convolved features using different dilation rates into a final feature representation. Although ASPP is able to generate multi-scale features, we argue the feature resolution in the scale-axis is not dense enough for the autonomous driving scenario. To this end, we propose Densely connected Atrous Spatial Pyramid Pooling (DenseASPP), which connects a set of atrous convolutional layers in a dense way, such that it generates multi-scale features that not only cover a larger scale range, but also cover that scale range densely, without significantly increasing the model size. We evaluate DenseASPP on the street scene benchmark Cityscapes[4] and achieve state-of-the-art performance.
Maoke Yang, Chi Zhang 0069, Zhiwei Li 0006, Kuiyuan Yang
CVPR4
2017 High throughput hardware architecture for accurate semi-global matching
abstract
As the most important step of a stereo vision system, stereo matching, which finds the correspondences in stereo image pairs, requires high-quality real-time depth computation. In this paper, a high accuracy and high throughput full-pipeline hardware architecture with disparity and row parallelism is proposed. In the semi-global aggregation stage, to improve the accuracy in discontinuous regions, adaptive weighted path costs are adopted, and, five aggregation paths are used without consuming external memory resources. The proposed hardware architecture is implemented on a Stratix V FPGA, which results in a throughput of 1280×960/197fps with 64 disparity levels at 156MHz.
Yan Li 0068, Zhiwei Li 0006, Song Chen 0001
ASP-DAC4
2017 Locality-Sensitive Deconvolution Networks with Gated Fusion for RGB-D Indoor Semantic Segmentation
abstract
This paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates large context to predict the label of each pixel, inherently limiting the segmentation precision of object boundaries. The other is about RGB-D fusion. Recent state-of-the-art methods generally fuse RGB and depth networks with equal-weight score fusion, regardless of the varying contributions of the two modalities on delineating different categories in different scenes. To address the two problems, we first propose a locality-sensitive DeconvNet (LS-DeconvNet) to refine the boundary segmentation over each modality. LS-DeconvNet incorporates locally visual and geometric cues from the raw RGB-D data into each DeconvNet, which is able to learn to upsample the coarse convolutional maps with large context whilst recovering sharp object boundaries. Towards RGB-D fusion, we introduce a gated fusion layer to effectively combine the two LS-DeconvNets. This layer can learn to adjust the contributions of RGB and depth over each pixel for high-performance object recognition. Experiments on the large-scale SUN RGB-D dataset and the popular NYU-Depth v2 dataset show that our approach achieves new state-of-the-art results for RGB-D indoor semantic segmentation.
Yanhua Cheng, Rui Cai 0002, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang
CVPR3
2016 Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency
abstract
In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouettes for multiple views on the fly. To obtain a set of accurate silhouettes, precise camera poses are required to propagate segmentation cues across views. To perform better localization, accurate silhouettes are needed to constrain camera poses. Therefore the two problems are intertwined with each other and require a joint treatment. Facilitated by the available depth, we introduce a simple but effective silhouette consistency energy term that binds traditional appearance-based multiview segmentation cost and RGB-D frame-to-frame matching cost together. Optimization of the problem w.r.t. binary segmentation masks and camera poses naturally fits in the graph cut minimization framework and the Gauss-Newton non-linear least-squares method respectively. Experiments show that the proposed approach achieves state-of-the-arts performance on both tasks of image segmentation and camera localization.
Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui
CVPR2
2016 Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition
Yanhua Cheng, Xin Zhao 0012, Rui Cai 0002, Zhiwei Li 0006, Kaiqi Huang, Yong Rui
IJCAI4
2015 Query Adaptive Similarity Measure for RGB-D Object Recognition
abstract
This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object pose and scale changes, as well as intra-class variations, and (2) effectively fusing RGB and depth cues is still an open problem. To address these problems, this paper first proposes a new similarity measure based on dense matching, through which objects in comparison are warped and aligned, to better tolerate variations. Towards RGB and depth fusion, we argue that a constant and golden weight doesn't exist. The two modalities have varying contributions when comparing objects from different categories. To capture such a dynamic characteristic, a group of matchers equipped with various fusion weights is constructed, to explore the responses of dense matching under different fusion configurations. All the response scores are finally merged following a learning-to-combination way, which provides quite good generalization ability in practice. The proposed approach win the best results on several public benchmarks, e.g., achieves 92.7% top-1 test accuracy on the Washington RGB-D object dataset, with a 5.1% improvement over the state-of-the-art.
Yanhua Cheng, Rui Cai 0002, Chi Zhang 0069, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang, Yong Rui
ICCV4
2015 MeshStereo: A Global Stereo Model with Mesh Alignment Regularization for View Interpolation
abstract
We present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D triangles with shared vertices. Lifting the 2D triangulation to 3D naturally generates a corresponding mesh. A technical difficulty is to properly split vertices to multiple copies when they appear at depth discontinuous boundaries. To deal with this problem, we formulate our objective as a two-layer MRF, with the upper layer modeling the splitting properties of the vertices and the lower layer optimizing a region-based stereo matching. Experiments on the Middlebury and the Herodion datasets demonstrate that our model is able to synthesize visually coherent new view angles with high PSNR, as well as outputting high quality disparity maps which rank at the first place on the new challenging high resolution Middlebury 3.0 benchmark.
Chi Zhang 0069, Zhiwei Li 0006, Yanhua Cheng, Rui Cai 0002, Hongyang Chao, Yong Rui
ICCV2
2014 As-Rigid-As-Possible Stereo under Second Order Smoothness Priors
Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui
ECCV (2)2
2013 Efficient 2D-to-3D Correspondence Filtering for Scalable 3D Object Recognition
abstract
3D model-based object recognition has been a noticeable research trend in recent years. Common methods find 2D-to-3D correspondences and make recognition decisions by pose estimation, whose efficiency usually suffers from noisy correspondences caused by the increasing number of target objects. To overcome this scalability bottleneck, we propose an efficient 2D-to-3D correspondence filtering approach, which combines a light-weight neighborhood-based step with a finer-grained pairwise step to remove spurious correspondences based on 2D/3D geometric cues. On a dataset of 300 3D objects, our solution achieves ~10 times speed improvement over the baseline, with a comparable recognition accuracy. A parallel implementation on a quad-core CPU can run at ~3fps for 1280×720 images.
Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001, Yong Rui
CVPR3
2012 Hierarchical Object Representations for Visual Recognition via Weakly Supervised Learning
Tianzhu Zhang 0001, Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Hanqing Lu
ACCV (1)3
2012 3D visual phrases for landmark recognition
abstract
In this paper, we study the problem of landmark recognition and propose to leverage 3D visual phrases to improve the performance. A 3D visual phrase is a triangular facet on the surface of a reconstructed 3D landmark model. In contrast to existing 2D visual phrases which are mainly based on co-occurrence statistics in 2D image planes, such 3D visual phrases explicitly characterize the spatial structure of a 3D object (landmark), and are highly robust to projective transformations due to viewpoint changes. We present an effective solution to discover, describe, and detect 3D visual phrases. The experiments on 10 landmarks have achieved promising results, which demonstrate that our approach provides a good balance between precision and recall of landmark recognition while reducing the dependence on post-verification to reject false positives.
Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001
CVPR3
2011 Rank-SIFT: Learning to rank repeatable local interest points
abstract
Scale-invariant feature transform (SIFT) has been well studied in recent years. Most related research efforts focused on designing and learning effective descriptors to characterize a local interest point. However, how to identify stable local interest points is still a very challenging problem. In this paper, we propose a set of differential features, and based on them we adopt a data-driven approach to learn a ranking function to sort local interest points according to their stabilities across images containing the same visual objects. Compared with the handcrafted rule-based method used by the standard SIFT algorithm, our algorithm substantially improves the stability of detected local interest point on a very challenging benchmark dataset, in which images were generated under very different imaging conditions. Experimental results on the Oxford and PASCAL databases further demonstrate the superior performance of the proposed algorithm on both object image retrieval and category recognition.
Rong Xiao 0003, Zhiwei Li 0006, Rui Cai 0002, Bao-Liang Lu, Lei Zhang 0001
CVPR3
2011 Contextual synonym dictionary for visual object retrieval
abstract
In this paper, we study the problem of visual object retrieval by introducing a dictionary of contextual synonyms to narrow down the semantic gap in visual word quantization. The basic idea is to expand a visual word in the query image with its synonyms to boost the retrieval recall. Unlike the existing work such as soft-quantization, which only focuses on the Euclidean (l2) distance in descriptor space, we utilize the visual words which are more likely to describe visual objects with the same semantic meaning by identifying the words with similar contextual distributions (i.e. contextual synonyms). We describe the contextual distribution of a visual word using the statistics of both co-occurrence and spatial information averaged over all the image patches having this visual word, and propose an efficient system implementation to construct the contextual synonym dictionary for a large visual vocabulary. The whole construction process is unsupervised and the synonym dictionary can be naturally integrated into a standard bag-of-feature image retrieval system. Experimental results on several benchmark datasets are quite promising. The contextual synonym dictionary-based expansion consistently outperforms the l2 distance-based soft-quantization, and advances the state-of-the-art performance remarkably.
Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001
ACM Multimedia3
2011 Query by document via a decomposition-based two-level retrieval approach
abstract
Retrieving similar documents from a large-scale text corpus according to a given document is a fundamental technique for many applications. However, most of existing indexing techniques have difficulties to address this problem due to special properties of a document query, e.g. high dimensionality, sparse representation and semantic issue. Towards addressing this problem, we propose a two-level retrieval solution based on a document decomposition idea. A document is decomposed to a compact vector and a few document specific keywords by a dimension reduction approach. The compact vector embodies the major semantics of a document, and the document specific keywords complement the discriminative power lost in dimension reduction process. We adopt locality sensitive hashing (LSH) to index the compact vectors, which guarantees to quickly find a set of related documents according to the vector of a query document. Then we re-rank documents in this set by their document
Linkai Weng, Zhiwei Li 0006, Rui Cai 0002, Yaoxue Zhang, Yue-Zhi Zhou, Laurence T. Yang, Lei Zhang 0001
SIGIR2
2010 Spatial-bag-of-features
abstract
In this paper, we study the problem of large scale image retrieval by developing a new class of bag-of-features to encode geometric information of objects within an image. Beyond existing orderless bag-of-features, local features of an image are first projected to different directions or points to generate a series of ordered bag-of-features, based on which different families of spatial bag-of-features are designed to capture the invariance of object translation, rotation, and scaling. Then the most representative features are selected based on a boosting-like method to generate a new bag-of-features-like vector representation of an image. The proposed retrieval framework works well in image retrieval task owing to the following three properties: 1) the encoding of geometric information of objects for capturing objects' spatial transformation, 2) the supervised feature selection and combination strategy for enhancing the discriminative power, and 3) the representation of bag-of-features for effective image matching and indexing for large scale image retrieval. Extensive experiments on 5000 Oxford building images and 1 million Panoramio images show the effectiveness and efficiency of the proposed features as well as the retrieval framework.
Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001
CVPR3
2010 Probabilistic models for supervised dictionary learning
abstract
Dictionary generation is a core technique of the bag-of-visual-words (BOV) models when applied to image categorization. Most of previous approaches generate dictionaries by unsupervised clustering techniques, e.g. k-means. However, the features obtained by such kind of dictionaries may not be optimal for image classification. In this paper, we propose a probabilistic model for supervised dictionary learning (SDLM) which seamlessly combines an unsuper-vised model (a Gaussian Mixture Model) and a supervised model (a logistic regression model) in a probabilistic framework. In the model, image category information directly affects the generation of a dictionary. A dictionary obtained by this approach is a trade-off between minimization of distortions of clusters and maximization of discriminative power of image-wise representations, i.e. histogram representations of images. We further extend the model to incorporate spatial information during the dictionary learning process in a spatial pyramid matching like manner. We extensively evaluated the two models on various benchmark dataset and obtained promising results.
Xiao-Chen Lian, Zhiwei Li 0006, Changhu Wang, Bao-Liang Lu, Lei Zhang 0001
CVPR2
2010 Max-Margin Dictionary Learning for Multiclass Image Categorization
Xiao-Chen Lian, Zhiwei Li 0006, Bao-Liang Lu, Lei Zhang 0001
ECCV (4)2
2010 MindFinder: interactive sketch-based image search on millions of images
abstract
In this paper, we showcase the MindFinder system, which is an interactive sketch-based image search engine. Different from existing work, most of which is limited to a small scale database or only enables single modality input, MindFinder is a sketch-based multimodal search engine for million-level database. It enables users to sketch major curves of the target image in their mind, and also supports tagging and coloring operations to better express their search intentions. Owning to a friendly interface, our system supports multiple actions, which help users to flexibly design their queries. After each operation, top returned images are updated in real time, based on which users could interactively refine their initial thoughts until ideal images are returned. The novelty of the MindFinder system includes the following two aspects: 1) A multimodal searching scheme is proposed to retrieve images which meet users' requirements not only in structure, but also in semantic meaning and color tone. 2) An indexing framework is designed to make MindFinder scalable in terms of database size, memory cost, and response time. By scaling up the database to more than two million images, MindFinder not only helps users to easily present whatever they are imagining, but also has the potential to retrieve the most desired images in their mind.
Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001
ACM Multimedia4
2010 MindFinder: image search by interactive sketching and tagging
abstract
In this technical demonstration, we showcase the MindFinder system $-$ a novel image search engine. Different from existing interactive image search engines, most of which only provide image-level relevance feedback, MindFinder enables users to sketch and tag query images at object level. By considering the image database as a huge repository, MindFinder is able to help users present and refine their initial thoughts in their mind, and finally turn thoughts to a beautiful image(s). Multiple actions are enabled for users to flexibly design their queries in a bilateral interactive manner by leveraging the whole image database, including tagging, refining query by dragging and dropping objects from search results, as well as editing objects. After each action, the search results will be updated in real time to provide users up-to-date materials to further formulate the query. By the deliberate but easy design of the query, MindFinder not only tries to enable users to present on the query panel whatever they are imagining, but also returns to users the most similar images to the picture in users' mind. By scaling up the image database to 10 million, MindFinder has the potential to reveal whatever in users' mind, that is where the name MindFinder comes from.
Changhu Wang, Zhiwei Li 0006, Lei Zhang 0001
WWW2
2009 Efficient indexing for large scale visual search
abstract
With the popularity of “bag of visual terms” representations of images, many text indexing techniques have been applied in large-scale image retrieval systems. However, due to a fundamental difference between an image query (e.g. 1500 visual terms) and a text query (e.g. 3-5 terms), the usages of some text indexing techniques, e.g. inverted list, are misleading. In this work, we develop a novel indexing technique for this problem. The basic idea is to decompose a document-like representation of an image into two components, one for dimension reduction and the other for residual information preservation. The computing of similarity of two images can be transferred to measuring similarities of their components. The decomposition has two major merits: (1) these components have good properties which enable them to be efficiently indexed and retrieved; (2) The decomposition has better generalization ability than other dimension reduction algorithms. The decomposition can be achieved by either a graphical model or a matrix factorization approach. Theoretic analysis and extensive experiments over a 2.3 million image database show that this framework is scalable to index large scale image database to support fast and accurate visual search.
Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma, Harry Shum
ICCV2
2009 LogisticLDA: Regularizing Latent Dirichlet Allocation by Logistic Regression
Jia-Cheng Guo, Bao-Liang Lu, Zhiwei Li 0006, Lei Zhang 0001
PACLIC3
2008 Delivering online advertisements inside images
abstract
We present in this paper a new channel to deliver online advertisements along with Web images and show a new business model to monetize billions of Web images. The idea is intuitively inspired by image displaying processes on the Web, which typically require people to wait a few seconds before they see full resolution images. This is due to large file sizes and limited network bandwidth. To utilize idle time and the display area, we propose an innovative method for non-intrusively embedding ads into images in a visually pleasant manner. To maintain a smooth user experience, we utilize the thumbnail of the full-resolution image because it is small and visually similar to the full-resolution image. At the client side, a rendering engine first enlarges and blurs the thumbnail, and then blends the pre-chosen ads information into the enlarged image. Based on this idea, we propose three typical scenarios that can adopt the proposed image-advertising mode. More importantly, we can encourage providers of images or other users to participate in our online image ads service by tagging or annotating images. We envision revenue sharing with the providers participating in our service, and we expect that a large number of users will actively submit, tag and annotate images using the system. We have implemented a prototype image ads system, and conducted a series of experiments and user studies to evaluate such a new advertisement channel. The experimental results and user studies show that the proposed online image ad delivery is a non-intrusive ads mode, and the proposed solution is practical. This work also opens multiple new research directions ranging from multimedia to web data mining
Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma
ACM Multimedia1
2008 Improving relevance judgment of web search results with image excerpts
abstract
Current web search engines return result pages containing mostly text summary even though the matched web pages may contain informative pictures. A text excerpt (i.e. snippet) is generated by selecting keywords around the matched query terms for each returned page to provide context for user's relevance judgment. However, in many scenarios, we found that the pictures in web pages, if selected properly, could be added into search result pages and provide richer contextual description because a picture is worth a thousand words. Such new summary is named as image excerpts. By well designed user study, we demonstrate image excerpts can help users make much quicker relevance judgment of search results for a wide range of query types. To implement this idea, we propose a practicable approach to automatically generate image excerpts in the result pages by considering the dominance of each picture in each web page and the relevance of the picture to the query. We also outline an efficient way to incorporate image excerpts in web search engines. Web search engines can adopt our approach by slightly modifying their index and inserting a few low cost operations in their workflow. Our experiments on a large web dataset indicate the performance of the proposed approach is very promising.
Zhiwei Li 0006, Shuming Shi 0001, Lei Zhang 0001
WWW1
2007 Improve Ranking by Using Image Information
Shuming Shi 0001, Zhiwei Li 0006, Ji-Rong Wen, Wei-Ying Ma
ECIR3
2007 Image Search Result Clustering and Re-Ranking via Partial Grouping
abstract
Image search result clustering has become an active research topic. However, due to the limitations of current image search engines, the search result always exhibits partial clustering character, which makes the traditional clustering assumption unreasonable. In this paper, we apply Bregman bubble clustering (BBC), which clusters only a fraction of the whole data set, to image search result clustering. We show that relevant and irrelevant images are less mixed in the clusters produced by BBC. Therefore, we are able to incorporate a cluster based relevance feedback scheme to the clustering result and improve the relevance ranking of the search result according to user's feedback. Experiments on animal images from Flickr demonstrate the effectiveness of our clustering and re-ranking algorithms.
Yang Hu 0006, Nenghai Yu, Zhiwei Li 0006, Mingjing Li
ICME3
2007 On Detection of Advertising Images
abstract
Online advertising has enjoyed exponential growth in recent years, and many advertisements appear in the form of images. Although it makes considerable profit, these advertisements tend to disturb the Internet surfing of normal users. Moreover, they always bring extra burden in indexing to commercial image search engines. Therefore, it is necessary to automatically detect those advertising images on the Web. In this paper, a classification based approach is proposed for advertising image detection, in which comprehensive features are exploited and effectively combined. Those features include visual content, link, text and visual layout in hosting Web pages. Promising experimental results are obtained on images collected from about 480 Web sites.
Zhiwei Li 0006, Nenghai Yu, Mingjing Li
ICME3
2007 Image Annotation in a Progressive Way
abstract
Automatic image annotation is crucial for keyword-based image retrieval because it can be used to improve the textual description of images efficiently. For this purpose, many methods have been developed. Due to the restrictions of computational complexity and small training set, the image annotation methods are usually based on the probability of individual word, instead of the joint probability of a set of words. Therefore the correlation between words is omitted. In this paper, we propose a method to approximate the joint probability of words in a progressive way. Given an image, the word with highest probability is first annotated. Then, the successive words are annotated by incorporating the information of previously annotated words. It can be seen as a "greedy" algorithm to calculate the joint probability of multiple words. The experiments show that the proposed progressive annotation method can effectively improve the annotation performance.
Zhiwei Li 0006, Nenghai Yu, Mingjing Li
ICME2
2007 Human behaviour consistent relevance feedback model for image retrieval
abstract
Due to the well known semantic gap, content based image retrieval is a difficult problem. To bridge it, relevance feedback as an effective solution has been extensively studied in literatures. However, existing methods follow a single-line searching philosophy, which may lead to a local optimum in search space. To address the problem, we propose a human behavior consistent relevance feedback model for image retrieval in this paper. Simulating human behaviors, the proposed model enable the user to perform relevance feedback in three manners: Follow up, Go back, and Restart. Each manner is a way for the user to provide the system with his or her opinions about search results. The accumulated feedback information can be used to refine the user query and regulate the similarity metric. We adopt the graph ranking algorithm to model the retrieval process. Experiments conducted on standard Corel dataset and Pascal VOC 2006 dataset demonstrate the effectiveness of the proposed mechanism.
Jing Liu 0001, Zhiwei Li 0006, Mingjing Li, Hanqing Lu, Songde Ma
ACM Multimedia2
2007 Dual cross-media relevance model for image annotation
abstract
Image annotation has been an active research topic in recent years due to its potential impact on both image understanding and web image retrieval. Existing relevance-model-based methods perform image annotation by maximizing the joint probability of images and words, which is calculated by the expectation over training images. However, the semantic gap and the dependence on training data restrict their performance and scalability. In this paper, a dual cross-media relevance model (DCMRM) is proposed for automatic image annotation, which estimates the joint probability by the expectation over words in a pre-defined lexicon. DCMRM involves two kinds of critical relations in image annotation. One is the word-to-image relation and the other is the word-to-word relation. Both relations can be estimated by using search techniques on the web data as well as available training data. Experiments conducted on the Corel dataset and a web image dataset demonstrate the effectiveness of the proposed model.
Jing Liu 0001, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma, Hanqing Lu, Songde Ma
ACM Multimedia4
2007 Bipartite graph reinforcement model for web image annotation
abstract
Automatic image annotation is an effective way for managing and retrieving abundant images on the internet. In this paper, a bipartite graph reinforcement model (BGRM) is proposed for web image annotation. Given a web image, a set of candidate annotations is extracted from its surrounding text and other textual information in the hosting web page. As this set is often incomplete, it is extended to include more potentially relevant annotations by searching and mining a large-scale image database. All candidates are modeled as a bipartite graph. Then a reinforcement algorithm is performed on the bipartite graph to re-rank the candidates. Only those with the highest ranking scores are reserved as the final annotations. Experimental results on real web images demonstrate the effectiveness of the proposed model.
Xiaoguang Rui, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma, Nenghai Yu
ACM Multimedia3
2007 Dual-Space Pyramid Matching for Medical Image Classification
Yang Hu 0006, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma
MMM (1)3
2007 Automatic Refinement of Keyword Annotations for Web Image Search
Zhiwei Li 0006, Mingjing Li
MMM (1)2
2006 Adaptive User Profile Model and Collaborative Filtering for Personalized News
Zhiwei Li 0006, Jinyi Yao, Zengqi Sun, Mingjing Li, Wei-Ying Ma
APWeb2
2006 Ranking Web News Via Homepage Visual Layout and Cross-Site Voting
Jinyi Yao, Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma
ECIR3
2006 Automatic Classification of Photographs and Graphics
abstract
In general, digital images can be classified into photographs and computer graphics. This taxonomy is very useful in many applications, such as Web image search. However, there are no effective methods to perform this classification automatically. In this paper, we manage to solve this problem from two aspects. At first, we propose some novel low-level features that can reveal perceptional differences between photographs and graphics. Then, we adopt an effective algorithm to perform the classification. The experiments conducted on a large-scale image database indicate the effectiveness of our algorithm
Yuanhao Chen, Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma
ICME2
2006 Large-Scale Duplicate Detection for Web Image Search
abstract
Finding visually identical images in large image collections is important for many applications such as intelligence propriety protection and search result presentation. Several algorithms have been reported in the literature, but they are not suitable for large image collections. In this paper, a novel algorithm is proposed to handle the situation, in which each image is compactly represented by a hash code. To detect duplicate images, only the hash codes are required. In addition, a very efficient search method is implemented to quickly group images with similar hash codes for fast detection. The experiments show that our algorithm can be both efficient and effective for duplicate detection in Web image search
Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma
ICME2
2005 Natural Image Retrieval with Sketches
abstract
In this paper, we present a method to retrieve natural images by sketch query. To measure the similarity between the sketch and an image, relevant regions are first located in that image through a multi-resolution search, and a normalized local shape similarity is proposed for image retrieval. Efficiency and other implementation issues are discussed. Experimental results show that it is an effective approach for content-based image retrieval
Jinyi Yao, Mingjing Li, Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma
ICME3
2005 Web object indexing using domain knowledge
abstract
A web object is defined to represent any meaningful object embedded in web pages (e.g. images, music) or pointed to by hyperlinks (e.g. downloadable files). In many cases, users would like to search for information of a certain 'object', rather than a web page containing the query terms. To facilitate web object searching and organizing, in this paper, we propose a novel approach to web object indexing, by discovering its inherent structure information with existed domain knowledge. In our approach, first, Layered LSI spaces are built for a better representation of the hierarchically structured domain knowledge, in order to emphasize the specific semantics and term space in each layer of the domain knowledge. Meanwhile, the web object representation is constructed by hyperlink analysis, and further pruned to remove the noises. Then an optimal matching between the web object and the domain knowledge is performed, in order to pick out the structure attributes of the web object from the knowledge. Finally, the obtained structure attributes are used to re-organize and index the web objects. Our approach also indicates a new promising way to use trust-worthy Deep Web knowledge to help organize dispersive information of Surface Web.
Muyuan Wang, Zhiwei Li 0006, Lie Lu, Wei-Ying Ma, Naiyao Zhang
KDD2
2005 Grouping WWW Image Search Results by Novel Inhomogeneous Clustering Method
abstract
In this paper, a novel inhomogeneous clustering method is proposed for grouping web images. It is used to re-organize the search result of web image search engines into a hierarchical structure so that the users can conveniently browse the search result. This method takes into account various features associated with web images, and treats them in different ways. For the surrounding text extracted from the containing web pages, co-clustering approach is adopted; for low-level features of the image content and other features, one-way clustering approach is adopted. The clustering results of different approaches are combined together to produce the final image groups. Experimental results demonstrate the effectiveness of the proposed method.
Zhiwei Li 0006, Gu Xu, Mingjing Li, Wei-Ying Ma, HongJiang Zhang
MMM1
2005 A probabilistic model for retrospective news event detection
abstract
Retrospective news event detection (RED) is defined as the discovery of previously unidentified events in historical news corpus. Although both the contents and time information of news articles are helpful to RED, most researches focus on the utilization of the contents of news articles. Few research works have been carried out on finding better usages of time information. In this paper, we do some explorations on both directions based on the following two characteristics of news articles. On the one hand, news articles are always aroused by events; on the other hand, similar articles reporting the same event often redundantly appear on many news sources. The former hints a generative model of news articles, and the latter provides data enriched environments to perform RED. With consideration of these characteristics, we propose a probabilistic model to incorporate both content and time information in a unified framework. This model gives new representations of both news articles and news events. Furthermore, based on this approach, we build an interactive RED system, HISCOVERY, which provides additional functions to present events, Photo Story and Chronicle.
Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma
SIGIR1
2004 Hierarchical clustering of WWW image search results using visual, textual and link information
abstract
We consider the problem of clustering Web image search results. Generally, the image search results returned by an image search engine contain multiple topics. Organizing the results into different semantic clusters facilitates users' browsing. In this paper, we propose a hierarchical clustering method using visual, textual and link analysis. By using a vision-based page segmentation algorithm, a web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. By using block-level link analysis techniques, an image graph can be constructed. We then apply spectral techniques to find a Euclidean embedding of the images which respects the graph structure. Thus for each image, we have three kinds of representations, i.e. visual feature based representation, textual feature based representation and graph based representation. Using spectral clustering techniques, we can cluster the search results into different semantic clusters. An image search example illustrates the potential of these techniques.
Deng Cai 0001, Xiaofei He 0001, Zhiwei Li 0006, Wei-Ying Ma, Ji-Rong Wen
ACM Multimedia3
2004 Intuitive and effective interfaces for WWW image search engines
abstract
Web image search engine has become an important tool to organize digital images on the Web. However, most commercial search engines still use a list presentation while little effort has been placed on improving their usability. How to present the image search results in a more intuitive and effective way is still an open question to be carefully studied. In this demo, we present iFind, a scalable Web image search engine, in which we integrated two kinds of search result browsing interfaces. User study results have proved that our interfaces are superior to traditional interfaces.
Zhiwei Li 0006, Xing Xie 0001, Hao Liu 0007, Xiaoou Tang, Mingjing Li, Wei-Ying Ma
ACM Multimedia1