Cai-Zhi Zhu

dblp:70/4054 · also Caizhi Zhu · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 64% Multimedia systems and quality of experience · 28% Image and video processing · 8%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › image retrieval › object retrieval
image object retrieval
0.212013
Query-Adaptive Asymmetrical Dissimilarities for Visual Object Retrieval · ICCV 2013
Information retrieval
image retrieval
0.212013
Query-Adaptive Asymmetrical Dissimilarities for Visual Object Retrieval · ICCV 2013
Information retrieval
retrieval models
0.212013
Query-Adaptive Asymmetrical Dissimilarities for Visual Object Retrieval · ICCV 2013
Multimedia analysis and retrieval › multimedia browsing
video browsing
0.112005
Natural video browsing · ACM Multimedia 2005
Multimedia systems and quality of experience
video quality assessment
0.112005
Spatio-temporal quality assessment for home videos · ACM Multimedia 2005
Multimedia analysis and retrieval
video summarization
0.112005
Spatio-temporal quality assessment for home videos · ACM Multimedia 2005
Multimedia analysis and retrieval
video retrieval
0.012005
Natural video browsing · ACM Multimedia 2005

Methods — techniques the papers use, named apart from their topics

inverted file · 0.2bag-of-words · 0.2user study · 0.1learning-based quality prediction · 0.1factor fusion · 0.1
YearPublicationVenuePosition
2014 Image Flows Visualization for Inter-media Comparison
abstract
To understand recent societal behavior, it is important to compare how multiple media are affected by real-world events and how each medium affects other media. This paper proposes a novel framework for inter-media comparison through visualizing images extracted from different types of media. We extract blog image clusters from our six-year blog archive and search for similar TV shots in each cluster from a broadcast news video archive by using image similarities. We then visualize such image flows on a timeline in 3D space to visually and interactively explore time sequential changes in influences among media resources and differences and/or similarities between them such as topics that become popular on only blogs or that become popular on blogs earlier than on TV.
Masahiko Itoh, Masashi Toyoda, Cai-Zhi Zhu, Shin'ichi Satoh 0001, Masaru Kitsuregawa
PacificVis3
2014 Multi-image aggregation for better visual object retrieval
abstract
We study how aggregating multiple images, on query or database side, impacts the performance of visual object retrieval in a Bag-of-Words framework. To this end, we first compare five different multi-image aggregation methods, and suggest selecting the average pooling method in most cases for its superior advantages in accuracy, speed, and memory footprint. Then we prove with experiments that more images generally yield better retrieval performance. What is more, we illustrate that simply aggregating query images without selection is far from optimal. Comprehensive experiments were conducted on three large-scale object retrieval datasets, and the new state-of the-art was achieved. This research can be leveraged in some real applications such as mobile search, where the retrieval performance will be improved once users snap multiple query images.
Cai-Zhi Zhu, Yu-Hui Huang, Shin'ichi Satoh 0001
ICASSP1
2014 A practical spatial re-ranking method for instance search from videos
abstract
Spatial re-ranking has proved to be successful in image retrieval. Yet no work on spatial re-ranking has been systematically reported for instance search from videos so far: As videos are composed of multiple frames, and frame-by-frame spatial verification is too prohibitive, till now we lack an efficient spatial re-ranking method designed for the purpose of video instance search. The effectiveness is unclear as well. This paper proposes a practical spatial re-ranking method for video instance search. We make two contributions to speed up the algorithm: One is to select the most representative image/frame for spatial verification, based on the available Bag-of-Words representation; the other is to efficiently build tentative matching feature pairs, which serve as the input of the RANSAC algorithm, by regarding features quantized to the same visual word as matches, thus avoid the costly nearest-neighbor search. These two modifications lead one order of magnitude speedup without compromising the performance. Another contribution is a ROI-originated RANSAC method, which improves the re-ranking performance significantly. Experiments were carried out on the TrecVid Instance Search 2013 dataset, and the new state-of-the-art performance was achieved by the proposed method, at a much faster speed.
Cai-Zhi Zhu, Qiang Zhu 0011, Shin'ichi Satoh 0001, Yu-tang Guo
ICIP2
2014 Tell Me about TV Commercials of This Product
Cai-Zhi Zhu, Siriwat Kasamwattanarote, Xiaomeng Wu, Shin'ichi Satoh 0001
MMM (1)1
2013 Query-Adaptive Asymmetrical Dissimilarities for Visual Object Retrieval
abstract
Visual object retrieval aims at retrieving, from a collection of images, all those in which a given query object appears. It is inherently asymmetric: the query object is mostly included in the database image, while the converse is not necessarily true. However, existing approaches mostly compare the images with symmetrical measures, without considering the different roles of query and database. This paper first measure the extent of asymmetry on large-scale public datasets reflecting this task. Considering the standard bag-of-words representation, we then propose new asymmetrical dissimilarities accounting for the different inlier ratios associated with query and database images. These asymmetrical measures depend on the query, yet they are compatible with an inverted file structure, without noticeably impacting search efficiency. Our experiments show the benefit of our approach, and show that the visual object retrieval task is better treated asymmetrically, in the spirit of state-of-the-art text retrieval.
Cai-Zhi Zhu, Hervé Jégou, Shin'ichi Satoh 0001
ICCV1
2013 Evaluation of visual object retrieval datasets
abstract
Recently visual object retrieval has being widely studied for its vast application prospect, and the progress is usually tracked by the performance achieved on benchmark datasets. Therefore, in order to faithfully evaluate object retrieval algorithms, the quality of datasets must be ensured by some means. In this paper, we propose a method to evaluate the quality of benchmarking visual object retrieval on two highly cited object retrieval datasets, the Oxford datasets and TrecVid instance search datasets. Our evaluation method leverages the essential differences between object retrieval and other similar image search, and digs out some unrevealed and rather interesting features from those datasets. To the best of our knowledge, this research has never been touched before. Our method is believed to be beneficial to dataset collection and algorithm evaluation of object retrieval. More importantly, we hope this work can attract more attentions on this topic in the community.
Cai-Zhi Zhu, Shin'ichi Satoh 0001
ICIP1
2013 Connect commercial films with realities
abstract
Broadcast TV program is a quite informative media resource which records our daily life over the time. While for emphasizing real-time reporting, those out-of-date video archives once were elaborately created with high quality are always left without being fully used. In this paper, many known state-of-the-art retrieval technologies are integrated into a commercial film retrieval system, which manages to index a huge commercial dataset archived from five TV channels within recent three years. The final purpose is to connect images queried by users with our archived broadcast video dataset via searching relevant commercials and accessing their broadcast information, such as air time and replay frequency. This system also serves as one part of our ongoing broadcast TV program reusing project.
Cai-Zhi Zhu, Siriwat Kasamwattanarote, Xiaomeng Wu, Shin'ichi Satoh 0001
ICMR1
2012 Large vocabulary quantization for searching instances from videos
abstract
A very promising application involving video collections is to search for relevant video segments from a video database when given few visual examples of the specific instance, e.g. a person, object, or place. However, this problem is difficult due to the lighting variations, different viewpoints, partial occlusion, and large changes in appearance. In this paper, we focus on a kind of restricted instance searching task, where the region of a specific instance to be searched for is manually labeled on each query image. We formulate this problem in a large vocabulary quantization based Bag-of-Words framework, while putting more research emphasis on investigating to what extent we can benefit from these labeled instance regions. The contribution of this paper mainly lies in two aspects: first, we proposed an algorithm for instance search that outperformed all submissions on the instance search dataset TRECVID 2011. Secondly, after thoroughly analyzing the experiment results, we show that our top performance is mainly due to similar scene retrieval, instead of the same instance search. This observation reveals that in the current dataset background is more dominated than instance, and it also suggests that a promising direction in which to further improve the current algorithm, which may also be the breakthrough for achieving this challenge, is to investigate more about how to truly take advantage of additional labeled instance regions. We believe our research opens a window for future new methods for searching instance.
Cai-Zhi Zhu, Shin'ichi Satoh 0001
ICMR1
2011 Efficient quantization of color sift for image classification
abstract
The local feature (e.g. SIFT) and Bag of Words (BOW) model play key roles in achieving a state-of-the-art performance for image classification. Although we realize that utilizing extra color information will undoubtedly boost the local feature, there still have not been any research that have carefully focused on how to efficiently transfer this color boosted local feature into a boosted BOW. In this paper, a channel-wise separate quantization is newly proposed to replace the traditional multi-channel combined quantization for color SIFT. Dozens of experiments conducted on common data sets have consistently proven that the latter is noticeably worse than the former. We attribute this to a larger quantization error induced by the unreasonable assumption of identical feature distributions in multiple channels. In addition, we compare the distinctness of the intensity and color components separately, and the experimental results are beyond our expectations. Our experiments also show that the accuracy is noticeably enhanced when we take the top N nearest neighbors in soft assignment into account.
Cai-Zhi Zhu, Shin'ichi Satoh 0001, Yu-tang Guo
ICIP2
2007 Home Video Visual Quality Assessment With Spatiotemporal Factors
abstract
Compared with the video programs taken by professionals, home videos are always with low quality content resulted from non-professional capture skills. In this paper, we present a novel spatiotemporal quality assessment scheme in terms of low-level content features for home videos. In contrast to existing frame-level-based quality assessment approaches, a type of temporal segment of video, subshot, is selected as the basic unit for quality assessment. A set of spatiotemporal visual artifacts, regarded as the key factors affecting the overall perceived quality (i.e., unstableness and jerkiness as temporal factors; infidelity, blurring, brightness, and orientation as spatial factors), are mined from each subshot based on particular characteristics of home videos. The relationship between the overall quality metric and these factors are exploited by three different methods, including user study-based, rule-based and learning-based. To validate the proposed scheme, we present a scalable quality-based home video summarization system from a novel perspective-achieving the best visual quality while simultaneously preserving the most informative content. A comparison user study between this system and the attention model-based video skimming approach demonstrated the effectiveness of the proposed quality assessment scheme
Tao Mei 0001, Xian-Sheng Hua 0001, Cai-Zhi Zhu, He-Qin Zhou, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2005 Spatio-temporal quality assessment for home videos
abstract
Compared with the video programs taken by professionals, home videos are always with low-quality content resulted from lack of professional capture skills. In this paper, we present a novel spatio-temporal quality assessment scheme in terms of low-level content features for home videos. In contrast to existing frame-level-based quality assessment approaches, a type of temporal segment of video, sub-shot, is selected as the basic unit for quality assessment. A set of spatio-temporal artifacts, regarded as the key factors affecting the overall perceived quality (i.e. unstableness, jerkiness, infidelity, blurring, brightness and orientation), are mined from each sub-shot based on the particular characteristics of home videos. The relationship between the overall quality metric and these factors are exploited by three different methods, including user study, factor fusion, and a learning-based scheme. To validate the proposed scheme, we present a scalable quality-based home video summarization system, aiming at achieving the best quality while simultaneously preserving the most informative content. A comparison user study between this system and the attention model based video skimming approach demonstrated the effectiveness of the proposed quality assessment scheme.
Tao Mei 0001, Cai-Zhi Zhu, He-Qin Zhou, Xian-Sheng Hua 0001
ACM Multimedia2
2005 Natural video browsing
abstract
In this demonstration, we show a novel system, Video Booklet, which enables nature personal video browsing and searching.
Cai-Zhi Zhu, Tao Mei 0001, Xian-Sheng Hua 0001
ACM Multimedia1