Yan Ke

dblp:13/6836 · DBLP profile ↗
← Back
43ranked-venue papers
22as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 9 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Theory of computation · 3 · 3 first-authorSecurity and privacy · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A reinforcement learning-based spatiotemporal dynamic multi-graph ensemble framework for multi-station air quality prediction
Yufei Qian, Yan Ke
Eng. Appl. Artif. Intell.3
2026 Reversible Data Hiding in Encrypted Images With Dual-Phase Embedding Based on Multi-Key Threshold Decryption
abstract
Existing reversible data hiding methods in encrypted images (RDH-EI) are primarily designed for point-to-point scenarios and not suitable for multiparty communication. To address this issue, a novel RDH-EI scheme based on multi-key (k,n)-threshold decryption (RDH-EITD) is proposed, in which a dual-phase embedding method is designed to support authentication for both the content owner and central server. In RDH-EITD, the private key of Paillier is split intonshares, which are then allocated tondistributed receivers. The image is first encrypted and then embedded with additional data to generate the marked ciphertext, which is uploaded to data server. The server can perform the 2nd-phase embedding. The dual-marked ciphertext is then distributed tonreceivers. Each receiver can extract the 2nd-phase embedded data but cannot decrypt the image individually. Only whenkout ofnreceivers submit their partially decrypted results, can the image be decrypted. Then the 1st-phase embedded data can be extracted and the original image can be recovered losslessly. Security reduction is employed to formally prove the semantic security of RDH-EITD. Experimental results demonstrate that RDH-EITD preserves (k,n)-threshold decryption, enabling resistance againstk− 1 collusion attacks and tolerance ofn−kfailures. The dual-phase embedding achieves embedding rates of λ − 9 bpp (bits per pixel) and 14 bpp respectively with security parameter λ. For an image of lengthLin one-to-ncommunication scenarios, the time complexity of RDH-EITD reachesO(L(λ3+k2)), outperforming existing Paillier-based point-to-point solutions.
Juanli Sun, Yan Ke, Minqing Zhang, Shijun Xiang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Genetic Algorithm-Based Deep Gradient Compression with Layer-Wise Adaptation for Distributed Training
abstract
Distributed deep learning faces a major bottleneck in communication efficiency due to the frequent exchange of gradient information among computing nodes. Gradient compression offers a viable solution to reduce communication overhead. However, existing methods often rely on uniform or limited adaptive strategies, which fail to fully exploit compression potential. To address this, we propose a layer-wise adaptive gradient compression method based on genetic algorithms. By dynamically adjusting compression parameters per layer using evolutionary search and incorporating multi-objective optimization, our approach achieves a better trade-off between compression rate and model accuracy. Experiments across various network architectures and datasets demonstrate that our method significantly improves communication efficiency without sacrificing accuracy.
Yan Ke
SMC2
2025 Multi-grained adaptive informer for multi-step solar irradiance forecasting
Xiang Ma 0004, Jing Huang 0005, Yan Ke
Expert Syst. Appl.4
2025 Energy consumption and carbon emission modeling and forecasting study with novel deep learning methods
Xiang Ma 0004, Jing Huang 0005, Yan Ke
Expert Syst. Appl.4
2025 Two-stage reversible data hiding in encrypted domain with public key embedding mechanism
Yan Ke, Jia Liu 0016, Yiliang Han
Signal Process.1
2025 BladeView: Toward Automatic Wind Turbine Inspection With Unmanned Aerial Vehicle
abstract
This paper presents a fully automatic method, BladeView, for drone-based wind turbine blade inspection using an Unmanned Aerial Vehicle (UAV). With the need for highly efficient blade inspection coupled with the rapid increase of wind turbines, existing methods provide limited automation on wind turbine parameter estimation, full blade coverage, and safety control. We introduce an Automatic Parameter Calculation (APC) algorithm and an Automatic Flight System (AFS) in BladeView to compute wind turbine parameters and inspection paths, respectively. Leveraging triangulation and linear fitting integration techniques, the APC automatically calculates the wind turbine parameters and estimates the relative angle and position between a drone and the turbine. Furthermore, with dynamic path finding and B-spline optimization, the AFS plans a path covering 3 blades within specified flight corridors, in compliance with the turbine parameters obtained from APC. Thus, the proposed BladeView can properly ensure an inspection’s automation, coverage, safety, and smoothness. The efficiency and usability of BladeView are validated through 100,000 flight simulations in the Gazebo simulation environment and 9,239 field runs at various wind farms, including offshore, near-shore, deserts, mountainous areas, farmlands, and suburbs.Note to Practitioners—The proposed BladeView is distinguished in three aspects: (1) It automatically adapts wind turbines with varying geometric properties and physical locations relative to the take-off point. (2) It dramatically improves the quality of collected data with optimal UAV speed and flight corridors. (3) It thoroughly covers all three blades of a Horizontal-Axis Wind Turbine (HAWT), including regions where defects frequently occur. Thus, BladeView is more efficient and robust than existing UAV-based methods for blade inspection, with only around 25 minutes per HAWT. Moreover, it does not require experienced pilots to fly the UAV and manual interventions are rarely needed. Extensive simulation and real-world experiments demonstrate the efficiency and usability of BladeView in various on- and offshore wind farms.
Huan Zhou 0002, Yan Ke, Marcin Grzegorzek, Zeyd Boukhers, John See
IEEE Trans Autom. Sci. Eng.4
2025 Federated Learning With Security Authentication and Traceability of Poisoning by Embedded Message Authentication Code
abstract
Federated learning (FL) allows for collaborative training without centralizing data, but concerns regarding model privacy leakage, intellectual property theft and poisoning attacks have hindered its development. To mitigate such risks, this paper proposes embedded message authentication code technology (EMAC) to integrate encryption, digital signatures, and watermark functions for model security. In EMAC, the authentication data is embedded into the model ciphertext using reversible data hiding after encryption. The marked ciphertext supports data extraction for subsequent authentication and lossless decryption for testing and training simultaneously. Based on EMAC, a novel FL with security authentication and traceability of poisoning (FL-SATP) is proposed, which integrates privacy protection, identity authentication and poisoning traceability into FL. The poisoner tracing is designed to detect and identify poisoners retrospectively based on the practical performance of trained or aggregated models, thus removing the malicious users' model and deterring poisoning behaviors. Theoretical analysis and experimental results demonstrate that FL-SATP could ensure the confidentiality of the model content, the availability of model function, and that when more than half of the users are benign, the proposed method can accurately and efficiently pinpoint all malicious poisoners in FL.
Yan Ke, Minqing Zhang, Jia Liu 0016, Yiliang Han, Wenchao Liu 0002
IEEE Trans. Dependable Secur. Comput.1
2024 Transforming GP-CNN Tree Search Into Trainable Architectures for Image Classification
abstract
Data-efficient image classification poses a challenge in achieving effectiveness with limited data, as evidenced by the current methods based on convolutional neural networks (CNNs) and genetic programming (GP). Existing works employing these two methods encounter limitations, such as a lack of flexibility and an inability to effectively explore the latent features of the data. To tackle these challenges, this paper introduces a genetic programming method for data-efficient image recognition, leveraging novel function sets, terminal sets, and program structures. This method transforms tree-based data structures in GP into trainable CNN architectures. Further, by employing block structures instead of single operations in the search space, the search space is reduced and the stability of the search structures enhanced. Comparative experiments with state-of-the-art neural network methods and GP-based methods on data-efficient classification datasets validate the GP-CNN method offering higher performance.
Yan Ke, Yue-Jiao Gong, Yun Li 0002
SMC2
2024 Skeleton Ground Truth Extraction: Methodology, Annotation Tool and Benchmarks
abstract
Abstract Skeleton Ground Truth (GT) is critical to the success of supervised skeleton extraction methods, especially with the popularity of deep learning techniques. Furthermore, we see skeleton GTs used not only for training skeleton detectors with Convolutional Neural Networks (CNN), but also for evaluating skeleton-related pruning and matching algorithms. However, most existing shape and image datasets suffer from the lack of skeleton GT and inconsistency of GT standards. As a result, it is difficult to evaluate and reproduce CNN-based skeleton detectors and algorithms on a fair basis. In this paper, we present a heuristic strategy for object skeleton GT extraction in binary shapes and natural images. Our strategy is built on an extended theory of diagnosticity hypothesis, which enables encoding human-in-the-loop GT extraction based on clues from the target’s context, simplicity, and completeness. Using this strategy, we developed a tool, SkeView, to generate skeleton GT of 17 existing shape and image datasets. The GTs are then structurally evaluated with representative methods to build viable baselines for fair comparisons. Experiments demonstrate that GTs generated by our strategy yield promising quality with respect to standard consistency, and also provide a balance between simplicity and completeness.
Bipin Indurkhya, John See, Yan Ke, Zeyd Boukhers, Marcin Grzegorzek
Int. J. Comput. Vis.5
2024 Implicit neural representation steganography by neuron pruning
Weina Dong, Jia Liu 0016, Lifeng Chen, Wenquan Sun, Xiaozhong Pan, Yan Ke
Multim. Syst.6
2024 Collaborative Intelligent Delivery With One Truck and Multiple Heterogeneous Drones in COVID-19 Pandemic Environment
abstract
The outbreak of COVID-19 has caused a serious impact on the traditional logistics industry. Considering that the truck-drone collaborative delivery system can both reduce the risk of COVID-19 propagation and deliver supplies in a cost effective and timely manner, this paper introduces the Multiple visits Travelling Salesman Problem with Multiple Heterogeneous Drones (MTSP-MHD). The model allows a truck to carry a fleet of heterogeneous multi-visit drones for cooperative deliveries, where the drones are capable of delivering to multiple customers on a single route and the flight is restricted by energy consumption and payload constraints. To solve MTSP-MHD, we develop an approach that combines K-Means$++$clustering, Nearest neighbor search and Greedy strategies (KNG) to construct feasible solutions. Meanwhile, an Improved Artificial Bee Colony algorithm combining Metropolis acceptance criterion of Simulated Annealing, Tabu list of Tabu Search, and Elite selection strategies (IABC-MTE) is proposed to enhance the quality of solutions. Particularly, three problem-specific neighborhood operators are adopted to search for new solutions. The massive experimental results indicate that IABC-MTE achieves significant improvements over other competitors, with average objective value reductions ranging from 1.81% to 29.16% and standard deviations reduced by 0.04 to 26.44. Finally, the influencing factors of the drone fleet, the performance of different drone fleets and delivery modes are evaluated in detail.
Yiwen Luo, Xiaoheng Deng, Yan Ke, Shaohua Wan 0001, Yurong Qian
IEEE Trans. Intell. Transp. Syst.4
2023 MDCN: Multi-scale Dilated Convolutional Enhanced Residual Network for Traffic Sign Detection
Yan Ke, Wanghao Mo, Ruyi Cao
ADMA (1)1
2023 PVDet: Towards pedestrian and vehicle detection on gigapixel-level images
Wanghao Mo, Hongyang Wei, Ruyi Cao, Yan Ke, Yiwen Luo
Eng. Appl. Artif. Intell.5
2023 Doing More With Moiré Pattern Detection in Digital Photos
abstract
Detecting moiré patterns in digital photographs is meaningful as it provides priors towards image quality evaluation and demoiréing tasks. In this paper, we present a simple yet efficient framework to extract moiré edge maps from images with moiré patterns. The framework includes a strategy for training triplet (natural image, moiré layer, and their synthetic mixture) generation, and a Moiré Pattern Detection Neural Network (MoireDet) for moiré edge map estimation. This strategy ensures consistent pixel-level alignments during training, accommodating characteristics of a diverse set of camera-captured screen images and real-world moiré patterns from natural images. The design of three encoders in MoireDet exploits both high-level contextual and low-level structural features of various moiré patterns. Through comprehensive experiments, we demonstrate the advantages of MoireDet: better identification precision of moiré images on two datasets, and a marked improvement over state-of-the-art demoiréing methods.
Yan Ke, Marcin Grzegorzek, John See
IEEE Trans. Image Process.3
2022 Image steganalysis based on attention augmented convolution
Minqing Zhang, Yan Ke, Xinliang Bi, Yongjun Kong
Multim. Tools Appl.3
2022 A Reversible Data Hiding Scheme in Encrypted Domain for Secret Image Sharing Based on Chinese Remainder Theorem
abstract
Schemes of reversible data hiding in encrypted domain (RDH-ED) based on symmetric or public key encryption are mainly applied in the scenarios of end-to-end communication. To provide security guarantees for the multi-party scenarios, a RDH-ED scheme for secret image sharing based on Chinese remainder theorem (CRT) is presented. In the application of ($t$,$n$) secret image sharing, an image is first shared into$n$different shares of ciphertext. Only when not less than$t$shares obtained, can the image be reconstructed. In our scheme, additional data could be embedded into the image shares. To realize data extraction from the image shares and the reconstructed image separably, two data hiding methods are proposed: one is homomorphic difference expansion in encrypted domain (HDE-ED) that supports data extraction from the reconstructed image by utilizing the addition homomorphism of CRT secret sharing; the other is difference expansion in image shares (DE-IS) that supports the data extraction from the marked shares before image reconstruction. Experimental results demonstrate that the proposed scheme could not only maintain the security and the threshold function of secret sharing system, but also obtain a better reversibility and efficiency compared with most existing RDH-ED algorithms. The maximum embedding rate of HDE-ED could reach 0.500 bits per pixel and the average embedding rate of DE-IS could reach 0.4652 bits per pixel.
Yan Ke, Minqing Zhang, Xinpeng Zhang 0004, Jia Liu 0016, Tingting Su, Xiaoyuan Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2020 PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
Kean Chen, Weiyao Lin, John See, Yan Ke
ECCV (5)6
2020 CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li 0009, Weiyao Lin, John See, Ning Xu 0007, Shugong Xu, Yan Ke
ECCV (16)6
2020 Fully Homomorphic Encryption Encapsulated Difference Expansion for Reversible Data Hiding in Encrypted Domain
abstract
This paper proposes a fully homomorphic encryption encapsulated difference expansion (FHEE-DE) scheme for reversible data hiding in encrypted domain (RDH-ED). The homomorphic circuits and ciphertext operations are elaborated. Key-switching and bootstrapping techniques are introduced to control the ciphertext extension and decryption failure of homomorphic encryption. A key-switching based least-significant-bit (KS-LSB) data hiding method has been designed to realize data extraction directly from the encrypted domain without the private key. In application, the user first encrypts the plaintext and uploads ciphertext to the server. The server embeds additional data into the ciphertext by performing FHEE-DE data hiding and KS-LSB data hiding. Additional data can be extracted directly from the marked ciphertext by the server without the private key. The user owns the private key and can decrypt the marked ciphertext to obtain the marked plaintext. Then additional data or plaintext can be obtained from the marked plaintext by using the standard DE extraction or recovery. The server could also implement FHEE-DE recovery or extraction on the marked ciphertext to return the ciphertext of original plaintext or additional data to the user. Experimental results demonstrate that the embedding capacity and reversibility of the proposed scheme are superior to existing RDH-ED methods, and fully separability is achieved without reducing the security of encryption.
Yan Ke, Minqing Zhang, Jia Liu 0016, Tingting Su, Xiaoyuan Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2019 Generative steganography with Kerckhoffs' principle
Yan Ke, Minqing Zhang, Jia Liu 0016, Tingting Su, Xiaoyuan Yang 0002
Multim. Tools Appl.1
2018 A multilevel reversible data hiding scheme in encrypted domain based on LWE
Yan Ke, Minqing Zhang, Jia Liu 0016, Tingting Su, Xiaoyuan Yang 0002
J. Vis. Commun. Image Represent.1
2016 Separable Multiple Bits Reversible Data Hiding in Encrypted Domain
Yan Ke, Minqing Zhang, Jia Liu 0016
IWDW1
2010 A pattern tree-based approach to learning URL normalization rules
abstract
Duplicate URLs have brought serious troubles to the whole pipeline of a search engine, from crawling, indexing, to result serving. URL normalization is to transform duplicate URLs to a canonical form using a set of rewrite rules. Nowadays URL normalization has attracted significant attention as it is lightweight and can be flexibly integrated into both the online (e.g. crawling) and the offline (e.g. index compression) parts of a search engine. To deal with a large scale of websites, automatic approaches are highly desired to learn rewrite rules for various kinds of duplicate URLs. In this paper, we rethink the problem of URL normalization from a global perspective and propose a pattern tree-based approach, which is remarkably different from existing approaches. Most current approaches learn rewrite rules by iteratively inducing local duplicate pairs to more general forms, and inevitably suffer from noisy training data and are practically inefficient. Given a training set of URLs partitioned into duplicate clusters for a targeted website, we develop a simple yet efficient algorithm to automatically construct a URL pattern tree. With the pattern tree, the statistical information from all the training samples is leveraged to make the learning process more robust and reliable. The learning process is also accelerated as rules are directly summarized based on pattern tree nodes. In addition, from an engineering perspective, the pattern tree helps select deployable rules by removing conflicts and redundancies. An evaluation on more than 70 million duplicate URLs from 200 websites showed that the proposed approach achieves very promising performance, in terms of both de-duping effectiveness and computational efficiency.
Rui Cai 0002, Jiang-Ming Yang, Yan Ke, Xiaodong Fan, Lei Zhang 0001
WWW4
2010 Volumetric Features for Video Event Detection
Yan Ke, Rahul Sukthankar, Martial Hebert
Int. J. Comput. Vis.1
2008 Fast Motion Consistency through Matrix Quantization
abstract
Determining the motion consistency between two video clips is a key component for many applications such as video event detection and human pose estimation. Shechtman and Irani recently proposed a method for measuring the motion consistency between two videos by representing the motion about each point with a space-time Harris matrix of spatial and temporal derivatives. A motion-consistency measure can be accurately estimated without explicitly calculating the optical flow from the videos, which could be noisy. However, the motion consistency calculation is computationally expensive and it must be evaluated between all possible pairs of points between the two videos. We propose a novel quantization method for the space-time Harris matrices that reduces the consistency calculation to a fast table lookup for any arbitrary consistency measure. We demonstrate that for the continuous rank drop consistency measure used by Shechtman and Irani, our quantization method is much faster and achieves the same accuracy as the existing approximation.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert, Yan Ke
BMVC4
2007 Spatio-temporal Shape and Flow Correlation for Action Recognition
abstract
This paper explores the use of volumetric features for action recognition. First, we propose a novel method to correlate spatio-temporal shapes to video clips that have been automatically segmented. Our method works on over-segmented videos, which means that we do not require background subtraction for reliable object segmentation. Next, we discuss and demonstrate the complementary nature of shape- and flow-based features for action recognition. Our method, when combined with a recent flow-based correlation technique, can detect a wide range of actions in video, as demonstrated by results on a long tennis video. Although not specifically designed for whole-video classification, we also show that our method's performance is competitive with current action classification techniques on a standard video classification dataset.
Yan Ke, Rahul Sukthankar, Martial Hebert
CVPR1
2007 Event Detection in Crowded Videos
abstract
Real-world actions occur often in crowded, dynamic environments. This poses a difficult challenge for current approaches to video event detection because it is difficult to segment the actor from the background due to distracting motion from other objects in the scene. We propose a technique for event recognition in crowded videos that reliably identifies actions in the presence of partial occlusion and background clutter. Our approach is based on three key ideas: (1) we efficiently match the volumetric representation of an event against oversegmented spatio-temporal video volumes; (2) we augment our shape-based features using flow; (3) rather than treating an event template as an atomic entity, we separately match by parts (both in space and time), enabling robustness against occlusions and actor variability. Our experiments on human actions, such as picking up a dropped object or waving in a crowd show reliable detection with few false positives.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV1
2006 The Design of High-Level Features for Photo Quality Assessment
abstract
We propose a principled method for designing high level features forphoto quality assessment. Our resulting system can classify between high quality professional photos and low quality snapshots. Instead of using the bag of low-level features approach, we first determine the perceptual factors that distinguish between professional photos and snapshots. Then, we design high level semantic features to measure the perceptual differences. We test our features on a large and diverse dataset and our system is able to achieve a classification rate of 72% on this difficult task. Since our system is able to achieve a precision of over 90% in low recall scenarios, we show excellent results in a web image search application.
Yan Ke, Xiaoou Tang
CVPR (1)1
2005 Computer Vision for Music Identification
abstract
We describe how certain tasks in the audio domain can be effectively addressed using computer vision approaches. This paper focuses on the problem of music identification, where the goal is to reliably identify a song given a few seconds of noisy audio. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem. By employing pairwise boosting on a large set of Viola-Jones features, our system learns compact, discriminative, local descriptors that are amenable to efficient indexing. During the query phase, we retrieve the set of song snippets that locally match the noisy sample and employ geometric verification in conjunction with an EM-based "occlusion" model to identify the song that is most consistent with the observed signal. We have implemented our algorithm in a practical system that can quickly and accurately recognize music from short audio samples in the presence of distortions such as poor recording quality and significant ambient noise. Our experiments demonstrate that this approach significantly outperforms the current state-of-the-art in content-based music identification.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (1)1
2005 Computer Vision for Music Identification: Video Demonstration
abstract
This paper describes a demonstration video for our music identification system. The goal of music identification is to reliably recognize a song from a small sample of noisy audio. This problem is challenging because the recording is often corrupted by noise and because the audio sample will only match a small portion of the target song. Additionally, a practical music identification system should scale (in both accuracy and speed) to databases containing hundreds of thousands of songs. Recently, the music identification problem has attracted considerable attention. However, the task remains unsolved, particularly for noisy real-world queries. We cast music identification into an equivalent sub-image retrieval framework: identify the portion of a spectrogram image from the database that best matches a given query snippet. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (2)1
2005 SOLAR: sound object localization and retrieval in complex audio environments
abstract
The ability to identify sounds in complex audio environments is highly useful for multimedia retrieval, security, and many mobile robotic applications, but very little work has been done in this area. We present the SOLAR system, a system capable of finding sound objects, such as dog barks or car horns, in complex audio data extracted from movies. SOLAR avoids the need for segmentation by scanning over the audio data in fixed increments and classifying each short audio window separately. SOLAR employs boosted decision tree classifiers to select suitable features for modeling each sound object and to discriminate between the object of interest and all other sounds. We demonstrate the effectiveness of our approach with experiments on thirteen sound object classes trained using only tens of positive examples and tested on hours of audio data extracted from popular movies.
Derek Hoiem, Yan Ke, Rahul Sukthankar
ICASSP (5)2
2005 Efficient Visual Event Detection Using Volumetric Features
abstract
This paper studies the use of volumetric features as an alternative to popular local descriptor approaches for event detection in video sequences. Motivated by the recent success of similar ideas in object detection on static images, we generalize the notion of 2D box features to 3D spatio-temporal volumetric features. This general framework enables us to do real-time video analysis. We construct a realtime event detector for each action of interest by learning a cascade of filters based on volumetric features that efficiently scans video sequences in space and time. This event detector recognizes actions that are traditionally problematic for interest point methods - such as smooth motions where insufficient space-time interest points are available. Our experiments demonstrate that the technique accurately detects actions on real-world sequences and is robust to changes in viewpoint, scale and action speed. We also adapt our technique to the related task of human action classification and confirm that it achieves performance comparable to a current interest point based human activity recognizer on a standard database of human activities.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV1
2005 Evaluating keypoint methods for content-based copyright protection of digital images
abstract
This paper evaluates the effectiveness of keypoint methods for content-based protection of digital images. These methods identify a set of "distinctive" regions (termed keypoints) in an image and encode them using descriptors that are robust to expected image transformations. To determine whether particular images were derived from a protected image, the keypoints for both images are generated and their descriptors matched. We describe a comprehensive set of experiments to examine how keypoint methods cope with three real-world challenges: (1) loss of keypoints due to cropping; (2) matching failures caused by approximate nearest-neighbor indexing schemes; (3) degraded descriptors due to significant image distortions. While keypoint methods perform very well in general, this paper identifies cases where the accuracy of such methods degrades.
Larry Huston, Rahul Sukthankar, Yan Ke
ICME3
2004 PCA-SIFT: A More Distinctive Representation for Local Image Descriptors
Yan Ke, Rahul Sukthankar
CVPR (2)1
2004 An efficient parts-based near-duplicate and sub-image retrieval system
abstract
We introduce a system for near-duplicate detection and sub-image retrieval. Such a system is useful for finding copyright violations and detecting forged images. We define near-duplicate as images altered with common transformations such as changing contrast, saturation, scaling, cropping, framing, etc. Our system builds a parts-based representation of images using distinctive local descriptors which give high quality matches even under severe transformations. To cope with the large number of features extracted from the images, we employ locality-sensitive hashing to index the local descriptors. This allows us to make approximate similarity queries that only examine a small fraction of the database. Although locality-sensitive hashing has excellent theoretical performance properties, a standard implementation would still be unacceptably slow for this application. We show that, by optimizing layout and access to the index data on disk, we can efficiently query indices containing millions of keypoints. Our system achieves near-perfect accuracy (100% precision at 99.85% recall) on the tests presented in Meng et al. [16], and consistently strong results on our own, significantly more challenging experiments. Query times are interactive even for collections of thousands of images.
Yan Ke, Rahul Sukthankar, Larry Huston
ACM Multimedia1
2003 IrisNet: An Architecture for Internet-scale Sensing Services
Suman Nath, Amol Deshpande, Yan Ke, Phillip B. Gibbons, Brad Karp, Srinivasan Seshan
VLDB3
2002 Enhancing foreign language tutors - In search of the golden speaker
Katharina Probst, Yan Ke, Maxine Eskénazi
Speech Commun.2
1990 A Journey into the Fourth Dimension
abstract
It is shown that by a simple (one-way) mapping from quaternions to complex numbers, the problem of generating a four-dimensional Mandelbrot set by iteration of a quadratic function in quaternions can be reduced to iteration of the same function in the complex domain, and thus, the function values in 4-D can be obtained by a simple table lookup. The computations are cut down by an order. Simple ways of displaying the fractal without shading and ways of fast ray tracing such a fractal using the table so generated are discussed. Further speedup in ray tracing can be achieved by estimates of a distance of a point from the Mandelbrot set. Animation is a key factor in visualizing 4-D objects. Three types of animation are attempted: translation in 4-D, rotation in 4-D, and fly-through in 3-D.>
Yan Ke, E. S. Panduranga
IEEE Visualization1
1989 An Efficient Algorithm for Link-Distance Problems
abstract
The link distance between two points inside a simple polygon P is defined to be the minimum number of edges required to form a polygonal path inside P that connects the points. A link furthest neighbor of a point p Ε P is a point of P whose link distance is the maximum from p. The link center of P is the collection of points whose link distances to their link furthest neighbors are minimized. We present an Ο(n log n) time and Ο(n) space algorithm for computing the link center of a simple polygon P, where n is the number of vertices of P. This improves the previous Ο(n2) time and space algorithm. Our algorithm essentially sweeps a chord through the polygon and spends Ο(log n) time at each step. We demonstrate that the output of the algorithm, a sequence of sets of chords, is a powerful tool for solving several other link distance problems.
Yan Ke
SCG1
1989 Computing the Kernel of a Point Set in a Polygon (Extended Abstract)
Yan Ke, Joseph O'Rourke
WADS1
1988 Lower Bounds on Moving a Ladder in Two and Three Dimensions
Yan Ke, Joseph O'Rourke
Discret. Comput. Geom.1
1987 Moving a Ladder in Three Dimensions: Upper and Lower Bounds
abstract
This paper summarizes two results in motion planning, the details of which are in two technical reports. The first establishes an Ω(n4) lower bound on moving a ladder (a line segment) in three dimensions in the presence of polyhedral obstacles with a total of n vertices. This bound is established via a complex arrangement of polygons in space that force a ladder to make Ω(n4 distinct moves between particular initial and final positions. The second report establishes an Ο (n6logn) upper bound by exhibiting an algorithm with that time complexity. The algorithm uses the cell decomposition approach pioneered by Schwartz and Sharir. We suspect that the lower bound is closer to the true complexity of the problem.
Yan Ke, Joseph O'Rourke
SCG1