Nenghai Yu

dblp:96/5144 · DBLP profile ↗
← Back
15ranked-venue papers in the field
0as first author
4since 2021 · last 2024
0000-0003-4417-9316ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 5Database Systems & Data Management · 4Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 2Knowledge Engineering, Semantic Web & Information Systems · 2
YearPublicationVenuePosition
2024 A Robust Database Watermarking Scheme That Preserves Statistical Characteristics
abstract
Database watermarking can be used for copyright verification and leakage traceability, effectively protecting the security of the database. However, the existing watermarking schemes commonly embed watermarks by modifying the original data, which changes the statistical characteristics and affects the statistical analysis of the database. Therefore, this paper proposes SCPW, aStatisticalCharacteristicsPreserving robust databaseWatermarking framework. First, we perform a theoretical analysis and propose a data modification scheme maintaining the statistical characteristics unchanged. Then, we establish the correspondence between the data and the watermarks that need to be embedded in it by grouping. Finally, the watermark message is embedded into the database through data verification and modification. Specifically, for data that needs to be watermarked, we first verify whether the potential watermark bits extracted from the data are the same as bits that need to be embedded. If they are the same, we regard this original data, usually a floating point number, as a “good number” and do not modify it. Otherwise, we modify the data until it becomes a “good number” using a data modification scheme that preserves the statistical characteristics proposed by the theoretical analysis. In addition, we also use the genetic algorithm to optimize the grouping results and increase the proportion of “good number”, thereby reducing the proportion of data that needs to be modified and further reducing distortion. To our best knowledge, SCPW is the first watermarking scheme that ensures the preservation of statistical characteristics, and the experimental results also prove its excellent ability to preserve statistical characteristics compared to existing schemes. Moreover, experiments also illustrate that our method is robust against a wide range of attacks. When under deletion attack (deletion rate = 90%), the bit error rate of watermark extraction is only 0.8%, which is more than 12% lower than the current best method.
Zhiwen Ren, Han Fang 0004, Jie Zhang 0073, Zehua Ma, Ronghao Lin, Weiming Zhang 0001, Nenghai Yu
IEEE Trans. Knowl. Data Eng.7
2023 Compressing the Trees of Canonical Binary AIFV Coding
abstract
Canonical binary AIFV coding [1] contains two trees T0and T1. We show the method to compress T0, and the method to compress T1is with a similar way. We provide a new method to store the number of leaves, master nodes and complete internal nodes in each layer and compactly encode the string of numbers according to the specific property between the nodes.
Sian-Jheng Lin, Nenghai Yu
DCC3
2022 Compressing the Tree of Canonical Huffman Coding
abstract
The codebook is important for canonical Huffman coding, which needs to contain the number of leaves in each layer of the canonical Huffman tree and the corresponding symbols. Specifically, as two conventional methods in [1], [2], only the number of leaves in each level of the canonical Huffman tree is needed to store. However, we provide a new method to store the number of internal nodes in each layer and compactly encode the string of numbers according to the specific property between the internal nodes.
Wei Yan 0014, Sian-Jheng Lin, Nenghai Yu
DCC4
2021 CDAE: Color decomposition-based adversarial examples for screen devices
Huanyu Bian, Hao Cui 0004, Kunlin Liu, Hang Zhou 0007, Dongdong Chen 0001, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu
Inf. Sci.8
2017 Sequence Generation with Target Attention
Yingce Xia, Tao Qin 0001, Nenghai Yu, Tie-Yan Liu
ECML/PKDD (1)4
2017 Semi-order preserving encryption
Weiming Zhang 0001, Nenghai Yu
Inf. Sci.3
2017 Large-Scale Online Feature Selection for Ultra-High Dimensional Sparse Data
abstract
Feature selection (FS) is an important technique in machine learning and data mining, especially for large-scale high-dimensional data. Most existing studies have been restricted to batch learning, which is often inefficient and poorly scalable when handling big data in real world. As real data may arrive sequentially and continuously, batch learning has to retrain the model for the new coming data, which is very computationally intensive. Online feature selection (OFS) is a promising new paradigm that is more efficient and scalable than batch learning algorithms. However, existing online algorithms usually fall short in their inferior efficacy. In this article, we present a novel second-order OFS algorithm that is simple yet effective, very fast and extremely scalable to deal with large-scale ultra-high dimensional sparse data streams. The basic idea is to exploit the second-order information to choose the subset of important features with high confidence weights. Unlike existing OFS methods that often suffer from extra high computational cost, we devise a novel algorithm with a MaxHeap-based approach, which is not only more effective than the existing first-order algorithms, but also significantly more efficient and scalable. Our extensive experiments validated that the proposed technique achieves highly competitive accuracy as compared with state-of-the-art batch FS methods, meanwhile it consumes significantly less computational cost that is orders of magnitude lower. Impressively, on a billion-scale synthetic dataset (1-billion dimensions, 1-billion non-zero features, and 1-million samples), the proposed algorithm takes less than 3 minutes to run on a single PC.
Steven C. H. Hoi, Tao Mei 0001, Nenghai Yu
ACM Trans. Knowl. Discov. Data4
2013 FoCUS: Learning to Crawl Web Forums
abstract
In this paper, we present Forum Crawler Under Supervision (FoCUS), a supervised web-scale forum crawler. The goal of FoCUS is to crawl relevant forum content from the web with minimal overhead. Forum threads contain information content that is the target of forum crawlers. Although forums have different layouts or styles and are powered by different forum software packages, they always have similar implicit navigation paths connected by specific URL types to lead users from entry pages to thread pages. Based on this observation, we reduce the web forum crawling problem to a URL-type recognition problem. And we show how to learn accurate and effective regular expression patterns of implicit navigation paths from automatically created training sets using aggregated results from weak page type classifiers. Robust page type classifiers can be trained from as few as five annotated forums and applied to a large set of unseen forums. Our test results show that FoCUS achieved over 98 percent effectiveness and 97 percent coverage on a large set of test forums powered by over 150 different forum software packages. In addition, the results of applying FoCUS on more than 100 community Question and Answer sites and Blog sites demonstrated that the concept of implicit navigation path could apply to other social media sites.
Jingtian Jiang, Xinying Song, Nenghai Yu, Chin-Yew Lin
IEEE Trans. Knowl. Data Eng.3
2012 Learning Bregman Distance Functions for Semi-Supervised Clustering
abstract
Learning distance functions with side information plays a key role in many data mining applications. Conventional distance metric learning approaches often assume that the target distance function is represented in some form of Mahalanobis distance. These approaches usually work well when data are in low dimensionality, but often become computationally expensive or even infeasible when handling high-dimensional data. In this paper, we propose a novel scheme of learning nonlinear distance functions with side information. It aims to learn a Bregman distance function using a nonparametric approach that is similar to Support Vector Machines. We emphasize that the proposed scheme is more general than the conventional approach for distance metric learning, and is able to handle high-dimensional data efficiently. We verify the efficacy of the proposed distance learning method with extensive experiments on semi-supervised clustering. The comparison with state-of-the-art approaches for learning distance functions with side information reveals clear advantages of the proposed technique.
Lei Wu 0017, Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Nenghai Yu
IEEE Trans. Knowl. Data Eng.5
2011 Distance metric learning from uncertain side information for automated photo tagging
abstract
Automated photo tagging is an important technique for many intelligent multimedia information systems, for example, smart photo management system and intelligent digital media library. To attack the challenge, several machine learning techniques have been developed and applied for automated photo tagging. For example, supervised learning techniques have been applied to automated photo tagging by training statistical classifiers from a collection of manually labeled examples. Although the existing approaches work well for small testbeds with relatively small number of annotation words, due to the long-standing challenge of object recognition, they often perform poorly in large-scale problems. Another limitation of the existing approaches is that they require a set of high-quality labeled data, which is not only expensive to collect but also time consuming. In this article, we investigate a social image based annotation scheme by exploiting implicit side information that is available for a large number of social photos from the social web sites. The key challenge of our intelligent annotation scheme is how to learn an effective distance metric based on implicit side information (visual or textual) of social photos. To this end, we present a novel “Probabilistic Distance Metric Learning” (PDML) framework, which can learn optimized metrics by effectively exploiting the implicit side information vastly available on the social web. We apply the proposed technique to photo annotation tasks based on a large social image testbed with over 1 million tagged photos crawled from a social photo sharing portal. Encouraging results show that the proposed technique is effective and promising for social photo based annotation tasks.
Lei Wu 0017, Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Nenghai Yu
ACM Trans. Intell. Syst. Technol.5
2011 A Privacy-Preserving Remote Data Integrity Checking Protocol with Data Dynamics and Public Verifiability
abstract
Remote data integrity checking is a crucial technology in cloud computing. Recently, many works focus on providing data dynamics and/or public verifiability to this type of protocols. Existing protocols can support both features with the help of a third-party auditor. In a previous work, Sebé et al. propose a remote data integrity checking protocol that supports data dynamics. In this paper, we adapt Sebé et al.'s protocol to support public verifiability. The proposed protocol supports public verifiability without help of a third-party auditor. In addition, the proposed protocol does not leak any private information to third-party verifiers. Through a formal analysis, we show the correctness and security of the protocol. After that, through theoretical analysis and experimental results, we demonstrate that the proposed protocol has a good performance.
Zhuo Hao, Sheng Zhong 0002, Nenghai Yu
IEEE Trans. Knowl. Data Eng.3
2010 BioSnowball: automated population of Wikis
abstract
Internet users regularly have the need to find biographies and facts of people of interest. Wikipedia has become the first stop for celebrity biographies and facts. However, Wikipedia can only provide information for celebrities because of its neutral point of view (NPOV) editorial policy. In this paper we propose an integrated bootstrapping framework named BioSnowball to automatically summarize the Web to generate Wikipedia-style pages for any person with a modest web presence. In BioSnowball, biography ranking and fact extraction are performed together in a single integrated training and inference process using Markov Logic Networks (MLNs) as its underlying statistical model. The bootstrapping framework starts with only a small number of seeds and iteratively finds new facts and biographies. As biography paragraphs on the Web are composed of the most important facts, our joint summarization model can improve the accuracy of both fact extraction and biography ranking compared to decoupled methods in the literature. Empirical results on both a small labeled data set and a real Web-scale data set show the effectiveness of BioSnowball. We also empirically show that BioSnowball outperforms the decoupled methods.
Xiaojiang Liu, Zaiqing Nie, Nenghai Yu, Ji-Rong Wen
KDD3
2009 Learning to tag
abstract
Social tagging provides valuable and crucial information for large-scale web image retrieval. It is ontology-free and easy to obtain; however, irrelevant tags frequently appear, and users typically will not tag all semantic objects in the image, which is also called semantic loss. To avoid noises and compensate for the semantic loss, tag recommendation is proposed in literature. However, current recommendation simply ranks the related tags based on the single modality of tag co-occurrence on the whole dataset, which ignores other modalities, such as visual correlation. This paper proposes a multi-modality recommendation based on both tag and visual correlation, and formulates the tag recommendation as a learning problem. Each modality is used to generate a ranking feature, and Rankboost algorithm is applied to learn an optimal combination of these ranking features from different modalities. Experiments on Flickr data demonstrate the effectiveness of this learning-based multi-modality recommendation strategy.
Lei Wu 0017, Linjun Yang, Nenghai Yu, Xian-Sheng Hua 0001
WWW3
2008 Can phrase indexing help to process non-phrase queries?
abstract
Modern web search engines, while indexing billions of web pages, are expected to process queries and return results in a very short time. Many approaches have been proposed for efficiently computing top-k query results, but most of them ignore one key factor in the ranking functions of commercial search engines - term-proximity, which is the metric of the distance between query terms in a document. When term-proximity is included in ranking functions, most of the existing top-k algorithms will become inefficient. To address this problem, in this paper we propose to build a compact phrase index to speed up the search process when incorporating the term-proximity factor. The compact phrase index can help more accurately estimate the score upper bounds of unknown documents. The size of the phrase index is controlled by including a small portion of phrases which are possibly helpful for improving search performance. Phrase index has been used to process phrase queries in existing work. It is, however, to the best of our knowledge, the first time that phrase index is used to improve the performance of generic queries. Experimental results show that, compared with the state-of-the-art top-k computation approaches, our approach can reduce average query processing time to 1/5 for typical setttings.
Mingjie Zhu, Shuming Shi 0001, Nenghai Yu, Ji-Rong Wen
CIKM3
2008 Maximum Margin Clustering with Pairwise Constraints
abstract
Maximum margin clustering (MMC), which extends the theory of support vector machine to unsupervised learning, has been attracting considerable attention recently. The existing approaches mainly focus on reducing the computational complexity of MMC. The accuracy of these methods, however, has not always been guaranteed. In this paper, we propose to incorporate additional side-information, which is in the form of pairwise constraints, into MMC to further improve its performance. A set of pairwise loss functions are introduced into the clustering objective function which effectively penalize the violation of the given constraints. We show that the resulting optimization problem can be easily solved via constrained concave-convex procedure (CCCP). Moreover, for constrained multi-class MMC, we present an efficient cutting-plane algorithm to solve the sub-problem in each iteration of CCCP. The experiments demonstrate that the pairwise constrained MMC algorithms considerably outperform the unconstrained MMC algorithms and two other clustering algorithms that exploit the same type of side-information.
Yang Hu 0006, Jingdong Wang 0001, Nenghai Yu, Xian-Sheng Hua 0001
ICDM3