VLDB 2026 Research / reviewers in the wild / expert
Mengyang Yu
dblp:97/11080
· DBLP profile ↗
33ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-1191-429XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11Systems, architecture and hardware · 2Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Representation and self-supervised learning · 18% Deep learning architectures and training · 17% Generative modeling · 14% | |
| Databases, data mining, and information retrieval
8 papers |
Information retrieval · 87% Data mining · 13% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 90% Multimedia analysis and retrieval · 10% | |
| Network and information security
3 papers |
Cryptographic primitives and cryptanalysis · 80% Digital forensics and information hiding · 13% Cryptographic protocols and secure computation · 7% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 82% Algorithms and data structures · 18% |
Topics — the 30 heaviest of 53, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cryptographic primitives and cryptanalysis › encryption
access control encryption |
0.9 | 1 | 2025 | Attribute-Based Access Control Encryption · IEEE Trans. Dependable Secur. Comput. 2025 |
Cryptographic primitives and cryptanalysis › functional encryption
attribute-based encryption |
0.9 | 1 | 2025 | Attribute-Based Access Control Encryption · IEEE Trans. Dependable Secur. Comput. 2025 |
Mathematical optimization › integer programming
binary optimization |
0.6 | 1 | 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Mathematical optimization
convergence analysis |
0.6 | 1 | 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Information retrieval › hashing
binary code learning |
0.5 | 2 | 2017 | Learning Short Binary Codes for Large-scale Image Retrieval · IEEE Trans. Image Process. 2017 Projection Bank: From High-Dimensional Data to Medium-Length Binary Codes · ICCV 2015 |
Information retrieval › hashing
hashing for nearest neighbor search |
0.5 | 2 | 2017 | Learning Short Binary Codes for Large-scale Image Retrieval · IEEE Trans. Image Process. 2017 Multiview Alignment Hashing for Efficient Image Search · IEEE Trans. Image Process. 2015 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
autoencoder representation learning |
0.4 | 1 | 2020 | Auto-Encoding Twin-Bottleneck Hashing · CVPR 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 1 | 2020 | On the Number of Linear Regions of Convolutional Neural Networks · ICML 2020 |
Machine learning › Deep learning architectures and training
neural network expressivity |
0.4 | 1 | 2020 | On the Number of Linear Regions of Convolutional Neural Networks · ICML 2020 |
Information retrieval › hashing
hashing-based retrieval |
0.4 | 1 | 2020 | Auto-Encoding Twin-Bottleneck Hashing · CVPR 2020 |
Information retrieval › hashing
unsupervised hashing |
0.4 | 1 | 2020 | Auto-Encoding Twin-Bottleneck Hashing · CVPR 2020 |
Image and video processing
image enhancement |
0.4 | 1 | 2020 | STAR: A Structure and Texture Aware Retinex Model · IEEE Trans. Image Process. 2020 |
Image and video processing
image restoration |
0.4 | 1 | 2020 | STAR: A Structure and Texture Aware Retinex Model · IEEE Trans. Image Process. 2020 |
Image and video processing › image enhancement
low-light image enhancement |
0.4 | 1 | 2020 | STAR: A Structure and Texture Aware Retinex Model · IEEE Trans. Image Process. 2020 |
Image and video processing › image enhancement
retinex |
0.4 | 1 | 2020 | STAR: A Structure and Texture Aware Retinex Model · IEEE Trans. Image Process. 2020 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.4 | 1 | 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit Test · NeurIPS 2019 |
Machine learning › Generative modeling
energy-based model |
0.4 | 1 | 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit Test · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
goodness-of-fit testing |
0.4 | 1 | 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit Test · NeurIPS 2019 |
Computer vision › Image recognition and object detection › image retrieval
sketch-based image retrieval |
0.3 | 1 | 2018 | Generative Domain-Migration Hashing for Sketch-to-Image Retrieval · ECCV (2) 2018 |
Data mining › predictive modeling › regression
multi-target regression |
0.3 | 1 | 2018 | Multi-Target Regression via Robust Low-Rank Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Data mining
dimensionality reduction |
0.3 | 1 | 2017 | Latent Structure Preserving Hashing · Int. J. Comput. Vis. 2017 |
Information retrieval
hashing |
0.3 | 1 | 2017 | Discretely Coding Semantic Rank Orders for Supervised Image Hashing · CVPR 2017 |
Information retrieval › similarity search
hashing for similarity search |
0.3 | 1 | 2017 | Latent Structure Preserving Hashing · Int. J. Comput. Vis. 2017 |
Information retrieval › image retrieval
large-scale image retrieval |
0.3 | 1 | 2017 | Learning Short Binary Codes for Large-scale Image Retrieval · IEEE Trans. Image Process. 2017 |
Information retrieval › hashing › similarity-preserving hashing
ranking-preserving hashing |
0.3 | 1 | 2017 | Discretely Coding Semantic Rank Orders for Supervised Image Hashing · CVPR 2017 |
Information retrieval
similarity search |
0.3 | 1 | 2017 | Latent Structure Preserving Hashing · Int. J. Comput. Vis. 2017 |
Information retrieval › hashing
supervised hashing |
0.3 | 1 | 2017 | Discretely Coding Semantic Rank Orders for Supervised Image Hashing · CVPR 2017 |
Computer vision › Video understanding and tracking
action recognition |
0.2 | 1 | 2016 | Kernelized Multiview Projection for Robust Action Recognition · Int. J. Comput. Vis. 2016 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2016 | Local Feature Discriminant Projection · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Computer vision › Video understanding and tracking › action recognition › human action recognition
multi-view action recognition |
0.2 | 1 | 2016 | Kernelized Multiview Projection for Robust Action Recognition · Int. J. Comput. Vis. 2016 |
Methods — techniques the papers use, named apart from their topics
alternating optimization · 2.0matrix perturbation · 1.1linear secret sharing · 0.9ciphertext-policy attribute-based encryption · 0.9twin bottleneck · 0.9gradient descent · 0.9autoencoder · 0.9logistic regression · 0.5vectorized least squares regression · 0.4linear region analysis · 0.4exponential filters · 0.4ReLU networks · 0.4goodness-of-fit test · 0.4energy-based model · 0.4matrix elastic nets · 0.3low-rank learning · 0.3kernel trick · 0.3generative domain-migration hashing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attribute-Based Access Control EncryptionabstractThe burgeoning complexity of communication necessitates a high demand for security. Access control encryption is a promising primitive to meet the security demand but the bulk of its constructions rely on formulating the access control policy with identities. Attribute-based access control policy in attribute-based encryption (ABE) is known to be more expressive without relying on enumerating identities. We propose a generic framework to build attribute-based access control encryption from ciphertext-policy ABE. Our instantiations prioritize different emphases on expressiveness and efficiency. The first instantiation supports multi-valued AND-gate access control structures, while the second supports the linear-secret-sharing access structure. Both are prototyped with efficiency validated empirically. Xiuhua Wang 0009, Mengyang Yu, Yinjia Pi, Peng Xu 0003, Shuai Wang 0033, Hai Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and ApplicationsabstractBinary optimization problems (BOPs) arise naturally in many fields, such as information retrieval, computer vision, and machine learning. Most existing binary optimization methods either use continuous relaxation which can cause large quantization errors, or incorporate a highly specific algorithm that can only be used for particular loss functions. To overcome these difficulties, we propose a novel generalized optimization method, named Alternating Binary Matrix Optimization (ABMO), for solving BOPs. ABMO can handle BOPs with/without orthogonality or linear constraints for a large class of loss functions. ABMO involves rewriting the binary, orthogonality and linear constraints for BOPs as an intersection of two closed sets, then iteratively dividing the original problems into several small optimization problems that can be solved as closed forms. To provide a strict theoretical convergence analysis, we add a sufficiently small perturbation and translate the original problem to an approximated problem whose feasible set is continuous. We not only provide rigorous mathematical proof for the convergence to a stationary and feasible point, but also derive the convergence rate of the proposed algorithm. The promising results obtained from four binary optimization tasks validate the superiority and the generality of ABMO compared with the state-of-the-art methods. Huan Xiong, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Fumin Shen, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Scaled Simplex Representation for Subspace ClusteringabstractThe self-expressive property of data points, that is, each data point can be linearly represented by the other data points in the same subspace, has proven effective in leading subspace clustering (SC) methods. Most self-expressive methods usually construct a feasible affinity matrix from a coefficient matrix, obtained by solving an optimization problem. However, the negative entries in the coefficient matrix are forced to be positive when constructing the affinity matrix via exponentiation, absolute symmetrization, or squaring operations. This consequently damages the inherent correlations among the data. Besides, the affine constraint used in these methods is not flexible enough for practical applications. To overcome these problems, in this article, we introduce a scaled simplex representation (SSR) for the SC problem. Specifically, the non-negative constraint is used to make the coefficient matrix physically meaningful, and the coefficient vector is constrained to be summed up to a scalar to make it more discriminative. The proposed SSR-based SC (SSRSC) model is reformulated as a linear equality-constrained problem, which is solved efficiently under the alternating direction method of multipliers framework. Experiments on benchmark datasets demonstrate that the proposed SSRSC algorithm is very efficient and outperforms the state-of-the-art SC methods on accuracy. The code can be found at https://github.com/csjunxu/SSRSC. Jun Xu 0019, Mengyang Yu, Ling Shao 0001, Wangmeng Zuo, Deyu Meng, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Auto-Encoding Twin-Bottleneck HashingabstractConventional unsupervised hashing methods usually take advantage of similarity graphs, which are either pre-computed in the high-dimensional space or obtained from random anchor points. On the one hand, existing methods uncouple the procedures of hash function learning and graph construction. On the other hand, graphs empirically built upon original data could introduce biased prior knowledge of data relevance, leading to sub-optimal retrieval performance. In this paper, we tackle the above problems by proposing an efficient and adaptive code-driven graph, which is updated by decoding in the context of an auto-encoder. Specifically, we introduce into our framework twin bottlenecks (i.e., latent variables) that exchange crucial information collaboratively. One bottleneck (i.e., binary codes) conveys the high-level intrinsic data structure captured by the code-driven graph to the other (i.e., continuous variables for low-level detail information), which in turn propagates the updated network feedback for the encoder to learn more discriminative binary codes. The auto-encoding learning objective literally rewards the code-driven graph to learn an optimal encoder. Moreover, the proposed model can be simply optimized by gradient descent without violating the binary constraints. Experiments on benchmarked datasets clearly show the superiority of our framework over the state-of-the-art hashing methods. Our source code can be found at https://github.com/ymcidence/TBH. Yuming Shen, Jie Qin 0004, Jiaxin Chen 0002, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Ling Shao 0001 |
CVPR | 4 |
| 2020 | On the Number of Linear Regions of Convolutional Neural NetworksabstractOne fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a ReLU NN can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of CNNs, and use them to derive the maximal and average numbers of linear regions for one-layer ReLU CNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer ReLU CNNs. Our results suggest that deeper CNNs have more powerful expressivity than their shallow counterparts, while CNNs have more expressivity than fully-connected NNs per parameter. Huan Xiong, Lei Huang 0015, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICML | 3 |
| 2020 | Deep quantization generative networks
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Lei Huang 0015, Mengyang Yu, Heng Tao Shen, Ling Shao 0001 |
Pattern Recognit. | 6 |
| 2020 | Heterogenous output regression network for direct face alignment
Xiantong Zhen, Mengyang Yu, Zehao Xiao, Lei Zhang 0093, Ling Shao 0001 |
Pattern Recognit. | 2 |
| 2020 | STAR: A Structure and Texture Aware Retinex ModelabstractRetinex theory is developed mainly to decompose an image into the illumination and reflectance components by analyzing local image derivatives. In this theory, larger derivatives are attributed to the changes in reflectance, while smaller derivatives are emerged in the smooth illumination. In this paper, we utilize exponentiated local derivatives (with an exponent γ) of an observed image to generate its structure map and texture map. The structure map is produced by been amplified with γ > 1, while the texture map is generated by been shrank with γ < 1. To this end, we design exponential filters for the local derivatives, and present their capability on extracting accurate structure and texture maps, influenced by the choices of exponents γ. The extracted structure and texture maps are employed to regularize the illumination and reflectance components in Retinex decomposition. A novel Structure and Texture Aware Retinex (STAR) model is further proposed for illumination and reflectance decomposition of a single image. We solve the STAR model by an alternating optimization algorithm. Each sub-problem is transformed into a vectorized least squares regression, with closed-form solutions. Comprehensive experiments on commonly tested datasets demonstrate that, the proposed STAR model produce better quantitative and qualitative performance than previous competing methods, on illumination and reflectance decomposition, low-light image enhancement, and color correction. The code is publicly available at https://github.com/csjunxu/STAR. Jun Xu 0019, Yingkun Hou, Dongwei Ren, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Haoqian Wang, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit TestabstractLearning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be interpreted as a GOF game between two generative networks, where one explicit generative network learns an energy-based distribution that fits the real data, and the other implicit generative network is trained by minimizing a GOF test statistic between the energy-based distribution and the generated data, such that the underlying distribution of the generated data is close to the energy-based distribution. We design a two-level alternative optimization procedure to train the explicit and implicit generative networks, such that the hyper-parameters can also be automatically learned. Experimental results show that DEAN achieves high quality generations compared to the state-of-the-art approaches. Lizhong Ding 0001, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Yong Liu 0018, Yu Li 0006, Ling Shao 0001 |
NeurIPS | 2 |
| 2018 | Generative Domain-Migration Hashing for Sketch-to-Image Retrieval
Jingyi Zhang 0005, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Ling Shao 0001, Heng Tao Shen, Luc Van Gool |
ECCV (2) | 5 |
| 2018 | DTRM: A new reputation mechanism to enhance data trustworthiness for high-performance cloud computingabstractCloud computing and the mobile Internet have been the two most influential information technology revolutions, which intersect in mobile cloud computing (MCC). The burgeoning MCC enables the large-scale collection and processing of big data, which demand trusted, authentic, and accurate data to ensure an important but often overlooked aspect of big data — data veracity. Troublesome internal attacks launched by internal malicious users is one key problem that reduces data veracity and remains difficult to handle. To enhance data veracity and thus improve the performance of big data computing in MCC, this paper proposes a Data Trustworthiness enhanced Reputation Mechanism (DTRM) which can be used to defend against internal attacks. In the DTRM, the sensitivity-level based data category, Metagraph theory based user group division, and reputation transferring methods are integrated into the reputation query and evaluation process. The extensive simulation results based on real datasets show that the DTRM outperforms existing classic reputation mechanisms under bad mouthing attacks and mobile attacks. Hui Lin 0007, Jia Hu 0001, Chuanfeng Xu, Jianfeng Ma 0001, Mengyang Yu |
Future Gener. Comput. Syst. | 5 |
| 2018 | Multi-Target Regression via Robust Low-Rank LearningabstractMulti-target regression has recently regained great popularity due to its capability of simultaneously learning multiple relevant regression tasks and its wide applications in data mining, computer vision and medical image analysis, while great challenges arise from jointly handling inter-target correlations and input-output relationships. In this paper, we propose Multi-layer Multi-target Regression (MMR) which enables simultaneously modeling intrinsic inter-target correlations and nonlinear input-output relationships in a general framework via robust low-rank learning. Specifically, the MMR can explicitly encode inter-target correlations in a structure matrix by matrix elastic nets (MEN); the MMR can work in conjunction with the kernel trick to effectively disentangle highly complex nonlinear input-output relationships; the MMR can be efficiently solved by a new alternating optimization algorithm with guaranteed convergence. The MMR leverages the strength of kernel methods for nonlinear feature learning and the structural advantage of multi-layer learning architectures for inter-target correlation modeling. More importantly, it offers a new multi-layer learning paradigm for multi-target regression which is endowed with high generality, flexibility and expressive ability. Extensive experimental evaluation on 18 diverse real-world datasets demonstrates that our MMR can achieve consistently high performance and outperforms representative state-of-the-art algorithms, which shows its great effectiveness and generality for multivariate prediction. Xiantong Zhen, Mengyang Yu, Xiaofei He 0001, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Multitarget Sparse Latent RegressionabstractMultitarget regression has recently generated intensive popularity due to its ability to simultaneously solve multiple regression tasks with improved performance, while great challenges stem from jointly exploring inter-target correlations and input-output relationships. In this paper, we propose multitarget sparse latent regression (MSLR) to simultaneously model intrinsic intertarget correlations and complex nonlinear input-output relationships in one single framework. By deploying a structure matrix, the MSLR accomplishes a latent variable model which is able to explicitly encode intertarget correlations via -norm-based sparse learning; the MSLR naturally admits a representer theorem for kernel extension, which enables it to flexibly handle highly complex nonlinear input-output relationships; the MSLR can be solved efficiently by an alternating optimization algorithm with guaranteed convergence, which ensures efficient multitarget regression. Extensive experimental evaluation on both synthetic data and six greatly diverse real-world data sets shows that the proposed MSLR consistently outperforms the state-of-the-art algorithms, which demonstrates its great effectiveness for multivariate prediction. Xiantong Zhen, Mengyang Yu, Feng Zheng 0001, Ilanit Ben Nachum, Mousumi Bhaduri, David T. Laidley, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Discretely Coding Semantic Rank Orders for Supervised Image HashingabstractLearning to hash has been recognized to accomplish highly efficient storage and retrieval for large-scale visual data. Particularly, ranking-based hashing techniques have recently attracted broad research attention because ranking accuracy among the retrieved data is well explored and their objective is more applicable to realistic search tasks. However, directly optimizing discrete hash codes without continuous-relaxations on a nonlinear ranking objective is infeasible by either traditional optimization methods or even recent discrete hashing algorithms. To address this challenging issue, in this paper, we introduce a novel supervised hashing method, dubbed Discrete Semantic Ranking Hashing (DSeRH), which aims to directly embed semantic rank orders into binary codes. In DSeRH, a generalized Adaptive Discrete Minimization (ADM) approach is proposed to discretely optimize binary codes with the quadratic nonlinear ranking objective in an iterative manner and is guaranteed to converge quickly. Additionally, instead of using 0/1 independent labels to form rank orders as in previous works, we generate the listwise rank orders from the high-level semantic word embeddings which can quantitatively capture the intrinsic correlation between different categories. We evaluate our DSeRH, coupled with both linear and deep convolutional neural network (CNN) hash functions, on three image datasets, i.e., CIFAR-10, SUN397 and ImageNet100, and the results manifest that DSeRH can outperform the state-of-the-art ranking-based hashing methods. Li Liu 0004, Ling Shao 0001, Fumin Shen, Mengyang Yu |
CVPR | 4 |
| 2017 | Fast action retrieval from videos via feature disaggregation
Jie Qin 0004, Li Liu 0004, Mengyang Yu, Yunhong Wang 0001, Ling Shao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | Latent Structure Preserving HashingabstractAiming at efficient similarity search, hash functions are designed to embed high-dimensional feature descriptors to low-dimensional binary codes such that similar descriptors will lead to binary codes with a short distance in the Hamming space. It is critical to effectively maintain the intrinsic structure and preserve the original information of data in a hashing algorithm. In this paper, we propose a novel hashing algorithm called Latent Structure Preserving Hashing (LSPH), with the target of finding a well-structured low-dimensional data representation from the original high-dimensional data through a novel objective function based on Nonnegative Matrix Factorization (NMF) with their corresponding Kullback-Leibler divergence of data distribution as the regularization term. Via exploiting the joint probabilistic distribution of data, LSPH can automatically learn the latent information and successfully preserve the structure of high-dimensional data. To further achieve robust performance with complex and nonlinear data, in this paper, we also contribute a more generalized multi-layer LSPH (ML-LSPH) framework, in which hierarchical representations can be effectively learned by a multiplicative up-propagation algorithm. Once obtaining the latent representations, the hash functions can be easily acquired through multi-variable logistic regression. Experimental results on three large-scale retrieval datasets, i.e., SIFT 1M, GIST 1M and 500 K TinyImage, show that ML-LSPH can achieve better performance than the single-layer LSPH and both of them outperform existing hashing techniques on large-scale data. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
Int. J. Comput. Vis. | 2 |
| 2017 | Learning Short Binary Codes for Large-scale Image RetrievalabstractLarge-scale visual information retrieval has become an active research area in this big data era. Recently, hashing/binary coding algorithms prove to be effective for scalable retrieval applications. Most existing hashing methods require relatively long binary codes (i.e., over hundreds of bits, sometimes even thousands of bits) to achieve reasonable retrieval accuracies. However, for some realistic and unique applications, such as on wearable or mobile devices, only short binary codes can be used for efficient image retrieval due to the limitation of computational resources or bandwidth on these devices. In this paper, we propose a novel unsupervised hashing approach called min-cost ranking (MCR) specifically for learning powerful short binary codes (i.e., usually the code length shorter than 100 b) for scalable image retrieval tasks. By exploring the discriminative ability of each dimension of data, MCR can generate one bit binary code for each dimension and simultaneously rank the discriminative separability of each bit according to the proposed cost function. Only top-ranked bits with minimum cost-values are then selected and grouped together to compose the final salient binary codes. Extensive experimental results on large-scale retrieval demonstrate that MCR can achieve comparative performance as the state-of-the-art hashing algorithms but with significantly shorter codes, leading to much faster large-scale retrieval. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Binary Set Embedding for Cross-Modal RetrievalabstractCross-modal retrieval is such a challenging topic that traditional global representations would fail to bridge the semantic gap between images and texts to a satisfactory level. Using local features from images and words from documents directly can be more robust for the scenario with large intraclass variations and small interclass discrepancies. In this paper, we propose a novel unsupervised binary coding algorithm called binary set embedding (BSE) to obtain meaningful hash codes for local features from the image domain and words from text domain. Understanding image features with the word vectors learned from the human language instead of the provided documents from data sets, BSE can map samples into a common Hamming space effectively and efficiently where each sample is represented by the sets of local feature descriptors from image and text domains. In particular, BSE explores relationship among local features in both feature level and image (text) level, which can balance the sensitivity of each other. Furthermore, a recursive orthogonalization procedure is applied to reduce the redundancy of codes. Extensive experiments demonstrate the superior performance of BSE compared with state-of-the-art cross-modal hashing methods using either image or text queries. Mengyang Yu, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Descriptor Learning via Supervised Manifold Regularization for Multioutput RegressionabstractMultioutput regression has recently shown great ability to solve challenging problems in both computer vision and medical image analysis. However, due to the huge image variability and ambiguity, it is fundamentally challenging to handle the highly complex input-target relationship of multioutput regression, especially with indiscriminate high-dimensional representations. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm for multioutput regression, which can establish discriminative and compact feature representations to improve the multivariate estimation performance. The SDL is formulated as generalized low-rank approximations of matrices with a supervised manifold regularization. The SDL is able to simultaneously extract discriminative features closely related to multivariate targets and remove irrelevant and redundant information by transforming raw features into a new low-dimensional space aligned to targets. The achieved discriminative while compact descriptor largely reduces the variability and ambiguity for multioutput regression, which enables more accurate and efficient multivariate estimation. We conduct extensive evaluation of the proposed SDL on both synthetic data and real-world multioutput regression tasks for both computer vision and medical image analysis. Experimental results have shown that the proposed SDL can achieve high multivariate estimation accuracy on all tasks and largely outperforms the algorithms in the state of the arts. Our method establishes a novel SDL framework for multioutput regression, which can be widely used to boost the performance in different applications. Xiantong Zhen, Mengyang Yu, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Kernelized Multiview Projection for Robust Action RecognitionabstractConventional action recognition algorithms adopt a single type of feature or a simple concatenation of multiple features. In this paper, we propose to better fuse and embed different feature representations for action recognition using a novel spectral coding algorithm called Kernelized Multiview Projection (KMP). Computing the kernel matrices from different features/views via time-sequential distance learning, KMP can encode different features with different weights to achieve a low-dimensional and semantically meaningful subspace where the distribution of each view is sufficiently smooth and discriminative. More crucially, KMP is linear for the reproducing kernel Hilbert space, which allows it to be competent for various practical applications. We demonstrate KMP’s performance for action recognition on five popular action datasets and the results are consistently superior to state-of-the-art techniques. Ling Shao 0001, Li Liu 0004, Mengyang Yu |
Int. J. Comput. Vis. | 3 |
| 2016 | Structure-Preserving Binary Representations for RGB-D Action RecognitionabstractIn this paper, we propose a novel binary local representation for RGB-D video data fusion with a structure-preserving projection. Our contribution consists of two aspects. Toacquire a general feature for the video data, we convert the problem to describing the gradient fields of RGB and depth information of video sequences. With the local fluxes of the gradient fields, which include the orientation and the magnitude of the neighborhood of each point, a new kind of continuous local descriptor called Local Flux Feature(LFF) is obtained. Then the LFFs from RGB and depth channels are fused into a Hamming space via the Structure Preserving Projection (SPP). Specifically, an orthogonal projection matrix is applied to preserve the pairwise structure with a shape constraint to avoid the collapse of data structure in the projected space. Furthermore, a bipartite graph structure of data is taken into consideration, which is regarded as a higher level connection between samples and classes than the pairwise structure of local features. Theextensive experiments show not only the high efficiency of binary codes and the effectiveness of combining LFFs from RGB-D channels via SPP on various action recognition benchmarks of RGB-D data, but also the potential power of LFF for general action recognition. Mengyang Yu, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Local Feature Discriminant ProjectionabstractIn this paper, we propose a novel subspace learning algorithm called Local Feature Discriminant Projection (LFDP) for supervised dimensionality reduction of local features. LFDP is able to efficiently seek a subspace to improve the discriminability of local features for classification. We make three novel contributions. First, the proposed LFDP is a general supervised subspace learning algorithm which provides an efficient way for dimensionality reduction of large-scale local feature descriptors. Second, we introduce the Differential Scatter Discriminant Criterion (DSDC) to the subspace learning of local feature descriptors which avoids the matrix singularity problem. Third, we propose a generalized orthogonalization method to impose on projections, leading to a more compact and less redundant subspace. Extensive experimental validation on three benchmark datasets including UIUC-Sports, Scene-15 and MIT Indoor demonstrates that the proposed LFDP outperforms other dimensionality reduction methods and achieves state-of-the-art performance for image classification. Mengyang Yu, Ling Shao 0001, Xiantong Zhen, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Unsupervised Local Feature Hashing for Image Similarity SearchabstractThe potential value of hashing techniques has led to it becoming one of the most active research areas in computer vision and multimedia. However, most existing hashing methods for image search and retrieval are based on global feature representations, which are susceptible to image variations such as viewpoint changes and background cluttering. Traditional global representations gather local features directly to output a single vector without the analysis of the intrinsic geometric property of local features. In this paper, we propose a novel unsupervised hashing method called unsupervised bilinear local hashing (UBLH) for projecting local feature descriptors from a high-dimensional feature space to a lower-dimensional Hamming space via compact bilinear projections rather than a single large projection matrix. UBLH takes the matrix expression of local features as input and preserves the feature-to-feature and image-to-image structures of local features simultaneously. Experimental results on challenging data sets including Caltech-256, SUN397, and Flickr 1M demonstrate the superiority of UBLH compared with state-of-the-art hashing methods. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Local Feature Binary Coding for Approximate Nearest Neighbor SearchabstractThe potential value of hashing techniques has led to it becoming one of the most active research areas in computer vision and multimedia. However, most existing hashing methods for image search and retrieval are based on global representations, e.g., GIST, which lack the analysis of the intrinsic geometric property of local features and heavily limit the effectiveness of the hash code. In this paper, we propose a novel supervised hashing method called Local Feature Binary Coding (LFBC) for projecting local feature descriptors from a high-dimensional feature space to a lower-dimensional Hamming space via compact bilinear projections rather than a single large projection matrix. LFBC takes the matrix expression of local features as input and preserves the feature-to-feature and image-to-class structures simultaneously. Experimental results on challenging datasets including Caltech-256, SUN397 and NUS-WIDE demonstrate the superiority of LFBC compared with state-of-the-art hashing methods. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
BMVC | 2 |
| 2015 | Latent Structure Preserving Hashing
Ziyun Cai, Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
BMVC | 3 |
| 2015 | Fast Action Retrieval from Videos via Feature DisaggregationabstractLearning based hashing methods, which aim at learning similarity-preserving binary codes for efficient nearest neighbor search, have been actively studied recently. A majority of the approaches address hashing problems for image collections. However, due to the extra temporal information, videos are usually represented by much higher dimensional (thousands or even more) features compared with images, causing high computational complexity for conventional hashing schemes. In this paper, we propose a simple and efficient hashing scheme for high-dimensional video data. This method, called Disaggregation Hashing, exploits the correlations among different feature dimensions. An intuitive feature disaggregation method is first proposed, followed by a novel hashing algorithm based on different feature clusters. We demonstrate the efficiency and effectiveness of our method by theoretical analysis and exploring its application on action retrieval from video databases. Extensive experiments show the superiority of our binary coding scheme over state-of-the-art hashing methods. Jie Qin 0004, Li Liu 0004, Mengyang Yu, Yunhong Wang 0001, Ling Shao 0001 |
BMVC | 3 |
| 2015 | Supervised descriptor learning for multi-output regressionabstractDescriptor learning has recently drawn increasing attention in computer vision, Existing algorithms are mainly developed for classification rather than for regression which however has recently emerged as a powerful tool to solve a broad range of problems, e.g., head pose estimation. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm to establish a discriminative and compact feature representation for multi-output regression. By formulating as generalized low-rank approximations of matrices with a supervised manifold regularization (SMR), the SDL removes irrelevant and redundant information from raw features by transforming into a low-dimensional space under the supervision of multivariate targets. The obtained discriminative while compact descriptor largely reduces the variability and ambiguity in multi-output regression, and therefore enables more accurate and efficient multivariate estimation. We demonstrate the effectiveness of the proposed SDL algorithm on a representative multi-output regression task: head pose estimation using the benchmark Pointing'04 dataset. Experimental results show that the SDL can achieve high pose estimation accuracy and significantly outperforms state-of-the-art algorithms by an error reduction up to 27.5%. The proposed SDL algorithm provides a general descriptor learning framework in a supervised way for multi-output regression which can largely boost the performance of existing multi-output regression tasks. Xiantong Zhen, Zhijie Wang 0003, Mengyang Yu, Shuo Li 0001 |
CVPR | 3 |
| 2015 | Projection Bank: From High-Dimensional Data to Medium-Length Binary CodesabstractRecently, very high-dimensional feature representations, e.g., Fisher Vector, have achieved excellent performance for visual recognition and retrieval. However, these lengthy representations always cause extremely heavy computational and storage costs and even become unfeasible in some large-scale applications. A few existing techniques can transfer very high-dimensional data into binary codes, but they still require the reduced code length to be relatively long to maintain acceptable accuracies. To target a better balance between computational efficiency and accuracies, in this paper, we propose a novel embedding method called Binary Projection Bank (BPB), which can effectively reduce the very high-dimensional representations to medium-dimensional binary codes without sacrificing accuracies. Instead of using conventional single linear or bilinear projections, the proposed method learns a bank of small projections via the max-margin constraint to optimally preserve the intrinsic data similarity. We have systematically evaluated the proposed method on three datasets: Flickr 1M, ILSVR2010 and UCF101, showing competitive retrieval and recognition accuracies compared with state-of-the-art approaches, but with a significantly smaller memory footprint and lower coding complexity. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
ICCV | 2 |
| 2015 | Multiview Alignment Hashing for Efficient Image SearchabstractHashing is a popular and efficient method for nearest neighbor search in large-scale data spaces by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. For most hashing methods, the performance of retrieval heavily depends on the choice of the high-dimensional feature descriptor. Furthermore, a single type of feature cannot be descriptive enough for different images when it is used for hashing. Thus, how to combine multiple representations for learning effective hashing functions is an imminent task. In this paper, we present a novel unsupervised multiview alignment hashing approach based on regularized kernel nonnegative matrix factorization, which can find a compact representation uncovering the hidden semantics and simultaneously respecting the joint probability distribution of data. In particular, we aim to seek a matrix factorization to effectively fuse the multiple information sources meanwhile discarding the feature redundancy. Since the raised problem is regarded as nonconvex and discrete, our objective function is then optimized via an alternate way with relaxation and converges to a locally optimal solution. After finding the low-dimensional representation, the hashing functions are finally obtained through multivariable logistic regression. The proposed method is systematically evaluated on three data sets: 1) Caltech-256; 2) CIFAR-10; and 3) CIFAR-20, and the results show that our method significantly outperforms the state-of-the-art multiview hashing techniques. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Cross-Modality Submodular Dictionary Learning for Information RetrievalabstractThis paper addresses the problem of joint modeling of multimedia components in different media forms. We consider the information retrieval task across both text and image documents, which includes retrieving relevant images that closely match the description in a text query and retrieving text documents that best explain the content of an image query. A greedy dictionary construction approach is introduced for learning an isomorphic feature space, to which cross-modality data can be adapted while data smoothness is guaranteed. The proposed objective function consists of two reconstruction error terms for both modalities and a Maximum Mean Discrepancy (MMD) term that measures the cross-modality discrepancy. Optimization of the reconstruction terms and the MMD term yields a compact and modality-adaptive dictionary pair. We formulate the joint combinatorial optimization problem by maximizing variance reduction over a candidate signal set while constraining the dictionary size and coefficients' sparsity. By exploiting the submodularity and the monotonicity property of the proposed objective function, the optimization problem can be solved by a highly efficient greedy algorithm, and is guaranteed to be at least a (e - 1)=/e≈0.632- approximation to the optimum. The proposed method achieves state-of-the-art performance on the Wikipedia dataset. Fan Zhu 0001, Ling Shao 0001, Mengyang Yu |
CIKM | 3 |
| 2012 | Comparison-based encryption for fine-grained access control in cloudsabstractAccess control is one of the most important security mechanisms in cloud computing. However, there has been little work that explores various comparison-based constraints for regulating data access in clouds. In this paper, we present an innovative comparison-based encryption scheme to facilitate fine-grained access control in cloud computing. By means of forward/backward derivation functions, we introduce comparison relation into attribute-based encryption to implement various range constraints on integer attributes, such as temporal and level attributes. Then, we present a new cryptosystem with dual decryption to reduce computational overheads on cloud clients, where the majority of decryption operations are executed in cloud servers. We also prove the security strength of our proposed scheme, and our experiment results demonstrate the efficiency of our methodology. Yan Zhu 0010, Hongxin Hu, Gail-Joon Ahn, Mengyang Yu, Hong-Jia Zhao |
CODASPY | 4 |
| 2012 | Efficient construction of provably secure steganography under ordinary covert channels
Yan Zhu 0010, Mengyang Yu, Hongxin Hu, Gail-Joon Ahn, Hong-Jia Zhao |
Sci. China Inf. Sci. | 2 |
| 2012 | Cooperative Provable Data Possession for Integrity Verification in Multicloud StorageabstractProvable data possession (PDP) is a technique for ensuring the integrity of data in storage outsourcing. In this paper, we address the construction of an efficient PDP scheme for distributed cloud storage to support the scalability of service and data migration, in which we consider the existence of multiple cloud service providers to cooperatively store and maintain the clients' data. We present a cooperative PDP (CPDP) scheme based on homomorphic verifiable response and hash index hierarchy. We prove the security of our scheme based on multiprover zero-knowledge proof system, which can satisfy completeness, knowledge soundness, and zero-knowledge properties. In addition, we articulate performance optimization mechanisms for our scheme, and in particular present an efficient method for selecting optimal parameter values to minimize the computation costs of clients and storage service providers. Our experiments show that our solution introduces lower computation and communication overheads in comparison with noncooperative approaches. Yan Zhu 0010, Hongxin Hu, Gail-Joon Ahn, Mengyang Yu |
IEEE Trans. Parallel Distributed Syst. | 4 |