Kerui Min

dblp:65/7181 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 68% Image recognition and object detection · 11% 3D vision · 11%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
extractive-abstractive summarization
0.312018
A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss · ACL (1) 2018
Natural language and speech › Language models and text generation
text summarization
0.312018
A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss · ACL (1) 2018
Computer vision › Image recognition and object detection
image retrieval
0.112010
Compact projection: Simple and efficient near neighbor search with practical memory requirements · CVPR 2010
Computer vision › 3D vision
similarity search
0.112010
Compact projection: Simple and efficient near neighbor search with practical memory requirements · CVPR 2010
Machine learning › Deep learning architectures and training
attention mechanism
0.112018
A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss · ACL (1) 2018
Algorithms and data structures › similarity search
nearest neighbor search
0.012010
Compact projection: Simple and efficient near neighbor search with practical memory requirements · CVPR 2010

Methods — techniques the papers use, named apart from their topics

inconsistency loss · 0.3end-to-end training · 0.3random projection · 0.2hashing · 0.2
YearPublicationVenuePosition
2018 A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss
abstract
We propose a unified model combining the strength of extractive and abstractive summarization.On the one hand, a simple extractive model can obtain sentence-level attention with high ROUGE scores but less readable.On the other hand, a more complicated abstractive model can obtain word-level dynamic attention to generate a more readable paragraph.In our model, sentence-level attention is used to modulate the word-level attention such that words in less attended sentences are less likely to be generated.Moreover, a novel inconsistency loss function is introduced to penalize the inconsistency between two levels of attentions.By end-to-end training our model with the inconsistency loss and original losses of extractive and abstractive models, we achieve state-of-theart ROUGE scores while being the most informative and readable summarization on the CNN/Daily Mail dataset in a solid human evaluation.
Wan Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min
ACL (1)4
2015 BosonNLP: An Ensemble Approach for Word Segmentation and POS Tagging
Kerui Min, Chenggang Ma, Tianmei Zhao
NLPCC1
2013 Joint topic-document modeling via low-dimensional sparse models
abstract
Topic modeling is a well-known approach for document analysis. In this paper, we propose a new model, and corresponding optimization algorithm for topic modeling. Experimental results on polarity classification demonstrate that the new model provides a more accurate characterization for document corpus, and archived higher classification accuracy compared to Latent Dirichlet Allocation (LDA).
Kerui Min, Yi Ma 0001
ICASSP1
2012 Principal Component Pursuit with reduced linear measurements
abstract
In this paper, we study the problem of decomposing a superposition of a low-rank matrix and a sparse matrix when a relatively few linear measurements are available. This problem arises in many data processing tasks such as aligning multiple images or rectifying regular texture, where the goal is to recover a low-rank matrix with a large fraction of corrupted entries in the presence of nonlinear domain transformation. We consider a natural convex heuristic to this problem which is a variant to the recently proposed Principal Component Pursuit. We prove that under suitable conditions, this convex program guarantees to recover the correct low-rank and sparse components despite reduced measurements. Our analysis covers both random and deterministic measurement models.
Arvind Ganesh, Kerui Min, John Wright 0001, Yi Ma 0001
ISIT2
2012 Compressive principal component pursuit
abstract
We consider the problem of recovering a target matrix that is a superposition of low-rank and sparse components, from a small set of linear measurements. This problem arises in compressed sensing of structured high-dimensional signals such as videos and hyperspectral images, as well as in the analysis of transformation invariant low-rank recovery. We analyze the performance of the natural convex heuristic for solving this problem, under the assumption that measurements are chosen uniformly at random. We prove that this heuristic exactly recovers low-rank and sparse terms, provided the number of observations exceeds the number of intrinsic degrees of freedom of the component signals by a polylogarithmic factor. Our analysis introduces several ideas that may be of independent interest for the more general problem of compressive sensing of superpositions of structured signals.
John Wright 0001, Arvind Ganesh, Kerui Min, Yi Ma 0001
ISIT3
2010 Decomposing background topics from keywords by principal component pursuit
abstract
Low-dimensional topic models have been proven very useful for modeling a large corpus of documents that share a relatively small number of topics. Dimensionality reduction tools such as Principal Component Analysis or Latent Semantic Indexing (LSI) have been widely adopted for document modeling, analysis, and retrieval. In this paper, we contend that a more pertinent model for a document corpus as the combination of an (approximately) low-dimensional topic model for the corpus and a sparse model for the keywords of individual documents. For such a joint topic-document model, LSI or PCA is no longer appropriate to analyze the corpus data. We hence introduce a powerful new tool called Principal Component Pursuit that can effectively decompose the low-dimensional and the sparse components of such corpus data. We give empirical results on data synthesized with a Latent Dirichlet Allocation (LDA) mode to validate the new model. We then show that for real document data analysis, the new tool significantly reduces the perplexity and improves retrieval performance compared to classical baselines.
Kerui Min, Zhengdong Zhang 0001, John Wright 0001, Yi Ma 0001
CIKM1
2010 Compact projection: Simple and efficient near neighbor search with practical memory requirements
abstract
Image similarity search is a fundamental problem in computer vision. Efficient similarity search across large image databases depends critically on the availability of compact image representations and good data structures for indexing them. Numerous approaches to the problem of generating and indexing image codes have been presented in the literature, but existing schemes generally lack explicit estimates of the number of bits needed to effectively index a given large image database. We present a very simple algorithm for generating compact binary representations of imagery data, based on random projections. Our analysis gives the first explicit bound on the number of bits needed to effectively solve the indexing problem. When applied to real image search tasks, these theoretical improvements translate into practical performance gains: experimental results show that the new method, while using significantly less memory, is several times faster than existing alternatives.
Kerui Min, Linjun Yang, John Wright 0001, Lei Wu 0017, Xian-Sheng Hua 0001, Yi Ma 0001
CVPR1
2009 The Closest Pair Problem under the Hamming Metric
Kerui Min, Ming-Yang Kao, Hong Zhu 0004
COCOON1