VLDB 2026 Research / reviewers in the wild / expert
Hassan H. Malik
dblp:28/5101
· DBLP profile ↗
8ranked-venue papers
8as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 8 first-authorArtificial intelligence and machine learning · 5 · 5 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 97% Indexing and storage engines · 3% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.1 | 2 | 2007 | Optimizing Frequency Queries for Data Mining Applications · ICDM 2007 High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2008 | Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008 |
Data mining › predictive modeling › classification
pattern classification |
0.1 | 1 | 2008 | Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008 |
Data mining › text mining
text classification |
0.1 | 1 | 2008 | Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.1 | 1 | 2007 | Optimizing Frequency Queries for Data Mining Applications · ICDM 2007 |
Data mining › pattern mining
support counting |
0.1 | 1 | 2007 | Optimizing Frequency Queries for Data Mining Applications · ICDM 2007 |
Data mining
clustering |
0.1 | 1 | 2006 | High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006 |
Data mining › clustering
document clustering |
0.1 | 1 | 2006 | High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006 |
Data mining › pattern mining
itemset mining |
0.1 | 1 | 2006 | High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006 |
Indexing and storage engines
bitmap index |
0.0 | 1 | 2007 | Optimizing Frequency Queries for Data Mining Applications · ICDM 2007 |
Data mining › pattern mining › itemset mining › frequent itemset mining
frequent closed itemset mining |
0.0 | 1 | 2006 | High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006 |
Methods — techniques the papers use, named apart from their topics
power law weighting · 0.1pattern mining · 0.1radix sort · 0.1hamming-distance reordering · 0.1gray code · 0.1WAH encoding · 0.1interestingness measures · 0.1hierarchical clustering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Accurate information extraction for quantitative financial eventsabstractIn this paper, we present a novel financial event extraction system that achieves very high extraction quality by combining the outcome of statistical classifiers with a set of rules. Using expert-annotated press releases as training data, and novel feature generation schemes, our system learns multiple binary classifiers for each "slot" in a financial event. At runtime, common parsing and search indexing methods are used to normalize incoming press releases and to identify candidate event "slots". Rules are applied on candidates that satisfy a combination of classifiers, and the system confidence on extracted events is estimated using a unique confidence model learned from training data. We present results of experiments performed on European corporate press releases for extracting dividend events, and show that our system achieves a precision of 96% and a recall of 79%. Hassan H. Malik, Vikas S. Bhardwaj, Huascar Fiorletta |
CIKM | 1 |
| 2011 | Exploring the corporate ecosystem with a semi-supervised entity graphabstractInvestment decisions in the financial markets require careful analysis of information available from multiple data sources. In this paper, we present Atlas, a novel entity-based information analysis and content aggregation platform that uses heterogeneous data sources to construct and maintain the "ecosystem" around tangible and logical entities such as organizations, products, industries, geographies, commodities and macroeconomic indicators. Entities are represented as vertices in a directed graph, and edges are generated using entity co-occurrences in unstructured documents and supervised information from structured data sources. Significance scores for the edges are computed using a method that combines supervised, unsupervised and temporal factors into a single score. Important entity attributes from the structured content and the entity neighborhood in the graph are automatically summarized as the entity "fingerprint". A highly interactive user interface provides exploratory access to the graph and supports common business use cases. We present results of experiments performed on five years of news and broker research data, and show that Atlas is able to accurately identify important and interesting connections in real-world entities. We also demonstrate that Atlas entity fingerprints are particularly useful in entity similarity queries, with a quality that rivals existing human maintained databases. Hassan H. Malik, Ian MacGillivray, Måns Olof-Ors, Siming Sun, Shailesh Saroha |
CIKM | 1 |
| 2011 | Single pass text classification by direct feature weighting
Hassan H. Malik, Dmitriy Fradkin, Fabian Mörchen |
Knowl. Inf. Syst. | 1 |
| 2010 | Hierarchical document clustering using local patterns
Hassan H. Malik, John R. Kender, Dmitriy Fradkin, Fabian Mörchen |
Data Min. Knowl. Discov. | 1 |
| 2008 | Classifying High-Dimensional Text and Web Data Using Very Short PatternsabstractIn this paper, we propose the "democratic classifier", a simple pattern-based classification algorithm that uses very short patterns for classification, and does not rely on the minimum support threshold. Borrowing ideas from democracy, our training phase allows each training instance to vote for an equal number of candidate size-2 patterns. The training instances select patterns by effectively balancing between local, class, and global significance of patterns. The selected patterns are simultaneously added to the model for all applicable classes and a novel power law based weighing scheme adjusts their weights with respect of each class. Results of experiments performed on 121 common text and Web datasets show that our algorithm almost always outperforms state of the art classification algorithms, without any parameter tuning. On 100 real-life Web datasets, the average absolute classification accuracy improvement was as great as 9.4% over SVM, Harmony, C4.5 and KNN. Also, our algorithm ran about 3.5 times faster than the fastest existing pattern-based classification algorithm. Hassan H. Malik, John R. Kender |
ICDM | 1 |
| 2007 | Optimizing Frequency Queries for Data Mining ApplicationsabstractData mining algorithms use various Trie and bitmap-based representations to optimize the support (i.e., frequency) counting performance. In this paper, we compare the memory requirements and support counting performance of FP Tree, and Compressed Patricia Trie against several novel variants of vertical bit vectors. First, borrowing ideas from the VLDB domain, we compress vertical bit vectors using WAH encoding. Second, we evaluate the Gray code rank- based transaction reordering scheme, and show that in practice, simple lexicographic ordering, obtained by applying LSB Radix sort, outperforms this scheme. Led by these results, we propose HDO, a novel Hamming-distance-based greedy transaction reordering scheme, and aHDO, a linear-time approximation to HDO. We present results of experiments performed on 15 common datasets with varying degrees of sparseness, and show that HDO- reordered, WAH encoded bit vectors can take as little as 5% of the uncompressed space, while aHDO achieves similar compression on sparse datasets. Finally, with results from over a billion database and data mining style frequency query executions, we show that bitmap-based approaches result in up to hundreds of times faster support counting, and HDO-WAH encoded bitmaps offer the best space-time tradeoff. Hassan H. Malik, John R. Kender |
ICDM | 1 |
| 2006 | High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting ItemsetsabstractHigh dimensionality remains a significant challenge for document clustering. Recent approaches used frequent itemsets and closed frequent itemsets to reduce dimensionality, and to improve the efficiency of hierarchical document clustering. In this paper, we introduce the notion of "closed interesting" itemsets (i.e. closed itemsets with high interestingness). We provide heuristics such as "super item" to efficiently mine these itemsets and show that they provide significant dimensionality reduction over closed frequent itemsets. Using "closed interesting" itemsets, we propose a new, sub-linearly scalable, hierarchical document clustering method that outperforms state of the art agglomerative, partitioning and frequent-itemset based methods both in terms of clustering quality and runtime performance, without requiring dataset specific parameter tuning. We evaluate twenty interestingness measures and show that when used to generate "closed interesting" itemsets, and to select parent nodes, mutual information, added value, Yule's Q and Chi- Square offer best clustering performance. Hassan H. Malik, John R. Kender |
ICDM | 1 |
| 2006 | Clustering web images using association rules, interestingness measures, and hypergraph partitionsabstractThis paper presents a new approach to cluster web images. Images are first processed to extract signal features such as color in HSV format and quantized orientation. Web pages referring to these images are processed to extract textual features (keywords) and feature reduction techniques such as stemming, stop word elimination, and Zipf's law are applied. All visual and textual features are used to generate association rules. Hypergraphs are generated from these rules, with features used as vertices and discovered associations as hyperedges. Twenty-two objective interestingness measures are evaluated on their ability to prune non-interesting rules and to assign weights to hyperedges. Then a hypergraph partitioning algorithm is used to generate clusters of features, and a simple scoring function is used to assign images to clusters. A tree-distance-based evaluation measure is used to evaluate the quality of image clustering with respect to manually generated ground truth. Our experiments indicate that combining textual and content-based features results in better clustering as compared to signal-only or text-only approaches. Online steps are done in real-time, which makes this approach practical for web images. Furthermore, we demonstrate that statistical interestingness measures such as Correlation Coefficient, Laplace, Kappa and J-Measure result in better clustering compared to traditional association rule interestingness measures such as Support and Confidence. Hassan H. Malik, John R. Kender |
ICWE | 1 |