VLDB 2026 Research / reviewers in the wild / expert
Zhaoyan Ming
dblp:13/7184
· DBLP profile ↗
39ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-6766-4579ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake DetectionabstractAdvanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it unlocks novel applications in fields like entertainment and education, its malicious use has sparked urgent ethical and societal concerns ranging fromidentity theftto thedissemination of misinformation. To tackle these challenges, feature analysis using frequency features has emerged as a promising direction for deepfake detection. However, one aspect that has been overlooked so far is that existing methods tend to concentrate on one or a few specific frequency domains, which risks overfitting to particular artifacts and significantly undermines their robustness when facing diverse forgery patterns. Another underexplored aspect we observe is that different features often attend to the same forged region, resulting in redundant feature representations and limiting the diversity of the extracted clues. This may undermine the ability of a model to capture complementary information across different facets, thereby compromising its generalization capability to diverse manipulations. In this paper, we seek to tackle these challenges from two aspects: (1) we propose a triple-branch network that jointly captures spatial and frequency features by learning from both original image and image reconstructed by different frequency channels, and (2) we mathematically derive feature decoupling and fusion losses grounded in the mutual information theory, which enhances the model to focus on task-relevant features across the original image and the image reconstructed by different frequency channels. Extensive experiments onsixlarge-scale benchmark datasets demonstrate that our method consistently achieves state-of-the-art performance. Our code is released athttps://github.com/injooker/Unveiling_Deepfake. Qihao Shen, Jiaxing Xuan, Zhenguang Liu, Sifan Wu 0001, Yutong Xie 0019, Zhaoyan Ming, Yingying Jiao, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | FairQuanti: Enhancing Fairness in Deep Neural Network Quantization via Neuron Role ContributionabstractThe increasing complexity of deep neural networks (DNNs) poses significant resource challenges for edge devices, prompting the development of compression technologies like model quantization. However, while improving model efficiency, quantization can introduce or perpetuate the original model’s bias. Existing debiasing methods for quantized models often incur additional costs. To address this issue, we propose FairQuanti , a novel quantization approach that leverages neuron role contribution to achieve fairness. By distinguishing between biased and normal neurons, FairQuanti employs mixed precision quantization to mitigate model bias during the quantization process. FairQuanti has four key differences from previous studies: (1) Neuron Roles - It formally defines biased and normal neuron roles, establishing a framework for feasible model quantization and bias mitigation; (2) Effectiveness - It introduces a fair quantization strategy that discriminatively quantizes neuron roles, balancing model accuracy and fairness through Bayesian optimization; (3) Generality - It applies to both structured and unstructured data across various quantization bit levels; (4) Robustness - It demonstrates resilience against adaptive attacks. Extensive experiments on five datasets (three structured and two unstructured) using five different models validate FairQuanti ’s superior performance against eight baseline methods. Specifically, fairness metrics such as demographic parity (DP) improve by approximately 1.03 times, and the demographic parity ratio (DPR) improves by approximately 1.51 times compared to the baselines, with an average accuracy loss of less than 7.5% at 8-bit quantization. FairQuanti presents a promising solution for deploying fair and efficient deep models on resource-constrained devices and holds potential for application in large language models to reduce size and computational demands while minimizing bias. Our source code is available at https://github.com/Caozq2/FairQuanti . Jinyin Chen, Zhiqi Cao, Haibin Zheng, Zhaoyan Ming, Yayu Zheng |
ACM Trans. Priv. Secur. | 5 |
| 2025 | Robust Visual Food Recognition for Enriching Nutrition Knowledge BasesabstractAcquiring nutrition information and health-related knowledge about food is a common need among individuals. However, using conventional food names as search queries often fails to yield accurate matches to entries within food nutrition knowledge bases (FoodnKB), which frequently utilize scientific or product names. In this study, we present a method for enriching FoodnKB entries with imagery and facilitating visual access to food-related knowledge through image recognition. We start with an official food nutrition database and propose a consensus-based approach using Large Language Models to identify visually discernible and directly edible foods, expanding food synonyms and harnessing diverse web-based food images for comprehensive visual representation. To minimize manual annotation of noisy web images, we introduce a cyclic training-based area under the margin metric (cAUM) approach that effectively distinguishes appropriate images, including rare instances, from noisy ones. Additionally, we design a generic accuracy gap (AccGap) algorithm to automatically estimate the noise ratio of the web-harnessed data. Our integrated cAUM and AccGap method demonstrates superior performance in noise detection and enhancement of image recognition accuracy compared to existing noise-robust frameworks. Furthermore, we successfully apply the visually enriched FoodnKB and food recognition capabilities within a smart nutritionist mobile application. Zhaoyan Ming, Zeyu Xie, Kui Su, Changzheng Yuan, Tat-Seng Chua |
IEEE Trans. Multim. | 1 |
| 2023 | HostNet: improved sequence representation in deep neural networks for virus-host predictionabstractBACKGROUND: The escalation of viruses over the past decade has highlighted the need to determine their respective hosts, particularly for emerging ones that pose a potential menace to the welfare of both human and animal life. Yet, the traditional means of ascertaining the host range of viruses, which involves field surveillance and laboratory experiments, is a laborious and demanding undertaking. A computational tool with the capability to reliably predict host ranges for novel viruses can provide timely responses in the prevention and control of emerging infectious diseases. The intricate nature of viral-host prediction involves issues such as data imbalance and deficiency. Therefore, developing highly accurate computational tools capable of predicting virus-host associations is a challenging and pressing demand. RESULTS: To overcome the challenges of virus-host prediction, we present HostNet, a deep learning framework that utilizes a Transformer-CNN-BiGRU architecture and two enhanced sequence representation modules. The first module, k-mer to vector, pre-trains a background vector representation of k-mers from a broad range of virus sequences to address the issue of data deficiency. The second module, an adaptive sliding window, truncates virus sequences of various lengths to create a uniform number of informative and distinct samples for each sequence to address the issue of data imbalance. We assess HostNet's performance on a benchmark dataset of "Rabies lyssavirus" and an in-house dataset of "Flavivirus". Our results show that HostNet surpasses the state-of-the-art deep learning-based method in host-prediction accuracies and F1 score. The enhanced sequence representation modules, significantly improve HostNet's training generalization, performance in challenging classes, and stability. CONCLUSION: HostNet is a promising framework for predicting virus hosts from genomic sequences, addressing challenges posed by sparse and varying-length virus sequence data. Our results demonstrate its potential as a valuable tool for virus-host prediction in various biological contexts. Virus-host prediction based on genomic sequences using deep neural networks is a promising approach to identifying their potential hosts accurately and efficiently, with significant impacts on public health, disease prevention, and vaccine development. Zhaoyan Ming, Xiangjun Chen, Shunlong Wang, Zhiming Yuan, Han Xia 0002 |
BMC Bioinform. | 1 |
| 2023 | GONE: A generic O(1) NoisE layer for protecting privacy of deep neural networks
Haibin Zheng, Jinyin Chen, Wenchang Shangguan, Zhaoyan Ming, Xing Yang 0004 |
Comput. Secur. | 4 |
| 2023 | Evil vs evil: using adversarial examples to against backdoor attack in federated learning
Tao Liu 0040, Haibin Zheng, Zhaoyan Ming, Jinyin Chen |
Multim. Syst. | 4 |
| 2023 | Coloring anime line art videos with transformation region enhancement networkabstractAutomatic colorization of anime line art videos aims to produce color frames given line art frames and reference color images, which is challenging due to various motions and geometric transformations across frame sequences. Existing methods usually utilize the feature maps of reference images directly and treat all the regions in an image equally. However, this may overlook the details of the regions undergoing geometric transformations . To emphasize the regions with significant transformations between the reference and target frames, we propose a Transformation Region Enhancement Network (TRE-Net) to exploit useful reference information and enhance the colorization of key transformation regions with Region Localization Module (RLM) and Feature Enhancement Module (FEM). Specifically, we propose Multi-scale Euclidean Distance Difference (Multi-scale EDD) Maps in RLM which effectively locate geometric transformation regions by contrasting the Euclidean Distance Maps of two line arts and aggregating representations at multiple scales of the network. In addition, FEM is devised to enhance feature learning in the regions with geometric transformation and to ensure proper color alignment. FEM learns locally enhanced features through an attention-gating operation at a low computational cost. With the well-represented key geometric transformation regions, our method exploits the multi-scale reference information well for color alignment, thus produces perceptually pleasing frames. Comprehensive experimental results show that our proposed method is superior to existing methods in terms of the overall quality of colorized anime line art videos. Ning Wang 0020, Muyao Niu, Zhi Dou, Zhihui Wang 0001, Zhiyong Wang 0001, Zhaoyan Ming, Bin Liu 0040 |
Pattern Recognit. | 6 |
| 2023 | CTL-DIFF: Control Information Diffusion in Social Network by Structure OptimizationabstractA critical side effect of online social networks’ flourishing is fast-spreading rumors on the Internet, making the information diffusion control on social networks a fundamental requirement. While information diffusion control has received extensive attention at the global level, there have been fewer user-level diffusion control studies under minimal budget. In this article, we study the information diffusion at the user level and propose a diffusion control method based on gradient information to generate an optimized network structure, namely ConTroL information DIFFusion (CTL-DIFF). CTL-DIFF targets a user through subtle modifications of its local network structure. It first selects the edges with the largest absolute gradient based on the prediction model to optimize the original network’s structure. It then employs several prediction methods to verify whether the target user’s social action status is controlled. CTL-DIFF achieves state-of-the-art control performance with a minimum budget, comparing with five baselines based on edge centrality strategies on four real-world datasets. We extend the diffusion control from user-level to global-level, comparing with four baselines on three datasets. Experimental results show that CTL-DIFF can effectively control information diffusion in the global social network by identifying and controlling the most influential users. Jinyin Chen, Lihong Chen, Zhongyuan Ruan, Zhaoyan Ming, Yi Liu 0024 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | FDA-GAN: Flow-Based Dual Attention GAN for Human Pose TransferabstractHuman pose transfer aims at transferring the appearance of the source person to the target pose. Existing methods utilizing flow-based warping for non-rigid human image generation have achieved great success. However, they fail to preserve the appearance details in synthesized images since the spatial correlation between the source and target is not fully exploited. To this end, we propose the Flow-based Dual Attention GAN (FDA-GAN) to apply occlusion- and deformation-aware feature fusion for higher generation quality. Specifically, deformable local attention and flow similarity attention, constituting the dual attention mechanism, can derive the output features responsible for deformable- and occlusion-aware fusion, respectively. Besides, to maintain the pose and global position consistency in transferring, we design a pose normalization network for learning adaptive normalization from the target pose to the source person. Both qualitative and quantitative results show that our method outperforms state-of-the-art models in public iPER and DeepFashion datasets. Kejie Huang, Dongxu Wei, Zhaoyan Ming, Haibin Shen |
IEEE Trans. Multim. | 4 |
| 2022 | Salient feature extractor for adversarial defense on deep neural networks
Ruoxi Chen, Jinyin Chen, Haibin Zheng, Qi Xuan 0001, Zhaoyan Ming, Wenrong Jiang |
Inf. Sci. | 5 |
| 2022 | ROBY: Evaluating the adversarial robustness of a deep model by its decision boundaries
Haibo Jin, Jinyin Chen, Haibin Zheng, Zhen Wang 0004, Jun Xiao 0001, Shanqing Yu, Zhaoyan Ming |
Inf. Sci. | 7 |
| 2021 | MASS: Multi-task anthropomorphic speech synthesis framework
Jinyin Chen, Linhui Ye, Zhaoyan Ming |
Comput. Speech Lang. | 3 |
| 2020 | Hierarchical Attention Network for Visually-Aware Food RecommendationabstractFood recommender systems play an important role in assisting users to identify the desired food to eat. Deciding what food to eat is a complex and multi-faceted process, which is influenced by many factors such as the ingredients, appearance of the recipe, the user's personal preference on food, and various contexts like what had been eaten in the past meals. This work formulates the food recommendation problem as predicting user preference on recipes based on three key factors that determine a user's choice on food, namely, 1) the user's (and other users') history; 2) the ingredients of a recipe; and 3) the descriptive image of a recipe. To address this challenging problem, this work develops a dedicated neural network-based solution Hierarchical Attention based Food Recommendation (HAFR) which is capable of: 1) capturing the collaborative filtering effect like what similar users tend to eat; 2) inferring a user's preference at the ingredient level; and 3) learning user preference from the recipe's visual images. To evaluate our proposed method, this work constructs a large-scale dataset consisting of millions of ratings from AllRecipes.com. Extensive experiments show that our method outperforms several competing recommender solutions like Factorization Machine and Visual Bayesian Personalized Ranking with an average improvement of 12%, offering promising results in predicting user preference on food. Xiaoyan Gao 0001, Fuli Feng, Xiangnan He 0001, Heyan Huang, Chong Feng 0001, Zhaoyan Ming, Tat-Seng Chua |
IEEE Trans. Multim. | 7 |
| 2019 | DietLens-Eout: Large Scale Restaurant Food Photo RecognitionabstractRestaurant dishes represent a significant portion of food that people consume in their daily life. While people are becoming health-conscious in their food intake, convenient restaurant food tracking becomes an essential task in wellness and fitness applications. Given the huge number of dishes (food categories) involved, it becomes extremely challenging for traditional food photo classification to be feasible in both algorithm design and training data availability. In this work, we present a demo that runs on restaurant dish images in a city of millions of residents and tens of thousand restaurants. We propose a rank-loss based convolutional neural network to optimize the image features representation. Context information such as GPS location of the recognition request is also used to further improve the performance. Our experimental results are highly promising. We have shown in our demo that the proposed algorithm is near ready to be deployed in real-world applications. Zhipeng Wei 0001, Jingjing Chen 0001, Zhaoyan Ming, Chong-Wah Ngo, Tat-Seng Chua, Fengfeng Zhou |
ICMR | 3 |
| 2019 | Mixed-dish Recognition with Contextual Relation NetworksabstractMixed dish is a food category that contains different dishes mixed in one plate, and is popular in Eastern and Southeast Asia. Recognizing individual dishes in a mixed dish image is important for health related applications, e.g. calculating the nutrition values. However, most existing methods that focus on single dish classification are not applicable to mixed-dish recognition. The new challenge in recognizing mixed-dish images are the complex ingredient combination and severe overlap among different dishes. In order to tackle these problems, we propose a novel approach called contextual relation networks (CR-Nets) that encodes the implicit and explicit contextual relations among multiple dishes using region-level features and label-level co-occurrence, respectively. This is inspired by the intuition that people are likely to choose dishes with common eating habits, e.g., with multiple nutrition but without repeating ingredients. In addition, we collect a large-scale dataset of mixed-dish images that contain $9,254$ mixed-dish images from $6$ school canteens in Singapore. Extensive experiments on both our dataset and a smaller-scale public dataset validate that our CR-Nets can achieve top performance for localizing the dishes and recognizing their food categories. Lixi Deng, Jingjing Chen 0001, Qianru Sun, Xiangnan He 0001, Sheng Tang, Zhaoyan Ming, Yongdong Zhang 0001, Tat-Seng Chua |
ACM Multimedia | 6 |
| 2018 | Food Photo Recognition for Dietary Tracking: System and Experiment
Zhaoyan Ming, Jingjing Chen 0001, Ciarán Forde, Chong-Wah Ngo, Tat-Seng Chua |
MMM (2) | 1 |
| 2017 | Product ranking using hierarchical aspect structures
Si Li 0001, Zhaoyan Ming, Yan Leng, Jun Guo 0002 |
J. Intell. Inf. Syst. | 2 |
| 2016 | Exploring heterogeneous features for query-focused summarization of categorized community answers
Wei Wei 0002, Zhaoyan Ming, Liqiang Nie, Guohui Li 0001, Jianjun Li 0010, Feida Zhu 0001, Tianfeng Shang, Changyin Luo |
Inf. Sci. | 2 |
| 2016 | Resolving local cuisines for tourists with multi-source social media contents
Zhaoyan Ming, Tat-Seng Chua |
Multim. Syst. | 1 |
| 2016 | Generating Incremental Length Summary Based on Hierarchical Topic Coverage MaximizationabstractDocument summarization is playing an important role in coping with information overload on the Web. Many summarization models have been proposed recently, but few try to adjust the summary length and sentence order according to application scenarios. With the popularity of handheld devices, presenting key information first in summaries of flexible length is of great convenience in terms of faster reading and decision-making and network consumption reduction. Targeting this problem, we introduce a novel task of generating summaries of incremental length. In particular, we require that the summaries should have the ability to automatically adjust the coverage of general-detailed information when the summary length varies. We propose a novel summarization model that incrementally maximizes topic coverage based on the document’s hierarchical topic model. In addition to the standard Rouge-1 measure, we define a new evaluation metric based on the similarity of the summaries’ topic coverage distribution in order to account for sentence order and summary length. Extensive experiments on Wikipedia pages, DUC 2007, and general noninverted writing style documents from multiple sources show the effectiveness of our proposed approach. Moreover, we carry out a user study on a mobile application scenario to show the usability of the produced summary in terms of improving judgment accuracy and speed, as well as reducing the reading burden and network traffic. Jintao Ye, Zhaoyan Ming, Tat-Seng Chua |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Capturing the Semantics of Key Phrases Using Multiple Languages for Question RetrievalabstractIn the age of Web 2.0, community user contributed questions and answers provide an important alternative for knowledge acquisition through web search. Question retrieval in current community-based question answering (CQA) services do not, in general, work well for long and complex queries, such as the questions. The main reasons are the verboseness in natural language queries and the word mismatch between the queries and the candidate questions in the CQA archive during retrieval. To address these two problems, existing solutions try to refine the search queries by distinguishing the key concepts in the queries and expanding the queries with relevant content. However, using the existing query refinement approaches can only identify the key and non-key concepts, while the differences between the key concepts are overlooked. Moreover, the existing query expansion approaches, not only overlook the weights of key concepts in the queries, but also fail to consider concept level expansion for them. In this paper, we explore a key concept identification approach for query refinement and a pivot language translation based approach to explore key concept paraphrasing. We further propose a new question retrieval model which can seamlessly integrate the key concepts and their paraphrases. The experimental results demonstrate that the integrated retrieval model significantly outperforms the state-of-the-art models in question retrieval. Weinan Zhang 0003, Zhaoyan Ming, Yu Zhang 0030, Ting Liu 0001, Tat-Seng Chua |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Volunteerism Tendency Prediction via Harvesting Multiple Social NetworksabstractVolunteers have always been extremely crucial and in urgent need for nonprofit organizations (NPOs) to sustain their continuing operations. However, it is expensive and time-consuming to recruit volunteers using traditional approaches. In the Web 2.0 era, abundant and ubiquitous social media data opens a door to the possibility of automatic volunteer identification. In this article, we aim to fully explore this possibility by proposing a scheme that is able to predict users’ volunteerism tendency from user-generated contents collected from multiple social networks based on a conceptual volunteering decision model. We conducted comprehensive experiments to investigate the effectiveness of our proposed scheme and further discussed its generalizibility and extendability. This novel interdisciplinary research will potentially inspire more promising and important human-centered applications. Xuemeng Song, Zhaoyan Ming, Liqiang Nie, Yi-Liang Zhao, Tat-Seng Chua |
ACM Trans. Inf. Syst. | 2 |
| 2015 | Exploring Key Concept Paraphrasing Based on Pivot Language Translation for Question RetrievalabstractQuestion retrieval in current community-based question answering (CQA) services does not, in general, work well for long and complex queries. One of the main difficulties lies in the word mismatch between queries and candidate questions. Existing solutions try to expand the queries at word level, but they usually fail to consider concept level enrichment. In this paper, we explore a pivot language translation based approach to derive the paraphrases of key concepts. We further propose a unified question retrieval model which integrates the keyconcepts and their paraphrases for the query question. Experimental results demonstrate that the paraphrase enhanced retrieval model significantly outperforms the state-of-the-art models in question retrieval. Weinan Zhang 0003, Zhaoyan Ming, Yu Zhang 0030, Ting Liu 0001, Tat-Seng Chua |
AAAI | 2 |
| 2015 | Tackling Data Sparseness in Recommendation using Social Media based Topic Hierarchy Modeling
Xingwei Zhu, Zhaoyan Ming, Yu Hao 0001, Xiaoyan Zhu 0001 |
IJCAI | 2 |
| 2015 | Resolving polysemy and pseudonymity in entity linking with comprehensive name and context modeling
Zhaoyan Ming, Tat-Seng Chua |
Inf. Sci. | 1 |
| 2014 | A Dynamic Reconstruction Approach to Topic Summarization of User-Generated-ContentabstractUser generated contents (UGCs) from various social media sites give analysts the opportunity to obtain a comprehensive and dynamic view of any topic from multiple heterogeneous information sources. Summarization provides a promising means of distilling the overview of the targeted topic by aggregating and condensing the related UGCs. However, the mass volume, uneven quality, and dynamics of UGCs, pose new challenges that are not addressed by existing multi-document summarization techniques. In this paper, we introduce a timely task of dynamic structural and textual summarization. We generate topic hierarchy from the UGCs as a high level overview and structural guide for exploring and organizing the content. To capture the evolution of events in the content, we propose a unified dynamic reconstruction approach to detect the update points and generate the time-sequence textual summary. To enhance the expressiveness of the reconstruction space, we further use the topic hierarchy to organize the UGCs and the hierarchical subtopics to augment the sentence representation. Experimental comparison with the state-of-the-art summarization models on a multi-source UGC dataset shows the superiority of our proposed methods. Moreover, we conducted a user study on our usability enhancement measures. It suggests that by disclosing some meta information of the summary generation process in the proposed framework, the time-sequence textual summaries can pair with the structural overview of the topic hierarchy to achieve interpretable and verifiable summarization. Zhaoyan Ming, Jintao Ye, Tat-Seng Chua |
CIKM | 1 |
| 2014 | Customized Organization of Social Media Contents using Focused Topic HierarchyabstractWith the popularity of social media platforms such as Facebook and Twitter, the amount of useful data in these sources is rapidly increasing, making them promising places for information acquisition. This research aims at the customized organization of a social media corpus using focused topic hierarchy. It organizes the contents into different structures to meet with users' different information needs (e.g., "iPhone 5 problem" or "iPhone 5 camera"). To this end, we introduce a novel function to measure the likelihood of a topic hierarchy, by which the users' information need can be incorporated into the process of topic hierarchy construction. Using the structure information within the generated topic hierarchy, we then develop a probability based model to identify the representative contents for topics to assist users in document retrieval on the hierarchy. Experimental results on real world data illustrate the effectiveness of our method and its superiority over state-of-the-art methods for both information organization and retrieval tasks. Xingwei Zhu, Zhaoyan Ming, Yu Hao 0001, Xiaoyan Zhu 0001, Tat-Seng Chua |
CIKM | 2 |
| 2014 | Discovering high quality answers in community question answering archives using a hierarchy of classifiers
Hapnes Toba, Zhaoyan Ming, Mirna Adriani, Tat-Seng Chua |
Inf. Sci. | 2 |
| 2013 | Topic hierarchy construction for the organization of multi-source user generated contentsabstractUser generated contents (UGCs) carry a huge amount of high quality information. However, the information overload and diversity of UGC sources limit their potential uses. In this research, we propose a framework to organize information from multiple UGC sources by a topic hierarchy which is automatically generated and updated using the UGCs. We explore the unique characteristics of UGCs like blogs, cQAs, microblogs, etc., and introduce a novel scheme to combine them. We also propose a graph-based method to enable incremental update of the generated topic hierarchy. Using the hierarchy, users can easily obtain a comprehensive, in-depth and up-to-date picture of their topics of interests. The experiment results demonstrate how information from multiple heterogeneous sources improves the resultant topic hierarchies. It also shows that the proposed method achieves better F1 scores in hierarchy generation as compared to the state-of-the-art methods. Xingwei Zhu, Zhaoyan Ming, Xiaoyan Zhu 0001, Tat-Seng Chua |
SIGIR | 2 |
| 2012 | Automatic labeling hierarchical topicsabstractRecently, statistical topic modeling has been widely applied in text mining and knowledge management due to its powerful ability. A topic, as a probability distribution over words, is usually difficult to be understood. A common, major challenge in applying such topic models to other knowledge management problem is to accurately interpret the meaning of each topic. Topic labeling, as a major interpreting method, has attracted significant attention recently. However, previous works simply treat topics individually without considering the hierarchical relation among topics, and less attention has been paid to creating a good hierarchical topic descriptors for a hierarchy of topics. In this paper, we propose two effective algorithms that automatically assign concise labels to each topic in a hierarchy by exploiting sibling and parent-child relations among topics. The experimental results show that the inter-topic relation is effective in boosting topic labeling accuracy and the proposed algorithms can generate meaningful topic labels that are useful for interpreting the hierarchical topics. Xianling Mao, Zhaoyan Ming, Zhengjun Zha, Tat-Seng Chua, Hongfei Yan, Xiaoming Li 0001 |
CIKM | 2 |
| 2012 | The Use of Dependency Relation Graph to Enhance the Term Weighting in Question Retrieval
Weinan Zhang 0003, Zhaoyan Ming, Yu Zhang 0030, Liqiang Nie, Ting Liu 0001, Tat-Seng Chua |
COLING | 2 |
| 2012 | SSHLDA: A Semi-Supervised Hierarchical Topic Model
Xianling Mao, Zhaoyan Ming, Tat-Seng Chua, Si Li 0001, Hongfei Yan, Xiaoming Li 0001 |
EMNLP-CoNLL | 2 |
| 2011 | Product comparison using comparative relationsabstractThis paper proposes a novel Product Comparison approach. The comparative relations between products are first mined from both user reviews on multiple review websites and community-based question answering pairs containing product comparison information. A unified graph model is then developed to integrate the resultant comparative relations for product comparison. Experiments on popular electronic products show that the proposed approach outperforms the state-of-the-art methods. Si Li 0001, Zhengjun Zha, Zhaoyan Ming, Meng Wang 0001, Tat-Seng Chua, Jun Guo 0002, Weiran Xu |
SIGIR | 3 |
| 2010 | Exploring domain-specific term weight in archived question searchabstractCommunity Question Answering services, e.g., Yahoo! Answers, have accumulated large archives of question answer (QA) pairs for information and answer retrieval. An effective question retrieval model is essential to increase the accessibility of the QA archives. QA archives are usually organized into categories and question search can be performed within the whole collection or within a certain category.. Zhaoyan Ming, Tat-Seng Chua, Gao Cong |
CIKM | 1 |
| 2010 | Vocabulary Filtering for Term Weighting in Archived Question Search
Zhaoyan Ming, Tat-Seng Chua |
PAKDD (1) | 1 |
| 2010 | Prototype hierarchy based clustering for the categorization and navigation of web collectionsabstractThis paper presents a novel prototype hierarchy based clustering (PHC) framework for the organization of web collections. It solves simultaneously the problem of categorizing web collections and interpreting the clustering results for navigation. By utilizing prototype hierarchies and the underlying topic structures of the collections, PHC is modeled as a multi-criterion optimization problem based on minimizing the hierarchy evolution, maximizing category cohesiveness and inter-hierarchy structural and semantic resemblance. The flexible design of metrics enables PHC to be a general framework for applications in various domains. In the experiments on categorizing 4 collections of distinct domains, PHC achieves 30% improvement in ¼F1 over the state-of-the-art techniques. Further experiments provide insights on performance variations with abstract and concrete domains, completeness of the prototype hierarchy, and effects of different combinations of optimization criteria. Zhaoyan Ming, Tat-Seng Chua |
SIGIR | 1 |
| 2010 | Segmentation of multi-sentence questions: towards effective question retrieval in cQA servicesabstractExisting question retrieval models work relatively well in finding similar questions in community-based question answering (cQA) services. However, they are designed for single-sentence queries or bag-of-word representations, and are not sufficient to handle multi-sentence questions complemented with various contexts. Segmenting questions into parts that are topically related could assist the retrieval system to not only better understand the user's different information needs but also fetch the most appropriate fragments of questions and answers in cQA archive that are relevant to user's query. In this paper, we propose a graph based approach to segmenting multi-sentence questions. The results from user studies show that our segmentation model outperforms traditional systems in question segmentation by over 30% in user's satisfaction. We incorporate the segmentation model into existing cQA question retrieval framework for more targeted question matching, and the empirical evaluation results demonstrate that the segmentation boosts the question retrieval performance by up to 12.93% in Mean Average Precision and 11.72% in Top One Precision. Our model comes with a comprehensive question detector equipped with both lexical and syntactic features. Zhaoyan Ming, Xia Ben Hu, Tat-Seng Chua |
SIGIR | 2 |
| 2009 | Video reference: question answering on YouTubeabstractCommunity-based question answering systems have become very popular for providing answers to a wide variety of "how-to" questions. However most such systems present only textual answers. In many cases, users would prefer visual answers such as videos which are more direct and intuitive. Guangda Li, Zhaoyan Ming, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2009 | A syntactic tree matching approach to finding similar questions in community-based qa servicesabstractWhile traditional question answering (QA) systems tailored to the TREC QA task work relatively well for simple questions, they do not suffice to answer real world questions. The community-based QA systems offer this service well, as they contain large archives of such questions where manually crafted answers are directly available. However, finding similar questions in the QA archive is not trivial. In this paper, we propose a new retrieval framework based on syntactic tree structure to tackle the similar question matching problem. We build a ground-truth set from Yahoo! Answers, and experimental results show that our method outperforms traditional bag-of-word or tree kernel based methods by 8.3% in mean average precision. It further achieves up to 50% improvement by incorporating semantic features as well as matching of potential answers. Our model does not rely on training, and it is demonstrated to be robust against grammatical errors as well. Zhaoyan Ming, Tat-Seng Chua |
SIGIR | 2 |