EDBT 2026 Demo / reviewers in the wild / expert
Ruofei Zhang
dblp:36/2351
· DBLP profile ↗
54ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0002-4063-0109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdDetector: Detecting Chinese Advertorials on Social Media Platforms with Textual and Social InformationabstractWith the widespread use of social media platforms and people’s increasing dependence on them, social media has emerged as one of the most important channels for advertorials. However, there is currently a lack of research on detecting advertorials on social media platforms. This research focuses on detecting advertorials, a type of advertisement that frequently conceals itself within normal articles, blurring the nature of advertising and deceiving users. To effectively carry out research on advertorial detection, we have constructed a multi-topic advertorial dataset in Chinese with rich social information. This dataset is obtained from the Chinese question-answering platform ZHIHU, and it is publicly available to facilitate further research. 1 Furthermore, we propose AdDetector, a novel dual-tower model that detects advertorials by jointly leveraging the article’s textual and social information. In addition, we use fine-grained sentence-level classification labels to improve the model’s generalization capability on previously unseen topic articles. Experiment results show that our model significantly improves the \(F_1\) score by 1.29% in the intra-domain advertorial detection setting and 1.52% in the transfer setting in comparison with several strong baselines. The extensive ablation studies and thorough performance analyses also validate the complementary and beneficial values of the novel components of AdDetector. We also make our source code publicly available to facilitate future studies. 2 This research provides crucial support for user protection and advertising management. Haitao Bai, Pinghui Wang, Ruofei Zhang, Zi Liang, Ziyang Zhou 0003, Zhou Su 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Exploring Intrinsic Alignments Within Text CorpusabstractRecent years have witnessed rapid advancements in the safety alignments of large language models (LLMs). Methods such as supervised instruction fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) have thus emerged as vital components in constructing LLMs. While these methods achieve robust and fine-grained alignment to human values, their practical application is still hindered by high annotation costs and incomplete human alignments. Besides, the intrinsic human values within training corpora have not been fully exploited. To address these issues, we propose ISAAC (Intrinsically Supervised Alignments by Assessing Corpus), a primary and coarse-grained safety alignment strategy for LLMs. ISAAC only relies on a prior assumption about the text corpus, and does not require preferences in RLHF or human responses selection in SFT. Specifically, it assumes a long-tail distribution of text corpus and employs a specialized sampling strategy to automatically sample high-quality responses. Theoretically, we prove that this strategy can improve the safety of LLMs under our assumptions. Empirically, our evaluations on mainstream LLMs show that ISAAC achieves a safety score comparable to current SFT solutions. Moreover, we conduct experiments on ISAAC for some RLHF-based LLMs, where we find that ISAAC can even improve the safety of these models under specific safety domains. These findings demonstrate that ISAAC can provide preliminary alignment to LLMs, thereby reducing the construction costs of existing human-feedback-based methods. Zi Liang, Pinghui Wang, Ruofei Zhang, Haibo Hu 0001, Qingqing Ye 0001, Nuo Xu 0012, Yaxin Xiao |
AAAI | 3 |
| 2025 | Cross-Modal 3D Representation with Multi-View Images and Point CloudsabstractThe advancement of 3D understanding and representation is a crucial step for the next phase of autonomous driving, robotics, augmented and virtual reality, 3D gaming and 3D e-commerce products. However, existing 3D semantic representation research has primarily focused on point clouds to perceive 3D objects and scenes, overlooking the rich visual details offered by multi-view images, thereby limiting the potential of 3D semantic representation. This paper introduces OpenView, a novel representation method that integrates both point clouds and multi-view images to form a unified 3D representation. OpenView comprises a unique fusion framework, sequence-independent modeling, a cross-modal fusion encoder, and a progressive hard learning strategy. Our experiments demonstrate that OpenView outperforms the state-of-the-art by 11.5% and 5.5% on the R@1 metric for cross-modal retrieval and the Top-1 metric for zero-shot classification tasks, respectively. Furthermore, we showcase some applications of OpenView: 3D retrieval, 3D captioning and hierarchical data clustering, highlighting its generality in the field of 3D representation learning. Ziyang Zhou 0003, Pinghui Wang, Zi Liang, Haitao Bai, Ruofei Zhang |
CVPR | 5 |
| 2025 | MutationGuard: A Graph and Temporal-Spatial Neural Method for Detecting Mutation Telecommunication FraudabstractTelecommunication fraud refers to deceptive activities in the field of communication services. This research focuses on a category of fraud identified as ''mutation telecommunication fraud". There is currently a lack of research on mutation telecommunication fraud detection, allowing this type of fraud to persist uncaught. We identify that detecting mutation fraud requires capturing multi-source patterns, including user communication graphs and temporal-spatial Voice of Call (VOC) features. Specifically, we introduce MutationGuard, which leverages Graph Neural Networks (GNN) to capture changes in user communication graphs. For VOC records, we map call start times onto a 3D cylindrical surface, thereby representing each VOC record in spatial coordinates and utilizing proposed LFFE and TCFE modules to capture local fraud behaviors and temporal behavior changes. The proposed neural modeling approach that facilitates multi-source information fusion constitutes a significant advancement in detecting mutation fraud. Experiment results reveal a significant improvement in the AUC score by 1.52% and the F1 score by 1.36% on the proposed telecommunication fraud dataset. Particularly, our method shows a significant improvement of 13.93% in accuracy on mutation fraud data. We also validate the effectiveness of our method on the publicly available Sichuan Telecommunication Fraud dataset. Haitao Bai, Pinghui Wang, Ruofei Zhang, Ziyang Zhou 0003, Juxiang Zeng, Yulou Su, Zhou Su 0001, Li-Zhen Cui 0001, Wei Wang 0012 |
IJCAI | 3 |
| 2024 | MERGE: Fast Private Text GenerationabstractThe drastic increase in language models' parameters has led to a new trend of deploying models in cloud servers, raising growing concerns about private inference for Transformer-based models. Existing two-party privacy-preserving techniques, however, only take into account natural language understanding (NLU) scenarios. Private inference in natural language generation (NLG), crucial for applications like translation and code completion, remains underexplored. In addition, previous privacy-preserving techniques suffer from convergence issues during model training and exhibit poor inference speed when used with NLG models due to the neglect of time-consuming operations in auto-regressive generations. To address these issues, we propose a fast private text generation framework for Transformer-based language models, namely MERGE. MERGE reuses the output hidden state as the word embedding to bypass the embedding computation and reorganize the linear operations in the Transformer module to accelerate the forward procedure. Extensive experiments show that MERGE achieves a 26.5x speedup to the vanilla encrypted model under the sequence length 512, and reduces 80% communication cost, with an up to 10x speedup to state-of-the-art approximated models. Zi Liang, Pinghui Wang, Ruofei Zhang, Nuo Xu 0012, Lifeng Xing, Haitao Bai, Ziyang Zhou 0003 |
AAAI | 3 |
| 2024 | PAIR: Pre-denosing Augmented Image Retrieval Model for Defending Adversarial PatchesabstractDeep neural networks are widely used in retrieval systems. However, they are notoriously vulnerable to attack. Among the various forms of adversarial attacks, the patch attack is one of the most threatening forms. This type of attack can introduce cognitive biases into the retrieval system by inserting deceptive patches into images. Despite the seriousness of this threat, there are still no well-established solutions in image retrieval systems. In this paper, we propose the Pre-denosing Augmented Image Retrieval (PAIR) model, a new approach designed to protect image retrieval systems against adversarial patch attacks. The core strategy of PAIR is to dynamically and randomly reconstruct entire images based on their semantic content. This purifies well-designed patch attacks while preserving the semantic integrity of the images. Furthermore, we present a novel training strategy that incorporates a semantic discriminator. This discriminator significantly improves PAIR's ability to capture real semantics and reconstruct images. Experiments show that PAIR significantly outperforms existing defense methods. It effectively reduces the success rate of two state-of-the-art patch attack methods to below 5%, achieving a 14% improvement over current leading methods. Moreover, in defending against global perturbation attacks, PAIR also achieves competitive results. Ziyang Zhou 0003, Pinghui Wang, Zi Liang, Ruofei Zhang, Haitao Bai |
ACM Multimedia | 4 |
| 2023 | SegFormer: A Topic Segmentation Model with Controllable Range of AttentionabstractTopic segmentation aims to reveal the latent structure of a document and divide it into multiple parts. However, current neural solutions are limited in the context modeling of sentences and feature representation of candidate boundaries. This causes the model to suffer from inefficient sentence context encoding and noise information interference. In this paper, we design a new text segmentation model SegFormer with unidirectional attention blocks to better model sentence representations. To alleviate the problem of noise information interference, SegFormer uses a novel additional context aggregator and a topic classification loss to guide the model to aggregate the information within the appropriate range. In addition, SegFormer applies an iterative prediction algorithm to search for optimal boundaries progressively. We evaluate SegFormer's generalization ability, multilingual ability, and application ability on multiple challenging real-world datasets. Experiments show that our model significantly improves the performance by 7.5% on the benchmark WIKI-SECTION compared to several strong baselines. The application of SegFormer to a real-world dataset to separate normal and advertisement segments in product marketing essays also achieves superior performance in the evaluation with other cutting-edge models. Haitao Bai, Pinghui Wang, Ruofei Zhang, Zhou Su 0001 |
AAAI | 3 |
| 2022 | SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad RelevanceabstractAd relevance modeling plays a critical role in online advertising systems including Microsoft Bing. To leverage powerful transformers like BERT in this low-latency setting, many existing approaches perform ad-side computations offline. While efficient, these approaches are unable to serve cold start ads, resulting in poor relevance predictions for such ads. This work aims to design a new, low-latency BERT via structured pruning to empower real-time online inference for cold start ads relevance on a CPU platform. Our challenge is that previous methods typically prune all layers of the transformer to a high, uniform sparsity, thereby producing models which cannot achieve satisfactory inference speed with an acceptable accuracy. Li Lyna Zhang, Youkow Homma, Yujing Wang 0002, Mao Yang 0004, Ruofei Zhang, Ting Cao 0003 |
CIKM | 6 |
| 2022 | Taming Sparsely Activated Transformer with Stochastic Experts
Simiao Zuo, Xiaodong Liu 0003, Jian Jiao 0007, Young Jin Kim 0006, Hany Hassan, Ruofei Zhang, Jianfeng Gao 0001, Tuo Zhao |
ICLR | 6 |
| 2022 | AdsCVLR: Commercial Visual-Linguistic Representation Modeling in Sponsored SearchabstractSponsored search advertisements (ads) appear next to search results when consumers look for products and services on search engines. As the fundamental basis of search ads, relevance modeling has attracted increasing attention due to the significant research challenges and tremendous practical value. In this paper, we address the problem of multi-modal modeling in sponsored search, which models the relevance between user query and commercial ads with multi-modal structured information. To solve this problem, we propose a transformer architecture with Ads data on Commercial Visual-Linguistic Representation (AdsCVLR) with contrastive learning that naturally extends the transformer encoder with the complementary multi-modal inputs, serving as a strong aggregator of image-text features. We also make a public advertising dataset, which includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on. Empirically, we evaluate the AdsCVLR model over the large industry dataset, and the experimental results of online/offline tests show the superiority of our method. Yongjie Zhu, Chunhui Han, Yuefeng Zhan, Bochen Pang, Zhaoju Li, Hao Sun 0015, Si Li 0001, Boxin Shi, Nan Duan 0001, Ruofei Zhang, Liangjie Zhang, Qi Zhang 0066 |
ACM Multimedia | 11 |
| 2022 | Enhancing Self-Attention with Knowledge-Assisted Attention MapsabstractJiangang Bai, Yujing Wang, Hong Sun, Ruonan Wu, Tianmeng Yang, Pengfei Tang, Defu Cao, Mingliang Zhang1, Yunhai Tong, Yaming Yang, Jing Bai, Ruofei Zhang, Hao Sun, Wei Shen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jiangang Bai, Yujing Wang 0002, Ruonan Wu, Tianmeng Yang, Defu Cao, Mingliang Zhang 0004, Yunhai Tong, Yaming Yang 0001, Jing Bai 0010, Ruofei Zhang, Hao Sun 0015 |
NAACL-HLT | 12 |
| 2021 | Improving Taxonomy-based Categorization with Categorical Graph Neural NetworksabstractIn search and retrieval, a critical subtask is the classification of user search queries into predefined categories. Traditional supervised multi-class classification algorithms usually treat each category independently. In practical applications however, the categories have implicit relationships. Categories are organized as a tree-based taxonomy, which can be viewed as a graph. In this work, we explore a novel and systematic way of leveraging semantic information for improving taxonomy-based categorization. We propose a class of graph-based network structures, which we call Categorical Graph Neural Networks (CaGNN). CaGNNs leverage relationship information between neighbor categories and overlay the semantic information for each category, thus improving the performance of query categorization. The CaGNN framework can integrate a baseline categorizer with any Graph Neural Network, such as the commonly used Graph Attention Network and Graph Convolutional Network. Over a query categorization dataset of 2k categories and another ad title categorization dataset of 5k categories, CaGNN improves categorizers’ performance significantly compared to a baseline Deep Neural Network model without the CaGNN structure. Notably top 3 prediction recall increases from 90.15% to 91.40% for the ad title categorization task, for which is quite significant at over 90% level for more than 5k categories. By inspecting the learned category embeddings and the flow of message passing, we show that CaGNN effectively encapsulates useful graph structural information. Online A/B testing result shows that an ad ranking model with CaGNN-based features has increased ad click-through rate by 1.81% and reduced defect rate by 2.64%. The model has been deployed to production. Tianchuan Du, Keng-hao Chang, Paul Liu 0001, Ruofei Zhang |
IEEE BigData | 4 |
| 2021 | HittER: Hierarchical Transformers for Knowledge Graph EmbeddingsabstractThis paper examines the challenging problem of learning representations of entities and relations in a complex multi-relational knowledge graph.We propose HittER, a Hierarchical Transformer model to jointly learn Entityrelation composition and Relational contextualization based on a source entity's neighborhood.Our proposed model consists of two different Transformer blocks: the bottom block extracts features of each entity-relation pair in the local neighborhood of the source entity and the top block aggregates the relational information from outputs of the bottom block.We further design a masked entity prediction task to balance information from the relational context and the source entity itself.Experimental results show that HittER achieves new stateof-the-art results on multiple link prediction datasets.We additionally propose a simple approach to integrate HittER into BERT and demonstrate its effectiveness on two Freebase factoid question answering datasets. Sanxing Chen, Xiaodong Liu 0003, Jianfeng Gao 0001, Jian Jiao 0007, Ruofei Zhang, Yangfeng Ji |
EMNLP (1) | 5 |
| 2021 | BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingabstractIn this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available. Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001 |
ICML | 10 |
| 2021 | EL-Attention: Memory Efficient Lossless Attention for GenerationabstractTransformer model with multi-head attention requires caching intermediate results for efficient inference in generation tasks. However, cache brings new memory-related costs and prevents leveraging larger batch size for faster speed. We propose memory-efficient lossless attention (called EL-attention) to address this issue. It avoids heavy operations for building multi-head keys and values, cache for them is not needed. EL-attention constructs an ensemble of attention results by expanding query while keeping key and value shared. It produces the same result as multi-head attention with less GPU memory and faster inference speed. We conduct extensive experiments on Transformer, BART, and GPT-2 for summarization and question generation tasks. The results show EL-attention speeds up existing models by 1.6x to 5.3x without accuracy loss. Jiusheng Chen, Weizhen Qi, Nikhil Bhendawade, Yeyun Gong, Nan Duan 0001, Ruofei Zhang |
ICML | 7 |
| 2021 | Mask Attention Networks: Rethinking and Strengthen TransformerabstractZhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001 |
NAACL-HLT | 8 |
| 2021 | GalaXC: Graph Neural Networks with Labelwise Attention for Extreme ClassificationabstractThis paper develops the GalaXC algorithm for Extreme Classification, where the task is to annotate a document with the most relevant subset of labels from an extremely large label set. Extreme classification has been successfully applied to several real world web-scale applications such as web search, product recommendation, query rewriting, etc. GalaXC identifies two critical deficiencies in leading extreme classification algorithms. First, existing approaches generally assume that documents and labels reside in disjoint sets, even though in several applications, labels and documents cohabit the same space. Second, several approaches, albeit scalable, do not utilize various forms of metadata offered by applications, such as label text and label correlations. To remedy these, GalaXC presents a framework that enables collaborative learning over joint document-label graphs at massive scales, in a way that naturally allows various auxiliary sources of information, including label metadata, to be incorporated. GalaXC also introduces a novel label-wise attention mechanism to meld high-capacity extreme classifiers with its framework. An efficient end-to-end implementation of GalaXC is presented that could be trained on a dataset with 50M labels and 97M training documents in less than 100 hours on 4 × V100 GPUs. This allowed GalaXC to not only scale to applications with several millions of labels, but also be up to 18% more accurate than leading deep extreme classifiers, while being upto 2-50 × faster to train and 10 × faster to predict on benchmark datasets. GalaXC is particularly well-suited to warm-start scenarios where predictions need to be made on data points with partially revealed label sets, and was found to be up to 25% more accurate than extreme classification algorithms specifically designed for warm start settings. In A/B tests conducted on the Bing search engine, GalaXC could improve the Click Yield (CY) and coverage by 1.52% and 1.11% respectively. Code for GalaXC is available at https://github.com/Extreme-classification/GalaXC Deepak Saini, Arnav Kumar Jain, Kushal Dave 0001, Jian Jiao 0007, Amit Singh 0003, Ruofei Zhang, Manik Varma |
WWW | 6 |
| 2021 | TextGNN: Improving Text Encoder via Graph Neural Network in Sponsored SearchabstractText encoders based on C-DSSM or transformers have demonstrated strong performance in many Natural Language Processing (NLP) tasks. Low latency variants of these models have also been developed in recent years in order to apply them in the field of sponsored search which has strict computational constraints. However these models are not the panacea to solve all the Natural Language Understanding (NLU) challenges as the pure semantic information in the data is not sufficient to fully identify the user intents. We propose the TextGNN model that naturally extends the strong twin tower structured encoders with the complementary graph information from user historical behaviors, which serves as a natural guide to help us better understand the intents and hence generate better language representations. The model inherits all the benefits of twin tower models such as C-DSSM and TwinBERT so that it can still be used in the low latency environment while achieving a significant performance gain than the strong encoder-only counterpart baseline models in both offline evaluations and online production system. In offline experiments, the model achieves a 0.14% overall increase in ROC-AUC with a 1% increased accuracy for long-tail low-frequency Ads, and in the online A/B testing, the model shows a 2.03% increase in Revenue Per Mille with a 2.32% decrease in Ad defect rate. Jason Zhu, Yanling Cui, Hao Sun 0015, Markus Pelger, Liangjie Zhang, Ruofei Zhang, Huasha Zhao |
WWW | 9 |
| 2020 | AutoADR: Automatic Model Design for Ad RelevanceabstractLarge-scale pre-trained models have attracted extensive attention in the research community and shown promising results on various tasks of natural language processing. However, these pre-trained models are memory and computation intensive, hindering their deployment into industrial online systems like Ad Relevance. Meanwhile, how to design an effective yet efficient model architecture is another challenging problem in online Ad Relevance. Recently, AutoML shed new lights on architecture design, but how to integrate it with pre-trained language models remains unsettled. In this paper, we propose AutoADR (Automatic model design for AD Relevance) --- a novel end-to-end framework to address this challenge, and share our experience to ship these cutting-edge techniques into online Ad Relevance system at Microsoft Bing. Specifically, AutoADR leverages a one-shot neural architecture search algorithm to find a tailored network architecture for Ad Relevance. The search process is simultaneously guided by knowledge distillation from a large pre-trained teacher model (e.g. BERT), while taking the online serving constraints (e.g. memory and latency) into consideration. We add the model designed by AutoADR as a sub-model into the production Ad Relevance model. This additional sub-model improves the Precision-Recall AUC (PR AUC) on top of the original Ad Relevance model by 2.65X of the normalized shipping bar. More importantly, adding this automatically designed sub-model leads to a statistically significant 4.6% Bad-Ad ratio reduction in online A/B testing. This model has been shipped into Microsoft Bing Ad Relevance Production model. Yaming Yang 0001, Yujing Wang 0002, Yunhai Tong, Jing Bai 0010, Ruofei Zhang |
CIKM | 10 |
| 2020 | TwinBERT: Distilling Knowledge to Twin-Structured Compressed BERT Models for Large-Scale RetrievalabstractPre-trained language models have achieved great success in a wide variety of natural language processing (NLP) tasks, while the superior performance comes with high demand in computational resources, which hinders the application in low-latency information retrieval (IR) systems. To address the problem, we present TwinBERT model, which has two improvements: 1) represent query and document separately using twin-structured encoders and 2) each encoder is a highly compressed BERT-like model with less than one third of the parameters. The former allows document embeddings to be pre-computed offline and cached in memory, which is different from BERT, where the two input sentences are concatenated and encoded together. The change saves large amount of computation time, however, it is still not sufficient for real-time retrieval considering the complexity of BERT model itself. To further reduce computational cost, a compressed multi-layer transformer encoder is proposed with special training strategies as a substitution of the original complex BERT encoder. Lastly, two versions of TwinBERT are developed to combine the query and keyword embeddings for retrieval and relevance tasks correspondingly. Both of them have met the real-time latency requirement and achieve close or on-par performance to BERT-Base model. Wenhao Lu, Jian Jiao 0007, Ruofei Zhang |
CIKM | 3 |
| 2020 | An Enhanced Knowledge Injection Model for Commonsense GenerationabstractCommonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics. Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang |
COLING | 9 |
| 2020 | Boosting Fashion Image Attributes Classification Performance with MT-GAN Training TechniqueabstractAutomatic understanding of the product images, particularly the image attributes such as pattern, color, category, material, department, etc., has become a prerequisite for many downstream applications in the e-commerce ecosystem such as visually similar product retrieval and recommendation, publisher website image content monetization. With the recent advancement of deep learning techniques, multi-task deep neural networks (MTDNN) have been widely adopted for multi-class multi-label fashion image classification tasks. In this work, we propose a training technique Multi-Task Generative Adversarial Network (MT-GAN) to improve fashion image classification performance with an additional image translation task. We evaluated the proposed technique on two real-world fashion image datasets and experimental results show that both image classification and image translation benefit from this joint training setting. Specifically the image classifier, even though trained with a smaller amount of data and fewer types of labeled information, manages to outperform existing image classification models such as WTBI and DARN by 54% and 20% respectively for category classification on DeepFashion benchmark. Two main contributors to the superior performance of the image classifier trained as the discriminator in the proposed MT-GAN framework are: 1) the classifier leverages additional amount of data and the regularization effect by taking on the task of a traditional discriminator. 2) The classifier takes not only the RGB image but also the conditional input image shared between the generator and itself as its input, boosting the classification accuracy and recall by 18%. The proposed MT-GAN training technique is GAN-agnostic and can be applied to various state-of-the-art designs of generator and discriminator as well. Qun Li 0003, Changbo Hu, Keng-hao Chang, Ruofei Zhang |
DSAA | 4 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 12 |
| 2020 | ProphetNet-Ads: A Looking Ahead Strategy for Generative Retrieval Models in Sponsored Search Engine
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Ruofei Zhang, Houqiang Li, Nan Duan 0001, Ming Zhou 0001 |
NLPCC (2) | 6 |
| 2019 | Generating Better Search Engine Text Advertisements with Deep Reinforcement LearningabstractDeep Reinforcement Learning has been applied in a number of fields to directly optimize non-differentiable reward functions, including in sequence to sequence settings using Self Critical Sequence Training (SCST). Previously, SCST has primarily been applied to bring conditional language models closer to the distribution of their training set, as in traditional neural machine translation and abstractive summarization. We frame the generation of search engine text ads as a sequence to sequence problem, and consider two related goals: to generate ads similar to those a human would write, and to generate ads with high click-through rates. We jointly train a model to minimize cross-entropy on an existing corpus of Landing Page/Text Ad pairs using typical sequence to sequence training techniques while also optimizing the expected click-through rate (CTR) as predicted by an existing oracle model using SCST. Through joint training we achieve a 6.7% increase in expected CTR without a meaningful drop in ROUGE score. Human experiments demonstrate that SCST training produces significantly more attractive ads without reducing grammatical quality. J. Weston Hughes, Keng-hao Chang, Ruofei Zhang |
KDD | 3 |
| 2018 | Rare Query Expansion Through Generative Adversarial Networks in Search AdvertisingabstractGenerative Adversarial Networks (GAN) have achieved great success in generating realistic synthetic data like images, tags, and sentences. We explore using GAN to generate bid keywords directly from query in sponsored search ads selection, especially for rare queries. Specifically, in the query expansion (query-keyword matching) scenario in search advertising, we train a sequence to sequence model as the generator to generate keywords, conditioned on the user query, and use a recurrent neural network model as the discriminator to play an adversarial game with the generator. By applying the trained generator, we can generate keywords directly from a given query, so that we can highly improve the effectiveness and efficiency of query-keyword matching based ads selection in search advertising. We trained the proposed model in the clicked query-keyword pair dataset from a commercial search advertising system. Evaluation results show that the generated keywords are more relevant to the given query compared with the baseline model and they have big potential to bring extra revenue improvement. Mu-Chu Lee, Bin Gao 0001, Ruofei Zhang |
KDD | 3 |
| 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level ParallelizationabstractRelevance ranking models based on additive ensembles of regression trees have shown quite good effectiveness in web search engines. In the era of big data, tree ensemble models grow large in both tree depth and ensemble size to provide even better search relevance and user experience. However, the computational cost for their scoring process is high, such that it becomes a challenging issue to apply the big tree ensemble models in a search engine which needs to answer thousands of queries per second. Although several works have been proposed to improve the scoring process, the challenge is still great especially when the model size grows large. In this paper, we present RapidScorer , a novel framework for speeding up the scoring process of industry-scale tree ensemble models, without hurting the quality of scoring results. RapidScorer introduces a modified run length encoding called epitome to the bitvector representation of the tree nodes. Epitome can greatly reduce the computation cost to traverse the tree ensemble, and work with several other proposed strategies to maximize the compactness of data units in memory. The achieved compactness makes it possible to fully utilize data parallelization to improve model scalability. Experiments on two web search benchmarks show that, RapidScorer achieves significant speed-up over the state-of-the-art methods: V-QuickScorer , ranging from 1.3x to 3.5x; QuickScorer , ranging from 2.1x to 25.0x; VPred , ranging from 2.3x to 18.3x; and XGBoost , ranging from 2.6x to 42.5x. Ting Ye, Hucheng Zhou, Will Y. Zou, Bin Gao 0001, Ruofei Zhang |
KDD | 5 |
| 2017 | DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural NetworksabstractInformation extraction and user intention identification is a central topic in modern query understanding and recommendation systems. In this paper, we propose DeepProbe, a generic information-directed interaction framework which is built around an attention-based sequence to sequence (seq2seq) recurrent neural network. DeepProbe can rephrase, evaluate, and even actively ask questions, leveraging the generative ability and likelihood estimation made possible by seq2seq models. DeepProbe makes decisions based on a derived uncertainty (entropy) measure conditioned on user inputs, possibly with multiple rounds of interactions. Three applications, namely a rewritter, a relevance scorer and a chatbot for ad recommendation, were built around DeepProbe, with the first two serving as precursory building blocks for the third. We first use the seq2seq model in DeepProbe to rewrite a user query into one of standard query form, which is submitted to an ordinary recommendation system. Secondly, we evaluate DeepProbe's seq2seq model-based relevance scoring. Finally, we build a chatbot prototype capable of making active user interactions, which can ask questions that maximize information gain, allowing for a more efficient user intention idenfication process. We evaluate first two applications by 1) comparing with baselines by BLEU and AUC, and 2) human judge evaluation. Both demonstrate significant improvements compared with current state-of-the-art systems, proving their values as useful tools on their own, and at the same time laying a good foundation for the ongoing chatbot application. Zi Yin, Keng-hao Chang, Ruofei Zhang |
KDD | 3 |
| 2016 | DeepIntent: Learning Attentions for Online Advertising with Recurrent Neural NetworksabstractIn this paper, we investigate the use of recurrent neural networks (RNNs) in the context of search-based online advertising. We use RNNs to map both queries and ads to real valued vectors, with which the relevance of a given (query, ad) pair can be easily computed. On top of the RNN, we propose a novel attention network, which learns to assign attention scores to different word locations according to their intent importance (hence the name DeepIntent). The vector output of a sequence is thus computed by a weighted sum of the hidden states of the RNN at each word according their attention scores. We perform end-to-end training of both the RNN and attention network under the guidance of user click logs, which are sampled from a commercial search engine. We show that in most cases the attention network improves the quality of learned vector representations, evaluated by AUC on a manually labeled dataset. Moreover, we highlight the effectiveness of the learned attention scores from two aspects: query rewriting and a modified BM25 metric. We show that using the learned attention scores, one is able to produce sub-queries that are of better qualities than those of the state-of-the-art methods. Also, by modifying the term frequency with the attention scores in a standard BM25 formula, one is able to improve its performance evaluated by AUC. Shuangfei Zhai, Keng-hao Chang, Ruofei Zhang, Zhongfei Zhang |
KDD | 3 |
| 2014 | Large-Scale Multi-Label Learning with Incomplete Label AssignmentsabstractMulti-label learning deals with the classification problems where each instance can be assigned with multiple labels simultaneously. Conventional multi-label learning approaches mainly focus on exploiting label correlations. It is usually assumed, explicitly or implicitly, that the label sets for training instances are fully labeled without any missing labels. However, in many real-world multi-label datasets, the label assignments for training instances can be incomplete. Some ground-truth labels can be missed by the labeler from the label set. This problem is especially typical when the number instances is very large, and the labeling cost is very high, which makes it almost impossible to get a fully labeled training set. In this paper, we study the problem of large-scale multi-label learning with incomplete label assignments. We propose an approach, called Mpu, based upon positive and unlabeled stochastic gradient descent and stacked models. Unlike prior works, our method can effectively and efficiently consider missing labels and label correlations simultaneously, and is very scalable, that has linear time complexities over the size of the data. Extensive experiments on two real-world multi-label datasets show that our Mpu model consistently outperform other commonly-used baselines. Xiangnan Kong, Zhaoming Wu, Li-Jia Li 0001, Ruofei Zhang, Philip S. Yu, Wei Fan 0001 |
SDM | 4 |
| 2013 | A Lipschitz Exploration-Exploitation Scheme for Bayesian Optimization
Ali Jalali, Javad Azimi, Xiaoli Z. Fern, Ruofei Zhang |
ECML/PKDD (1) | 4 |
| 2012 | Visual appearance of display ads and its effect on click through rateabstractOne of the most important categories of online advertising is display advertising which provides publishers with significant revenue. Similar to other categories, the main goal in display advertising is to maximize user response rate for advertising campaigns, such as click through rates (CTR) or conversion rates. Previous studies have tried to optimize these parameters using objectives such as behavioral targeting. However, there is no published work so far to address the effect of the visual appearance of ads (creatives) on user response rate via a systematic data-driven approach. In this paper, we quantitatively study the relationship between the visual appearance and performance of creatives using large scale data in the world's largest display ads exchange system, RightMedia. We designed a set of 43 visual features, some of which are novel and others are inspired by related work. We extracted these features from real creatives served on RightMedia. We also designed and conducted a series of experiments to evaluate the effectiveness of visual features for CTR prediction, ranking and performance classification. Based on the evaluation results, we selected a subset of features that have the highest impact on CTR. We believe that the findings presented in this paper will be very useful for the online advertising industry in designing high-performance creatives. It also provides the research community with the first ever data set, initial insights into visual appearance's effect on user response propensity, and evaluation benchmarks for further study. Javad Azimi, Ruofei Zhang, Yang Zhou 0033, Vidhya Navalpakkam, Jianchang Mao, Xiaoli Z. Fern |
CIKM | 2 |
| 2012 | Multiview hierarchical bayesian regression model andapplication to online advertisingabstractWith the development of Web applications, large scale data are popular; and they are not only getting richer, but also ubiquitously interconnected with users and other objects in various ways, which brings about multi-view data with implicit structure. In this paper, we propose a novel hierarchical Bayesian mixture regression model, which discovers and then exploits the relationships among multiple views of the data to perform various machine learning tasks. A stochastic EM inference and learning algorithm is derived; and a parallel implementation in Hadoop MapReduce [9] paradigm is developed to scale up the learning. We apply the developed model and algorithm on click-through-rate (CTR) prediction and campaign targeting recommendation in online advertising to measure its effectiveness. The experiments on both synthetic data and large scale ads serving data from a real world online advertising exchange demonstrate the superior CTR prediction accuracy of our method compared to existing state-of-the-art methods. The results also show that our model can recommend high performance targeting features for online advertising campaigns. Tianbing Xu, Ruofei Zhang |
CIKM | 2 |
| 2012 | Multimedia features for click prediction of new ads in display advertisingabstractNon-guaranteed display advertising (NGD) is a multi-billion dollar business that has been growing rapidly in recent years. Advertisers in NGD sell a large portion of their ad campaigns using performance dependent pricing models such as cost-per-click (CPC) and cost-per-action (CPA). An accurate prediction of the probability that users click on ads is a crucial task in NGD advertising because this value is required to compute the expected revenue. State-of-the-art prediction algorithms rely heavily on historical information collected for advertisers, users and publishers. Click prediction of new ads in the system is a challenging task due to the lack of such historical data. The objective of this paper is to mitigate this problem by integrating multimedia features extracted from display ads into the click prediction models. Multimedia features can help us capture the attractiveness of the ads with similar contents or aesthetics. In this paper we evaluate the use of numerous multimedia features (in addition to commonly used user, advertiser and publisher features) for the purposes of improving click prediction in ads with no history. We provide analytical results generated over billions of samples and demonstrate that adding multimedia features can significantly improve the accuracy of click prediction for new ads, compared to a state-of-the-art baseline model. Haibin Cheng, Roelof van Zwol, Javad Azimi, Eren Manavoglu, Ruofei Zhang, Yang Zhou 0033, Vidhya Navalpakkam |
KDD | 5 |
| 2011 | Bid landscape forecasting in online ad exchange marketplaceabstractDisplay advertising has been a significant source of revenue for publishers and ad networks in online advertising ecosystem. One important business model in online display advertising is Ad Exchange marketplace, also called non-guaranteed delivery (NGD), in which advertisers buy targeted page views and audiences on a spot market through real-time auction. In this paper, we describe a bid landscape forecasting system in NGD marketplace for any advertiser campaign specified by a variety of targeting attributes. In the system, the impressions that satisfy the campaign targeting attributes are partitioned into multiple mutually exclusive samples. Each sample is one unique combination of quantified attribute values. We develop a divide-and-conquer approach that breaks down the campaign-level forecasting problem. First, utilizing a novel star-tree data structure, we forecast the bid for each sample using non-linear regression by gradient boosting decision trees. Then we employ a mixture-of-log-normal model to generate campaign-level bid distribution based on the sample-level forecasted distributions. The experiment results of a system developed with our approach show that it can accurately forecast the bid distributions for various campaigns running on the world's largest NGD advertising exchange system, outperforming two baseline methods in term of forecasting errors. Ruofei Zhang, Wei Li 0010, Jianchang Mao |
KDD | 2 |
| 2011 | Internet multimedia advertising: techniques and technologiesabstractThe explosive growth of multimedia data on the Internet creates huge opportunities for multimedia advertising. In this tutorial, we present the techniques and technologies for Internet multimedia advertising. The tutorial aims at bringing together recent insights from the research on multimedia advertising that addresses the theoretical fundamentals, solution concepts, and the issues related to the development of modern multimedia advertising schemes. Tao Mei 0001, Ruofei Zhang, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2011 | A stochastic learning-to-rank algorithm and its application to contextual advertisingabstractThis paper is concerned with the problem of learning a model to rank objects (Web pages, ads and etc.). We propose a framework where the ranking model is both optimized and evaluated using the same information retrieval measures such as Normalized Discounted Cumulative Gain (NDCG) and Mean Average Precision (MAP). The main difficulty in direct optimization of NDCG and MAP is that these measures depend on the rank of objects and are not differentiable. Most learning-to-rank methods that attempt to optimize NDCG or MAP approximate such measures so that they can be differentiable. In this paper, we propose a simple yet effective stochastic optimization algorithm to directly minimize any loss function, which can be defined on NDCG or MAP for the learning-to-rank problem. The algorithm employs Simulated Annealing along with Simplex method for its parameter search and finds the global optimal parameters. Experiment results using NDCG-Annealing algorithm, an instance of the proposed algorithm, on LETOR benchmark data sets show that the proposed algorithm is both effective and stable when compared to the baselines provided in LETOR 3.0. In addition, we applied the algorithm for ranking ads in contextual advertising. Our method has shown to significantly improve relevance in offline evaluation and business metrics in online tests in a real large-scale advertising serving system. To scale our computations, we parallelize the algorithm in a MapReduce framework running on Hadoop. Maryam Karimzadehgan, Wei Li 0010, Ruofei Zhang, Jianchang Mao |
WWW | 3 |
| 2010 | Exploitation and exploration in a performance based contextual advertising systemabstractThe dynamic marketplace in online advertising calls for ranking systems that are optimized to consistently promote and capitalize better performing ads. The streaming nature of online data inevitably makes an advertising system choose between maximizing its expected revenue according to its current knowledge in short term (exploitation) and trying to learn more about the unknown to improve its knowledge (exploration), since the latter might increase its revenue in the future. The exploitation and exploration (EE) tradeoff has been extensively studied in the reinforcement learning community, however, not been paid much attention in online advertising until recently. In this paper, we develop two novel EE strategies for online advertising. Specifically, our methods can adaptively balance the two aspects of EE by automatically learning the optimal tradeoff and incorporating confidence metrics of historical performance. Within a deliberately designed offline simulation framework we apply our algorithms to an industry leading performance based contextual advertising system and conduct extensive evaluations with real online event log data. The experimental results and detailed analysis reveal several important findings of EE behaviors in online advertising and demonstrate that our algorithms perform superiorly in terms of ad reach and click-through-rate (CTR). Wei Li 0010, Ruofei Zhang, Jianchang Mao, Rong Jin 0001 |
KDD | 3 |
| 2009 | Towards developing a unified multimodal image retrieval frameworkabstractBuilt upon the previous work on automatic image annotation and multimodal image retrieval, in this paper we present a unified multimodal image retrieval framework, called UPMIR. The contributions of this paper include: (1) the development of the UPMIR framework; and (2) extensive evaluations of UPMIR including evaluations against a state-of-the-art image retrieval system in a large scale, visually and semantically diverse database crawled from the Web to demonstrate that UPMIR not only has more effective and efficient retrieval performance, but also facilitates more enhanced retrieval modalities. Zhongfei Zhang, Ruofei Zhang |
ICME | 3 |
| 2009 | Learning to Rank by Optimizing NDCG MeasureabstractLearning to rank is a relatively new field of study, aiming to learn a ranking function from a set of training data with relevancy labels. The ranking algorithms are often evaluated using Information Retrieval measures, such as Normalized Discounted Cumulative Gain [1] and Mean Average Precision [2]. Until recently, most learning to rank algorithms were not using a loss function related to the above mentioned evaluation measures. The main difficulty in direct optimization of these measures is that they depend on the ranks of documents, not the numerical values output by the ranking function. We propose a probabilistic framework that addresses this challenge by optimizing the expectation of NDCG over all the possible permutations of documents. A relaxation strategy is used to approximate the average of NDCG over the space of permutation, and a bound optimization approach is proposed to make the computation efficient. Extensive experiments show that the proposed algorithm outperforms state-of-the-art ranking algorithms on several benchmark data sets. Hamed Valizadegan, Rong Jin 0001, Ruofei Zhang, Jianchang Mao |
NIPS | 3 |
| 2007 | Effective Image Retrieval Based on Hidden Concept Discovery in Image DatabaseabstractThis paper addresses content-based image retrieval in general, and in particular, focuses on developing a hidden semantic concept discovery methodology to address effective semantics-intensive image retrieval. In our approach, each image in the database is segmented into regions associated with homogenous color, texture, and shape features. By exploiting regional statistical information in each image and employing a vector quantization method, a uniform and sparse region-based representation is achieved. With this representation, a probabilistic model based on statistical-hidden-class assumptions of the image database is obtained, to which the expectation-maximization technique is applied to analyze semantic concepts hidden in the database. An elaborated retrieval algorithm is designed to support the probabilistic model. The semantic similarity is measured through integrating the posterior probabilities of the transformed query image, as well as a constructed negative example, to the discovered semantic concepts. The proposed approach has a solid statistical foundation; the experimental evaluations on a database of 10000 general-purposed images demonstrate its promise and effectiveness. Ruofei Zhang, Zhongfei Zhang |
IEEE Trans. Image Process. | 1 |
| 2006 | BALAS: Empirical Bayesian learning in the relevance feedback for image retrieval
Ruofei Zhang, Zhongfei Zhang |
Image Vis. Comput. | 1 |
| 2006 | A probabilistic semantic model for image annotation and multi-modal image retrieval
Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
Multim. Syst. | 1 |
| 2005 | A Probabilistic Semantic Model for Image Annotation and Multi-Modal Image RetrievaabstractThis paper addresses automatic image annotation problem and its application to multi-modal image retrieval. The contribution of our work is three-fold. (1) We propose a probabilistic semantic model in which the visual features and the textual words are connected via a hidden layer which constitutes the semantic concepts to be discovered to explicitly exploit the synergy among the modalities. (2) The association of visual features and textual words is determined in a Bayesian framework such that the confidence of the association can be provided. (3) Extensive evaluation on a large-scale, visually and semantically diverse image collection crawled from Web is reported to evaluate the prototype system based on the model. In the proposed probabilistic model, a hidden concept layer which connects the visual feature and the word layer is discovered by fitting a generative model to the training image and annotation words through an Expectation-Maximization (EM) based iterative learning procedure. The evaluation of the prototype system on 17,000 images and 7,736 automatically extracted annotation words from crawled Web pages for multi-modal image retrieval has indicated that the proposed semantic model and the developed Bayesian framework are superior to a state-of-the-art peer system in the literature. Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
ICCV | 1 |
| 2005 | Image Database Classification based on Concept Vector ModelabstractAutomatic semantic classification of image databases is very useful for users' searching and browsing, but it is at the same time a very challenging research problem as well. In this paper, we develop a hidden semantic concept discovery methodology to address effective semantics-intensive image database classification. Each image is segmented into regions and then a uniform and sparse region-based representation is obtained. With this representation a probabilistic model based on statistical-hidden-class assumptions of the image database is proposed, to which the Expectation-Maximization (EM) technique is applied to analyze semantic concepts hidden in the database. Two methods are proposed to make use of the semantic concepts discovered from the probabilistic model for unsupervised and supervised image database classifications, respectively, based on the automatically learned concept vectors. It is shown that the concept vectors are more reliable and robust and thus promising than the low level features through the theoretic analysis and the experimental evaluations on a database of 10,000 general-purpose images. Ruofei Zhang, Zhongfei Zhang |
ICME | 1 |
| 2005 | FAST: Toward more effective and efficient image retrieval
Ruofei Zhang, Zhongfei Zhang |
Multim. Syst. | 1 |
| 2004 | Hidden Semantic Concept Discovery in Region Based Image Retrieval
Ruofei Zhang, Zhongfei Zhang |
CVPR (2) | 1 |
| 2004 | Stretching Bayesian Learning in the Relevance Feedback of Image Retrieval
Ruofei Zhang, Zhongfei Zhang |
ECCV (3) | 1 |
| 2004 | Exploiting the cognitive synergy between different media modalities in multimodal information retrievalabstractThis is a position paper reporting an on-going collaboration project between SUNY Binghamton, USA, and Waseda University, Japan, on multimodal information retrieval through exploiting the cognitive synergy across the different modalities of the information, to facilitate an effective retrieval. Specifically, we focus on image retrieval in the applications where imagery data appear along with collateral text. It is noted that these applications are ubiquitous. We have proposed the synergistic indexing scheme (SIS) to explicitly exploit the synergy between the information of imagery and text modalities. Since the synergy we have exploited between the information of imagery and text modalities is subjective and depends on specific cognitive context, we call this type of synergy as cognitive synergy. We have reported part of the empirical evaluation and are in the process of fully implementing the SIS prototype for an extensive evaluation. Zhongfei Zhang, Ruofei Zhang, Jun Ohya |
ICME | 2 |
| 2004 | Semantic repository modeling in image databaseabstractThis work is about content based image database retrieval, focusing on developing a classification based methodology to address semantics-intensive image retrieval. With self organization map based image feature grouping, a visual dictionary is created for color, texture, and shape feature attributes, respectively. Labeling each training image with the keywords in the visual dictionary, a classification tree is built. Based on the statistical properties of the feature space we define a structure, called /spl alpha/-semantics graph, to discover the hidden semantic relationships among the semantic repositories embodied in the image database. With the /spl alpha/-semantics graph, each semantic repository is modeled as a unique fuzzy set to explicitly address the semantic uncertainty and the semantic overlap existing among the repositories in the feature space. A retrieval algorithm combining the classification tree with the fuzzy set models to deliver semantically relevant image retrieval is provided. The experimental evaluations have demonstrated that the proposed approach models the semantic relationships effectively and outperforms a state-of-the-art content based image retrieval system in the literature both in effectiveness and efficiency. Ruofei Zhang, Zhongfei Zhang, Zhongyuan Qin |
ICME | 1 |
| 2004 | A data mining approach to modeling relationships among categories in image collectionabstractThis paper proposes a data mining approach to modeling relationships among categories in image collection. In our approach, with image feature grouping, a visual dictionary is created for color, texture, and shape feature attributes respectively. Labeling each training image with the keywords in the visual dictionary, a classification tree is built. Based on the statistical properties of the feature space we define a structure, called α-Semantics Graph, to discover the hidden semantic relationships among the semantic categories embodied in the image collection. With the α-Semantics Graph, each semantic category is modeled as a unique fuzzy set to explicitly address the semantic uncertainty and semantic overlap among the categories in the feature space. The model is utilized in the semantics-intensive image retrieval application. An algorithm using the classification accuracy measures is developed to combine the built classification tree with the fuzzy set modeling method to deliver semantically relevant image retrieval for a given query image. The experimental evaluations have demonstrated that the proposed approach models the semantic relationships effectively and the image retrieval prototype system utilizing the derived model is promising both in effectiveness and efficiency. Ruofei Zhang, Zhongfei Zhang, Sandeep Khanzode |
KDD | 1 |
| 2003 | A 3D Modeling Scheme for Cerebral Vasculature from MRA DatasetsabstractThis paper proposes an integrative approach that facilitates physicians to semi automatically obtain a 3-D symbolic representation of cerebral vasculature from 3-D magnetic resonance angiography (MRA) datasets. In this approach, firstly vessels are segmented by morphology method followed by 3-D parallel thinning to obtain the one voxel wide skeleton. Then a novel method employing general tree and its combinations is introduced to depict the 3-D geometrical structure of the vasculature. With the generated tree, post processing, such as traversal and visualization, is implemented. The method has been tested on both synthetic images and real images; the results are promising. A system based on this approach provides a useful visualization tool of the intracerebral vasculature for clinic applications. Zhongyuan Qin, Xuanqin Mou, Ruofei Zhang |
CBMS | 3 |
| 2003 | A Unified Fuzzy Feature Indexing Scheme for Region Based Online Image QueryingabstractThis paper describes a novel indexing and retrieval methodology integrating color, texture and shape information for content-based image retrieval in online image databases. This methodology, called PicSearcher, applies unsupervised image segmentation to partition an image into a set of regions, then fuzzy color histogram as well as fuzzy texture and shape properties of each region is calculated to be part of their signatures. The fuzzification procedures resolve the recognition uncertainty stemming from color quantization and human perception of colors. At the same time, this unified fuzzy scheme incorporates the segmentation-related uncertainties into the retrieval algorithm. Then an adaptive and effective measure for the overall similarity between images is developed by integrating properties of all the regions in the image. An implemented prototype system of PicSearcher has demonstrated a promising retrieval performance for an online test database containing 10,000 general-purpose color images, as compared with its peer systems in the literature. 1. Ruofei Zhang, Zhongfei Zhang, Jian Yao 0003 |
Web Intelligence | 1 |
| 2002 | A Clustering Based Approach to Efficient Image RetrievalabstractThis paper addresses the issue of effective and efficient content based image retrieval by presenting a novel indexing and retrieval methodology that integrates color, texture, and shape information for the indexing and retrieval, and applies these features in regions obtained through unsupervised segmentation, as opposed to applying them to the whole image domain. In order to address the typical color feature "inaccuracy" problem in the literature, fuzzy logic is applied to the traditional color histogram to solve for the problem to a certain degree. The similarity is defined through a balanced combination between global and regional similarity measures incorporating all the features. In order to further improve the retrieval efficiency, a secondary clustering technique is developed and employed to significantly save query processing time without compromising the retrieval precision. An implemented prototype system has demonstrated a promising retrieval performance for a test database containing 2000 general-purpose color images, as compared with its peer systems in the literature. Ruofei Zhang, Zhongfei Zhang |
ICTAI | 1 |