Huan-Bo Luan

dblp:55/5812 · also Huanbo Luan · DBLP profile ↗
← Back
72ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-3938-119XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding and Grounding
Fengbin Zhu, Xiang Yao Ng, Haohui Wu, Wenjie Wang 0007, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Tat-Seng Chua
MMM (1)8
2025 Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in Finance
abstract
Finance decision-making often relies on in-depth data analysis across various data sources, including financial tables, news articles, stock prices, etc. In this work, we introduce FINTMMBench, the first comprehensive benchmark for evaluating temporal-aware multi-modal Retrieval-Augmented Generation (RAG) systems in finance. Built from heterologous data of NASDAQ 100 companies, FINTMMBench offers three significant advantages. 1) Multi-modal Corpus: It encompasses a hybrid of financial tables, news articles, daily stock prices, and visual technical charts as the corpus. 2) Temporal-aware Questions: Each question requires the retrieval and interpretation of its relevant data over a specific time period, including daily, weekly, monthly, quarterly, and annual periods. 3) Diverse Financial Analysis Tasks: The questions involve 10 different financial analysis tasks designed by domain experts, including information extraction, trend analysis, sentiment analysis and event detection, etc. We further propose a novel TMMHybridRAG method, which first leverages a multi-modal LLM to convert data from other modalities (e.g., tabular, visual and time-series data) into textual format and then incorporates temporal information in each node when constructing graphs and dense indexes. Its effectiveness has been validated in extensive experiments, but notable gaps remain, highlighting the challenges presented by our FINTMMBench. The benchmark and source code will be made publicly available.
Fengbin Zhu, Liangming Pan, Wenjie Wang 0007, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia7
2025 FinIR: The 2nd Workshop on Financial Information Retrieval in the Era of Generative AI
abstract
Recent advancements in Generative AI, such as Large Language Models (LLMs), have demonstrated remarkable success across various general tasks. Extensive studies have explored leveraging generative models in finance, but significant challenges persist. This half-day workshop explores potential approaches and research directions to address these challenges by equipping generative models with advanced Information Retrieval (IR) models. Specifically, this workshop seeks to provide a platform for discussing innovative ideas that facilitate the advancement of IR technology to enrich generative models in finance from four key perspectives: (i) financial IR techniques (ii) financial IR benchmarking and evaluation (iii) financial systems and agents/assistants (iv) and trustworthiness, privacy and security when applying financial IR and generative models. This workshop aims to deepen understanding, accelerate progress, and support the advancement of IR technology to enhance generative models to address financial challenges.
Fengbin Zhu, Yunshan Ma 0002, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Guangnan Ye, Shuo Zhang 0006, Dhagash Mehta, Pingping Chen 0004, Bing Xiang, Tat-Seng Chua
SIGIR5
2022 Data augmentation for low-resource languages NMT guided by constrained sampling
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
Int. J. Intell. Syst.3
2021 Segment, Mask, and Predict: Augmenting Chinese Word Segmentation with Self-Supervision
abstract
Mieradilijiang Maimaiti, Yang Liu, Yuanhang Zheng, Gang Chen, Kaiyu Huang, Ji Zhang, Huanbo Luan, Maosong Sun. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Mieradilijiang Maimaiti, Yang Liu 0005, Yuanhang Zheng, Gang Chen 0039, Ji Zhang 0011, Huan-Bo Luan, Maosong Sun 0001
EMNLP (1)7
2021 Self-Supervised Quality Estimation for Machine Translation
abstract
Yuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti, Huanbo Luan, Maosong Sun, Qun Liu, Yang Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Yuanhang Zheng, Zhixing Tan, Meng Zhang 0019, Mieradilijiang Maimaiti, Huan-Bo Luan, Maosong Sun 0001, Qun Liu 0001, Yang Liu 0005
EMNLP (1)5
2021 Metaphor identification: A contextual inconsistency based neural sequence labeling approach
Xin Chen 0070, Zhen Hai, Suge Wang, Deyu Li 0001, Chao Wang 0049, Huan-Bo Luan
Neurocomputing6
2021 Improving Data Augmentation for Low-Resource NMT Guided by POS-Tagging and Paraphrase Embedding
abstract
Data augmentation is an approach for several text generation tasks. Generally, in the machine translation paradigm, mainly in low-resource language scenarios, many data augmentation methods have been proposed. The most used approaches for generating pseudo data mainly lay in word omission, random sampling, or replacing some words in the text. However, previous methods barely guarantee the quality of augmented data. In this work, we try to build the data by using paraphrase embedding and POS-Tagging. Namely, we generate the fake monolingual corpus by replacing the main four POS-Tagging labels, such as noun, adjective, adverb, and verb, based on both the paraphrase table and their similarity. We select the bigger corpus size of the paraphrase table with word level and obtain the word embedding of each word in the table, then calculate the cosine similarity between these words and tagged words in the original sequence. In addition, we exploit the ranking algorithm to choose highly similar words to reduce semantic errors and leverage the POS-Tagging replacement to mitigate syntactic error to some extent. Experimental results show that our augmentation method consistently outperforms all previous SOTA methods on the low-resource language pairs in seven language pairs from four corpora by 1.16 to 2.39 BLEU points.
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Zegao Pan, Maosong Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Learning to Generate Explainable Plots for Neural Story Generation
abstract
Story generation is an important natural language processing task that aims to generate coherent stories automatically. While the use of neural networks has proven effective in improving story generation, how to learn to generate an explainable high-level plot still remains a major challenge. In this article, we propose a latent variable model for neural story generation. The model treats an outline, which is a natural language sentence explainable to humans, as a latent variable to represent a high-level plot that bridges the input and output. We adopt an external summarization model to guide the latent variable model to learn how to generate outlines from training data. Experiments show that our approach achieves significant improvements over state-of-the-art methods in both automatic and human evaluations.
Gang Chen 0039, Yang Liu 0005, Huan-Bo Luan, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Neural Machine Translation With Explicit Phrase Alignment
abstract
While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads to three problems: it is hard to (1) interpret the translation process, (2) impose lexical constraints, and (3) impose structural constraints. These problems not only increase the difficulty of designing new architectures for neural machine translation, but also limit its applications in practice. To alleviate these problems, we propose to introduce explicit phrase alignment into the translation process of arbitrary neural machine translation models. The key idea is to build a search space similar to that of phrase-based statistical machine translation for neural machine translation where phrase alignment is readily available. We design a new decoding algorithm that can easily impose lexical and structural constraints. Experiments show that our approach makes the translation process of neural machine translation more interpretable without sacrificing translation quality. In addition, our approach achieves significant improvements in lexically and structurally constrained translation tasks.
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Neural Diffusion Model for Microscopic Cascade Study
abstract
The study of information diffusion or cascade has attracted much attention over the last decade. Most related works target on studying cascade-level macroscopic properties such as the final size of a cascade. Existing microscopic cascade models which focus on user-level modeling either make strong assumptions on how a user gets infected by a cascade or limit themselves to a specific scenario where “who infected whom” information is explicitly labeled. The strong assumptions oversimplify the complex diffusion mechanism and prevent these models from better fitting real-world cascade data. Also, the methods which focus on specific scenarios cannot be generalized to a general setting where the diffusion graph is unobserved. To overcome the drawbacks of previous works, we propose a Neural Diffusion Model (NDM) for general microscopic cascade study. NDM makes relaxed assumptions and employs deep learning techniques including attention mechanism and convolutional network for cascade modeling. Both advantages enable our model to go beyond the limitations of previous methods, better fit the diffusion data and generalize to unseen cascades. Experimental results on diffusion identification task over four realistic cascade datasets show that our model can achieve a relative improvement up to 26 percent against the best performing baseline in terms of F1 score.
Cheng Yang 0002, Maosong Sun 0001, Shiyi Han, Zhiyuan Liu 0001, Huan-Bo Luan
IEEE Trans. Knowl. Data Eng.6
2020 Modeling Voting for System Combination in Machine Translation
abstract
System combination is an important technique for combining the hypotheses of different machine translation systems to improve translation performance. Although early statistical approaches to system combination have been proven effective in analyzing the consensus between hypotheses, they suffer from the error propagation problem due to the use of pipelines. While this problem has been alleviated by end-to-end training of multi-source sequence-to-sequence models recently, these neural models do not explicitly analyze the relations between hypotheses and fail to capture their agreement because the attention to a word in a hypothesis is calculated independently, ignoring the fact that the word might occur in multiple hypotheses. In this work, we propose an approach to modeling voting for system combination in machine translation. The basic idea is to enable words in hypotheses from different systems to vote on words that are representative and should get involved in the generation process. This can be done by quantifying the influence of each voter and its preference for each candidate. Our approach combines the advantages of statistical and neural methods since it can not only analyze the relations between hypotheses but also allow for end-to-end training. Experiments show that our approach is capable of better taking advantage of the consensus between hypotheses and achieves significant improvements over state-of-the-art baselines on Chinese-English and English-German machine translation tasks.
Xuancheng Huang, Zhixing Tan, Derek F. Wong, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
IJCAI5
2020 Graph Random Neural Networks for Semi-Supervised Learning on Graphs
abstract
We study the problem of semi-supervised learning on graphs, for which graph neural networks (GNNs) have been extensively explored. However, most existing GNNs inherently suffer from the limitations of over-smoothing, non-robustness, and weak-generalization when labeled nodes are scarce. In this paper, we propose a simple yet effective framework—GRAPH RANDOM NEURAL NETWORKS (GRAND)—to address these issues. In GRAND, we first design a random propagation strategy to perform graph data augmentation. Then we leverage consistency regularization to optimize the prediction consistency of unlabeled nodes across different data augmentations. Extensive experiments on graph benchmark datasets suggest that GRAND significantly outperforms state-of- the-art GNN baselines on semi-supervised node classification. Finally, we show that GRAND mitigates the issues of over-smoothing and non-robustness, exhibiting better generalization behavior than existing GNNs. The source code of GRAND is publicly available at https://github.com/Grand20/grand.
Wenzheng Feng, Jie Zhang 0078, Yuxiao Dong, Yu Han 0001, Huan-Bo Luan, Qian Xu 0005, Qiang Yang 0001, Evgeny Kharlamov, Jie Tang 0001
NeurIPS5
2019 Learning to Copy for Automatic Post-Editing
abstract
Xuancheng Huang, Yang Liu, Huanbo Luan, Jingfang Xu, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xuancheng Huang, Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
EMNLP/IJCNLP (1)3
2019 Improving Back-Translation with Uncertainty-based Confidence Estimation
abstract
Shuo Wang, Yang Liu, Chao Wang, Huanbo Luan, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuo Wang 0013, Yang Liu 0005, Chao Wang 0049, Huan-Bo Luan, Maosong Sun 0001
EMNLP/IJCNLP (1)4
2019 Stance Influences Your Thoughts: Psychology-Inspired Social Media Analytics
Weizhi Ma, Zhen Wang 0040, Min Zhang 0006, Huan-Bo Luan, Yiqun Liu 0001, Shaoping Ma
NLPCC (1)5
2019 Multi-Round Transfer Learning for Low-Resource NMT Using Multiple High-Resource Languages
abstract
Neural machine translation (NMT) has made remarkable progress in recent years, but the performance of NMT suffers from a data sparsity problem since large-scale parallel corpora are only readily available for high-resource languages (HRLs). In recent days, transfer learning (TL) has been used widely in low-resource languages (LRLs) machine translation, while TL is becoming one of the vital directions for addressing the data sparsity problem in low-resource NMT. As a solution, a transfer learning method in NMT is generally obtained via initializing the low-resource model (child) with the high-resource model (parent). However, leveraging the original TL to low-resource models is neither able to make full use of highly related multiple HRLs nor to receive different parameters from the same parents. In order to exploit multiple HRLs effectively, we present a language-independent and straightforward multi-round transfer learning (MRTL) approach to low-resource NMT. Besides, with the intention of reducing the differences between high-resource and low-resource languages at the character level, we introduce a unified transliteration method for various language families, which are both semantically and syntactically highly analogous with each other. Experiments on low-resource datasets show that our approaches are effective, significantly outperform the state-of-the-art methods, and yield improvements of up to 5.63 BLEU points.
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Improving the Transformer Translation Model with Document-Level Context
abstract
Although the Transformer translation model (Vaswani et al., 2017) has achieved state-ofthe-art performance in a variety of translation tasks, how to use document-level context to deal with discourse phenomena problematic for Transformer still remains a challenge.In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder.As large-scale document-level parallel corpora are usually not available, we introduce a two-step training method to take full advantage of abundant sentence-level parallel corpora and limited document-level parallel corpora.Experiments on the NIST Chinese-English datasets and the IWSLT French-English datasets show that our approach improves over Transformer significantly. 1
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Min Zhang 0005, Yang Liu 0005
EMNLP2
2018 Cross-Domain Depression Detection via Harvesting Social Media
abstract
Depression detection is a significant issue for human well-being. In previous studies, online detection has proven effective in Twitter, enabling proactive care for depressed users. Owing to cultural differences, replicating the method to other social media platforms, such as Chinese Weibo, however, might lead to poor performance because of insufficient available labeled (self-reported depression) data for model training. In this paper, we study an interesting but challenging problem of enhancing detection in a certain target domain (e.g. Weibo) with ample Twitter data as the source domain. We first systematically analyze the depression-related feature patterns across domains and summarize two major detection challenges, namely isomerism and divergency. We further propose a cross-domain Deep Neural Network model with Feature Adaptive Transformation & Combination strategy (DNN-FATC) that transfers the relevant information across heterogeneous domains. Experiments demonstrate improved performance compared to existing heterogeneous transfer methods or training directly in the target domain (over 3.4% improvement in F1), indicating the potential of our model to enable depression detection via social media for more countries with different cultural settings.
Tiancheng Shen, Jia Jia 0001, Guangyao Shen, Fuli Feng, Xiangnan He 0001, Huan-Bo Luan, Jie Tang 0001, Thanassis Tiropanis, Tat-Seng Chua, Wendy Hall 0001
IJCAI6
2017 Bilingual Lexicon Induction from Non-Parallel Data with Minimal Supervision
abstract
Building bilingual lexica from non-parallel data is a long-standing natural language processing research problem that could benefit thousands of resource-scarce languages which lack parallel data. Recent advances of continuous word representations have opened up new possibilities for this task, e.g. by establishing cross-lingual mapping between word embeddings via a seed lexicon. The method is however unreliable when there are only a limited number of seeds, which is a reasonable setting for resource-scarce languages. We tackle the limitation by introducing a novel matching mechanism into bilingual word representation learning. It captures extra translation pairs exposed by the seeds to incrementally improve the bilingual word embeddings. In our experiments, we find the matching mechanism to substantially improve the quality of the bilingual vector space, which in turn allows us to induce better bilingual lexica with seeds as few as 10.
Meng Zhang 0019, Haoruo Peng, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
AAAI4
2017 Visualizing and Understanding Neural Machine Translation
abstract
While neural machine translation (NMT) has made remarkable progress in recent years, it is hard to interpret its internal workings due to the continuous representations and non-linearity of neural networks.In this work, we propose to use layer-wise relevance propagation (LRP) to compute the contribution of each contextual word to arbitrary hidden states in the attention-based encoderdecoder framework.We show that visualization with LRP helps to interpret the internal workings of NMT and analyze translation errors.
Yanzhuo Ding, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACL (1)3
2017 Adversarial Training for Unsupervised Bilingual Lexicon Induction
abstract
Word embeddings are well known to capture linguistic regularities of the language on which they are trained.Researchers also observe that these regularities can transfer across languages.However, previous endeavors to connect separate monolingual word embeddings typically require cross-lingual signals as supervision, either in the form of parallel corpus or seed lexicon.In this work, we show that such cross-lingual connection can actually be established without any form of supervision.We achieve this end by formulating the problem as a natural adversarial game, and investigating techniques that are crucial to successful training.We carry out evaluation on the unsupervised bilingual lexicon induction task.Even though this task appears intrinsically cross-lingual, we are able to demonstrate encouraging performance without any cross-lingual clues.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACL (1)3
2017 Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization
abstract
Although neural machine translation has made significant progress recently, how to integrate multiple overlapping, arbitrary prior knowledge sources remains a challenge.In this work, we propose to use posterior regularization to provide a general framework for integrating prior knowledge into neural machine translation.We represent prior knowledge sources as features in a log-linear model, which guides the learning process of the neural translation model.Experiments on Chinese-English translation show that our approach leads to significant improvements.
Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
ACL (1)3
2017 Earth Mover's Distance Minimization for Unsupervised Bilingual Lexicon Induction
abstract
Cross-lingual natural language processing hinges on the premise that there exists invariance across languages. At the word level, researchers have identified such invariance in the word embedding semantic spaces of different languages. However, in order to connect the separate spaces, cross-lingual supervision encoded in parallel data is typically required. In this paper, we attempt to establish the cross-lingual connection without relying on any cross-lingual supervision. By viewing word embedding spaces as distributions, we propose to minimize their earth mover's distance, a measure of divergence between distributions. We demonstrate the success on the unsupervised bilingual lexicon induction task. In addition, we reveal an interesting finding that the earth mover's distance shows potential as a measure of language difference.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
EMNLP3
2017 Image-embodied Knowledge Representation Learning
abstract
Entity images could provide significant visual information for knowledge representation learning. Most conventional methods learn knowledge representations merely from structured triples, ignoring rich visual information extracted from entity images. In this paper, we propose a novel Image-embodied Knowledge Representation Learning model (IKRL), where knowledge representations are learned with both triple facts and images. More specifically, we first construct representations for all images of an entity with a neural image encoder. These image representations are then integrated into an aggregated image-based representation via an attention-based method. We evaluate our IKRL models on knowledge graph completion and triple classification. Experimental results demonstrate that our models outperform all baselines on both tasks, which indicates the significance of visual information for knowledge representations and the capability of our models in learning knowledge representations with images.
Ruobing Xie, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001
IJCAI3
2017 PIC2DISH: A Customized Cooking Assistant System
abstract
The art of cooking is always fascinating. Nevertheless, reproducing a delicious dish that one has never encountered before is not easy. Even if the name of dish is known and the corresponding recipe could be retrieved, the right ingredients for cooking the dish may not be available due to factors such as geography region or season. Furthermore, knowing how to cut, cook and control timing may be challenging for one whose has no cooking experience. In this paper, an all-around cooking assistant mobile app, named Pic2Dish, is developed to help users who would like to cook a dish but neither know the name of dish nor has cooking skill. Basically, by inputting a picture of the dish and the list of ingredients at hand, Pic2Dish automatically recognizes the dish name and recommends a customized recipe together with video clips to guide user on how to cook the dish. Importantly, the recommended recipe is modified from a retrieved recipe that best matches the given dish, with missing ingredients being replaced with the available ingredients that match dish context and taste. The whole process involves the recognition of dishes with convolutional neural network, classification of key and non-key ingredients, and context analysis of ingredient relationship and their cooking/cutting methods. The user studies, which recruit real users to cook dishes by using Pic2Dish, shows the usefulness of the app.
Yongsheng An, Jingjing Chen 0001, Chong-Wah Ngo, Jia Jia 0001, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia6
2017 Understanding and Predicting Usefulness Judgment in Web Search
abstract
Usefulness judgment measures the user-perceived amount of useful information for the search task in the current search context. Understanding and predicting usefulness judgment are crucial for developing user-centric evaluation methods and providing contextualize results according to the search context. With a dataset collected in a laboratory user study, we systematically investigate the effects of a variety of content, context, and behavior factors on usefulness judgments and find that while user behavior factors are most important in determining usefulness judgments, content and context factors also have significant effects on it. We further adopt these factors as features to build prediction models for usefulness judgments. An AUC score of 0.909 in binary usefulness classification and a Pearson's correlation coefficient of 0.694 in usefulness regression demonstrate the effectiveness of our models. Our study sheds light on the understanding of the dynamics of the user-perceived usefulness of documents in a search session and provides implications for the evaluation and design of Web search engines.
Jiaxin Mao, Yiqun Liu 0001, Huan-Bo Luan, Min Zhang 0006, Shaoping Ma, Hengliang Luo
SIGIR3
2017 Neural Parse Combination
Liner Yang, Maosong Sun 0001, Yong Cheng 0003, Zhenghao Liu 0001, Huan-Bo Luan, Yang Liu 0005
J. Comput. Sci. Technol.6
2017 PRISM: Profession Identification in Social Media
abstract
Profession is an important social attribute of people. It plays a crucial role in commercial services such as personalized recommendation and targeted advertising. In practice, profession information is usually unavailable due to privacy and other reasons. In this article, we explore the task of identifying user professions according to their behaviors in social media. The task confronts the following challenges that make it non-trivial: how to incorporate heterogeneous information of user behaviors, how to effectively utilize both labeled and unlabeled data, and how to exploit community structure. To address these challenges, we present a framework called Profession Identification in Social Media. It takes advantage of both personal information and community structure of users in the following aspects: (1) We present a cascaded two-level classifier with heterogeneous personal features to measure the confidence of users belonging to different professions. (2) We present a multi-training process to take advantages of both labeled and unlabeled data to enhance classification performance. (3) We design a profession identification method synthetically considering the confidences from personal features and community structure. We collect a real-world dataset to conduct experiments, and experimental results demonstrate the significant effectiveness of our method compared with other baseline methods. By applying prediction on large-scale users, we also analyze characteristics of microblog users, finding that there are significant diversities among users of different professions in demographics, social network structures, and linguistic styles.
Cunchao Tu, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001
ACM Trans. Intell. Syst. Technol.3
2017 Compact Indexing and Judicious Searching for Billion-Scale Microblog Retrieval
abstract
In this article, we study the problem of efficient top-kdisjunctive query processing in a huge microblog dataset. In terms of compact indexing, we categorize the keywords into rare terms and common terms based on inverse document frequency (idf) and propose tailored block-oriented organization to save memory consumption. In terms of fast searching, we classify the queries into three types based on term category and judiciously design an efficient search algorithm for each type. We conducted extensive experiments on a billion-scale Twitter dataset and examined the performance with both simple and more advanced ranking functions. The results showed that with much smaller index size, our search algorithm achieves a factor of 2--3 times faster speedup over state-of-the-art solutions in both ranking scenarios.
Dongxiang Zhang, Liqiang Nie, Huan-Bo Luan, Kian-Lee Tan, Tat-Seng Chua, Heng Tao Shen
ACM Trans. Inf. Syst.3
2016 Learning to Appreciate the Aesthetic Effects of Clothing
abstract
How do people describe clothing? The words like “formal”or "casual" are usually used. However, recent works often focus on recognizing or extracting visual features (e.g., sleeve length, color distribution and clothing pattern) from clothing images accurately. How can we bridge the gap between the visual features and the aesthetic words? In this paper, we formulate this task to a novel three-level framework: visual features(VF) - image-scale space (ISS) - aesthetic words space(AWS). Leveraging the art-field image-scale space served as an intermediate layer, we first propose a Stacked Denoising Autoencoder Guided by CorrelativeLabels (SDAE-GCL) to map the visual features to the image-scale space; and then according to the semantic distances computed byWordNet::Similarity, we map the most often used aesthetic words in online clothing shops to the image-scale space too. Employing upper body menswear images downloaded from several global online clothing shops as experimental data, the results indicate that the proposed three-level framework can help to capture the subtle relationship between visual features and aesthetic words better compared to several baselines. To demonstrate that our three-level framework and its implementation methods are universally applicable, we finally present some interesting analyses on the fashion trend of menswear in the last 10 years.
Jia Jia 0001, Guangyao Shen, Tao He 0016, Zhiyuan Liu 0001, Huan-Bo Luan
AAAI6
2016 Moodee: An Intelligent Mobile Companion for Sensing Your Stress from Your Social Media Postings
abstract
In this demo, we build a practical mobile application, Moodee,to help detect and release users’ psychological stress byleveraging users’ social media data in online social networks,and provide an interactive user interface to present users’and friends’ psychological stress states in an visualized andintuitional way.Given users’ online social media data as input, Moodee intelligentlyand automatically detects users’ stress states. Moreover,Moodee would recommend users with different linksto help release their stress. The main technology of this demois a novel hybrid model - a factor graph model combinedwith Deep Neural Network, which can leverage social mediacontent and social interaction information for stress detection.We think that Moodee can be helpful to people’s mentalhealth, which is a vital problem in modern world.
Huijie Lin, Jia Jia 0001, Enze Zhou, Jingtian Fu, Yejun Liu, Huan-Bo Luan
AAAI7
2016 Representation Learning of Knowledge Graphs with Entity Descriptions
abstract
Representation learning (RL) of knowledge graphs aims to project both entities and relations into a continuous low-dimensional space. Most methods concentrate on learning representations with knowledge triples indicating relations between entities. In fact, in most knowledge graphs there are usually concise descriptions for entities, which cannot be well utilized by existing methods. In this paper, we propose a novel RL method for knowledge graphs taking advantages of entity descriptions. More specifically, we explore two encoders, including continuous bag-of-words and deep convolutional neural models to encode semantics of entity descriptions. We further learn knowledge representations with both triples and descriptions. We evaluate our method on two tasks, including knowledge graph completion and entity classification. Experimental results on real-world datasets show that, our method outperforms other baselines on the two tasks, especially under the zero-shot setting, which indicates that our method is capable of building representations for novel entities according to their descriptions. The source code of this paper can be obtained from https://github.com/xrb92/DKRL.
Ruobing Xie, Zhiyuan Liu 0001, Jia Jia 0001, Huan-Bo Luan, Maosong Sun 0001
AAAI4
2016 Building Earth Mover's Distance on Bilingual Word Embeddings for Machine Translation
abstract
Following their monolingual counterparts, bilingual word embeddings are also on the rise. As a major application task, word translation has been relying on the nearest neighbor to connect embeddings cross-lingually. However, the nearest neighbor strategy suffers from its inherently local nature and fails to cope with variations in realistic bilingual word embeddings. Furthermore, it lacks a mechanism to deal with many-to-many mappings that often show up across languages. We introduce Earth Mover's Distance to this task by providing a natural formulation that translates words in a holistic fashion, addressing the limitations of the nearest neighbor. We further extend the formulation to a new task of identifying parallel sentences, which is useful for statistical machine translation systems, thereby expanding the application realm of bilingual word embeddings. We show encouraging performance on both tasks.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001, Tatsuya Izuha
AAAI3
2016 Discrete Image Hashing Using Large Weakly Annotated Photo Collections
abstract
We address the problem of image hashing by learning binary codes from large and weakly supervised photo collections. Due to the explosive growth of user generated media on the Web, this problem is becoming critical for large-scale visual applications like image retrieval. While most existing hashing methods fail to address this challenge well, our method shows promising improvement due to the following two key advantages.First, we formulate a novel hashing objective that can effectively mine implicit weak supervision by collaborative filtering. Second, we propose a discrete hashing algorithm, offered with efficient optimization, to overcome the inferior optimizations in obtaining binary codes from real-valued solutions. In this way, our method can be considered as a weakly-supervised discrete hashing framework which jointly learns image semantics and their corresponding binary codes. Through training on one million weakly annotated images, our experimental results demonstrate that image retrieval using the proposed hashing method outperforms the other state-of-the-art ones on image and video benchmarks.
Hanwang Zhang, Na Zhao 0004, Xindi Shang, Huan-Bo Luan, Tat-Seng Chua
AAAI4
2016 Neural Relation Extraction with Selective Attention over Instances
abstract
Distant supervised relation extraction has been widely used to find novel relational facts from text.However, distant supervision inevitably accompanies with the wrong labelling problem, and these noisy data will substantially hurt the performance of relation extraction.To alleviate this issue, we propose a sentence-level attention-based model for relation extraction.In this model, we employ convolutional neural networks to embed the semantics of sentences.Afterwards, we build sentence-level attention over multiple instances, which is expected to dynamically reduce the weights of those noisy instances.Experimental results on real-world datasets show that, our model can make full use of all informative sentences and effectively reduce the influence of wrong labelled instances.Our model achieves significant and consistent improvements on relation extraction as compared with baselines.The source code of this paper can be obtained from https: //github.com/thunlp/NRE.
Yankai Lin 0001, Shiqi Shen, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001
ACL (1)4
2016 Agreement-based Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora
abstract
We introduce an agreement-based approach to learning parallel lexicons and phrases from non-parallel corpora.The basic idea is to encourage two asymmetric latent-variable translation models (i.e., source-to-target and target-to-source) to agree on identifying latent phrase and word alignments.The agreement is defined at both word and phrase levels.We develop a Viterbi EM algorithm for jointly training the two unidirectional models efficiently.Experiments on the Chinese-English dataset show that agreementbased learning significantly improves both alignment and translation performance.
Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan, Heng Yu 0006
ACL (1)4
2016 Inducing Bilingual Lexica From Non-Parallel Data With Earth Mover's Distance Regularization
abstract
Being able to induce word translations from non-parallel data is often a prerequisite for cross-lingual processing in resource-scarce languages and domains. Previous endeavors typically simplify this task by imposing the one-to-one translation assumption, which is too strong to hold for natural languages. We remove this constraint by introducing the Earth Mover’s Distance into the training of bilingual word embeddings. In this way, we take advantage of its capability to handle multiple alternative word translations in a natural form of regularization. Our approach shows significant and consistent improvements across four language pairs. We also demonstrate that our approach is particularly preferable in resource-scarce settings as it only requires a minimal seed lexicon.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Yiqun Liu 0001, Maosong Sun 0001
COLING3
2016 Online Collaborative Learning for Open-Vocabulary Visual Classifiers
abstract
We focus on learning open-vocabulary visual classifiers, which scale up to a large portion of natural language vocabulary (e.g., over tens of thousands of classes). In particular, the training data are large-scale weakly labeled Web images since it is difficult to acquire sufficient well-labeled data at this category scale. In this paper, we propose a novel online learning paradigm towards this challenging task. Different from traditional N-way independent classifiers that generally fail to handle the extremely sparse and inter-related labels, our classifiers learn from continuous label embeddings discovered by collaboratively decomposing the sparse image-label matrix. Leveraging on the structure of the proposed collaborative learning formulation, we develop an efficient online algorithm that can jointly learn the label embeddings and visual classifiers. The algorithm can learn over 30,000 classes of 1,000 training images within 1 second on a standard GPU. Extensively experimental results on four benchmarks demonstrate the effectiveness of our method.
Hanwang Zhang, Xindi Shang, Wenzhuo Yang, Huan Xu 0001, Huan-Bo Luan, Tat-Seng Chua
CVPR5
2016 Predicting Search User Examination with Visual Saliency
abstract
Predicting users' examination of search results is one of the key concerns in Web search related studies. With more and more heterogeneous components federated into search engine result pages (SERPs), it becomes difficult for traditional position-based models to accurately predict users' actual examination patterns. Therefore, a number of prior works investigate the connection between examination and users' explicit interaction behaviors (e.g.~click-through, mouse movement). Although these works gain much success in predicting users' examination behavior on SERPs, they require the collection of large scale user behavior data, which makes it impossible to predict examination behavior on newly-generated SERPs. To predict user examination on SERPs containing heterogenous components without user interaction information, we propose a new prediction model based on visual saliency map and page content features. Visual saliency, which is designed to measure the likelihood of a given area to attract human visual attention, is used to predict users' attention distribution on heterogenous search components. With an experimental search engine, we carefully design a user study in which users' examination behavior (eye movement) is recorded. Examination prediction results based on this collected data set demonstrate that visual saliency features significantly improve the performance of examination model in heterogeneous search environments. We also found that saliency features help predict internal examination behavior within vertical results.
Yiqun Liu 0001, Zeyang Liu 0004, Ke Zhou 0003, Meng Wang 0001, Huan-Bo Luan, Chao Wang 0049, Min Zhang 0006, Shaoping Ma
SIGIR5
2016 Discrete Collaborative Filtering
abstract
We address the efficiency problem of Collaborative Filtering (CF) by hashing users and items as latent vectors in the form of binary codes, so that user-item affinity can be efficiently calculated in a Hamming space. However, existing hashing methods for CF employ binary code learning procedures that most suffer from the challenging discrete constraints. Hence, those methods generally adopt a two-stage learning scheme composed of relaxed optimization via discarding the discrete constraints, followed by binary quantization. We argue that such a scheme will result in a large quantization loss, which especially compromises the performance of large-scale CF that resorts to longer binary codes. In this paper, we propose a principled CF hashing framework called Discrete Collaborative Filtering (DCF), which directly tackles the challenging discrete optimization that should have been treated adequately in hashing. The formulation of DCF has two advantages: 1) the Hamming similarity induced loss that preserves the intrinsic user-item similarity, and 2) the balanced and uncorrelated code constraints that yield compact yet informative binary codes. We devise a computationally efficient algorithm with a rigorous convergence proof of DCF. Through extensive experiments on several real-world benchmarks, we show that DCF consistently outperforms state-of-the-art CF hashing techniques, e.g, though using only 8 bits, DCF is even significantly better than other methods using 128 bits.
Hanwang Zhang, Fumin Shen, Wei Liu 0005, Xiangnan He 0001, Huan-Bo Luan, Tat-Seng Chua
SIGIR5
2016 Listwise Ranking Functions for Statistical Machine Translation
abstract
Decision rules play an important role in the tuning and decoding steps of statistical machine translation. The traditional decision rule selects the candidate with the greatest potential from a candidate space by examining each candidate individually. However, viewing each candidate as independent imposes a serious limitation on the translation task. We instead view the problem from a ranking perspective that naturally allows the consideration of an entire list of candidates as a whole through the adoption of a listwise ranking function. Our shift from a pointwise to a listwise perspective proves to be a simple yet powerful extension to current modeling that allows arbitrary pairwise functions to be incorporated as features, whose weights can be estimated jointly with traditional ones. We further demonstrate that our formulation encompasses the minimum Bayes risk (MBR) approach, another decision rule that considers restricted listwise information, as a special case. Experiments show that our approach consistently outperforms the baseline and MBR methods across the considered test sets.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Deep Fusion of Multiple Semantic Cues for Complex Event Recognition
abstract
We present a deep learning strategy to fuse multiple semantic cues for complex event recognition. In particular, we tackle the recognition task by answering how to jointly analyze human actions (who is doing what), objects (what), and scenes (where). First, each type of semantic features (e.g., human action trajectories) is fed into a corresponding multi-layer feature abstraction pathway, followed by a fusion layer connecting all the different pathways. Second, the correlations of how the semantic cues interacting with each other are learned in an unsupervised cross-modality autoencoder fashion. Finally, by fine-tuning a large-margin objective deployed on this deep architecture, we are able to answer the question on how the semantic cues of who, what, and where compose a complex event. As compared with the traditional feature fusion methods (e.g., various early or late strategies), our method jointly learns the essential higher level features that are most effective for fusion and recognition. We perform extensive experiments on two real-world complex event video benchmarks, MED'11 and CCV, and demonstrate that our method outperforms the best published results by 21% and 11%, respectively, on an event recognition task.
Xishan Zhang, Hanwang Zhang, Yongdong Zhang 0001, Yang Yang 0002, Meng Wang 0001, Huan-Bo Luan, Jintao Li 0001, Tat-Seng Chua
IEEE Trans. Image Process.6
2016 Learning from Collective Intelligence: Feature Learning Using Social Images and Tags
abstract
Feature representation for visual content is the key to the progress of many fundamental applications such as annotation and cross-modal retrieval. Although recent advances in deep feature learning offer a promising route towards these tasks, they are limited in application domains where high-quality and large-scale training data are expensive to obtain. In this article, we propose a novel deep feature learning paradigm based on social collective intelligence, which can be acquired from the inexhaustible social multimedia content on the Web, in particular, largely social images and tags. Differing from existing feature learning approaches that rely on high-quality image-label supervision, our weak supervision is acquired by mining the visual-semantic embeddings from noisy, sparse, and diverse social image collections. The resultant image-word embedding space can be used to (1) fine-tune deep visual models for low-level feature extractions and (2) seek sparse representations as high-level cross-modal features for both image and text. We offer an easy-to-use implementation for the proposed paradigm, which is fast and compatible with any state-of-the-art deep architectures. Extensive experiments on several benchmarks demonstrate that the cross-modal features learned by our paradigm significantly outperforms others in various applications such as content-based retrieval, classification, and image captioning.
Hanwang Zhang, Xindi Shang, Huan-Bo Luan, Meng Wang 0001, Tat-Seng Chua
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Modeling Relation Paths for Representation Learning of Knowledge Bases
abstract
Representation learning of knowledge bases aims to embed both entities and relations into a low-dimensional space.Most existing methods only consider direct relations in representation learning.We argue that multiple-step relation paths also contain rich inference patterns between entities, and propose a path-based representation learning model.This model considers relation paths as translations between entities for representation learning, and addresses two key challenges: (1) Since not all relation paths are reliable, we design a path-constraint resource allocation algorithm to measure the reliability of relation paths.(2) We represent relation paths via semantic composition of relation embeddings.Experimental results on real-world datasets show that, as compared with baselines, our model achieves significant and consistent improvements on knowledge base completion and relation extraction from text.The source code of this paper can be obtained from https://github.com/mrlyk423/relation_extraction.
Yankai Lin 0001, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001, Siwei Rao
EMNLP3
2015 Generalized Agreement for Bidirectional Word Alignment
abstract
While agreement-based joint training has proven to deliver state-of-the-art alignment accuracy, the produced word alignments are usually restricted to one-toone mappings because of the hard constraint on agreement.We propose a general framework to allow for arbitrary loss functions that measure the disagreement between asymmetric alignments.The loss functions can not only be defined between asymmetric alignments but also between alignments and other latent structures such as phrase segmentations.We use a Viterbi EM algorithm to train the joint model since the inference is intractable.Experiments on Chinese-English translation show that joint training with generalized agreement achieves significant improvements over two state-ofthe-art alignment methods.
Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan, Heng Yu 0006
EMNLP4
2015 Online Learning of Interpretable Word Embeddings
abstract
Word embeddings encode semantic meanings of words into low-dimension word vectors.In most word embeddings, one cannot interpret the meanings of specific dimensions of those word vectors.Nonnegative matrix factorization (NMF) has been proposed to learn interpretable word embeddings via non-negative constraints.However, NMF methods suffer from scale and memory issue because they have to maintain a global matrix for learning.To alleviate this challenge, we propose online learning of interpretable word embeddings from streaming text data.Experiments show that our model consistently outperforms the state-of-the-art word embedding methods in both representation ability and interpretability.The source code of this paper can be obtained from http: //github.com/skTim/OIWE.
Hongyin Luo, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001
EMNLP3
2015 Consistency-Aware Search for Word Alignment
abstract
As conventional word alignment search algorithms usually ignore the consistency constraint in translation rule extraction, improving alignment accuracy does not necessarily increase translation quality.We propose to use coverage, which reflects how well extracted phrases can recover the training data, to enable word alignment to model consistency and correlate better with machine translation.This can be done by introducing an objective that maximizes both alignment model score and coverage.We introduce an efficient algorithm to calculate coverage on the fly during search.Experiments show that our consistency-aware search algorithm significantly outperforms both generative and discriminative alignment approaches across various languages and translation models.
Shiqi Shen, Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan
EMNLP4
2015 Joint Learning of Character and Word Embeddings
Xinxiong Chen, Lei Xu 0040, Zhiyuan Liu 0001, Maosong Sun 0001, Huan-Bo Luan
IJCAI5
2015 Iterative Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora
Meiping Dong, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001, Tatsuya Izuha, Dakun Zhang
IJCAI3
2015 Learning Features from Large-Scale, Noisy and Social Image-Tag Collection
abstract
Feature representation for multimedia content is the key to the progress of many fundamental multimedia tasks. Although recent advances in deep feature learning offer a promising route towards these tasks, they are limited in application to domains where high-quality and large-scale training data are hard to obtain. In this paper, we propose a novel deep feature learning paradigm based on large, noisy and social image-tag collections, which can be acquired from the inexhaustible social multimedia content on the Web. Instead of learning features from high-quality image-label supervision, we propose to learn from the image-word semantic relations, in a way of seeking a unified image-word embedding space, where the pairwise feature similarities preserve the semantic relations in the original image-word pairs. We offer an easy-to-use implementation for the proposed paradigm, which is fast and compatible for integrating into any state-of-the-art deep architectures. Experiments on NUSWIDE benchmark demonstrate that the features learned by our method significantly outperforms other state-of-the-art ones.
Hanwang Zhang, Xindi Shang, Huan-Bo Luan, Yang Yang 0002, Tat-Seng Chua
ACM Multimedia3
2015 Microblog Sentiment Analysis with Emoticon Space Model
Yiqun Liu 0001, Huan-Bo Luan, Jiashen Sun, Xuan Zhu 0006, Min Zhang 0006, Shaoping Ma
J. Comput. Sci. Technol.3
2015 Enhancing Video Event Recognition Using Automatically Constructed Semantic-Visual Knowledge Base
abstract
The task of recognizing events from video has attracted a lot of attention in recent years. However, due to the complex nature of user-defined events, the use of purely audio- visual content analysis without domain knowledge has been found to be grossly inadequate. In this paper, we propose to construct a semantic-visual knowledge base to encode the rich event-centric concepts and their relationships from the well- established lexical databases, including FrameNet, as well as the concept-specific visual knowledge from ImageNet. Based on this semantic-visual knowledge bases, we design an effective system for video event recognition. Specifically, in order to narrow the semantic gap between the high-level complex events and low-level visual representations, we utilize the event-centric semantic concepts encoded in the knowledge base as the intermediate-level event representation, which offers both human-perceivable and machine-interpretable semantic clues for event recognition. In addition, in order to leverage the abundant ImageNet images, we propose a robust transfer learning model to learn the noise- resistant concept classifiers for videos. Extensive experiments on various real-world video datasets demonstrate the superiority of our proposed system as compared to the state-of-the-art approaches.
Xishan Zhang, Yang Yang 0002, Yongdong Zhang 0001, Huan-Bo Luan, Jintao Li 0001, Hanwang Zhang, Tat-Seng Chua
IEEE Trans. Multim.4
2014 Brand Data Gathering From Live Social Media Streams
abstract
Social media streams, such as Twitter, Facebook, and Sina Weibo, have become essential real-time information resources with a wide range of users and applications. The rapidly increasing amount of live information in social media streams has important societal and marketing values for large corporations and government organizations. There is a strong need for effective techniques for data gathering and content analysis. This problem is particularly challenging due to the short and conversational nature of posts, the huge data volume, and the increasing heterogeneous multimedia content in social media streams. Moreover, as the focus of "conversation" often shifts quickly in social media space, the traditional keywords based approach to gather data with respect to a target brand is grossly inadequate. To address these problems, we propose a multi-faceted brand tracking method that gathers relevant data based on not just evolving keywords, but also social factors (users, relations and locations) as well as visual contents as increasing number of social media posts are in multimedia form. For evaluation, we set up a large scale microblog dataset (Brand-Social-Net) on brand/product information, containing 3 million microblogs with over 1.2 million images for 100 famous brands. Experiments on this dataset have demonstrated that the proposed framework is able to gather a more complete set of relevant brand-related data from live social media streams. We have released this dataset to promote social media research.
Yue Gao 0002, Fanglin Wang, Huan-Bo Luan, Tat-Seng Chua
ICMR3
2014 One of a Kind: User Profiling by Social Curation
abstract
Social Curation Service (SCS) is a new type of emerging social media platform, where users can select, organize and keep track of multimedia contents they like. In this paper, we take advantage of this great opportunity and target at the very starting point in social media: user profiling, which supports fundamental applications such as personalized search and recommendation. As compared to other profiling methods in conventional Social Network Services (SNS), our work benefits from the two distinguishable characteristics of SCS: a) organized multimedia user-generated contents, and b) content-centric social network. Based on these two characteristics, we are able to deploy the state-of-the-art multimedia analysis techniques to establish content-based user profiles by extracting user preferences and their social relations. First, we automatically construct a content-based user preference ontology and learn the ontological models to generate comprehensive user profiles. In particular, we propose a new deep learning strategy called multi-task convolutional neural network (mtCNN) to learn profile models and profile-related visual features simultaneously. Second, we propose to model the multi-level social relations offered by SCS to refine the user profiles in a low-rank recovery framework. To the best of our knowledge, our work is the first that explores how social curation can help in content-based social media technologies, taking user profiling as an example. Extensive experiments on 1,293 users and 1.5 million images collected from Pinterest in fashion domain demonstrate that recommendation methods based on the proposed user profiles are considerably more effective than other state-of-the-art recommendation strategies.
Xue Geng, Hanwang Zhang, Yang Yang 0002, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia5
2014 Start from Scratch: Towards Automatically Identifying, Modeling, and Naming Visual Attributes
abstract
Higher-level semantics such as visual attributes are crucial for fundamental multimedia applications. We present a novel attribute discovery approach that can automatically identify, model and name attributes from an arbitrary set of image and text pairs that can be easily gathered on the Web. Different from conventional attribute discovery methods, our approach does not rely on any pre-defined vocabularies and human labeling. Therefore, we are able to build a large visual knowledge base without any human efforts. The discovery is based on a novel deep architecture, named Independent Component Multimodal Autoencoder (ICMAE), that can continually learn shared higher-level representations across the visual and textual modalities. With the help of the resultant representations encoding strong visual and semantic evidences, we propose to (a) identify attributes and their corresponding high-quality training images, (b) iteratively model them with maximum compactness and comprehensiveness, and (c) name the attribute models with human understandable words. To date, the proposed system has discovered 1,898 attributes over 1.3 million pairs of image and text. Extensive experiments on various real-world multimedia datasets demonstrate the quality and effectiveness of the discovered attributes, facilitating multimedia applications such as image annotation and retrieval as compared to the state-of-the-art approaches.
Hanwang Zhang, Yang Yang 0002, Huan-Bo Luan, Shuicheng Yan, Tat-Seng Chua
ACM Multimedia3
2014 Social-oriented visual image search
Peng Cui 0001, Huan-Bo Luan, Wenwu Zhu 0001, Shiqiang Yang, Qi Tian 0001
Comput. Vis. Image Underst.3
2014 Single/cross-camera multiple-person tracking by graph matching
Weizhi Nie, Anan Liu, Yuting Su 0001, Huan-Bo Luan, Zhaoxuan Yang, Liujuan Cao, Rongrong Ji
Neurocomputing4
2014 Social-Sensed Image Search
abstract
Although Web search techniques have greatly facilitate users’ information seeking, there are still quite a lot of search sessions that cannot provide satisfactory results, which are more serious in Web image search scenarios. How to understand user intent from observed data is a fundamental issue and of paramount significance in improving image search performance. Previous research efforts mostly focus on discovering user intent either from clickthrough behavior in user search logs (e.g., Google), or from social data to facilitate vertical image search in a few limited social media platforms (e.g., Flickr). This article aims to combine the virtues of these two information sources to complement each other, that is, sensing and understanding users’ interests from social media platforms and transferring this knowledge to rerank the image search results in general image search engines. Toward this goal, we first propose a novel social-sensed image search framework, where both social media and search engine are jointly considered. To effectively and efficiently leverage these two kinds of platforms, we propose an example-based user interest representation and modeling method, where we construct a hybrid graph from social media and propose a hybrid random-walk algorithm to derive the user-image interest graph. Moreover, we propose a social-sensed image reranking method to integrate the user-image interest graph from social media and search results from general image search engines to rerank the images by fusing their social relevance and visual relevance. We conducted extensive experiments on real-world data from Flickr and Google image search, and the results demonstrated that the proposed methods can significantly improve the social relevance of image search results while maintaining visual relevance well.
Peng Cui 0001, Wenwu Zhu 0001, Huan-Bo Luan, Tat-Seng Chua, Shiqiang Yang
ACM Trans. Inf. Syst.4
2014 Memory recall based video search: Finding videos you have seen before based on your memory
abstract
We often remember images and videos that we have seen or recorded before but cannot quite recall the exact venues or details of the contents. We typically have vague memories of the contents, which can often be expressed as a textual description and/or rough visual descriptions of the scenes. Using these vague memories, we then want to search for the corresponding videos of interest. We call this “Memory Recall based Video Search” (MRVS). To tackle this problem, we propose a video search system that permits a user to input his/her vague and incomplete query as a combination of text query, a sequence of visual queries, and/or concept queries. Here, a visual query is often in the form of a visual sketch depicting the outline of scenes within the desired video, while each corresponding concept query depicts a list of visual concepts that appears in that scene. As the query specified by users is generally approximate or incomplete, we need to develop techniques to handle this inexact and incomplete specification by also leveraging on user feedback to refine the specification. We utilize several innovative approaches to enhance the automatic search. First, we employ a visual query suggestion model to automatically suggest potential visual features to users as better queries. Second, we utilize a color similarity matrix to help compensate for inexact color specification in visual queries. Third, we leverage on the ordering of visual queries and/or concept queries to rerank the results by using a greedy algorithm. Moreover, as the query is inexact and there is likely to be only one or few possible answers, we incorporate an interactive feedback loop to permit the users to label related samples which are visually similar or semantically close to the relevant sample. Based on the labeled samples, we then propose optimization algorithms to update visual queries and concept weights to refine the search results. We conduct experiments on two large-scale video datasets: TRECVID 2010 and YouTube. The experimental results demonstrate that our proposed system is effective for MRVS tasks.
Jin Yuan 0002, Yi-Liang Zhao, Huan-Bo Luan, Meng Wang 0001, Tat-Seng Chua
ACM Trans. Multim. Comput. Commun. Appl.3
2013 Learning attribute-aware dictionary for image classification and search
abstract
Bag-of-visual words (BoW) model has recently been well advocated for image classification and search. However, one critical limitation of existing BoW model is the lack of semantic information. To alleviate the impact of this issue, it is imperative to construct semantic-aware visual dictionary. In this paper, we propose a novel approach for learning visual word dictionary embedding intermediate-level semantics. Specifically, we first introduce an Attribute aware Dictionary Learning(AttrDL) scheme to learn multiple sub-dictionaries with specific semantic meanings. We divide training images into different sets and each represents a specific attribute. For each image set, an attribute-aware sub-vocabulary is learned. Hence, these resulting sub-vocabularies are more discriminative for semantics than the traditional vocabularies. Second, to get semantic-aware and discriminative BoW representation with the learned sub-vocabularies, we adopt the idea of L21-norm regularized sparse coding and recode the resulting sparse representation of each image. Experimental results show that the proposed scheme outperforms the state-of-the-art algorithms in both image classification and search tasks.
Zhengjun Zha, Huan-Bo Luan, Shiliang Zhang, Qi Tian 0001
ICMR3
2013 Stereotime: a wireless 2D and 3D switchable video communication system
abstract
Mobile 3D video communication, especially with 2D and 3D compatible, is a new paradigm for both video communication and 3D video processing. Current techniques face challenges in mobile devices when bundled constraints such as computation resource and compatibility should be considered. In this work, we present a wireless 2D and 3D switchable video communication to handle the previous challenges, and name it as Stereotime. The methods of Zig-Zag fast object segmentation, depth cues detection and merging, and texture-adaptive view generation are used for 3D scene reconstruction. We show the functionalities and compatibilities on 3D mobile devices in WiFi network environment.
You Yang 0002, Qiong Liu 0001, Yue Gao 0002, Binbin Xiong, Li Yu 0003, Huan-Bo Luan, Rongrong Ji, Qi Tian 0001
ACM Multimedia6
2013 Social Visual Image Ranking for Web Image Search
Peng Cui 0001, Huan-Bo Luan, Wenwu Zhu 0001, Shiqiang Yang, Qi Tian 0001
MMM (1)3
2013 NExT-Live: A Live Observatory on Social Media
Huan-Bo Luan, Dejun Hou, Tat-Seng Chua
MMM (2)1
2013 Detecting Group Activities With Multi-Camera Context
abstract
Human group activities detection in multi-camera CCTV surveillance videos is a pressing demand on smart surveillance. Previous works on this topic are mainly based on camera topology inference that is hard to apply to real-world unconstrained surveillance videos. In this paper, we propose a new approach for multi-camera group activities detection. Our approach simultaneously exploits intra-camera and inter-camera contexts without topology inference. Specifically, a discriminative graphical model with hidden variables is developed. The intra-camera and inter-camera contexts are characterized by the structure of hidden variables. By automatically optimizing the structure, the contexts are effectively explored. Furthermore, we propose a new spatiotemporal feature, named vigilant area (VA), to characterize the quantity and appearance of the motion in an area. This feature is effective for group activity representation and is easy to extract from a dynamic and crowded scene. We evaluate the proposed VA feature and discriminative graphical model extensively on two real-world multi-camera surveillance video data sets, including a public corpus consisting of 2.5 h of videos and a 468-h video collection, which, to the best of our knowledge, is the largest video collection ever used in human activity detection. The experimental results demonstrate the effectiveness of our approach.
Zhengjun Zha, Hanwang Zhang, Meng Wang 0001, Huan-Bo Luan, Tat-Seng Chua
IEEE Trans. Circuits Syst. Video Technol.4
2012 View-based 3D object retrieval by bipartite graph matching
abstract
Bipartite graph matching has been investigated in multiple view matching for 3D object retrieval. However, existing methods employ one-to-one vertex matching scheme while more than two views may share close semantic meanings in practice. In this work, we propose a bipartite graph matching method to measure the distance between two objects based on multiple views. In the proposed method, representative views are first selected by using view clustering for each object, and the corresponding weights are given based on the cluster results. A bipartite graph is constructed by using the two groups of representative views from two compared objects. To calculate the similarity between two objects, the bipartite graph is first partitioned to several subsets, and the views in the same sub-set are with high possibility to be with similar semantic meanings. The distances between two objects within individual subsets are then assembled through the graph to obtain the final similarity. Experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm.
Yue Wen, Yue Gao 0002, Richang Hong, Huan-Bo Luan, Qiong Liu 0001, Jialie Shen 0001, Rongrong Ji
ACM Multimedia4
2012 Attribute feedback
abstract
This demonstration presents a new interactive Content Based Image Retrieval (CBIR) system, termed Attribute Feedback (AF). Unlike traditional relevance feedback purely founded on low-level features, AF system shapes user's search intents more precisely and quickly by collecting feedbacks on intermediate-level semantic attribute. At each interaction iteration, the AF system first determines the most informative binary attributes for feedbacks and then augments the binary attribute feedbacks by a new type of attributes, "affinity attributes", each of which is learnt offline to describe the distance/similarity between user's envisioned image(s) and a retrieved image with respect to the corresponding affinity attribute. Based on the feedbacks on binary and affinity attributes, the images in corpus are further re-ranked towards better fitting user's search intents. The experimental results on two real-world image datasets have demonstrated the superiority of the AF system over other state-of-the-art relevance feedback based CBIR approaches.
Hanwang Zhang, Zhengjun Zha, Jingwen Bian, Yue Gao 0002, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia5
2012 Video Browser Showdown by NUS
Jin Yuan 0002, Huan-Bo Luan, Dejun Hou, Han Zhang 0010, Yantao Zheng, Zhengjun Zha, Tat-Seng Chua
MMM2
2011 Tag-based social image search with visual-text joint hypergraph learning
abstract
Tag-based social image search has attracted great interest and how to order the search results based on relevance level is a research problem. Visual content of images and tags have both been investigated. However, existing methods usually employ tags and visual content separately or sequentially to learn the image relevance. This paper proposes a tag-based image search with visual-text joint hypergraph learning. We simultaneously investigate the bag-of-words and bag-of-visual-words representations of images and accomplish the relevance estimation with a hypergraph learning approach. Each textual or visual word generates a hyperedge in the constructed hypergraph. We conduct experiments with a real-world data set and experimental results demonstrate the effectiveness of our approach.
Yue Gao 0002, Meng Wang 0001, Huan-Bo Luan, Jialie Shen 0001, Shuicheng Yan, Dacheng Tao
ACM Multimedia3
2011 VisionGo: Towards video retrieval with joint exploration of human and computer
Huan-Bo Luan, Yantao Zheng, Meng Wang 0001, Tat-Seng Chua
Inf. Sci.1
2007 Interactive Spatio-Temporal Visual Map Model for Web Video Retrieval
abstract
The massive amount of multimedia information especially video available on the Web requires a more precise and interactive retrieval. Current operational video retrieval systems do not make use of the implicit visual features but rely only on textual metadata supplied by the user during uploading. This greatly affects the retrieval performance as the metadata may not be comprehensive or consistent. In this paper, we describe the use of a spatio-temporal visual map (STVM) model to supplement Web video retrieval. This is done by employing the spatio-temporal visual similarity to rerank the text-retrieval results and find new results. Experimental results on a dynamic Web video corpus show significant improvement based on STVM model, with good usability scores based on human users.
Huan-Bo Luan, Shouxun Lin, Sheng Tang, Shi-Yong Neo, Tat-Seng Chua
ICME1
2007 Segregated feedback with performance-based adaptive sampling for interactive news video retrieval
abstract
Existing video research incorporates the use of relevance feedback based on user-dependent interpretations to improve the retrieval results. In this paper, we segregate the process of relevance feedback into 2 distinct facets: (a) recall-directed feedback; and (b) precision-directed feedback. The recall-directed facet employs general features such as text and high level features (HLFs) to maximize efficiency and recall during feedback, making it very suitable for large corpuses. The precision-directed facet on the other hand uses many other multimodal features in an active learning environment for improved accuracy. Combined with a performance-based adaptive sampling strategy, this process continuously re-ranks a subset of instances as the user annotates. Experiments done using TRECVID 2006 dataset show that our approach is efficient and effective.
Huan-Bo Luan, Shi-Yong Neo, Hai-Kiat Goh, Yongdong Zhang 0001, Shouxun Lin, Tat-Seng Chua
ACM Multimedia1