EDBT 2026 Demo / reviewers in the wild / expert
Yong Rui
dblp:r/YongRui
· DBLP profile ↗
211ranked-venue papers
24as first author
33since 2021 · last 2026
0000-0002-9142-5914ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 156 · 22 first-author · 21 since 2021Artificial intelligence and machine learning · 64 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 23 · 1 since 2021Computer networks · 17 · 4 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Dual Representations for Semi-Supervised Partial Label LearningabstractSemi-supervised partial label learning (SSPLL) aims to improve the generalization performance of partial label (PL) classifiers by effectively leveraging unlabeled data. Nevertheless, the inherent ambiguity in supervision, where the ground-truth label of a PL example is hidden within a set of candidate labels, poses significant challenges. The presence of false positive labels potentially misleads model's judgment, resulting in pronounced confirmation bias. To address these issues, we propose a novel approach named CODUAL, which jointly learns a pair of dual representations for each instance: the predictive class distribution and the low-dimensional embedding. The dual representations interact and progress collaboratively during training. On one hand, in the embedding space the class prototypes are derived via solving a tailored empirical distance minimization problem and employed to smooth the pseudo-targets of unlabeled instances. On the other hand, the refined class distributions regularize the embedding space via encouraging instances with similar pseudo-targets to exhibit similar embeddings. Through an in-depth analysis, we provide-to the best of our knowledge-the first theoretical explanation of how collaborative dual representations facilitate more effective use of unlabeled data for disambiguation. Extensive experiments over benchmark datasets validate the superiority of our proposed approach. Wei-Xuan Bao, Yong Rui, Min-Ling Zhang |
AAAI | 2 |
| 2026 | Classifier-induced Reciprocal Points for Multi-label Open-set RecognitionabstractMulti-label learning is a practical machine learning paradigm dealing with instances associated with multiple labels simultaneously. Most existing multi-label learning studies are designed under the closed-world assumption, i.e. a fixed size of label space. However, it encounters significant difficulties in open-set scenarios, where test data may contain unknown labels absent from the training set to be recognized. Existing method typically tackles this challenging problem through sub-labeling approximations and prototype-based comparisons, which often overlooks the implicit information carried by unknown labels. To address this, we propose a novel framework CREM, i.e. Classifier-induced REciprocal point for Multi-label open-set recognition, which rethinks the above problem from the reciprocal point perspective. Specifically, reciprocal points are formulated by explicitly constraining the opposition feature space to a learnable bounded margin. Then reciprocal points can be induced through the classifier with the instance-wise bias eliminated. Subsequently, a unified optimization framework is introduced to jointly facilitate the classifier and reciprocal points induction. Extensive experiments demonstrate the effectiveness and superiority of the proposed CREM approach in the multi-label open-set recognition paradigm. Yong Rui, Min-Ling Zhang |
AAAI | 2 |
| 2026 | DivControl: Knowledge Diversion for Controllable Image GenerationabstractDiffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with entangled representations, resulting in poor generalization and high adaptation costs for novel conditions. To this end, we propose DivControl, a decomposable pretraining framework for unified controllable generation and efficient adaptation. DivControl factorizes ControlNet via SVD into basic components—pairs of singular vectors—which are disentangled into condition-agnostic learngenes and condition-specific tailors through knowledge diversion during multi-condition training. Knowledge diversion is implemented via a dynamic gate that performs soft routing over tailors based on the semantics of condition instructions, enabling zero-shot generalization and parameter-efficient adaptation to novel conditions. To further improve condition fidelity and training efficiency, we introduce a representation alignment loss that aligns condition embeddings with early diffusion features. Extensive experiments demonstrate that DivControl achieves state-of-the-art controllability with 36.4× less training cost, while simultaneously improving average performance on basic conditions. It also delivers strong zero-shot and few-shot performance on unseen conditions, demonstrating superior scalability, modularity, and transferability. Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001 |
AAAI | 5 |
| 2026 | Continual novel class discovery under domain shift with entropy-based selection and representation evolution
Feifei Shi, Xiangyang Li 0002, Shuqiang Jiang, Yong Rui |
Multim. Syst. | 4 |
| 2026 | Semi-Supervised Unconstrained Head Pose Estimation in the WildabstractExisting research on unconstrained in-the-wild head pose estimation suffers from the flaws of its datasets, which consist of either numerous samples by non-realistic synthesis or constrained collection, or small-scale natural images yet with plausible manual annotations. This makes fully-supervised solutions compromised due to the reliance on generous labels. To alleviate it, we propose the first semi-supervised unconstrained head pose estimation method SemiUHPE, which can leverage abundant easily available unlabeled head images. Technically, we choose semi-supervised rotation regression and adapt it to the error-sensitive and label-scarce problem of unconstrained head pose. Our method is based on the observation that the aspect-ratio invariant cropping of wild heads is superior to previous landmark-based affine alignment given that landmarks of unconstrained human heads are usually unavailable, especially for underexplored non-frontal heads. Instead of using a pre-fixed threshold to filter out pseudo labeled heads, we propose dynamic entropy based filtering to adaptively remove unlabeled outliers as training progresses by updating the threshold in multiple stages. We then revisit the design of weak-strong augmentations and improve it by devising two novel head-oriented strong augmentations, termed pose-irrelevant cut-occlusion and pose-altering rotation consistency respectively. Extensive experiments and ablation studies show that SemiUHPE outperforms its counterparts greatly on public benchmarks under both the front-range and full-range settings. Furthermore, our proposed method is also beneficial for solving other closely related problems, including generic object rotation regression and 3D head reconstruction, demonstrating good versatility and extensibility. Huayi Zhou 0001, Fei Jiang 0006, Yong Rui, Hongtao Lu 0001, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | A Bootstrap Pipeline for Chat-Based Image Retrieval with Effective Question GenerationabstractChat-based image retrieval uses Large Language Models (LLM) to guide user input to enable more specific and precise search results, where LLM can enhance this process by asking user retrieval-oriented questions eliciting additional details about the target image. Despite the potential of this approach, no specialized Questioner model has been developed for this task due to the following significant challenges: (a) the difficulty of determining the optimal questions to ask; (b) the lack of a suitable protocol for fair model comparison; and (c) the notable scarcity of dialog-to-image retrieval data. To address these challenges, two fundamental principles are developed in this article to ensure the simplicity and effectiveness of the generated questions while enabling a fair comparison and accurate estimation of data quality and model performance. A bootstrap training methodology is introduced to collect retrieval-oriented dialog data and concurrently train the Questioner and the image Retriever. Under a fair comparison protocol, our extensive experiments have demonstrated that our proposed method can not only address the critical data gap but also achieve state-of-the-art results, which substantially surpass GPT-4o and GPT-4-Turbo through the fine-tuning of an 8B model. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2025 | Implicit Relative Labeling-Importance Aware Multi-Label Metric LearningabstractMulti-label metric learning, as an extension of metric learning to multi-label scenarios, aims to learn better similarity metrics for objects with rich semantics. Existing multi-label metric learning approaches employ the common assumption of equal labeling-importance, i.e., all associated labels are considered relevant to the training instance, while there is no differentiation in the relative importance of their semantics. However, this common assumption does not reflect the fact that the importance of each relevant label is generally different, even though such importance information is not directly accessible from the training examples. In this paper, we claim that it is beneficial to leverage the implicit Relative LabelingImportance (RLI) information to facilitate multi-label metric learning. Specifically, the manifold structure within the feature space is exploited by local linear reconstruction, and then the RLIs are recovered by transferring such structure to the label space. Subsequently, a discrimiative multi-label metric learning framework is introduced to align the predictive modeling outputs with the recovered RLIs, under which instances with similar RLI are implicitly pulled closer to each other, while those with dissimilar RLI are pushed further apart. Comprehensive experiments on benchmark multi-label datasets validate the superiority of our proposed approach in learning effective similarity metrics between multi-label examples. Junxiang Mao, Yong Rui, Min-Ling Zhang |
AAAI | 2 |
| 2025 | KIND: Knowledge Integration and Diversion for Training Decomposable ModelsabstractPre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pre-trained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target tasks.
To address this, we propose **KIND**, a novel pre-training method designed to construct decomposable models.
KIND integrates knowledge by incorporating Singular Value Decomposition (SVD) as a structural constraint, with each basic component represented as a combination of a column vector, singular value, and row vector from $U$, $\Sigma$, and $V^\top$ matrices.
These components are categorized into **learngenes** for encapsulating class-agnostic knowledge and \textbf{tailors} for capturing class-specific knowledge, with knowledge diversion facilitated by a class gate mechanism during training.
Extensive experiments demonstrate that models pre-trained with KIND can be decomposed into learngenes and tailors, which can be adaptively recombined for diverse resource-constrained deployments.
Moreover, for tasks with large domain shifts, transferring only learngenes with task-agnostic knowledge, when combined with randomly initialized tailors, effectively mitigates domain shifts.
Code will be made available at https://github.com/Te4P0t/KIND. Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001 |
ICML | 5 |
| 2025 | Building Embodied EvoAgent: A Brain-inspired Paradigm for Bridging Multimodal Large Models and World ModelsabstractEmbodied artificial intelligence has rapidly developed under the impetus of multimodal learning, robotics, and cognitive science, demonstrating great potential in fields such as navigation and manipulation. However, building embodied agents that can robustly operate in diverse and dynamic environments still faces challenges, such as handling partial observability and environmental adaptability. Multimodal large language models (MLLMs) are vital for embodied intelligence due to their ability to process multimodal information, but they encounter difficulties in understanding spatial environments and performing dynamic decisions and evolution. Inspired by the functional specialization of the left and right hemispheres of the human brain, this paper proposes a brain-inspired learning and evolution paradigm for embodied agents. The method designs an embodied context-augmented MLLM to simulate the language processing and logical analysis capabilities of the left hemisphere, responsible for understanding instructions and visual scenes. At the same time, it constructs a perceptual context-guided world model based on the recurrent state space model to simulate the spatial perception and holistic thinking functions of the right hemisphere, capturing environmental dynamics and predicting future states. By simulating the communication function of the corpus callosum, we propose dynamic communication slots for efficient information exchange between MLLMs and the world model, which also allows the agent to quickly adapt to dynamic environments without requiring extensive computational resources. Experiments show that the proposed paradigm significantly improves the performance of embodied agents in a series of tasks and enhances their generalization ability in zero-shot tasks through embodied exploration experience and online evolution. Our project page is available at https://feliciaxyao.github.io/EvoAgent/. Junyu Gao 0002, Xuan Yao 0001, Yong Rui, Changsheng Xu |
ACM Multimedia | 3 |
| 2025 | Learning Preference Distributions: A Label-Side Paradigm for Explainable Reward ModelsabstractThe reward model is a critical component in training powerful large language models. However, current methods largely overlook the inherent subjectivity and variability in human preferences. Typically, this issue is indirectly addressed through model-side approaches, such as ensemble methods to estimate uncertainty from multiple predictions or uncertainty-aware regression to predict mean and variance. These indirect approaches fail to capture the intrinsic distributional characteristics and inter-rater disagreements present in human judgments. In contrast, we propose a direct, label-side solution by explicitly modeling human preference distributions. We recover missing information from scalar ratings to construct meaningful distribution labels. By employing Label Distribution Learning (LDL), each dimension of the resulting multidimensional discrete distribution explicitly corresponds to a specific preference score, naturally reflecting the subjective and multidimensional nature of human evaluations. Our approach improves explainability and confidence estimation, while also enabling more effective data selection and sample-efficient test-time adaptation. Empirical results demonstrate that our method not only achieves state-of-the-art performance but also provides a robust framework for uncertainty quantification and nuanced preference modeling. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
SMC | 7 |
| 2025 | Collective domain adversarial learning for unsupervised domain adaptation
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
Frontiers Comput. Sci. | 7 |
| 2025 | Epistemic graph: A plug-and-play module for hybrid representation learningabstractIn recent years, deep models have achieved remarkable success in various vision tasks. However, their performance heavily relies on large training datasets. In contrast, humans exhibit hybrid learning, seamlessly integrating structured knowledge for cross-domain recognition or relying on a smaller amount of data samples for few-shot learning. Motivated by this human-like epistemic process, we aim to extend hybrid learning to computer vision tasks by integrating structured knowledge with data samples for more effective representation learning. Nevertheless, this extension faces significant challenges due to the substantial gap between structured knowledge and deep features learned from data samples, encompassing both dimensions and knowledge granularity. In this paper, a novel Epistemic Graph Layer (EGLayer) is introduced to enable hybrid learning, enhancing the exchange of information between deep features and a structured knowledge graph. Our EGLayer is composed of three major parts, including a local graph module, a query aggregation model, and a novel correlation alignment loss function to emulate human epistemic ability. Serving as a plug-and-play module that can replace the standard linear classifier, EGLayer significantly improves the performance of deep models. Extensive experiments demonstrate that EGLayer can greatly enhance representation learning for the tasks of cross-domain recognition and few-shot learning, and the visualization of knowledge graphs can aid in model interpretation. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Yong Rui |
Neurocomputing | 7 |
| 2025 | DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain GeneralizationabstractTraditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 11 |
| 2025 | Multimodal Food LearningabstractFood-centered study has received more attention in the multimedia community for its profound impact on our survival, nutrition and health, pleasure, and enjoyment. Our experience of food is typically multi-sensory: We see food objects, smell its odors, taste its flavors, feel its texture, and hear sounds when chewing. Therefore, multimodal food learning is vital in food-centered study, which aims to relate information from multiple food modalities to support various multimedia tasks, ranging from recognition, retrieval, generation, recommendation, and interaction, enabling applications in different fields like healthcare and agriculture. However, there is no surveys on this topic to our knowledge. To fill this gap, this article formalizes multimodal food learning and comprehensively surveys its typical tasks, technical achievements, existing datasets, and applications to provide the blueprint with researchers and practitioners. Based on the current state of the art, we identify both open research issues and promising research directions, such as multimodal food learning benchmark construction, multimodal food foundation model construction, and multimodality diet estimation. We also point out that closer cooperation from researchers between multimedia and food science can handle some existing challenges and meanwhile open up more new opportunities to advance the fast development of multimodal food learning. This is the first comprehensive survey in this topic and we anticipate about 170 reviewed research articles can benefit academia and industry in this community and beyond. Weiqing Min, Xingjian Hong, Yuxin Liu 0009, Leyi Xu, Yilin Wang 0012, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 10 |
| 2024 | Disentangled Partial Label LearningabstractPartial label learning (PLL) induces a multi-class classifier from training examples each associated with a set of candidate labels, among which only one is valid. The formation of real-world data typically arises from heterogeneous entanglement of series latent explanatory factors, which are considered intrinsic properties for discriminating between different patterns. Though learning disentangled representation is expected to facilitate label disambiguation for partial-label (PL) examples, few existing works were dedicated to addressing this issue. In this paper, we make the first attempt towards disentangled PLL and propose a novel approach named TERIAL, which makes predictions according to derived disentangled representation of instances and label embeddings. The TERIAL approach formulates the PL examples as an undirected bipartite graph where instances are only connected with their candidate labels, and employs a tailored neighborhood routing mechanism to yield disentangled representation of nodes in the graph. Specifically, the proposed routing mechanism progressively infers the explanatory factors that contribute to the edge between adjacent nodes and augments the representation of the central node with factor-aware embedding information propagated from specific neighbors simultaneously via iteratively analyzing the promising subspace clusters formed by the node and its neighbors. The estimated labeling confidence matrix is also introduced to accommodate unreliable links owing to the inherent ambiguity of PLL. Moreover, we theoretically prove that the neighborhood routing mechanism will converge to the point estimate that maximizes the marginal likelihood of observed PL training examples. Comprehensive experiments over various datasets demonstrate that our approach outperforms the state-of-the-art counterparts. Wei-Xuan Bao, Yong Rui, Min-Ling Zhang |
AAAI | 2 |
| 2024 | See Widely, Think Wisely: Toward Designing a Generative Multi-agent System to Burst Filter BubblesabstractThe proliferation of AI-powered search and recommendation systems has accelerated the formation of “filter bubbles” that reinforce people’s biases and narrow their perspectives. Previous research has attempted to address this issue by increasing the diversity of information exposure, which is often hindered by a lack of user motivation to engage with. In this study, we took a human-centered approach to explore how Large Language Models (LLMs) could assist users in embracing more diverse perspectives. We developed a prototype featuring LLM-powered multi-agent characters that users could interact with while reading social media content. We conducted a participatory design study with 18 participants and found that multi-agent dialogues with gamification incentives could motivate users to engage with opposing viewpoints. Additionally, progressive interactions with assessment tasks could promote thoughtful consideration. Based on these findings, we provided design implications with future work outlooks for leveraging LLMs to help users burst their filter bubbles. Yu Zhang 0124, Jingwei Sun 0005, Cen Yao, Mingming Fan 0001, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui |
CHI | 9 |
| 2024 | Structure Matters: Tackling the Semantic Discrepancy in Diffusion Models for Image InpaintingabstractDenoising diffusion probabilistic models (DDPMs) for image inpainting aim to add the noise to the texture of the image during the forward process and recover the masked regions with the unmasked ones of the texture via the reverse denoising process. Despite the meaningful semantics gen-eration, the existing arts suffer from the semantic discrep-ancy between the masked and unmasked regions, since the semantically dense unmasked texture fails to be completely degraded while the masked regions turn to the pure noise in diffusion process, leading to the large discrepancy between them. In this paper, we aim to answer how the unmasked se-mantics guide the texture denoising process; together with how to tackle the semantic discrepancy, to facilitate the con-sistent and meaningful semantics generation. To this end, we propose a novel structure-guided diffusion model for image inpainting named StrDiffusion, to reformulate the conventional texture denoising process under the structure guidance to derive a simplified denoising objective for im-age inpainting, while revealing: 1) the semantically sparse structure is beneficial to tackle the semantic discrepancy in the early stage, while the dense texture generates the rea-sonable semantics in the late stage; 2) the semantics from the unmasked regions essentially offer the time-dependent structure guidance for the texture denoising process, ben-efiting from the time-dependent sparsity of the structure semantics. For the denoising process, a structure-guided neural network is trained to estimate the simplified denoising objective by exploiting the consistency of the denoised structure between masked and unmasked regions. Besides, we devise an adaptive resampling strategy as aformal criterion as whether the structure is competent to guide the texture denoising process, while regulate their semantic corre-lations. Extensive experiments validate the merits of StrDif-fusion over the state-of-the-arts. Our code is available at https://github.com/htyjers/StrDiffusion. Haipeng Liu 0004, Yang Wang 0023, Biao Qian, Meng Wang 0001, Yong Rui |
CVPR | 5 |
| 2024 | Towards Workplace Metaverse: A Human-Centered Approach for Designing and Evaluating XR Virtual DisplaysabstractWork is becoming more and more flexible nowadays. It can take place in the office, at home, or even on the go; tasks may encompass activities in the form of documents, multimedia, or 3D models in virtual space. Under such circumstances, personal computers (PC), the most widely used productivity devices for work today, cannot well address the diverse needs such as screen size, privacy, and flexibility to display diverse content formats (e.g., 2D to 3D), due to their fixed hardware specs. We believe the solution lies in Extended Reality (XR). In this article, we explored how XR glasses can be used for PC’s virtual extended displays, and conducted user interviews and usability tests to propose a systematic user experience design and evaluation framework. We discovered that the design space encompasses four dimensions general placement, display specs, operating system integration, and interaction behaviors) and summarized users’ corresponding preferences. We proposed a quality-of-experience (QoE) evaluation framework for XR virtual displays consisting of visual quality, visual fatigue and discomfort, as well as immersiveness, and identified clarity as the most significant factor that affects user satisfaction. Our design and evaluation frameworks could serve as a resource for both practitioners and scholars with an interest in the design and evaluation of virtual displays. Yu Zhang 0124, Jingwei Sun 0005, Qicheng Ding, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui |
Int. J. Hum. Comput. Interact. | 7 |
| 2024 | Unpacking the Gap Box Against Data-Free Knowledge DistillationabstractData-free knowledge distillation (DFKD) improves the student model (S) by mimicking the class probability from a pre-trained teacher model (T) without training data. Under such setting, an ideal scenario is that T can help generate ”good” samples from a generator (G) to maximally benefit S. However, existing arts suffer from the non-ideal generated samples under the disturbance of the gap (i.e., either too large or small) between the class probabilities of T and S; for example, the generated samples with too large gap may exhibitexcessiveinformation for S, while too small gap leads to thelimitedknowledge in the samples, resulting into the poor generalization. Meanwhile, they fail to judge the ”goodness” of the generated samples for S since thefixedT is not necessarily ideal. In this paper, we aim to answerwhat is inside the gap box; together withhow to yield ”good” generated samples for DFKD?To this end, we propose aGap-SensitiveSampleGeneration (GapSSG) approach, by revisiting the empirical distilled risk from a data-free perspective, which confirms the existence of an ideal teacher (T$^*$), while theoretically implying: (1) the gap disturbance originates from themismatchbetween T and T$^*$, hence the class probabilities of T enable the approximation to those of T$^*$; and (2) ”good” samples should maximally benefit S via T's class probabilities, owing to unknown T$^*$. To this end, we unpack the gap box between T and S as two findings:inherentgap to perceive T and T$^*$;derivedgap to monitor S and T$^*$. Benefiting from thederivedgap that focuses on the adaptability of generated sample to S, we attempt to track student's training route (a series of training epochs) to capture the category distribution of S; upon which, a regulatory factor is further devised to approximate T$^*$overinherentgap, so as to generate ”good” samples to S. Furthermore, during the distillation process, a sample-balanced strategy comes up to tackle the overfitting and missing knowledge issues between the generated partial and critical samples by training G. The theoretical and empirical studies verify the advantages of GapSSG over the state-of-the-arts.Our code is available athttps://github.com/hfutqian/GapSSG. Yang Wang 0023, Biao Qian, Haipeng Liu 0004, Yong Rui, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Convolution-Enhanced Bi-Branch Adaptive Transformer With Cross-Task Interaction for Food Category and Ingredient RecognitionabstractRecently, visual food analysis has received more and more attention in the computer vision community due to its wide application scenarios, e.g., diet nutrition management, smart restaurant, and personalized diet recommendation. Considering that food images are unstructured images with complex and unfixed visual patterns, mining food-related semantic-aware regions is crucial. Furthermore, the ingredients contained in food images are semantically related to each other due to the cooking habits and have significant semantic relationships with food categories under the hierarchical food classification ontology. Therefore, modeling the long-range semantic relationships between ingredients and the categories-ingredients semantic interactions is beneficial for ingredient recognition and food analysis. Taking these factors into consideration, we propose a multi-task learning framework for food category and ingredient recognition. This framework mainly consists of a food-orient Transformer named Convolution-Enhanced Bi-Branch Adaptive Transformer (CBiAFormer) and a multi-task category-ingredient recognition network called Structural Learning and Cross-Task Interaction (SLCI). In order to capture the complex and unfixed fine-grained patterns of food images, we propose a query-aware data-adaptive attention mechanism called Bi-Branch Adaptive Attention (BiA-Attention) in CBiAFormer, which consists of a local fine-grained branch and a global coarse-grained branch to mine local and global semantic-aware regions for different input images through an adaptive candidate key/value sets assignment for each query. Additionally, a convolutional patch embedding module is proposed to extract the fine-grained features which are neglected by Transformers. To fully utilize the ingredient information, we propose SLCI, which consists of cross-layer attention to model the semantic relationships between ingredients and two cross-task interaction modules to mine the semantic interactions between categories and ingredients. Extensive experiments show that our method achieves competitive performance on three mainstream food datasets (ETH Food-101, Vireo Food-172, and ISIA Food-200). Visualization analyses of CBiAFormer and SLCI on two tasks prove the effectiveness of our method. Codes will be released upon publication. Code and models are available at https://github.com/Liuyuxinict/CBiAFormer. Yuxin Liu 0009, Weiqing Min, Shuqiang Jiang, Yong Rui |
IEEE Trans. Image Process. | 4 |
| 2024 | Domain-Aware Graph Network for Bridging Multi-Source Domain AdaptationabstractDomain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG. Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 9 |
| 2024 | Toward Egocentric Compositional Action Anticipation with Adaptive Semantic DebiasingabstractPredicting the unknown from the first-person perspective is expected as a necessary step toward machine intelligence, which is essential for practical applications including autonomous driving and robotics. As a human-level task, egocentric action anticipation aims at predicting an unknown action seconds before it is performed from the first-person viewpoint. Egocentric actions are usually provided as verb-noun pairs; however, predicting the unknown action may be trapped in insufficient training data for all possible combinations. Therefore, it is crucial for intelligent systems to use limited known verb-noun pairs to predict new combinations of actions that have never appeared, which is known as compositional generalization. In this article, we are the first to explore the egocentric compositional action anticipation problem, which is more in line with real-world settings but neglected by existing studies. Whereas prediction results are prone to suffer from semantic bias considering the distinct difference between training and test distributions, we further introduce a general and flexible adaptive semantic debiasing framework that is compatible with different deep neural networks. To capture and mitigate semantic bias, we can imagine one counterfactual situation where no visual representations have been observed and only semantic patterns of observation are used to predict the next action. Instead of the traditional counterfactual analysis scheme that reduces semantic bias in a mindless way, we devise a novel counterfactual analysis scheme to adaptively amplify or penalize the effect of semantic experience by considering the discrepancy both among categories and among examples. We also demonstrate that the traditional counterfactual analysis scheme is a special case of the devised adaptive counterfactual analysis scheme. We conduct experiments on three large-scale egocentric video datasets. Experimental results verify the superiority and effectiveness of our proposed solution. Weiqing Min, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Learning From Biased Soft LabelsabstractSince the advent of knowledge distillation, many researchers have been intrigued by the $\textit{dark knowledge}$ hidden in the soft labels generated by the teacher model. This prompts us to scrutinize the circumstances under which these soft labels are effective. Predominant existing theories implicitly require that the soft labels are close to the ground-truth labels. In this paper, however, we investigate whether biased soft labels are still effective. Here, bias refers to the discrepancy between the soft labels and the ground-truth labels. We present two indicators to measure the effectiveness of the soft labels. Based on the two indicators, we propose moderate conditions to ensure that, the biased soft label learning problem is both $\textit{classifier-consistent}$ and $\textit{Empirical Risk Minimization}$ (ERM) $\textit{learnable}$, which can be applicable even for large-biased soft labels. We further design a heuristic method to train Skillful but Bad Teachers (SBTs), and these teachers with accuracy less than 30\% can teach students to achieve accuracy over 90\% on CIFAR-10, which is comparable to models trained on the original data. The proposed indicators adequately measure the effectiveness of the soft labels generated in this process. Moreover, our theoretical framework can be adapted to elucidate the effectiveness of soft labels in three weakly-supervised learning paradigms, namely incomplete supervision, partial label learning and learning with additive noise. Experimental results demonstrate that our indicators can measure the effectiveness of biased soft labels generated by teachers or in these weakly-supervised learning paradigms. Ning Xu 0009, Xu Yang 0021, Xin Geng 0001, Yong Rui |
NeurIPS | 6 |
| 2023 | Balanced masking strategy for multi-label image classification
Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
Neurocomputing | 6 |
| 2023 | Graph Attention Transformer Network for Multi-label Image ClassificationabstractMulti-label classification aims to recognize multiple objects or attributes from images. The key to solving this issue relies on effectively characterizing the inter-label correlations or dependencies, which bring the prevailing graph neural network. However, current methods often use the co-occurrence probability of labels based on the training set as the adjacency matrix to model this correlation, which is greatly limited by the dataset and affects the model’s generalization ability. This article proposes a Graph Attention Transformer Network, a general framework for multi-label image classification by mining rich and effective label correlation. First, we use the cosine similarity value of the pre-trained label word embedding as the initial correlation matrix, which can represent richer semantic information than the co-occurrence one. Subsequently, we propose the graph attention transformer layer to transfer this adjacency matrix to adapt to the current domain. Our extensive experiments have demonstrated that our proposed methods can achieve highly competitive performance on three datasets. Shikai Chen, Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Semi-Supervised 3D Medical Image Segmentation Via Boundary-Aware Consistent Hidden Representation LearningabstractThis paper proposes a novel Boundary-aware Consistent Hidden Representation Learning Network (BA-CHRLN), which contains two branches for semi-supervised 3D medical image segmentation. Inspired by the contrastive learning, the two branches share the same encoder and each has its individual decoder, namely supervised decoder and unsupervised one. A stop-gradient operation is also utilized to prevent collapsing of solutions. Taking the unlabeled images as references, BA-CHRLN imposes the consistency by applying a perturbation on the high-level hidden feature representations, which significantly improves the encoder’s representation and the network’s robustness. A boundary-aware map is further introduced to capture the organ’s boundary without any prior knowledge and additional parameters. Experiments on the Left Atrium (LA) benchmark dataset demonstrate the effectiveness of the BA-CHRLN. Linhu Liu, Jiang Tian, Xiangqian Cheng, Zhongchao Shi, Jianping Fan 0007, Yong Rui |
ICIP | 6 |
| 2022 | Delving Globally into Texture and Structure for Image InpaintingabstractImage inpainting has achieved remarkable progress and inspired abundant methods, where the critical bottleneck is identified as how to fulfill the high-frequency structure and low-frequency texture information on the masked regions with semantics. To this end, deep models exhibit powerful superiority to capture them, yet constrained on the local spatial regions. In this paper, we delve globally into texture and structure information to well capture the semantics for image inpainting. As opposed to the existing arts trapped on the independent local patches, the texture information of each patch is reconstructed from all other patches across the whole image, to match the coarsely filled information, especially the structure information over the masked regions. Unlike the current decoder-only transformer within the pixel level for image inpainting, our model adopts the transformer pipeline paired with both encoder and decoder. On one hand, the encoder captures the texture semantic correlations of all patches across image via self-attention module. On the other hand, an adaptive patch vocabulary is dynamically established in the decoder for the filled patches over the masked regions. Building on this, a structure-texture matching attention module anchored on the known regions comes up to marry the best of these two worlds for progressive inpainting via a probabilistic diffusion process. Our model is orthogonal to the fashionable arts, such as Convolutional Neural Networks (CNNs), Attention and Transformer model, from the perspective of texture and structure information for image inpainting. The extensive experiments over the benchmarks validate its superiority. Our code is available here Haipeng Liu 0004, Yang Wang 0023, Meng Wang 0001, Yong Rui |
ACM Multimedia | 4 |
| 2022 | Self-Supervised Graph Neural Network for Multi-Source Domain AdaptationabstractDomain adaptation (DA) tries to tackle the scenarios when the test data does not fully follow the same distribution of the training data, and multi-source domain adaptation (MSDA) is very attractive for real world applications. By learning from large-scale unlabeled samples, self-supervised learning has now become a new trend in deep learning. It is worth noting that both self-supervised learning and multi-source domain adaptation share a similar goal: they both aim to leverage unlabeled data to learn more expressive representations. Unfortunately, traditional multi-task self-supervised learning faces two challenges: (1) the pretext task may not strongly relate to the downstream task, thus it could be difficult to learn useful knowledge being shared from the pretext task to the target task; (2) when the same feature extractor is shared between the pretext task and the downstream one and only different prediction heads are used, it is ineffective to enable inter-task information exchange and knowledge sharing. To address these issues, we propose a novel Self-Supervised Graph Neural Network (SSG), where a graph neural network is used as the bridge to enable more effective inter-task information exchange and knowledge sharing. More expressive representation is learned by adopting a mask token strategy to mask some domain information. Our extensive experiments have demonstrated that our proposed SSG method has achieved state-of-the-art results over four multi-source domain adaptation datasets, which have shown the effectiveness of our proposed SSG method from different aspects. Feng Hou, Yangzhou Du, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
ACM Multimedia | 7 |
| 2022 | Hierarchical Deep Click Feature Prediction for Fine-Grained Image RecognitionabstractThe click feature of an image, defined as the user click frequency vector of the image on a predefined word vocabulary, is known to effectively reduce the semantic gap for fine-grained image recognition. Unfortunately, user click frequency data are usually absent in practice. It remains challenging to predict the click feature from the visual feature, because the user click frequency vector of an image is always noisy and sparse. In this paper, we devise a Hierarchical Deep Word Embedding (HDWE) model by integrating sparse constraints and an improved RELU operator to address click feature prediction from visual features. HDWE is a coarse-to-fine click feature predictor that is learned with the help of an auxiliary image dataset containing click information. It can therefore discover the hierarchy of word semantics. We evaluate HDWE on three dog and one bird image datasets, in which Clickture-Dog and Clickture-Bird are utilized as auxiliary datasets to provide click data, respectively. Our empirical studies show that HDWE has 1) higher recognition accuracy, 2) a larger compression ratio, and 3) good one-shot learning ability and scalability to unseen categories. Jun Yu 0002, Min Tan 0005, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Hierarchical User Intent Graph Network for Multimedia RecommendationabstractUnderstanding user preference on item context is the key to acquire a high-quality multimedia recommendation. Typically, the pre-existing features of items are derived from pre-trained models (e.g. visual features of micro-videos extracted from some neural networks), and then introduced into the recommendation framework (e.g. collaborative filtering) to capture user preference. However, we argue that such a paradigm is insufficient to output satisfactory user representations, which hardly profile personal interests well. The key reason is that present works largely leave user intents untouched, then failing to encode such informative representation of users. In this work, we aim to learn multi-level user intents from the co-interacted patterns of items, so as to obtain high-quality representations of users and items and further enhance the recommendation performance. Towards this end, we develop a novel framework,Hierarchical User Intent Graph Network, which exhibits user intents in a hierarchical graph structure, from the fine-grained to coarse-grained intents. In particular, we get the multi-level user intents by recursively performing two operations: 1) intra-level aggregation, which distills the signal pertinent to user intents from co-interacted item graphs; and 2) inter-level aggregation, which constitutes the supernode in higher levels to model coarser-grained user intents via gathering the nodes’ representations in the lower ones. Then, we refine the user and item representations as a distribution over the discovered intents, instead of simple pre-existing features. To demonstrate the effectiveness of our model, we conducted extensive experiments on three public datasets. Our model achieves significant improvements over the state-of-the-art methods, including MMGCN and DisenGCN. Furthermore, by visualizing the item representations, we provide the semantics of user intents. Yinwei Wei, Xiang Wang 0010, Xiangnan He 0001, Liqiang Nie, Yong Rui, Tat-Seng Chua |
IEEE Trans. Multim. | 5 |
| 2021 | What If We Could Not See? Counterfactual Analysis for Egocentric Action AnticipationabstractEgocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacity. In the factual case, we can predict the upcoming action based on visual features and semantic labels from past observation. Imagining one counterfactual situation where no visual representation had been observed, we would obtain a counterfactual predicted action only using past semantic labels. In this way, we can reduce the side-effect caused by semantic labels via a comparison between factual and counterfactual outcomes, which moves a step towards unbiased prediction for egocentric action anticipation. We conduct experiments on two large-scale egocentric video datasets. Qualitative and quantitative results validate the effectiveness of our proposed CA-EAA. Weiqing Min, Shuqiang Jiang, Yong Rui |
IJCAI | 6 |
| 2021 | MMPT'21: International Joint Workshop on Multi-Modal Pre-Training for Multimedia UnderstandingabstractPre-training has been an emerging topic that provides a way to learn strong representation in many fields (e.g., natural language processing, computing vision). In the last few years, we have witnessed many research works on multi-modal pre-training which have achieved state-of-the-art performances on many multimedia tasks (e.g., image-text retrieval, video localization, speech recognition). In this workshop, we aim to gather peer researchers on related topics for more insightful discussion. We also intend to attract more researchers to explore and investigate more opportunities of designing and using innovative pre-training models for multimedia tasks. Bei Liu 0001, Jianlong Fu, Shizhe Chen, Qin Jin, Alex Hauptmann 0001, Yong Rui |
ICMR | 6 |
| 2021 | HoloBoard: a Large-format Immersive Teaching Board based on pseudo HoloGraphicsabstractIn this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices. Jiangtao Gong, Teng Han, Siling Guo, Jiannan Li, Siyu Zha, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002, Yong Rui |
UIST | 9 |
| 2020 | Label Distribution Learning on Auxiliary Label Space Graphs for Facial Expression RecognitionabstractMany existing studies reveal that annotation inconsistency widely exists among a variety of facial expression recognition (FER) datasets. The reason might be the subjectivity of human annotators and the ambiguous nature of the expression labels. One promising strategy tackling such a problem is a recently proposed learning paradigm called Label Distribution Learning (LDL), which allows multiple labels with different intensity to be linked to one expression. However, it is often impractical to directly apply label distribution learning because numerous existing datasets only contain one-hot labels rather than label distributions. To solve the problem, we propose a novel approach named Label Distribution Learning on Auxiliary Label Space Graphs(LDL-ALSG) that leverages the topological information of the labels from related but more distinct tasks, such as action unit recognition and facial landmark detection. The underlying assumption is that facial images should have similar expression distributions to their neighbours in the label space of action unit recognition and facial landmark detection. Our proposed method is evaluated on a variety of datasets and outperforms those state-of-the-art methods consistently with a huge margin. Shikai Chen, Yuedong Chen, Zhongchao Shi, Xin Geng 0001, Yong Rui |
CVPR | 6 |
| 2020 | Selecting Useful Knowledge from Previous Tasks for Future Learning in a Single NetworkabstractContinual learning can learn new tasks incrementally while avoiding catastrophic forgetting. Recent work has shown that packing multiple tasks into a single network incrementally by iterative pruning and re-training network is a promising method. We build upon this idea and propose an improved version of PackNet. Specifically, we propose a novel gradient-based threshold method to reuse the knowledge of the previous tasks selectively when learning new tasks. Our experiments on a variety of classification tasks and different network architectures demonstrate that our method obtains competitive results when compared to PackNet. Feifei Shi, Peng Wang 0095, Zhongchao Shi, Yong Rui |
ICPR | 4 |
| 2020 | An Egocentric Action Anticipation Framework via Fusing Intuition and AnalysisabstractIn this paper, we focus on egocentric action anticipation from videos, which enables various applications, such as helping intelligent wearable assistants understand users' needs and enhance their capabilities in the interaction process. It requires intelligent systems to observe from the perspective of the first person and predict an action before it occurs. Owing to the uncertainty of future, it is insufficient to perform action anticipation relying on visual information especially when there exists salient visual difference between past and future. In order to alleviate this problem, which we call visual gap in this paper, we propose one novel Intuition-Analysis Integrated (IAI) framework inspired by psychological research, which mainly consists of three parts: Intuition-based Prediction Network (IPN), Analysis-based Prediction Network (APN) and Adaptive Fusion Network (AFN). To imitate the implicit intuitive thinking process, we model IPN as an encoder-decoder structure and introduce one procedural instruction learning strategy implemented by textual pre-training. On the other hand, we allow APN to process information under designed rules to imitate the explicit analytical thinking, which is divided into three steps: recognition, transitions and combination. Both the procedural instruction learning strategy in IPN and the transition step of APN are crucial to improving the anticipation performance via mitigating the visual gap problem. Considering the complementarity of intuition and analysis, AFN adopts attention fusion to adaptively integrate predictions from IPN and APN to produce the final anticipation results. We conduct experiments on the largest egocentric video dataset. Qualitative and quantitative evaluation results validate the effectiveness of our IAI framework, and demonstrate the advantage of bridging visual gap by utilizing multi-modal information, including both visual features of observed segments and sequential instructions of actions. Weiqing Min, Yong Rui, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2020 | CDbin: Compact Discriminative Binary Descriptor Learned With Efficient Neural NetworkabstractAs an important computer vision task, image matching requires efficient and discriminative local descriptors. Most of the existing descriptors like SIFT and ORB are hand-crafted; therefore it is necessary to study more optimized descriptors through end-to-end learning. This paper proposes the compact binary descriptors learned with a lightweight Convolutional Neural Network (CNN), which is efficient for training and testing. Specifically, we propose a CNN with no larger than five layers for descriptor learning. The resulting descriptors, i.e., Compact Discriminative binary descriptors (CDbin) are optimized with four complementary loss functions, i.e., 1) triplet loss to ensure the discriminative power; 2) quantization loss to decrease the quantization error; 3) correlation loss to ensure the feature compactness; and 4) even-distribution loss to enrich the embedded information. The extensive experiments on two image patch datasets and three image retrieval datasets show that the CDbin exhibits competitive performance compared with the existing descriptors. For example, the 64-bit CDbin substantially outperforms the 256-bit ORB and 1024-bit SIFT on Hpatches dataset. Although generated by a shallow CNN, CDbin also outperforms several recent deep descriptors. Jianming Ye, Shiliang Zhang, Tiejun Huang 0001, Yong Rui |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Self-Supervised Agent Learning for Unsupervised Cross-Domain Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) has better scalability and practicability than supervised Re-ID in the actual deployment. However, it is difficult to learn a discriminative Re-ID model without annotations. To address the above issue, we propose an end-to-end Self-supervised Agent Learning (SAL) algorithm by exploiting a set of agents as a bridge to reduce domain gaps for unsupervised cross-domain person Re- ID. The proposed SAL model enjoys several merits. First, to the best of our knowledge, this is the first work to exploit selfsupervised learning for unsupervised person Re-ID. Second, our model has designed three effective learning mechanisms including supervised label learning in source domain, similarity consistency learning in target domain, and self-supervised learning in cross domain, which can learn domain-invariant yet discriminative representations through the principled lens of agent learning by reducing domain discrepancy adaptively. Extensive experimental results on three standard benchmarks demonstrate that the proposed SAL performs favorably against state-of-the-art unsupervised person Re-ID methods. Kongzhu Jiang, Tianzhu Zhang 0001, Yongdong Zhang 0001, Feng Wu 0001, Yong Rui |
IEEE Trans. Image Process. | 5 |
| 2019 | Toward efficient indexing structure for scalable content-based music retrievalabstractWith advancement of various information processing and storage techniques, the scale of digital music collections has been growing at very fast speed during recent decades. To support high-quality content-based retrieval over such a large volume of music data, how to develop indexing structure with good effectiveness, efficiency and scalability becomes an important research issue. However, existing techniques mainly focus on improving query efficiency. Very few approaches have been proposed to address issues related to scalability and accuracy. In this study, we address the problem via introducing a novel indexing technique called effective music indexing framework (EMIF) to facilitate scalable and accurate music retrieval. It is designed based on a “classification-and-indexing” principle and consists of two main functionality modules: (1) music classification—a novel semantic-sensitive classification to identify an input song’s category and (2) indexing module—multiple local indexing structures, one for each semantic category to reduce query response time significantly. In particular, the classification model combining linear discriminative mixture model (LDMM) and advanced score fusion scheme has been applied to estimate category of music accurately. Layered architecture enables EMIF to enjoy superior scalability and efficiency. To evaluate the approach, a set of experimental studies has been carried out using two large music test collections and the results demonstrate various advantages of EMIF over state-of-the-art approaches including efficiency, scalability and effectiveness. Jialie Shen 0001, Tao Mei 0001, Qiang Qu 0001, Dacheng Tao, Yong Rui |
Multim. Syst. | 5 |
| 2019 | Image Recognition by Predicted User Click Feature With Multidomain Multitask Transfer Deep NetworkabstractThe click feature of an image, defined as a user click count vector based on click data, has been demonstrated to be effective for reducing the semantic gap for image recognition. Unfortunately, most of the traditional image recognition datasets do not contain click data. To address this problem, researchers have begun to develop a click prediction model using assistant datasets containing click information and have adapted this predictor to a common click-free dataset for different tasks. This method can be customized to our problem, but it has two main limitations: 1) the predicted click feature often performs badly in the recognition task since the prediction model is constructed independently of the subsequent recognition problem and 2) transferring the predictor from one dataset to another is challenging due to the large cross-domain diversity. In this paper, we devise a multitask and multidomain deep network with varied modals (MTMDD-VM) to formulate image recognition and click prediction tasks in a unified framework. Datasets with and without click information are integrated in the training. Furthermore, a nonlinear word embedding with a position-sensitive loss function is designed to discover the visual click correlation. We evaluate the proposed method on three public dog breed image datasets, and we utilize the Clickture-Dog dataset as the auxiliary dataset that provides click data. The experimental results show that: 1) the nonlinear word embedding and position-sensitive loss function largely enhance the predicted click feature in the recognition task, realizing a 32% improvement in accuracy; 2) the multitask learning framework improves accuracies in both image recognition and click prediction; and 3) the unified training using the combined dataset with and without click data further improves the performance. Compared with the state-of-the-art methods, the proposed approach not only performs much better in accuracy but also achieves good scalability and one-shot learning ability. Min Tan 0005, Jun Yu 0002, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2019 | Unified Spatio-Temporal Attention Networks for Action Recognition in VideosabstractRecognizing actions in videos is not a trivial task because video is an information-intensive media and includes multiple modalities. Moreover, on each modality, an action may only appear at some spatial regions, or only part of the temporal video segments may contain the action. A valid question is how to locate the attended spatial areas and selective video segments for action recognition. In this paper, we devise a general attention neural cell, called AttCell, that estimates the attention probability not only at each spatial location but also for each video segment in a temporal sequence. With AttCell, a unified Spatio-Temporal Attention Networks (STAN) is proposed in the context of multiple modalities. Specifically, STAN extracts the feature map of one convolutional layer as the local descriptors on each modality and pools the extracted descriptors with the spatial attention measured by AttCell as a representation of each segment. Then, we concatenate the representation on each modality to seek a consensus on the temporal attention, a priori, to holistically fuse the combined representation of video segments to the video representation for recognition. Our model differs from conventional deep networks, which focus on the attention mechanism, because our temporal attention provides a principled and global guidance across different modalities and video segments. Extensive experiments are conducted on four public datasets; UCF101, CCV, THUMOS14, and Sports-1M; our STAN consistently achieves superior results over several state-of-the-art techniques. More remarkably, we validate and demonstrate the effectiveness of our proposal when capitalizing on the different number of modalities. Dong Li 0019, Ting Yao 0003, Ling-Yu Duan, Tao Mei 0001, Yong Rui |
IEEE Trans. Multim. | 5 |
| 2019 | Learning Click-Based Deep Structure-Preserving Embeddings with Visual AttentionabstractOne fundamental problem in image search is to learn the ranking functions (i.e., the similarity between query and image). Recent progress on this topic has evolved through two paradigms: the text-based model and image ranker learning. The former relies on image surrounding texts, making the similarity sensitive to the quality of textual descriptions. The latter may suffer from the robustness problem when human-labeled query-image pairs cannot represent user search intent precisely. We demonstrate in this article that the preceding two limitations can be well mitigated by learning a cross-view embedding that leverages click data. Specifically, a novel click-based Deep Structure-Preserving Embeddings with visual Attention (DSPEA) model is presented, which consists of two components: deep convolutional neural networks followed by image embedding layers for learning visual embedding, and a deep neural networks for generating query semantic embedding. Meanwhile, visual attention is incorporated at the top of the convolutional neural network to reflect the relevant regions of the image to the query. Furthermore, considering the high dimension of the query space, a new click-based representation on a query set is proposed for alleviating this sparsity problem. The whole network is end-to-end trained by optimizing a large margin objective that combines cross-view ranking constraints with in-view neighborhood structure preservation constraints. On a large-scale click-based image dataset with 11.7 million queries and 1 million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods and achieves, to date, the best reported NDCG@25 of 52.21%. Yehao Li, Yingwei Pan, Ting Yao 0003, Hongyang Chao, Yong Rui, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2018 | Sequence-to-Sequence Learning via Shared Latent RepresentationabstractSequence-to-sequence learning is a popular research area in deep learning, such as video captioning and speech recognition. Existing methods model this learning as a mapping process by first encoding the input sequence to a fixed-sized vector, followed by decoding the target sequence from the vector. Although simple and intuitive, such mapping model is task-specific, unable to be directly used for different tasks. In this paper, we propose a star-like framework for general and flexible sequence-to-sequence learning, where different types of media contents (the peripheral nodes) could be encoded to and decoded from a shared latent representation (SLR) (the central node). This is inspired by the fact that human brain could learn and express an abstract concept in different ways. The media-invariant property of SLR could be seen as a high-level regularization on the intermediate vector, enforcing it to not only capture the latent representation intra each individual media like the auto-encoders, but also their transitions like the mapping models. Moreover, the SLR model is content-specific, which means it only needs to be trained once for a dataset, while used for different tasks. We show how to train a SLR model via dropout and use it for different sequence-to-sequence tasks. Our SLR model is validated on the Youtube2Text and MSR-VTT datasets, achieving superior performance on video-to-sentence task, and the first sentence-to-video results. Xu Shen 0001, Xinmei Tian 0001, Jun Xing, Yong Rui, Dacheng Tao |
AAAI | 4 |
| 2018 | Hierarchical semantic image matching using CNN feature pyramid
Wei Yu 0004, Xiaoshuai Sun, Kuiyuan Yang, Yong Rui, Hongxun Yao |
Comput. Vis. Image Underst. | 4 |
| 2018 | Automatic Generation of Social Event Storyboard From Image Click-Through DataabstractRecent studies have shown that a noticeable percentage of Web search traffic is about social events. While traditional Web sites can only show human-edited events, in this paper, we present a novel system to automatically detect events from search log data and generate storyboards where the events are arranged chronologically. We chose image search log as the resource for event mining, as search logs can directly reflect people's interests. To discover events from log data, we present a smooth nonnegative matrix factorization framework, which combines the information of query semantics, temporal correlations, search logs, and time continuity. Moreover, we consider the time factor to be an important element, since different events will develop in different time tendencies. In addition, to provide a media-rich and visually appealing storyboard, each event is associated with a set of representative photos arranged along a timeline. These relevant photos are automatically selected from image search results by analyzing image content features. We use celebrities as our test domain, which takes a large percentage of image search traffic. Experiments consisting of Web search traffic on 200 celebrities, for a period of six months, show very encouraging results compared with handcrafted editorial storyboards. Tao Mei 0001, Rui Cai 0002, Houqiang Li, Yong Rui |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Multimodal Deep Embedding via Hierarchical Grounded Compositional SemanticsabstractFor a number of important problems, isolated semantic representations of individual syntactic words or visual objects do not suffice, but instead a compositional semantic representation is required; for example, a literal phrase or a set of spatially concurrent objects. In this paper, we aim to harness the existing image-sentence databases to exploit the compositional nature of image-sentence data for multimodal deep embedding. In particular, we propose an approach called hierarchical-alike (bottom-up two layers) multimodal grounded compositional semantics (hiMoCS) learning. The proposed hiMoCS systemically captures the compositional semantic connotation of multimodal data in the setting of hierarchical-alike deep learning by modeling the inherent correlations between two modalities of collaboratively grounded semantics, such as the textual entity (with its describing attribute) and visual object, the phrase (e.g., subject-verb-object triplet), and spatially concurrent objects. We argue that hiMoCS is more appropriate to reflect the multimodal compositional semantics of the image and its narrative textual sentence, which are strongly coupled. We evaluate hiMoCS on the several benchmark data sets and show that the utilization of the hiMoCS (textual entities and visual objects, textual phrase, and spatially concurrent objects) achieves a much better performance than only using the flat grounded compositional semantics. Yueting Zhuang, Jun Song 0004, Fei Wu 0001, Xi Li 0001, Zhongfei Zhang, Yong Rui |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Scalable Content-Aware Collaborative Filtering for Location RecommendationabstractLocation recommendation plays an essential role in helping people find attractive places. Though recent research has studied how to recommend locations with social and geographical information, few of them addressed the cold-start problem of new users. Because mobility records are often shared on social networks, semantic information can be leveraged to tackle this challenge. A typical method is to feed them into explicit-feedback-based content-aware collaborative filtering, but they require drawing negative samples for better learning performance, as users’ negative preference is not observable in human mobility. However, prior studies have empirically shown sampling-based methods do not perform well. To this end, we propose a scalable Implicit-feedback-based Content-aware Collaborative Filtering (ICCF) framework to incorporate semantic content and to steer clear of negative sampling. We then develop an efficient optimization algorithm, scaling linearly with data size and feature size, and quadratically with the dimension of latent space. We further establish its relationship with graph Laplacian regularized matrix factorization. Finally, we evaluate ICCF with a large-scale LBSN dataset in which users have profiles and textual content. The results show that ICCF outperforms several competing baselines, and that user information is not only effective for improving recommendations but also coping with cold-start scenarios. Defu Lian, Yong Ge 0001, Nicholas Jing Yuan, Xing Xie 0001, Tao Zhou 0001, Yong Rui |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2018 | You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food AnalysisabstractCuisine is a style of cooking and usually associated with a specific geographic region. Recipes from different cuisines shared on the web are an indicator of culinary cultures in different countries. Therefore, analysis of these recipes can lead to deep understanding of food from the cultural perspective. In this paper, we perform the first cross-region recipe analysis by jointly using the recipe ingredients, food images, and attributes such as the cuisine and course (e.g., main dish and dessert). For that solution, we propose a culinary culture analysis framework to discover the topics of ingredient bases and visualize them to enable various applications. We first propose a probabilistic topic model to discover cuisine-course specific topics. The manifold ranking method is then utilized to incorporate deep visual features to retrieve food images for topic visualization. At last, we applied the topic modeling and visualization method for three applications: 1) multimodal cuisine summarization with both recipe ingredients and images, 2) cuisine-course pattern analysis including topic-specific cuisine distribution and cuisine-specific course distribution of topics, and 3) cuisine recommendation for both cuisine-oriented and ingredient-oriented queries. Through these three applications, we can analyze the culinary cultures at both macro and micro levels. We conduct the experiment on a recipe database Yummly-66K with 66,615 recipes from 10 cuisines in Yummly. Qualitative and quantitative evaluation results have validated the effectiveness of topic modeling and visualization, and demonstrated the advantage of the framework in utilizing rich recipe information to analyze and interpret the culinary cultures from different regions. Weiqing Min, Bing-Kun Bao, Shuhuan Mei, Yong Rui, Shuqiang Jiang |
IEEE Trans. Multim. | 5 |
| 2018 | User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric LearningabstractWe present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks. Min Tan 0005, Jun Yu 0002, Zhou Yu 0001, Fei Gao 0006, Yong Rui, Dacheng Tao |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2017 | Multi-level Attention Networks for Visual Question Answering
Dongfei Yu, Jianlong Fu, Tao Mei 0001, Yong Rui |
CVPR | 4 |
| 2017 | Beyond the Words: Predicting User Personality from Heterogeneous InformationabstractAn incisive understanding of user personality is not only essential to many scientific disciplines, but also has a profound business impact on practical applications such as digital marketing, personalized recommendation, mental diagnosis, and human resources management. Previous studies have demonstrated that language usage in social media is effective in personality prediction. However, except for single language features, a less researched direction is how to leverage the heterogeneous information on social media to have a better understanding of user personality. In this paper, we propose a Heterogeneous Information Ensemble framework, called HIE, to predict users' personality traits by integrating heterogeneous information including self-language usage, avatar, emoticon, and responsive patterns. In our framework, to improve the performance of personality prediction, we have designed different strategies extracting semantic representations to fully leverage heterogeneous information on social media. We evaluate our methods with extensive experiments based on a real-world data covering both personality survey results and social media usage from thousands of volunteers. The results reveal that our approaches significantly outperform several widely adopted state-of-the-art baseline methods. To figure out the utility of HIE in a real-world interactive setting, we also present DiPsy, a personalized chatbot to predict user personality through heterogeneous information in digital traces and conversation logs. Honghao Wei, Nicholas Jing Yuan, Chuan Cao, Hao Fu 0015, Xing Xie 0001, Yong Rui, Wei-Ying Ma |
WSDM | 7 |
| 2017 | LEGO-MM: LEarning Structured Model by Probabilistic loGic Ontology Tree for MultiMediaabstractRecent advances in multimedia ontology have resulted in a number of concept models, e.g., large-scale concept for multimedia and Mediamill 101, which are accessible and public to other researchers. However, most current research effort still focuses on building new concepts from scratch, very few work explores the appropriate method to construct new concepts upon the existing models already in the warehouse. To address this issue, we propose a new framework in this paper, termed LEarning Structured Model by Probabilistic loGic Ontology Tree for MultiM edia (LEGO1-MM), which can seamlessly integrate both the new target training examples and the existing primitive concept models to infer the more complex concept models. LEGO-MM treats the primitive concept models as the lego toy to potentially construct an unlimited vocabulary of new concepts. Specifically, we first formulate the logic operations to be the lego connectors to combine the existing concept models hierarchically in probabilistic logic ontology trees. Then, we incorporate new target training information simultaneously to efficiently disambiguate the underlying logic tree and correct the error propagation. Extensive experiments are conducted on a large vehicle domain data set from ImageNet. The results demonstrate that LEGO-MM has significantly superior performance over the existing state-of-the-art methods, which build new concept models from scratch. Jinhui Tang 0001, Shiyu Chang, Guo-Jun Qi, Qi Tian 0001, Yong Rui, Thomas S. Huang |
IEEE Trans. Image Process. | 5 |
| 2017 | Robust Spammer Detection in Microblogs: Leveraging User CarefulnessabstractMicroblogging Web sites, such as Twitter and Sina Weibo, have become popular platforms for socializing and sharing information in recent years. Spammers have also discovered this new opportunity to unfairly overpower normal users with unsolicited content, namely social spams. Although it is intuitive for everyone to follow legitimate users, recent studies show that both legitimate users and spammers follow spammers for different reasons. Evidence of users seeking spammers on purpose is also observed. We regard this behavior as useful information for spammer detection. In this article, we approach the problem of spammer detection by leveraging the “carefulness” of users, which indicates how careful a user is when she is about to follow a potential spammer. We propose a framework to measure the carefulness and develop a supervised learning algorithm to estimate it based on known spammers and legitimate users. We illustrate how the robustness of the detection algorithms can be improved with aid of the proposed measure. Evaluation on two real datasets from Sina Weibo and Twitter with millions of users are performed, as well as an online test on Sina Weibo. The results show that our approach indeed captures the carefulness, and it is effective for detecting spammers. In addition, we find that our measure is also beneficial for other applications, such as link prediction. Hao Fu 0015, Xing Xie 0001, Yong Rui, Neil Zhenqiang Gong, Guangzhong Sun, Enhong Chen |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Search by Screenshots for Universal Article Clipping in Mobile AppsabstractTo address the difficulty in clipping articles from various mobile applications (apps), we propose a novel framework called UniClip, which allows a user to snap a screen of an article to save the whole article in one place. The key task of the framework is search by screenshots , which has three challenges: (1) how to represent a screenshot; (2) how to formulate queries for effective article retrieval; and (3) how to identify the article from search results. We solve these by (1) segmenting a screenshot into structural units called blocks, (2) formulating effective search queries by considering the role of each block, and (3) aggregating the search result lists of multiple queries. To improve efficiency, we also extend our approach with learning-to-rank techniques so that we can find the desired article with only one query. Experimental results show that our approach achieves high retrieval performance ( F 1 = 0.868), which outperforms baselines based on keyword extraction and chunking methods. Learning-to-rank models improve our approach without learning by about 6%. A user study conducted to investigate the usability of UniClip reveals that ours is preferred by 21 out of 22 participants for its simplicity and effectiveness. Kazutoshi Umemoto, Ruihua Song, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui |
ACM Trans. Inf. Syst. | 6 |
| 2017 | Enhancing Person Re-identification in a Self-Trained SubspaceabstractDespite the promising progress made in recent years, person re-identification (re-ID) remains a challenging task due to the complex variations in human appearances from different camera views. For this challenging problem, a large variety of algorithms have been developed in the fully supervised setting, requiring access to a large amount of labeled training data. However, the main bottleneck for fully supervised re-ID is the limited availability of labeled training samples. To address this problem, we propose a self-trained subspace learning paradigm for person re-ID that effectively utilizes both labeled and unlabeled data to learn a discriminative subspace where person images across disjoint camera views can be easily matched. The proposed approach first constructs pseudo-pairwise relationships among unlabeled persons using the k-nearest neighbors algorithm. Then, with the pseudo-pairwise relationships, the unlabeled samples can be easily combined with the labeled samples to learn a discriminative projection by solving an eigenvalue problem. In addition, we refine the pseudo-pairwise relationships iteratively, which further improves learning performance. A multi-kernel embedding strategy is also incorporated into the proposed approach to cope with the non-linearity in a person’s appearance and explore the complementation of multiple kernels. In this way, the performance of person re-ID can be greatly enhanced when training data are insufficient. Experimental results on six widely used datasets demonstrate the effectiveness of our approach, and its performance can be comparable to the reported results of most state-of-the-art fully supervised methods while using much fewer labeled data. Xun Yang 0001, Meng Wang 0001, Richang Hong, Qi Tian 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2017 | Saliency Detection on Light Field: A Multi-Cue ApproachabstractSaliency detection has recently received increasing research interest on using high-dimensional datasets beyond two-dimensional images. Despite the many available capturing devices and algorithms, there still exists a wide spectrum of challenges that need to be addressed to achieve accurate saliency detection. Inspired by the success of the light-field technique, in this article, we propose a new computational scheme to detect salient regions by integrating multiple visual cues from light-field images. First, saliency prior maps are generated from several light-field features based on superpixel-level intra-cue distinctiveness, such as color, depth, and flow inherited from different focal planes and multiple viewpoints. Then, we introduce the location prior to enhance the saliency maps. These maps will finally be merged into a single map using a random-search-based weighting strategy. Besides, we refine the object details by employing a two-stage saliency refinement to obtain the final saliency map. In addition, we present a more challenging benchmark dataset for light-field saliency analysis, named HFUT-Lytro , which consists of 255 light fields with a range from 53 to 64 images generated from each light-field image, therein spanning multiple occurrences of saliency detection challenges such as occlusions, cluttered background, and appearance changes. Experimental results show that our approach can achieve 0.6--6.7% relative improvements over state-of-the-art methods in terms of the F-measure and Precision metrics, which demonstrates the effectiveness of the proposed approach. Jun Zhang 0017, Meng Wang 0001, Liang Lin 0004, Xun Yang 0001, Jun Gao 0006, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2016 | Jointly Modeling Embedding and Translation to Bridge Video and LanguageabstractAutomatically describing video content with natural language is a fundamental challenge of computer vision. Re-current Neural Networks (RNNs), which models sequence dynamics, has attracted increasing attention on visual interpretation. However, most existing approaches generate a word locally with the given previous words and the visual content, while the relationship between sentence semantics and visual content is not holistically exploited. As a result, the generated sentences may be contextually correct but the semantics (e.g., subjects, verbs or objects) are not true. This paper presents a novel unified framework, named Long Short-Term Memory with visual-semantic Embedding (LSTM-E), which can simultaneously explore the learning of LSTM and visual-semantic embedding. The former aims to locally maximize the probability of generating the next word given previous words and visual content, while the latter is to create a visual-semantic embedding space for enforcing the relationship between the semantics of the entire sentence and visual content. The experiments on YouTube2Text dataset show that our proposed LSTM-E achieves to-date the best published performance in generating natural sentences: 45.3% and 31.0% in terms of BLEU@4 and METEOR, respectively. Superior performances are also reported on two movie description datasets (M-VAD and MPII-MD). In addition, we demonstrate that LSTM-E outperforms several state-of-the-art techniques in predicting Subject-Verb-Object (SVO) triplets. Yingwei Pan, Tao Mei 0001, Ting Yao 0003, Houqiang Li, Yong Rui |
CVPR | 5 |
| 2016 | MSR-VTT: A Large Video Description Dataset for Bridging Video and LanguageabstractWhile there has been increasing interest in the task of describing video with natural language, current computer vision algorithms are still severely limited in terms of the variability and complexity of the videos and their associated language that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on specific fine-grained domains with limited videos and simple descriptions. While researchers have provided several benchmark datasets for image captioning, we are not aware of any large-scale video description dataset with comprehensive categories yet diverse video content. In this paper we present MSR-VTT (standing for "MSRVideo to Text") which is a new large-scale video benchmark for video understanding, especially the emerging task of translating video to text. This is achieved by collecting 257 popular queries from a commercial video search engine, with 118 videos for each query. In its current version, MSR-VTT provides 10K web video clips with 41.2 hours and 200K clip-sentence pairs in total, covering the most comprehensive categories and diverse visual content, and representing the largest dataset in terms of sentence and vocabulary. Each clip is annotated with about 20 natural sentences by 1,327 AMT workers. We present a detailed analysis of MSR-VTT in comparison to a complete set of existing datasets, together with a summarization of different state-of-the-art video-to-text approaches. We also provide an extensive evaluation of these approaches on this dataset, showing that the hybrid Recurrent Neural Networkbased approach, which combines single-frame and motion representations with soft-attention pooling strategy, yields the best generalization capability on MSR-VTT. Tao Mei 0001, Ting Yao 0003, Yong Rui |
CVPR | 4 |
| 2016 | Highlight Detection with Pairwise Deep Ranking for First-Person Video SummarizationabstractThe emergence of wearable devices such as portable cameras and smart glasses makes it possible to record life logging first-person videos. Browsing such long unstructured videos is time-consuming and tedious. This paper studies the discovery of moments of user's major or special interest (i.e., highlights) in a video, for generating the summarization of first-person videos. Specifically, we propose a novel pairwise deep ranking model that employs deep learning techniques to learn the relationship between high-light and non-highlight video segments. A two-stream network structure by representing video segments from complementary information on appearance of video frames and temporal dynamics across frames is developed for video highlight detection. Given a long personal video, equipped with the highlight detection model, a highlight score is assigned to each segment. The obtained highlight segments are applied for summarization in two ways: video time-lapse and video skimming. The former plays the highlight (non-highlight) segments at low (high) speed rates, while the latter assembles the sequence of segments with the highest scores. On 100 hours of first-person videos for 15 unique sports categories, our highlight detection achieves the improvement over the state-of-the-art RankSVM method by 10.5% in terms of accuracy. Moreover, our approaches produce video summary with better quality by a user study from 35 human subjects. Ting Yao 0003, Tao Mei 0001, Yong Rui |
CVPR | 3 |
| 2016 | Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette ConsistencyabstractIn this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouettes for multiple views on the fly. To obtain a set of accurate silhouettes, precise camera poses are required to propagate segmentation cues across views. To perform better localization, accurate silhouettes are needed to constrain camera poses. Therefore the two problems are intertwined with each other and require a joint treatment. Facilitated by the available depth, we introduce a simple but effective silhouette consistency energy term that binds traditional appearance-based multiview segmentation cost and RGB-D frame-to-frame matching cost together. Optimization of the problem w.r.t. binary segmentation masks and camera poses naturally fits in the graph cut minimization framework and the Gauss-Newton non-linear least-squares method respectively. Experiments show that the proposed approach achieves state-of-the-arts performance on both tasks of image segmentation and camera localization. Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui |
CVPR | 5 |
| 2016 | Network MorphismabstractWe present a systematic study on how to morph a well-trained neural network to a new one so that its network function can be completely preserved. We define this as network morphism in this research. After morphing a parent network, the child network is expected to inherit the knowledge from its parent network and also has the potential to continue growing into a more powerful one with much shortened training time. The first requirement for this network morphism is its ability to handle diverse morphing types of networks, including changes of depth, width, kernel size, and even subnet. To meet this requirement, we first introduce the network morphism equations, and then develop novel morphing algorithms for all these morphing types for both classic and convolutional neural networks. The second requirement is its ability to deal with non-linearity in a network. We propose a family of parametric-activation functions to facilitate the morphing of any continuous non-linear activation neurons. Experimental results on benchmark datasets and typical neural networks demonstrate the effectiveness of the proposed network morphism scheme. Changhu Wang, Yong Rui, Chang Wen Chen |
ICML | 3 |
| 2016 | Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition
Yanhua Cheng, Xin Zhao 0012, Rui Cai 0002, Zhiwei Li 0006, Kaiqi Huang, Yong Rui |
IJCAI | 6 |
| 2016 | Learning Deep Intrinsic Video Representation by Exploring Temporal Coherence and Graph Structure
Yingwei Pan, Yehao Li, Ting Yao 0003, Tao Mei 0001, Houqiang Li, Yong Rui |
IJCAI | 6 |
| 2016 | Deep Semantic-Preserving and Ranking-Based Hashing for Image Retrieval
Ting Yao 0003, Fuchen Long, Tao Mei 0001, Yong Rui |
IJCAI | 4 |
| 2016 | Action Recognition by Learning Deep Multi-Granular Spatio-Temporal Video RepresentationabstractRecognizing actions in videos is a challenging task as video is an information-intensive media with complex variations. Most existing methods have treated video as a flat data sequence while ignoring the intrinsic hierarchical structure of the video content. In particular, an action may span different granularities in this hierarchy including, from small to large, a single frame, consecutive frames (motion), a short clip, and the entire video. In this paper, we present a novel framework to boost action recognition by learning a deep spatio-temporal video representation at hierarchical multi-granularity. Specifically, we model each granularity as a single stream by 2D (for frame and motion streams) or 3D (for clip and video streams) convolutional neural networks (CNNs). The framework therefore consists of multi-stream 2D or 3D CNNs to learn both the spatial and temporal representations. Furthermore, we employ the Long Short-Term Memory (LSTM) networks on the frame, motion, and clip streams to exploit long-term temporal dynamics. With a softmax layer on the top of each stream, the classification scores can be predicted from all the streams, followed by a novel fusion scheme based on the multi-granular score distribution. Our networks are learned in an end-to-end fashion. On two video action benchmarks of UCF101 and HMDB51, our framework achieves promising performance compared with the state-of-the-art. Qing Li 0003, Zhaofan Qiu, Ting Yao 0003, Tao Mei 0001, Yong Rui, Jiebo Luo 0001 |
ICMR | 5 |
| 2016 | Video ChatBot: Triggering Live Social Interactions by Automatic Video CommentingabstractWe demonstrate a video chatbot, which can generate human-level emotional comments referring to the videos shared by users and trigger a conversation with users. Our video chatbot performs a large-scale similar video search to find visually similar videos w.r.t. a given video using approximate nearest-neighbor search. Then, the comments associated with the searched similar videos are ranked by learning a deep multi-view embedding space for modeling video content, visual sentiment and textual comments. The top ranked comments are selected as responses to the given video and trigger the succeeding text-based chat between users and the chatbot. The demonstration is conducted on a newly collected dataset with over 102K videos and 10.6M comments. Moreover, our video chatbot has great potential to increase live social interactions. Yehao Li, Ting Yao 0003, Tao Mei 0001, Yong Rui |
ACM Multimedia | 5 |
| 2016 | Share-and-Chat: Achieving Human-Level Video Commenting by Search and Multi-View EmbeddingabstractVideo has become a predominant social media for the booming live interactions. Automatic generation of emotional comments to a video has great potential to significantly increase user engagement in many socio-video applications (e.g., chat bot). Nevertheless, the problem of video commenting has been overlooked by the research community. The major challenges are that the generated comments are to be not only as natural as those from human beings, but also relevant to the video content. We present in this paper a novel two-stage deep learning-based approach to automatic video commenting. Our approach consists of two components. The first component, similar video search, efficiently finds the visually similar videos w.r.t. a given video using approximate nearest-neighbor search based on the learned deep video representations, while the second dynamic ranking effectively ranks the comments associated with the searched similar videos by learning a deep multi-view embedding space. For modeling the emotional view of videos, we incorporate visual sentiment, video content, and text comments into the learning of the embedding space. On a newly collected dataset with over 102K videos and 10.6M comments, we demonstrate that our approach outperforms several state-of-the-art methods and achieves human-level video commenting. Yehao Li, Ting Yao 0003, Tao Mei 0001, Hongyang Chao, Yong Rui |
ACM Multimedia | 5 |
| 2016 | Image2Text: A Multimodal Image CaptionerabstractIn this work, we showcase the Image2Text system, which is a real-time captioning system that can generate human-level natural language description for any input image. We formulate the problem of image captioning as a multimodal translation task. Analogous to machine translation, we present a sequence-to-sequence recurrent neural networks (RNN) model for image caption generation. Different from most existing work where the whole image is represented by a convolutional neural networks (CNN) feature, we propose to represent the input image as a sequence of detected objects to serve as the source sequence of the RNN model. Based on the captioning framework, we develop a user-friendly system to automatically generated human-level captions for users. The system also enables users to detect salient objects in an image, and retrieve similar images and corresponding descriptions from a database. Changhu Wang, Fuchun Sun 0001, Yong Rui |
ACM Multimedia | 4 |
| 2016 | Who Will Reply to/Retweet This Tweet?: The Dynamics of Intimacy from Online Social InteractionsabstractFriendships are dynamic. Previous studies have converged to suggest that social interactions, in both online and offline social networks, are diagnostic reflections of friendship relations (also called social ties). However, most existing approaches consider a social tie as either a binary relation, or a fixed value (named tie strength). In this paper, we investigate the dynamics of dyadic friend relationships through online social interactions, in terms of a variety of aspects, such as reciprocity, temporality, and contextuality. In turn, we propose a model to predict repliers and retweeters given a particular tweet posted at a certain time in a microblog-based social network. More specifically, we have devised a learning-to-rank approach to train a ranker that considers elaborate user-level and tweet-level features (like sentiment, self-disclosure, and responsiveness) to address these dynamics. In the prediction phase, a tweet posted by a user is deemed a query and the predicted repliers/retweeters are retrieved using the learned ranker. We have collected a large dataset containing 73.3 million dyadic relationships with their interactions (replies and retweets). Extensive experimental results based on this dataset show that by incorporating the dynamics of friendship relations, our approach significantly outperforms state-of-the-art models in terms of multiple evaluation metrics, such as MAP, NDCG and Topmost Accuracy. In particular, the advantage of our model is even more promising in predicting the exact sequence of repliers/retweeters considering their orders. Furthermore, the proposed approach provides emerging implications for many high-value applications in online social networks. Nicholas Jing Yuan, Xing Xie 0001, Chin-Yew Lin, Yong Rui |
WSDM | 6 |
| 2016 | Exploiting Dining Preference for Restaurant RecommendationabstractThe wide adoption of location-based services provide the potential to understand people's mobility pattern at an unprecedented level, which can also enable food-service industry to accurately predict consumers' dining behavior. In this paper, based on users' dining implicit feedbacks (restaurant visit via check-ins), explicit feedbacks (restaurant reviews) as well as some meta data (e.g., location, user demographics, restaurant attributes), we aim at recommending each user a list of restaurants for his next dining. Implicit and Explicit feedbacks of dining behavior exhibit different characteristics of user preference. Therefore, in our work, user's dining preference mainly contains two parts: implicit preference coming from check-in data (implicit feedbacks) and explicit preference coming from rating and review data (explicit feedbacks). For implicit preference, we first apply a probabilistic tensor factorization model (PTF) to capture preference in a latent subspace. Then, in order to incorporate contextual signals from meta data, we extend PTF by proposing an Implicit Preference Model (IPM), which can simultaneously capture users'/restaurants'/time' preference in the collaborative filtering and dining preference in a specific context (e.g., spatial distance preference, environmental preference). For explicit preference, we propose Explicit Preference Model (EPM) by combining matrix factorization with topic modeling to discover the user preference embedded both in rating score and text content. Finally, we design a unified model termed as Collective Implicit Explicit Preference Model (CIEPM) to combine implicit and explicit preference together for restaurant recommendation. To evaluate the performance of our system, we conduct extensive experiments with large-scale datasets covering hundreds of thousands of users and restaurants. The results reveal that our system is effective for restaurant recommendation. Nicholas Jing Yuan, Kai Zheng 0001, Defu Lian, Xing Xie 0001, Yong Rui |
WWW | 6 |
| 2016 | UniClip: Leveraging Web Search for Universal Clipping of Articles on MobileabstractIn this paper we address the difficulty of clipping articles from mobile apps. We propose a service called UniClip that allows a user to save the full content of an article by snapping a screenshot part of it. UniClip leverages a huge amount of indexed web data to mine the article by starting with a snapped screenshot. We propose approaches to solve three challenges: (1) how to represent a screenshot; (2) how to formulate effective queries for retrieving a full article; and (3) how to rank the best URL at the top from multiple search result lists. Experimental results indicate that our approach is effective in achieving as high an $$F_1$$ F 1 measure as 0.905, which outperforms the best of three baseline methods by 18 points. Ruihua Song, Kazutoshi Umemoto, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui |
Data Sci. Eng. | 6 |
| 2016 | Building Hierarchical Representations for Oracle Character and Sketch RecognitionabstractIn this paper, we study oracle character recognition and general sketch recognition. First, a data set of oracle characters, which are the oldest hieroglyphs in China yet remain a part of modern Chinese characters, is collected for analysis. Second, typical visual representations in shape- and sketch-related works are evaluated. We analyze the problems suffered when addressing these representations and determine several representation design criteria. Based on the analysis, we propose a novel hierarchical representation that combines a Gabor-related low-level representation and a sparse-encoder-related mid-level representation. Extensive experiments show the effectiveness of the proposed representation in both oracle character recognition and general sketch recognition. The proposed representation is also complementary to convolutional neural network (CNN)-based models. We introduce a solution to combine the proposed representation with CNN-based models, and achieve better performances over both approaches. This solution has beaten humans at recognizing general sketches. Jun Guo 0024, Changhu Wang, Edgar Roman-Rangel, Hongyang Chao, Yong Rui |
IEEE Trans. Image Process. | 5 |
| 2016 | Learning of Multimodal Representations With Random Walks on the Click GraphabstractIn multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. With the click data collected from the users' searching behavior, existing approaches take either one-to-one paired data (text-image pairs) or ranking examples (text-query-image and/or image-query-text ranking lists) as training examples, which do not make full use of the click data, particularly the implicit connections among the data objects. In this paper, we treat the click data as a large click graph, in which vertices are images/text queries and edges indicate the clicks between an image and a query. We consider learning a multimodal representation from the perspective of encoding the explicit/implicit relevance relationship between the vertices in the click graph. By minimizing both the truncated random walk loss as well as the distance between the learned representation of vertices and their corresponding deep neural network output, the proposed model which is named multimodal random walk neural network (MRW-NN) can be applied to not only learn robust representation of the existing multimodal data in the click graph, but also deal with the unseen queries and images to support cross-modal retrieval. We evaluate the latent representation learned by MRW-NN on a public large-scale click log data set Clickture and further show that MRW-NN achieves much better cross-modal retrieval performance on the unseen queries/images than the other state-of-the-art methods. Fei Wu 0001, Jun Song 0004, Shuicheng Yan, Zhongfei Zhang, Yong Rui, Yueting Zhuang |
IEEE Trans. Image Process. | 6 |
| 2016 | Monet: A System for Reliving Your Memories by Theme-Based Photo StorytellingabstractWith the ever-increasing use of smartphones and digital cameras, people are now able to take photos anywhere and anytime. Most of these photos simply end up stored in the cloud without further interaction. This occurs because we lack intelligent services to organize these personal photos well. Therefore, there is an urgent need for such a system to enable people to relive their memories by turning their photos into stories. This paper presents a storytelling system named Monet, which automatically creates interesting stories from personal photos by mimicking cinematic knowledge based on a set of predesigned editing styles. The system consists of two stages: photo summarization, which selects a subset of the “best” photos to represent a photo collection, and story remixing, which generates a stylish music video from the selected photos. During photo summarization, photos are grouped into events based on multimodal features (time and location). The “best” photos are then selected according to visual quality, event representativeness, and diversity. The second stage, story remixing, automatically selects an appropriate theme-dependent editing style based on the photo content. Each selected photo is converted to a video clip by applying a virtual camera with appropriate motions. A series of video effects, color filters, shapes, and transitions are then applied to the video clips according to cinematic rules. The generated video is finally multiplexed with a music clip to generate the story. Evaluations show that our system achieves superior performance to state-of-the-art photo event detection and story generation systems. Xu Shen 0001, Tao Mei 0001, Xinmei Tian 0001, Nenghai Yu, Yong Rui |
IEEE Trans. Multim. | 6 |
| 2016 | Automatic Generation of Visual-Textual Presentation LayoutabstractVisual-textual presentation layout (e.g., digital magazine cover, poster, Power Point slides, and any other rich media), which combines beautiful image and overlaid readable texts, can result in an eye candy touch to attract users’ attention. The designing of visual-textual presentation layout is therefore becoming ubiquitous in both commercially printed publications and online digital magazines. However, handcrafting aesthetically compelling layouts still remains challenging for many small businesses and amateur users. This article presents a system to automatically generate visual-textual presentation layouts by investigating a set of aesthetic design principles, through which an average user can easily create visually appealing layouts. The system is attributed with a set of topic-dependent layout templates and a computational framework integrating high-level aesthetic principles (in a top-down manner) and low-level image features (in a bottom-up manner). The layout templates, designed with prior knowledge from domain experts, define spatial layouts, semantic colors, harmonic color models, and font emotion and size constraints. We formulate the typography as an energy optimization problem by minimizing the cost of text intrusion, the utility of visual space, and the mismatch of information importance in perception and semantics, constrained by the automatically selected template and further preserving color harmonization. We demonstrate that our designs achieve the best reading experience compared with the reimplementation of parts of existing state-of-the-art designs through a series of user studies. Xuyong Yang, Tao Mei 0001, Ying-Qing Xu, Yong Rui, Shipeng Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2015 | Scalable Visual Instance Mining with Instance GraphabstractIn this paper we address the problem of visual instance mining, which is to automatically discover frequently appearing visual instances from a large collection of images. We propose a scalable mining method by leveraging the graph structure with images as vertices. Different from most existing work that focused on either instance-level similarities or image-level context properties, our graph captures both information. The instance-level information is integrated during the construction of a weighted and undirected instance graph based on the similarity between augmented local features, while the image-level context is explored with a greedy breadth-first search algorithm to discover clusters of visual instances from the graph. This method is capable of mining challenging small visual instances with diverse variations. We evaluated our method on two fully annotated datasets and outperformed the state of the arts on both datasets with higher recalls. We also applied our method on a one-million Flickr dataset and proved its scalability. Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010 |
BMVC | 4 |
| 2015 | Automatically Solving Number Word Problems by Semantic Parsing and ReasoningabstractThis paper presents a semantic parsing and reasoning approach to automatically solving math word problems.A new meaning representation language is designed to bridge natural language text and math expressions.A CFG parser is implemented based on 9,600 semi-automatically created grammar rules.We conduct experiments on a test set of over 1,500 number word problems (i.e., verbally expressed number problems) and yield 95.4% precision and 60.2% recall. Shuming Shi 0001, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, Yong Rui |
EMNLP | 5 |
| 2015 | Mining consumer impulsivity from offline and online behaviorabstractConsumer impulsivity is a psychological feature characterizing the impulsive buying tendency. In this paper, by bridging consumer behavior with perceived stimuli on social networks, we present a computational framework, termed Consumer Impulsivity Model (CIM), for exploring a consumer's impulsivity in both offline and online context: consumption-related location visit indicating consumption patterns in the physical realm, and online shopping behavior indicating economic activities on the Internet. To demonstrate the effectiveness of CIM, we conduct extensive experiments, with a large dataset we have collected from thousands of consumers. The results show that 1) for 103 subjects, the inferred consumer impulsivity has a positive Pearson correlation with survey results in the situation of product and product category, respectively. 2) females inferred impulsivity is higher than males on average in the situation of product and product category, respectively. Age has a negative Pearson correlation with inferred impulsivity in the situation of POI, POI category and product category, respectively. 3) for next behavior prediction, our model defeats several presented baselines. These results suggest that our framework CIM offers a powerful paradigm for 1) presenting an effective measurement for consumer impulsivity. 2) uncovering the correlation between consumer impulsivity and demographic factors and 3) revealing that the introduction of impulsivity is effective in predicting consumer behavior. Nicholas Jing Yuan, Kai Zheng 0001, Defu Lian, Xing Xie 0001, Yong Rui |
UbiComp | 6 |
| 2015 | Query Adaptive Similarity Measure for RGB-D Object RecognitionabstractThis paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object pose and scale changes, as well as intra-class variations, and (2) effectively fusing RGB and depth cues is still an open problem. To address these problems, this paper first proposes a new similarity measure based on dense matching, through which objects in comparison are warped and aligned, to better tolerate variations. Towards RGB and depth fusion, we argue that a constant and golden weight doesn't exist. The two modalities have varying contributions when comparing objects from different categories. To capture such a dynamic characteristic, a group of matchers equipped with various fusion weights is constructed, to explore the responses of dense matching under different fusion configurations. All the response scores are finally merged following a learning-to-combination way, which provides quite good generalization ability in practice. The proposed approach win the best results on several public benchmarks, e.g., achieves 92.7% top-1 test accuracy on the Washington RGB-D object dataset, with a 5.1% improvement over the state-of-the-art. Yanhua Cheng, Rui Cai 0002, Chi Zhang 0069, Zhiwei Li 0006, Xin Zhao 0012, Kaiqi Huang, Yong Rui |
ICCV | 7 |
| 2015 | Relaxing from Vocabulary: Robust Weakly-Supervised Deep Learning for Vocabulary-Free Image TaggingabstractThe development of deep learning has empowered machines with comparable capability of recognizing limited image categories to human beings. However, most existing approaches heavily rely on human-curated training data, which hinders the scalability to large and unlabeled vocabularies in image tagging. In this paper, we propose a weakly-supervised deep learning model which can be trained from the readily available Web images to relax the dependence on human labors and scale up to arbitrary tags (categories). Specifically, based on the assumption that features of true samples in a category tend to be similar and noises tend to be variant, we embed the feature map of the last deep layer into a new affinity representation, and further minimize the discrepancy between the affinity representation and its low-rank approximation. The discrepancy is finally transformed into the objective function to give relevance feedback to back propagation. Experiments show that we can achieve a performance gain of 14.0% in terms of a semantic-based relevance metric in image tagging with 63,043 tags from the WordNet, against the typical deep model trained on the ImageNet 1,000 vocabulary set. Jianlong Fu, Tao Mei 0001, Jinqiao Wang, Hanqing Lu, Yong Rui |
ICCV | 6 |
| 2015 | MeshStereo: A Global Stereo Model with Mesh Alignment Regularization for View InterpolationabstractWe present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D triangles with shared vertices. Lifting the 2D triangulation to 3D naturally generates a corresponding mesh. A technical difficulty is to properly split vertices to multiple copies when they appear at depth discontinuous boundaries. To deal with this problem, we formulate our objective as a two-layer MRF, with the upper layer modeling the splitting properties of the vertices and the lower layer optimizing a region-based stereo matching. Experiments on the Middlebury and the Herodion datasets demonstrate that our model is able to synthesize visually coherent new view angles with high PSNR, as well as outputting high quality disparity maps which rank at the first place on the new challenging high resolution Middlebury 3.0 benchmark. Chi Zhang 0069, Zhiwei Li 0006, Yanhua Cheng, Rui Cai 0002, Hongyang Chao, Yong Rui |
ICCV | 6 |
| 2015 | Content-Aware Collaborative Filtering for Location Recommendation Based on Human Mobility DataabstractLocation recommendation plays an essential role in helping people find places they are likely to enjoy. Though some recent research has studied how to recommend locations with the presence of social network and geographical information, few of them addressed the cold-start problem, specifically, recommending locations for new users. Because the visits to locations are often shared on social networks, rich semantics (e.g., tweets) that reveal a person's interests can be leveraged to tackle this challenge. A typical way is to feed them into traditional explicit-feedback content-aware recommendation methods (e.g., LibFM). As a user's negative preferences are not explicitly observable in most human mobility data, these methods need draw negative samples for better learning performance. However, prior studies have empirically shown that sampling-based methods don't perform as well as a method that considers all unvisited locations as negative but assigns them a lower confidence. To this end, we propose an Implicit-feedback based Content-aware Collaborative Filtering (ICCF) framework to incorporate semantic content and steer clear of negative sampling. For efficient parameter learning, we develop a scalable optimization algorithm, scaling linearly with the data size and the feature size. Furthermore, we offer a good explanation to ICCF, such that the semantic content is actually used to refine user similarity based on mobility. Finally, we evaluate ICCF with a large-scale LBSN dataset where users have profiles and text content. The results show that ICCF outperforms LibFM of the best configuration, and that user profiles and text content are not only effective at improving recommendation but also helpful for coping with the cold-start problem. Defu Lian, Yong Ge 0001, Nicholas Jing Yuan, Xing Xie 0001, Tao Zhou 0001, Yong Rui |
ICDM | 7 |
| 2015 | On the selection of trending image from the webabstractThe recommendation of trending images has become a popular feature used by commercial search engines to attract public attention. By browsing through trending images, search engine users can discover trending events at a glance. However, the selection of trending images is very challenging and remains an open issue. Most existing work is highly dependent on editorial efforts, though some preliminarily identify a few plain features for trending images. In this paper, we investigate a set of perceptual factors that can distinguish trending images from common ones. We propose a set of trending-aware features based on several common criteria, which reflect the characteristics of trending images. We further construct a manually labeled dataset based on a commercial search engine's query log over a two-week timespan. We evaluate our proposed method on this dataset and the results demonstrate its effectiveness. Dongfei Yu, Xinmei Tian 0001, Tao Mei 0001, Yong Rui |
ICME | 4 |
| 2015 | Offline Sketch Parsing via Shapeness Estimation
Changhu Wang, Liqing Zhang 0001, Yong Rui |
IJCAI | 4 |
| 2015 | Regularity and Conformity: Location Prediction Using Heterogeneous Mobility DataabstractMobility prediction enables appealing proactive experiences for location-aware services and offers essential intelligence to business and governments. Recent studies suggest that human mobility is highly regular and predictable. Additionally, social conformity theory indicates that people's movements are influenced by others. However, existing approaches for location prediction fail to organically combine both the regularity and conformity of human mobility in a unified model, and lack the capacity to incorporate heterogeneous mobility datasets to boost prediction performance. To address these challenges, in this paper we propose a hybrid predictive model integrating both the regularity and conformity of human mobility as well as their mutual reinforcement. In addition, we further elevate the predictive power of our model by learning location profiles from heterogeneous mobility datasets based on a gravity model. We evaluate the proposed model using several city-scale mobility datasets including location check-ins, GPS trajectories of taxis, and public transit data. The experimental results validate that our model significantly outperforms state-of-the-art approaches for mobility prediction in terms of multiple metrics such as accuracy and percentile rank. The results also suggest that the predictability of human mobility is time-varying, e.g., the overall predictability is higher on workdays than holidays while predicting users' unvisited locations is more challenging for workdays than holidays. Yingzi Wang, Nicholas Jing Yuan, Defu Lian, Linli Xu 0002, Xing Xie 0001, Enhong Chen, Yong Rui |
KDD | 7 |
| 2015 | EMIF: Towards a Scalable and Effective Indexing Framework for Large Scale Music RetrievalabstractIn this article, we present a novel indexing technique called EMIF (Effective Music Indexing Framework) to facilitate scalable and accurate content based music retrieval. It is designed based on a "classification-and-indexing" principle and consists of two main functionality layers: 1) a novel semantic-sensitive classification to identify input music's category and 2) multiple indexing structures - one local indexing structure corresponds to one semantic category. EMIF's layered architecture not only enables superior search accuracy but also reduces query response time significantly. To evaluate the system, a set of comprehensive experimental studies have been carried out using large test collection and EMIF demonstrates promising performance over state-of-the-art approaches. Jialie Shen 0001, Tao Mei 0001, Dacheng Tao, Xuelong Li 0001, Yong Rui |
ICMR | 5 |
| 2015 | Predicting Smartphone Adoption in Social Networks
Le Wu 0001, Nicholas Jing Yuan, Enhong Chen, Xing Xie 0001, Yong Rui |
PAKDD (1) | 6 |
| 2015 | Resorting Relevance Evidences to Cumulative Citation Recommendation for Knowledge Base Acceleration
Jingang Wang, Lejian Liao, Lerong Ma, Chin-Yew Lin, Yong Rui |
WAIM | 6 |
| 2015 | Tagging Personal Photos with Transfer Deep LearningabstractThe advent of mobile devices and media cloud services has led to the unprecedented growing of personal photo collections. One of the fundamental problems in managing the increasing number of photos is automatic image tagging. Existing research has predominantly focused on tagging general Web images with a well-labelled image database, e.g., ImageNet. However, they can only achieve limited success on personal photos due to the domain gaps between personal photos and Web images. These gaps originate from the differences in semantic distribution and visual appearance. To deal with these challenges, in this paper, we present a novel transfer deep learning approach to tag personal photos. Specifically, to solve the semantic distribution gap, we have designed an ontology consisting of a hierarchical vocabulary tailored for personal photos. This ontology is mined from $10,000$ active users in Flickr with 20 million photos and 2.7 million unique tags. To deal with the visual appearance gap, we discover the intermediate image representations and ontology priors by deep learning with bottom-up and top-down transfers across two domains, where Web images are the source domain and personal photos are the target. Moreover, we present two modes (single and batch-modes) in tagging and find that the batch-mode is highly effective to tag photo collections. We conducted personal photo tagging on 7,000 real personal photos and personal photo search on the MIT-Adobe FiveK photo dataset. The proposed tagging approach is able to achieve a performance gain of $12.8\%$ and $4.5\%$ in terms of [email protected], against the state-of-the-art hand-crafted feature-based and deep learning-based methods, respectively. Jianlong Fu, Tao Mei 0001, Kuiyuan Yang, Hanqing Lu, Yong Rui |
WWW | 5 |
| 2015 | Multi-order visual phrase for scalable partial-duplicate visual search
Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
Multim. Syst. | 5 |
| 2015 | Image Tag Refinement With View-Dependent Concept RepresentationsabstractImage tag refinement is the task of refining initial tags of an image such that the refined tags can better reflect the content of the image and, therefore, can help users better access that image. The quality of tag refinement depends on the quality of concept representations that build a mapping from concepts to visual images. While good progress was made in the past decade on tag refinement, the previous approaches only achieved a limited success due to their limited concept representations. In this paper, we show that the visual appearances of a concept consist of both a generic view and a specific view, and therefore we can comprehensively represent a concept by two components. To ensure a clean concept representation, this representation is learned on clean click-through data, where noises are greatly reduced. In the framework, a coarse-to-fine image tag refinement is proposed, which: (1) first generates an efficient star graph to find candidate tags but missing in the initial tag list of an input image and (2) guided by this view-dependent concept representation, formulates a probabilistic objective function to eliminate irrelevant tags. Extensive experiments on two widely used standard data sets (MIRFlickr-25K and NUS-WIDE-270K) demonstrate the effectiveness of our approach. Jianlong Fu, Jinqiao Wang, Yong Rui, Xin-Jing Wang, Tao Mei 0001, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Learning to Rank Using User Clicks and Visual Features for Image RetrievalabstractThe inconsistency between textual features and visual contents can cause poor image search results. To solve this problem, click features, which are more reliable than textual information in justifying the relevance between a query and clicked images, are adopted in image ranking model. However, the existing ranking model cannot integrate visual features, which are efficient in refining the click-based search results. In this paper, we propose a novel ranking model based on the learning to rank framework. Visual features and click features are simultaneously utilized to obtain the ranking model. Specifically, the proposed approach is based on large margin structured output learning and the visual consistency is integrated with the click features through a hypergraph regularizer term. In accordance with the fast alternating linearization method, we design a novel algorithm to optimize the objective function. This algorithm alternately minimizes two different approximations of the original objective function by keeping one function unchanged and linearizing the other. We conduct experiments on a large-scale dataset collected from the Microsoft Bing image search engine, and the results demonstrate that the proposed learning to rank models based on visual features and user clicks outperforms state-of-the-art algorithms. Jun Yu 0002, Dacheng Tao, Meng Wang 0001, Yong Rui |
IEEE Trans. Cybern. | 4 |
| 2015 | Where2Stand: A Human Position Recommendation System for Souvenir PhotographyabstractPeople often take photographs at tourist sites and these pictures usually have two main elements: a person in the foreground and scenery in the background. This type of “souvenir photo” is one of the most common photos clicked by tourists. Although algorithms that aid a user-photographer in taking a well-composed picture of a scene exist [Ni et al. 2013], few studies have addressed the issue of properly positioning human subjects in photographs. In photography, the common guidelines of composing portrait images exist. However, these rules usually do not consider the background scene. Therefore, in this article, we investigate human-scenery positional relationships and construct a photographic assistance system to optimize the position of human subjects in a given background scene, thereby assisting the user in capturing high-quality souvenir photos. We collect thousands of well-composed portrait photographs to learn human-scenery aesthetic composition rules. In addition, we define a set of negative rules to exclude undesirable compositions. Recommendation results are achieved by combining the first learned positive rule with our proposed negative rules. We implement the proposed system on an Android platform in a smartphone. The system demonstrates its efficacy by producing well-composed souvenir photos. Yinting Wang, Mingli Song, Dacheng Tao, Yong Rui, Jiajun Bu, Ah Chung Tsoi, Shaojie Zhuo, Ping Tan 0002 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Super Fast Event Recognition in Internet VideosabstractTechniques for recognizing high-level events in consumer videos on the Internet have many applications. Systems that produced state-of-the-art recognition performance usually contain modules requiring extensive computation, such as the extraction of the temporal motion trajectories, which cannot be deployed on large-scale datasets. In this paper, we provide a comprehensive study on efficient methods in this area and identify technical options for super fast event recognition in Internet videos. We start from analyzing a multimodal baseline that has produced good performance on popular benchmarks, by systematically evaluating each component in terms of both computational cost and contribution to recognition accuracy. After that, we identify alternative features, classifiers, and fusion strategies that can all be efficiently computed. In addition, we also provide a study on the following interesting question: for event recognition in Internet videos, what is the minimum number of visual and audio frames needed to obtain a comparable accuracy to that of using all the frames? Results on two rigorously designed datasets indicate that similar results can be maintained by using only a small portion of the visual frames. We also find that, different from the visual frames, the soundtracks contain little redundant information and thus sampling is always harmful. Integrating all the findings, our suggested recognition system is 2,350-fold faster than a baseline approach with even higher recognition accuracies. It recognizes 20 classes on a 120-second video sequence in just 1.78 seconds, using a regular desktop computer. Yu-Gang Jiang 0001, Qi Dai 0001, Tao Mei 0001, Yong Rui, Shih-Fu Chang |
IEEE Trans. Multim. | 4 |
| 2015 | Partial-Duplicate Clustering and Visual Pattern Discovery on Web Scale Image DatabaseabstractIn this paper, we study the problem of discovering visual patterns and partial-duplicate images, which is fundamental to visual concept representation and image parsing, but very challenging when the database is extremely large, such as billions of images indexed by a commercial search engine. Although extensive research with sophisticated algorithms has been conducted for either partial-duplicate clustering or visual pattern discovery, most of them can not be easily extended to this scale, since both are clustering problems in nature and require pairwise comparisons. To tackle this computational challenge, we introduce a novel and highly parallelizable framework to discover partial-duplicate images and visual patterns in a unified way in distributed computing systems. We emphasize the nested property of local features, and propose the generalized nested feature (GNF) as a mid-level representation for regions and local patterns. Initial coarse clusters are then discovered by GNFs, upon which$n$-gram GNF is defined to represent co-occurrent visual patterns. After that, efficient merging and refining algorithms are used to get the partial-duplicate clusters, and logical combinations of probabilistic GNF models are leveraged to represent the visual patterns of partially duplicate images. Extensive experiments show the parallelizable property and effectiveness of the algorithms on both partial-duplicate clustering and visual pattern discovery. With 2000 machines, it costs about eight and 400 minutes to process one million and 40 million images respectively, which is quite efficient compared to previous methods. Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010 |
IEEE Trans. Multim. | 4 |
| 2015 | Mining Latent Attributes From Click-Through Logs for Image RecognitionabstractAttribute-based image representation, which represents an image by projecting it into a space spanned by attributes, has attracted increasing attention from both computer vision and multimedia communities for its compactness and potential to bridge the semantic gap. While many works focus on learning attribute models and utilizing them in image recognition and retrieval, few touch on the problem of how to effectively construct a vocabulary of attributes, which is an essential part of effective attribute-based representation. Most existing approaches define the attribute vocabulary by human experts or through existing ontology, which is often limited in coverage of general concept space. In this paper, we propose automatically constructing the attribute vocabulary by mining latent topics from the click-through log of a commercial image search engine. These attributes are referred to as latent topic attributes (LTA), which take advantage of tens of millions of interactions between user submitted queries and images, thereby providing better coverage for the concept space than existing approaches. The mining of latent topics from the click log is formulated as a matrix factorization problem, and further improved by weighted terms-based matrix factorization to address the extreme sparsity of the click-through matrix. Both qualitative results of the mined LTA and quantitative results on the standard image recognition benchmark demonstrate the mined LTA's effectiveness. Yi-Jie Lu, Linjun Yang, Kuiyuan Yang, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2015 | Learning Cross Space Mapping via DNN Using Large Scale Click-Through LogsabstractThe gap between low-level visual signals and high-level semantics has been progressively bridged by continuous development of deep neural network (DNN). With recent progress of DNN, almost all image classification tasks have achieved new records of accuracy. To extend the ability of DNN to image retrieval tasks, we proposed a unified DNN model for image-query similarity calculation by simultaneously modeling image and query in one network. The unified DNN is named the cross space mapping (CSM) model, which contains two parts, a convolutional part and a query-embedding part. The image and query are mapped to a common vector space via these two parts respectively, and image-query similarity is naturally defined as an inner product of their mappings in the space. To ensure good generalization ability of the DNN, we learn weights of the DNN from a large number of click-through logs which consists of 23 million clicked image-query pairs between 1 million images and 11.7 million queries. Both the qualitative results and quantitative results on an image retrieval evaluation task with 1000 queries demonstrate the superiority of the proposed method. Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
IEEE Trans. Multim. | 5 |
| 2015 | Retargeting Semantically-Rich PhotosabstractSemantically-rich photos contain a rich variety of semantic objects (e.g., pedestrians and bicycles). Retargeting these photos is a challenging task since each semantic object has fixed geometric characteristics. Shrinking these objects simultaneously during retargeting is prone to distortion. In this paper, we propose to retarget semantically-rich photos by detecting photo semantics from image tags, which are predicted by a multi-label SVM. The key technique is a generative model termed latent stability discovery (LSD). It can robustly localize various semantic objects in a photo by making use of the predicted noisy image tags. Based on LSD, a feature fusion algorithm is proposed to detect salient regions at both the low-level and high-level. These salient regions are linked into a path sequentially to simulate human visual perception . Finally, we learn the prior distribution of such paths from aesthetically pleasing training photos. The prior enforces the path of a retargeted photo to be maximally similar to those from the training photos. In the experiment, we collect 217 1600 ×1200 photos, each containing over seven salient objects. Comprehensive user studies demonstrate the competitiveness of our method. Meng Wang 0001, Liqiang Nie, Liang Hong 0001, Yong Rui, Qi Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | A Distributed Approach Toward Discriminative Distance Metric LearningabstractDistance metric learning (DML) is successful in discovering intrinsic relations in data. However, most algorithms are computationally demanding when the problem size becomes large. In this paper, we propose a discriminative metric learning algorithm, develop a distributed scheme learning metrics on moderate-sized subsets of data, and aggregate the results into a global solution. The technique leverages the power of parallel computation. The algorithm of the aggregated DML (ADML) scales well with the data size and can be controlled by the partition. We theoretically analyze and provide bounds for the error induced by the distributed treatment. We have conducted experimental evaluation of the ADML, both on specially designed tests and on practical image annotation tasks. Those tests have shown that the ADML achieves the state-of-the-art performance at only a fraction of the cost incurred by most existing methods. Jun Li 0010, Xun Lin, Xiaoguang Rui, Yong Rui, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Learning Word Representation Considering Proximity and AmbiguityabstractDistributed representations of words (aka word embedding) have proven helpful in solving natural language processing (NLP) tasks. Training distributed representations of words with neural networks has lately been a major focus of researchers in the field. Recent work on word embedding, the Continuous Bag-of-Words (CBOW) model and the Continuous Skip-gram (Skip-gram) model, have produced particularly impressive results, significantly speeding up the training process to enable word representation learning from large-scale data. However, both CBOW and Skip-gram do not pay enough attention to word proximity in terms of model or word ambiguity in terms of linguistics. In this paper, we propose Proximity-Ambiguity Sensitive (PAS) models (i.e. PAS CBOW and PAS Skip-gram) to produce high quality distributed representations of words considering both word proximity and ambiguity. From the model perspective, we introduce proximity weights as parameters to be learned in PAS CBOW and used in PAS Skip-gram. By better modeling word proximity, we reveal the strength of pooling-structured neural networks in word representation learning. The proximity-sensitive pooling layer can also be applied to other neural network applications that employ pooling layers. From the linguistics perspective, we train multiple representation vectors per word. Each representation vector corresponds to a particular group of POS tags of the word. By using PAS models, we achieved a 16.9% increase in accuracy over state-of-the-art models. Zaiqing Nie, Yong Yu 0001, Yong Rui |
AAAI | 5 |
| 2014 | Sketch Recognition with Natural Correction and EditingabstractIn this paper, we target at the problem of sketch recognition. We systematically study how to incorporate users' correction and editing into isolated and full sketch recognition. This is a natural and necessary interaction in real systems such as Visio where very similar shapes exist. First, a novel algorithm is proposed to mine the prior shape knowledge for three editing modes. Second, to differentiate visually similar shapes, a novel symbol recognition algorithm is introduced by leveraging the learnt shape knowledge. Then, a novel editing detection algorithm is proposed to facilitate symbol recognition. Furthermore, both of the symbol recognizer and the editing detector are systematically incorporated into the full sketch recognition. Finally, based on the proposed algorithms, a real-time sketch recognition system is built to recognize hand-drawn flowcharts and diagrams with flexible interactions. Extensive experiments show the effectiveness of the proposed algorithms. Changhu Wang, Liqing Zhang 0001, Yong Rui |
AAAI | 4 |
| 2014 | What Visual Attributes Characterize an Object Class?
Jianlong Fu, Jinqiao Wang, Xin-Jing Wang, Yong Rui, Hanqing Lu |
ACCV (1) | 4 |
| 2014 | DNN Flow: DNN Feature Pyramid based Image Matching
Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
BMVC | 5 |
| 2014 | As-Rigid-As-Possible Stereo under Second Order Smoothness Priors
Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui |
ECCV (2) | 5 |
| 2014 | Unsupervised Template Mining for Semantic Category UnderstandingabstractWe propose an unsupervised approach to constructing templates from a large collection of semantic category names, and use the templates as the semantic representation of categories.The main challenge is that many terms have multiple meanings, resulting in a lot of wrong templates.Statistical data and semantic knowledge are extracted from a web corpus to improve template generation.A nonlinear scoring function is proposed and demonstrated to be effective.Experiments show that our approach achieves significantly better results than baseline methods.As an immediate application, we apply the extracted templates to the cleaning of a category collection and see promising results (precision improved from 81% to 89%). Lei Shi 0015, Shuming Shi 0001, Chin-Yew Lin, Yidong Shen, Yong Rui |
EMNLP | 5 |
| 2014 | Large-margin Weakly Supervised Dimensionality ReductionabstractThis paper studies dimensionality reduction in a weakly supervised setting, in which the preference relationship between examples is indicated by weak cues. A novel framework is proposed that integrates two aspects of the large margin principle (angle and distance), which simultaneously encourage angle consistency between preference pairs and maximize the distance between examples in preference pairs. Two specific algorithms are developed: an alternating direction method to learn a linear transformation matrix and a gradient boosting technique to optimize a non-linear transformation directly in the function space. Theoretical analysis demonstrates that the proposed large margin optimization criteria can strengthen and improve the robustness and generalization performance of preference learning algorithms on the obtained low-dimensional subspace. Experimental results on real-world datasets demonstrate the significance of studying dimensionality reduction in the weakly supervised setting and the effectiveness of the proposed framework. Chang Xu 0002, Dacheng Tao, Chao Xu 0006, Yong Rui |
ICML | 4 |
| 2014 | GeoMF: joint geographical modeling and matrix factorization for point-of-interest recommendationabstractPoint-of-Interest (POI) recommendation has become an important means to help people discover attractive locations. However, extreme sparsity of user-POI matrices creates a severe challenge. To cope with this challenge, viewing mobility records on location-based social networks (LBSNs) as implicit feedback for POI recommendation, we first propose to exploit weighted matrix factorization for this task since it usually serves collaborative filtering with implicit feedback better. Besides, researchers have recently discovered a spatial clustering phenomenon in human mobility behavior on the LBSNs, i.e., individual visiting locations tend to cluster together, and also demonstrated its effectiveness in POI recommendation, thus we incorporate it into the factorization model. Particularly, we augment users' and POIs' latent factors in the factorization model with activity area vectors of users and influence area vectors of POIs, respectively. Based on such an augmented model, we not only capture the spatial clustering phenomenon in terms of two-dimensional kernel density estimation, but we also explain why the introduction of such a phenomenon into matrix factorization helps to deal with the challenge from matrix sparsity. We then evaluate the proposed algorithm on a large-scale LBSN dataset. The results indicate that weighted matrix factorization is superior to other forms of factorization models and that incorporating the spatial clustering phenomenon into matrix factorization improves recommendation performance. Defu Lian, Xing Xie 0001, Guangzhong Sun, Enhong Chen, Yong Rui |
KDD | 6 |
| 2014 | CeleBrowser: An example of browsing big data on small deviceabstractIn this demonstration, we demonstrate a mobile-based celebrity video browsing system called CeleBrowser. Using this system, users can interactively switch among four views: people-centric, timeline-centric, month-centric and topic-centric, for browsing celebrity-related hot videos. A peculiarity of the demonstration is to highlight the advantage of multi-perspective information organization and presentation in engaging users for exploratory browsing of large number of Web videos on a device with small screen. Technology-wise the demonstration shows how query logs collected for six months from two vertical search engines are leveraged for mining hot events and videos of celebrities. Song Tan, Chong-Wah Ngo, Yong Rui |
ICMR | 4 |
| 2014 | SmartVisio: Interactive Sketch Recognition with Natural Correction and EditingabstractIn this work, we introduce the SmartVisio system for interactive hand-drawn shape/diagram recognition. Different from existing work, SmartVisio is a real-time sketch recognition system based on Visio, to recognize hand-drawn flowchart/diagram with flexible interactions. This system enables a user to draw shapes or diagrams on the Visio interface, and then the hand-drawn shapes are automatically converted to formal shapes in real-time. To satisfy the interaction needs from common users, we propose an algorithm to detect a user's correction and editing during drawing, and then recognize in real time. We also propose a novel symbol recognition algorithm to better recognize or differentiate some visually similar shapes. By enabling users' natural correction/editing on various shapes, our system makes flowchart/diagram production much more natural and easier. Changhu Wang, Liqing Zhang 0001, Yong Rui |
ACM Multimedia | 4 |
| 2014 | Click-through-based cross-view learning for image searchabstractOne of the fundamental problems in image search is to rank image documents according to a given textual query. Existing search engines highly depend on surrounding texts for ranking images, or leverage the query-image pairs annotated by human labelers to train a series of ranking functions. However, there are two major limitations: 1) the surrounding texts are often noisy or too few to accurately describe the image content, and 2) the human annotations are resourcefully expensive and thus cannot be scaled up. We demonstrate in this paper that the above two fundamental challenges can be mitigated by jointly exploring the cross-view learning and the use of click-through data. The former aims to create a latent subspace with the ability in comparing information from the original incomparable views (i.e., textual and visual views), while the latter explores the largely available and freely accessible click-through data (i.e., ``crowdsourced" human intelligence) for understanding query. Specifically, we propose a novel cross-view learning method for image search, named Click-through-based Cross-view Learning (CCL), by jointly minimizing the distance between the mappings of query and image in the latent subspace and preserving the inherent structure in each original space. On a large-scale click-based image dataset, CCL achieves the improvement over Support Vector Machine-based method by 4.0\% in terms of relevance, while reducing the feature dimension by several orders of magnitude (e.g., from thousands to tens). Moreover, the experiments also demonstrate the superior performance of CCL to several state-of-the-art subspace learning techniques. Yingwei Pan, Ting Yao 0003, Tao Mei 0001, Houqiang Li, Chong-Wah Ngo, Yong Rui |
SIGIR | 6 |
| 2014 | Learning to personalize trending image search suggestionabstractTrending search suggestion is leading a new paradigm of image search, where user's exploratory search experience is facilitated with the automatic suggestion of trending queries. Existing image search engines, however, only provide general suggestions and hence cannot capture user's personal interest. In this paper, we move one step forward to investigate personalized suggestion of trending image searches according to users' search behaviors. To this end, we propose a learning-based framework including two novel components. The first component, i.e., trending-aware weight-regularized matrix factorization (TA-WRMF), is able to suggest personalized trending search queries by learning user preference from many users as well as auxiliary common searches. The second component associates the most representative and trending image with each suggested query. The personalized suggestion of image search consists of a trending textual query and its associated trending image. The combined textual-visual queries not only are trending (bursty) and personalized to user's search preference, but also provide the compelling visual aspect of these queries. We evaluate our proposed learning-based framework on a large-scale search logs with 21 million users and 41 million queries in two weeks from a commercial image search engine. The evaluations demonstrate that our system achieve about 50% gain compared with state-of-the-art in terms of query prediction accuracy. Chun-Che Wu, Tao Mei 0001, Winston H. Hsu, Yong Rui |
SIGIR | 4 |
| 2014 | Preface: Internet multimedia computing and service
Shuqiang Jiang, Changsheng Xu, Yong Rui, Alberto Del Bimbo, Hongxun Yao |
Multim. Tools Appl. | 3 |
| 2014 | High-Order Distance-Based Multiview Stochastic Learning in Image ClassificationabstractHow do we find all images in a larger set of images which have a specific content? Or estimate the position of a specific object relative to the camera? Image classification methods, like support vector machine (supervised) and transductive support vector machine (semi-supervised), are invaluable tools for the applications of content-based image retrieval, pose estimation, and optical character recognition. However, these methods only can handle the images represented by single feature. In many cases, different features (or multiview data) can be obtained, and how to efficiently utilize them is a challenge. It is inappropriate for the traditionally concatenating schema to link features of different views into a long vector. The reason is each view has its specific statistical property and physical interpretation. In this paper, we propose a high-order distance-based multiview stochastic learning (HD-MSL) method for image classification. HD-MSL effectively combines varied features into a unified representation and integrates the labeling information based on a probabilistic framework. In comparison with the existing strategies, our approach adopts the high-order distance obtained from the hypergraph to replace pairwise distance in estimating the probability matrix of data distribution. In addition, the proposed approach can automatically learn a combination coefficient for each view, which plays an important role in utilizing the complementary information of multiview data. An alternative optimization is designed to solve the objective functions of HD-MSL and obtain different views on coefficients and classification scores simultaneously. Experiments on two real world datasets demonstrate the effectiveness of HD-MSL in image classification. Jun Yu 0002, Yong Rui, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2014 | Click Prediction for Web Image Reranking Using Multimodal Sparse CodingabstractImage reranking is effective for improving the performance of a text-based image search. However, existing reranking algorithms are limited for two main reasons: 1) the textual meta-data associated with images is often mismatched with their actual visual content and 2) the extracted visual features do not accurately describe the semantic similarities between images. Recently, user click information has been used in image reranking, because clicks have been shown to more accurately describe the relevance of retrieved images to search queries. However, a critical problem for click-based methods is the lack of click data, since only a small number of web images have actually been clicked on by users. Therefore, we aim to solve this problem by predicting image clicks. We propose a multimodal hypergraph learning-based sparse coding method for image click prediction, and apply the obtained click data to the reranking of images. We adopt a hypergraph to build a group of manifolds, which explore the complementarity of different features through a group of weights. Unlike a graph that has an edge between two vertices, a hyperedge in a hypergraph connects a set of vertices, and helps preserve the local smoothness of the constructed sparse codes. An alternating optimization procedure is then performed, and the weights of different modalities and the sparse codes are simultaneously obtained. Finally, a voting strategy is used to describe the predicted click as a binary event (click or no click), from the images' corresponding sparse codes. Thorough empirical studies on a large-scale database including nearly 330 K images demonstrate the effectiveness of our approach for click prediction when compared with several other methods. Additional image reranking experiments on real-world data show the use of click prediction is beneficial to improving the performance of prominent graph-based image reranking algorithms. Jun Yu 0002, Yong Rui, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2014 | Cascade Category-Aware Visual SearchabstractIncorporating image classification into image retrieval system brings many attractive advantages. For instance, the search space can be narrowed down by rejecting images in irrelevant categories of the query. The retrieved images can be more consistent in semantics by indexing and returning images in the relevant categories together. However, due to their different goals on recognition accuracy and retrieval scalability, it is hard to efficiently incorporate most image classification works into large-scale image search. To study this problem, we propose cascade category-aware visual search, which utilizes weak category clue to achieve better retrieval accuracy, efficiency, and memory consumption. To capture the category and visual clues of an image, we first learn category-visual words, which are discriminative and repeatable local features labeled with categories. By identifying category-visual words in database images, we are able to discard noisy local features and extract image visual and category clues, which are hence recorded in a hierarchical index structure. Our retrieval system narrows down the search space by: 1) filtering the noisy local features in query; 2) rejecting irrelevant categories in database; and 3) preforming discriminative visual search in relevant categories. The proposed algorithm is tested on object search, landmark search, and large-scale similar image search on the large-scale LSVRC10 data set. Although the category clue introduced is weak, our algorithm still shows substantial advantages in retrieval accuracy, efficiency, and memory consumption than the state-of-the-art. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
IEEE Trans. Image Process. | 5 |
| 2014 | USB: Ultrashort Binary Descriptor for Fast Visual Matching and RetrievalabstractCurrently, many local descriptors have been proposed to tackle a basic issue in computer vision: duplicate visual content matching. These descriptors either are represented as high-dimensional vectors relatively expensive to extract and compare or are binary codes limited in robustness. Bag-of-visual words (BoWs) model compresses local features into a compact representation that allows for fast matching and scalable indexing. However, the codebook training, high-dimensional feature extraction, and quantization significantly degrade the flexibility and efficiency of BoWs model. In this paper, we study an alternative to current local descriptors and BoWs model by extracting the ultrashort binary descriptor (USB) and a compact auxiliary spatial feature from each keypoint detected in images. A typical USB is a 24-bit binary descriptor, hence it directly quantizes visual clues of image keypoints to about 16 million unique IDs. USB allows fast image matching and indexing and avoids the expensive codebook training and feature quantization in BoWs model. The spatial feature complementarily captures the spatial configuration in neighbor region of each keypoint, hence is used to filter mismatched USBs in a cascade verification. In image matching task, USB shows promising accuracy and nearly one-order faster speed than SIFT. We also test USB in retrieval tasks on UKbench, Oxford5K, and 1.2 million distractor images. Comparisons with recent retrieval methods manifest the competitive accuracy, memory consumption, and significantly better efficiency of our approach. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
IEEE Trans. Image Process. | 5 |
| 2014 | Topic-Sensitive Influencer Mining in Interest-Based Social Media Networks via Hypergraph LearningabstractSocial media is emerging as a new mainstream means of interacting around online media. Social influence mining in social networks is therefore of critical importance in real-world applications such as friend suggestion and photo recommendation. Social media is inherently multimodal, including rich types of user contributed content and social link information. Most of the existing research suffers from two limitations: 1) only utilizing the textual information, and/or 2) only analyzing the generic influence but ignoring the more important topic-level influence. To address these limitations, in this paper we develop a novel Topic-Sensitive Influencer Mining (TSIM) framework in interest-based social media networks. Specifically, we take Flickr as the study platform. People in Flickr interact with each other through images. TSIM aims to find topical influential users and images. The influence estimation is determined with a hypergraph learning approach. In the hypergraph, the vertices represent users and images, and the hyperedges are utilized to capture multi-type relations including visual-textual content relations among images, and social links between users and images. Algorithmwise, TSIM first learns the topic distribution by leveraging user-contributed images, and then infers the influence strength under different topics for each node in the hypergraph. Extensive experiments on a real-world dataset of more than 50 K images and 70 K comment/favorite links from Flickr have demonstrated the effectiveness of our proposed framework. In addition, we also report promising results of friend suggestion and photo recommendation via TSIM on the same dataset. Quan Fang, Changsheng Xu, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2014 | Exploiting Click Constraints and Multi-view Features for Image Re-rankingabstractImage re-ranking is effective in improving performance of text-based image searches. However, improvements from existing re-ranking algorithms are limited by two factors: one is that the associated textual information of images often mismatches their actual visual contents; the other is that a visual's features cannot accurately describe the semantic similarities between images. In this paper, we adopt click data to bridge the semantic gap. We propose a novel multi-view hypergraph-based learning (MHL) method that adaptively integrates click data with varied visual features. In particular, MHL considers pairwise discriminative constraints from click data to maximally distinguish images with high click counts from images with no click counts, and a semantic manifold is constructed. It then adopts hypergraph learning to build multiple manifolds from varied visual features. Finally, MHL integrates the semantic manifold with visual manifolds through an iterative optimization procedure. The weights of different manifolds and the re-ranking score are simultaneously obtained after using this optimization strategy. We conduct experiments on real world datasets and the results demonstrate that MHL outperforms state-of-the-art image re-ranking methods. Jun Yu 0002, Yong Rui |
IEEE Trans. Multim. | 2 |
| 2014 | Bilateral Correspondence Model for Words-and-Pictures Association in Multimedia-Rich MicroblogsabstractNowadays, the amount of multimedia contents in microblogs is growing significantly. More than 20% of microblogs link to a picture or video in certain large systems. The rich semantics in microblogs provides an opportunity to endow images with higher-level semantics beyond object labels. However, this raises new challenges for understanding the association between multimodal multimedia contents in multimedia-rich microblogs. Disobeying the fundamental assumptions of traditional annotation, tagging, and retrieval systems, pictures and words in multimedia-rich microblogs are loosely associated and a correspondence between pictures and words cannot be established. To address the aforementioned challenges, we present the first study analyzing and modeling the associations between multimodal contents in microblog streams, aiming to discover multimodal topics from microblogs by establishing correspondences between pictures and words in microblogs. We first use a data-driven approach to analyze the new characteristics of the words, pictures, and their association types in microblogs. We then propose a novel generative model called the Bilateral Correspondence Latent Dirichlet Allocation (BC-LDA) model. Our BC-LDA model can assign flexible associations between pictures and words and is able to not only allow picture-word co-occurrence with bilateral directions, but also single modal association. This flexible association can best fit the data distribution, so that the model can discover various types of joint topics and generate pictures and words with the topics accordingly. We evaluate this model extensively on a large-scale real multimedia-rich microblogs dataset. We demonstrate the advantages of the proposed model in several application scenarios, including image tagging, text illustration, and topic discovery. The experimental results demonstrate that our proposed model can significantly and consistently outperform traditional approaches. Peng Cui 0001, Lexing Xie, Wenwu Zhu 0001, Yong Rui, Shiqiang Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2014 | Up-Fusion: An Evolving Multimedia Fusion MethodabstractThe amount of multimedia data on the Internet has increased exponentially in the past few decades and this trend is likely to continue. Multimedia content inherently has multiple information sources, therefore effective fusion methods are critical for data analysis and understanding. So far, most of the existing fusion methods are static with respect to time, making it difficult for them to handle the evolving multimedia content. To address this issue, in recent years, several evolving fusion methods were proposed, however, their requirements are difficult to meet, making them useful only in limited applications. In this article, we propose a novel evolving fusion method based on the online portfolio selection theory. The proposed method takes into account the correlation among different information sources and evolves the fusion model when new multimedia data is added. It performs effectively on both crisp and soft decisions without requiring additional context information. Extensive experiments on concept detection and human detection tasks over the TRECVID dataset and surveillance data have been conducted and significantly better performance has been obtained. Xiangyu Wang 0002, Yong Rui, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Efficient 2D-to-3D Correspondence Filtering for Scalable 3D Object Recognitionabstract3D model-based object recognition has been a noticeable research trend in recent years. Common methods find 2D-to-3D correspondences and make recognition decisions by pose estimation, whose efficiency usually suffers from noisy correspondences caused by the increasing number of target objects. To overcome this scalability bottleneck, we propose an efficient 2D-to-3D correspondence filtering approach, which combines a light-weight neighborhood-based step with a finer-grained pairwise step to remove spurious correspondences based on 2D/3D geometric cues. On a dataset of 300 3D objects, our solution achieves ~10 times speed improvement over the baseline, with a comparable recognition accuracy. A parallel implementation on a quad-core CPU can run at ~3fps for 1280×720 images. Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001, Yong Rui |
CVPR | 7 |
| 2013 | Multimedia LEGO: Learning Structured Model by Probabilistic Logic Ontology TreeabstractRecent advances in Multimedia research have generated a large collection of concept models, e.g., LSCOM and Media mill 101, which become accessible to other researchers. While most current research effort still focuses on building new concepts from scratch, little effort has been made on constructing new concepts upon the existing models already in the warehouse. To address this issue, we develop a new framework in this paper, termed LEGO, to seamlessly integrate both the new target training examples and the existing primitive concept models. LEGO treats the primitive concept models as a lego toy to potentially construct an unlimited vocabulary of new concepts. Specifically, LEGO first formulates the logic operations to be the lego connectors to combine existing concept models hierarchically in probabilistic logic ontology trees. LEGO then simultaneously incorporates new target training information to efficiently disambiguate the underlying logic tree and correct the error propagation. We present extensive experimental results on a large vehicle domain data set from Image Net, and demonstrate significantly superior performance over existing state-of-the-art approaches which build new concept models from scratch. Shiyu Chang, Guo-Jun Qi, Jinhui Tang 0001, Qi Tian 0001, Yong Rui, Thomas S. Huang |
ICDM | 5 |
| 2013 | Clickage: towards bridging semantic and intent gaps via mining click logs of search enginesabstractThe semantic gap between low-level visual features and high-level semantics has been investigated for decades but still remains a big challenge in multimedia. When "search" became one of the most frequently used applications, "intent gap", the gap between query expressions and users' search intents, emerged. Researchers have been focusing on three approaches to bridge the semantic and intent gaps: 1) developing more representative features, 2) exploiting better learning approaches or statistical models to represent the semantics, and 3) collecting more training data with better quality. However, it remains a challenge to close the gaps. In this paper, we argue that the massive amount of click data from commercial search engines provides a data set that is unique in the bridging of the semantic and intent gap. Search engines generate millions of click data (a.k.a. image-query pairs), which provide almost "unlimited" yet strong connections between semantics and images, as well as connections between users' intents and queries. To study the intrinsic properties of click data and to investigate how to effectively leverage this huge amount of data to bridge semantic and intent gap is a promising direction to advance multimedia research. In the past, the primary obstacle is that there is no such dataset available to the public research community. This changes as Microsoft has released a new large-scale real-world image click data to public. This paper presents preliminary studies on the power of large-scale click data with a variety of experiments, such as building large-scale concept detectors, tag processing, search, definitive tag detection, intent analysis, etc., with the goal to inspire deeper researches based on this dataset. Xian-Sheng Hua 0001, Linjun Yang, Jingdong Wang 0001, Jing Wang 0068, Kuansan Wang, Yong Rui, Jin Li 0001 |
ACM Multimedia | 7 |
| 2013 | Cross-media semantic representation via bi-directional learning to rankabstractIn multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. Existing approaches take either one-to-one paired data or uni-directional ranking examples (i.e., utilizing only text-query-image ranking examples or image-query-text ranking examples) as training examples, which do not make full use of bi-directional ranking examples (bi-directional ranking means that both text-query-image and image-query-text ranking examples are utilized in the training period) to achieve a better performance. In this paper, we consider learning a cross-media representation model from the perspective of optimizing a listwise ranking problem while taking advantage of bi-directional ranking examples. We propose a general cross-media ranking algorithm to optimize the bi-directional listwise ranking loss with a latent space embedding, which we call Bi-directional Cross-Media Semantic Representation Model (Bi-CMSRM). The latent space embedding is discriminatively learned by the structural large margin learning for optimization with certain ranking criteria (mean average precision in this paper) directly. We evaluate Bi-CMSRM on the Wikipedia and NUS-WIDE datasets and show that the utilization of the bi-directional ranking examples achieves a much better performance than only using the uni-directional ranking examples. Fei Wu 0001, Zhongfei Zhang, Shuicheng Yan, Yong Rui, Yueting Zhuang |
ACM Multimedia | 5 |
| 2013 | Pairwise constraints based multiview features fusion for scene classification
Jun Yu 0002, Dacheng Tao, Yong Rui, Jun Cheng 0002 |
Pattern Recognit. | 3 |
| 2013 | Hierarchical affective content analysis in arousal and valence dimensions
Min Xu 0001, Changsheng Xu, Xiangjian He, Jesse S. Jin, Suhuai Luo, Yong Rui |
Signal Process. | 6 |
| 2013 | View-Based Discriminative Probabilistic Modeling for 3D Object Retrieval and RecognitionabstractIn view-based 3D object retrieval and recognition, each object is described by multiple views. A central problem is how to estimate the distance between two objects. Most conventional methods integrate the distances of view pairs across two objects as an estimation of their distance. In this paper, we propose a discriminative probabilistic object modeling approach. It builds probabilistic models for each object based on the distribution of its views, and the distance between two objects is defined as the upper bound of the Kullback-Leibler divergence of the corresponding probabilistic models. 3D object retrieval and recognition is accomplished based on the distance measures. We first learn models for each object by the adaptation from a set of global models with a maximum likelihood principle. A further adaption step is then performed to enhance the discriminative ability of the models. We conduct experiments on the ETH 3D object dataset, the National Taiwan University 3D model dataset, and the Princeton Shape Benchmark. We compare our approach with different methods, and experimental results demonstrate the superiority of our approach. Meng Wang 0001, Yue Gao 0002, Ke Lu 0002, Yong Rui |
IEEE Trans. Image Process. | 4 |
| 2013 | Image search - from thousands to billions in 20 yearsabstractThis article presents a comprehensive review and analysis on image search in the past 20 years, emphasizing the challenges and opportunities brought by the astonishing increase of dataset scales from thousands to billions in the same time period, which was witnessed first-hand by the authors as active participants in this research area. Starting with a retrospective review of three stages of image search in the history, the article highlights major breakthroughs around the year 2000 in image search features, indexing methods, and commercial systems, which marked the transition from stage two to stage three. Subsequent sections describe the image search research from four important aspects: system framework, feature extraction and image representation, indexing, and big data's potential. Based on the review, the concluding section discusses open research challenges and suggests future research directions in effective visual representation, image knowledge base construction, implicit user feedback and crowdsourcing, mobile image search, and creative multimedia interfaces. Lei Zhang 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Large-scale multilabel propagation based on efficient sparse graph constructionabstractWith the popularity of photo-sharing websites, the number of web images has exploded into unseen magnitude. Annotating such large-scale data will cost huge amount of human resources and is thus unaffordable. Motivated by this challenging problem, we propose a novel sparse graph based multilabel propagation (SGMP) scheme for super large scale datasets. Both the efficacy and accuracy of the image annotation are further investigated under different graph construction strategies, where Gaussian noise and non-Gaussian sparse noise are simultaneously considered in the formulations of these strategies. Our proposed approach outperforms the state-of-the-art algorithms by focusing on: (1) For large-scale graph construction, a simple yet efficient LSH (Locality Sensitive Hashing)-based sparse graph construction scheme is proposed to speed up the construction. We perform the multilabel propagation on this hashing-based graph construction, which is derived with LSH approach followed by sparse graph construction within the individual hashing buckets; (2) To further improve the accuracy, we propose a novel sparsity induced scalable graph construction scheme, which is based on a general sparse optimization framework. Sparsity essentially implies a very strong prior: for large scale optimization, the values of most variables shall be zeros when the solution reaches the optimum. By utilizing this prior, the solutions of large-scale sparse optimization problems can be derived by solving a series of much smaller scale subproblems; (3) For multilabel propagation, different from the traditional algorithms that propagate over individual label independently, our proposed propagation first encodes the label information of an image as a unit label confidence vector and naturally imposes inter-label constraints and manipulates labels interactively. Then, the entire propagation problem is formulated on the concept of Kullback-Leibler divergence defined on probabilistic distributions, which guides the propagation of the supervision information. Extensive experiments on the benchmark dataset NUS-WIDE with 270k images and its lite version NUS-WIDE-LITE with 56k images well demonstrate the effectiveness and scalability of the proposed multi-label propagation scheme. Yadong Mu, Hairong Liu, Shuicheng Yan, Yong Rui, Tat-Seng Chua |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2013 | Towards decrypting attractiveness via multi-modality cuesabstractDecrypting the secret of beauty or attractiveness has been the pursuit of artists and philosophers for centuries. To date, the computational model for attractiveness estimation has been actively explored in computer vision and multimedia community, yet with the focus mainly on facial features. In this article, we conduct a comprehensive study on female attractiveness conveyed by single/multiple modalities of cues, that is, face, dressing and/or voice, and aim to discover how different modalities individually and collectively affect the human sense of beauty. To extensively investigate the problem, we collect the Multi-Modality Beauty (M2B) dataset, which is annotated with attractiveness levels converted from manualk-wise ratings and semantic attributes of different modalities. Inspired by the common consensus that middle-level attribute prediction can assist higher-level computer vision tasks, we manually labeled many attributes for each modality. Next, a tri-layer Dual-supervised Feature-Attribute-Task (DFAT) network is proposed to jointly learn the attribute model and attractiveness model of single/multiple modalities. To remedy possible loss of information caused by incomplete manual attributes, we also propose a novel Latent Dual-supervised Feature-Attribute-Task (LDFAT) network, where latent attributes are combined with manual attributes to contribute to the final attractiveness estimation. The extensive experimental evaluations on the collected M2B dataset well demonstrate the effectiveness of the proposed DFAT and LDFAT networks for female attractiveness prediction. Tam V. Nguyen 0002, Si Liu 0001, Bingbing Ni, Yong Rui, Shuicheng Yan |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2012 | Sense beauty via face, dressing, and/or voiceabstractDiscovering the secret of beauty has been the pursuit of artists and philosophers for centuries. Nowadays, the computational model for beauty estimation has been actively explored in computer science community, yet with the focus mainly on facial features. In this work, we perform a comprehensive study of female attractiveness conveyed by single/multiple modalities of cues, i.e., face, dressing and/or voice, and aim to uncover how different modalities individually and collectively affect the human sense of beauty. To this end, we collect the first Multi-Modality Beauty (M2B) dataset in the world for female attractiveness study, which is thoroughly annotated with attractiveness levels converted from manual k-wise ratings and semantic attributes of different modalities. A novel Dual-supervised Feature-Attribute-Task (DFAT) network is proposed to jointly learn the beauty estimation models of single/multiple modalities as well as the attribute estimation models. The DFAT network differentiates itself by its supervision in both attribute and task layers. Several interesting beauty-sense observations over single/multiple modalities are reported, and the extensive experimental evaluations on the collected M2B dataset well demonstrate the effectiveness of the proposed DFAT network for female attractiveness estimation. Tam V. Nguyen 0002, Si Liu 0001, Bingbing Ni, Yong Rui, Shuicheng Yan |
ACM Multimedia | 5 |
| 2012 | Towards indexing representative images on the webabstractEven after 20 years of research on real-world image retrieval, there is still a big gap between what search engines can provide and what users expect to see. To bridge this gap, we present an image knowledge base, ImageKB, a graph representation of structured entities, categories, and representative images, as a new basis for practical image indexing and search. ImageKB is automatically constructed via a both bottom-up and top-down, scalable approach that efficiently matches 2 billion web images onto an ontology with millions of nodes. Our approach consists of identifying duplicate image clusters from billions of images, obtaining a candidate set of entities and their images, discovering definitive texts to represent an image and identifying representative images for an entity. To date, ImageKB contains 235.3M representative images corresponding to 0.52M entities, much larger than the state-of-the-art alternative ImageNet that contains 14.2M images for 0.02M synsets. Compared to existing image databases, ImageKB reflects the distributions of both images on the web and users' interests, contains rich semantic descriptions for images and entities, and can be widely used for both text to image search and image to text understanding. Xin-Jing Wang, Zheng Xu 0002, Lei Zhang 0001, Ce Liu 0001, Yong Rui |
ACM Multimedia | 5 |
| 2012 | Annotating web images using NOVA: NOn-conVex group spArsityabstractAs image feature vector is large, selecting the right features plays a fundamental role in Web image annotation. Most existing approaches are either based on individual feature selection, which leads to local optima, or using a convex penalty, which leads to inconsistency. To address these difficulties, in this paper we propose a new sparsity-based approach NOVA (NOn-conVex group spArsity). To the best of our knowledge, NOVA is the first to introduce non-convex penalty for group selection in high-dimensional heterogeneous features space. Because it is a group-sparsity approach, it approximately reaches global optima. Because it uses non-convex penalty, it achieves the consistency. We demonstrate the superior performance of NOVA via three means. First, we present theoretical proof that NOVA is consistent, satisfying un-biasness, sparsity and continuity. Second, we show NOVA converges to the true underlying model by using a ground-truth-available generative-model simulation. Third, we report extensive experimental results on three diverse and widely-used data sets Kodak, MSRA-MM 2.0, and NUS-WIDE. We also compare NOVA against the state-of-the-art approaches, and report superior experimental results. Fei Wu 0001, Yong Rui, Shuicheng Yan, Yueting Zhuang |
ACM Multimedia | 3 |
| 2012 | Location Discriminative Vocabulary Coding for Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Junsong Yuan 0001, Yong Rui, Wen Gao 0001 |
Int. J. Comput. Vis. | 6 |
| 2012 | Sparse transfer learning for interactive video search rerankingabstractVisual reranking is effective to improve the performance of the text-based video search. However, existing reranking algorithms can only achieve limited improvement because of the well-known semantic gap between low-level visual features and high-level semantic concepts. In this article, we adopt interactive video search reranking to bridge the semantic gap by introducing user's labeling effort. We propose a novel dimension reduction tool, termed sparse transfer learning (STL), to effectively and efficiently encode user's labeling information. STL is particularly designed for interactive video search reranking. Technically, it (a) considers the pair-wise discriminative information to maximally separate labeled query relevant samples from labeled query irrelevant ones, (b) achieves a sparse representation for the subspace to encodes user's intention by applying the elastic net penalty, and (c) propagates user's labeling information from labeled samples to unlabeled samples by using the data distribution knowledge. We conducted extensive experiments on the TRECVID 2005, 2006 and 2007 benchmark datasets and compared STL with popular dimension reduction algorithms. We report superior performance by using the proposed STL-based interactive video search reranking. Xinmei Tian 0001, Dacheng Tao, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Cross-Domain Human Action RecognitionabstractConventional human action recognition algorithms cannot work well when the amount of training videos is insufficient. We solve this problem by proposing a transfer topic model (TTM), which utilizes information extracted from videos in the auxiliary domain to assist recognition tasks in the target domain. The TTM is well characterized by two aspects: 1) it uses the bag-of-words model trained from the auxiliary domain to represent videos in the target domain; and 2) it assumes each human action is a mixture of a set of topics and uses the topics learned from the auxiliary domain to regularize the topic estimation in the target domain, wherein the regularization is the summation of Kullback-Leibler divergences between topic pairs of the two domains. The utilization of the auxiliary domain knowledge improves the generalization ability of the learned topic model. Experiments on Weizmann and KTH human action databases suggest the effectiveness of the proposed TTM for cross-domain human action recognition. Wei Bian 0003, Dacheng Tao, Yong Rui |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Towards cross-category knowledge propagation for learning visual conceptsabstractIn recent years, knowledge transfer algorithms have become one of most the active research areas in learning visual concepts. Most of the existing learning algorithms focuses on leveraging the knowledge transfer process which is specific to a given category. However, in many cases, such a process may not be very effective when a particular target category has very few samples. In such cases, it is interesting to examine, whether it is feasible to use cross-category knowledge for improving the learning process by exploring the knowledge in correlated categories. Such a task can be quite challenging due to variations in semantic similarities and differences between categories, which could either help or hinder the cross-category learning process. In order to address this challenge, we develop a cross-category label propagation algorithm, which can directly propagate the inter-category knowledge at instance level between the source and the target categories. Furthermore, this algorithm can automatically detect conditions under which the transfer process can be detrimental to the learning process. This provides us a way to know when the transfer of cross-category knowledge is both useful and desirable. We present experimental results on real image and video data sets in order to demonstrate the effectiveness of our approach. Guo-Jun Qi, Charu C. Aggarwal, Yong Rui, Qi Tian 0001, Shiyu Chang, Thomas S. Huang |
CVPR | 3 |
| 2011 | Towards multi-semantic image annotation with graph regularized exclusive group lassoabstractTo bridge the semantic gap between low level feature and human perception, most of the existing algorithms aim mainly at annotating images with concepts coming from only one semantic space, e.g. cognitive or affective. The naive combination of the outputs from these spaces will implicitly force the conditional independence and ignore the correlations among the spaces. In this paper, to exploit the comprehensive semantic of images, we propose a general framework for harmoniously integrating the above multiple semantics, and investigating the problem of learning to annotate images with training images labeled in two or more correlated semantic spaces, such as fascinating nighttime, or exciting cat. This kind of semantic annotation is more oriented to real world search scenario. Our proposed approach outperforms the baseline algorithms by making the following contributions. 1) Unlike previous methods that annotate images within only one semantic space, our proposed multi-semantic annotation associates each image with labels from multiple semantic spaces. 2) We develop a multi-task linear discriminative model to learn a linear mapping from features to labels. The tasks are correlated by imposing the exclusive group lasso regularization for competitive feature selection, and the graph Laplacian regularization to deal with insufficient training sample issue. 3) A Nesterov-type smoothing approximation algorithm is presented for efficient optimization of our model. Extensive experiments on NUS-WIDEEmotive dataset (56k images) with 8×81 emotive cognitive concepts and Object&Scene datasets from NUS-WIDE well validate the effectiveness of the proposed approach. Xiao-Tong Yuan, Shuicheng Yan, Jinhui Tang 0001, Yong Rui, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2011 | Towards low bit rate mobile visual search with multiple-channel codingabstractIn this paper, we propose a multiple-channel coding scheme to extract compact visual descriptors for low bit rate mobile visual search. Different from previous visual search scenarios that send the query image, we make use of the ever growing mobile computational capability to directly extract compact visual descriptors at the mobile end. Meanwhile, stepping forward from the state-of-the-art compact descriptor extractions, we exploit the rich contextual cues at the mobile end (such as GPS tags for mobile visual search and 2D barcodes or RFID tags for mobile product search), together with the visual statistics at the reference database, to learn multiple coding channels. Therefore, we describe the query with one of many forms of high-dimensional visual signature, which is subsequently mapped to one or more channels and compressed. The compression function within each channel is learnt based on a novel robust PCA scheme, with specific consideration to preserve the retrieval ranking capability of the original signature. We have deployed our scheme on both iPhone4 and HTC DESIRE 7 to search ten million landmark images in a low bit rate setting. Quantitative comparisons to the state-of-the-arts demonstrate our significant advantages in descriptor compactness (with orders of magnitudes improvement) and retrieval mAP in mobile landmark, product, and CD/book cover search. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Yong Rui, Shih-Fu Chang, Wen Gao 0001 |
ACM Multimedia | 5 |
| 2011 | Up-fusion: an evolving multimedia decision fusion methodabstractThe amount of multimedia data available on the Internet has increased exponentially in the past few decades and is likely to keep on increasing. Given that a multimedia system has multiple information sources, fusion methods are critical for its analysis and understanding. However, most of the traditional fusion methods are static with respect to time. To address this, in recent years, several evolving fusion methods have been proposed. However, they can only be used in limited scenarios. For example, the context aware fusion methods need the context information to update the fusion model, but the context may not always be available in many applications. In this paper, a new evolving fusion method is proposed based on the online portfolio selection theory. The proposed method takes the correlation among different information sources into account, and evolves the fusion model when new multimedia data is added. It can deal with either crisp or soft decisions without requiring additional context information. Extensive experiments on concept detection task over TRECVID dataset have been conducted, and promising results have been obtained. Xiangyu Wang 0002, Yong Rui, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2010 | Video based 3D reconstruction using spatio-temporal attention analysisabstract3D reconstruction has been widely used in many important applications. While extensive research has been done in 3D reconstruction, several key issues are still open and the precision of the recovered regions is still far from satisfaction. In this paper, we propose a novel approach to selecting regions of interest in video frames by analyzing multiple spatio-temporal characteristics and reconstructing 3D objects based on the selected regions. Firstly, the static, location and motion attention are extracted from video frames to generate saliency maps. Then, all the video frames are clustered and a candidate set of key frames is extracted based on the saliency maps, where the key frames are extracted according to the constraints in terms of geometry and visibility. Finally, the 3D structure of the attention region is recovered using the selected key frames and the generated saliency maps. The experiments on real-world indoor and outdoor scenes demonstrate that the proposed approach is both more accurate (better attention regions) and computationally more efficient. Xian Xiao, Changsheng Xu, Yong Rui |
ICME | 3 |
| 2010 | Unified tag analysis with multi-edge graphabstractImage tags have become a key intermediate vehicle to organize, index and search the massive online image repositories. Extensive research has been conducted on different yet related tag analysis tasks, e.g., tag refinement, tag-to-region assignment, and automatic tagging. In this paper, we propose a new concept of multi-edge graph, through which a unified solution is derived for the different tag analysis tasks. Specifically, each vertex of the graph is first characterized by a unique image. Then each image is encoded as a region bag with multiple image segmentations, and the thresholding of the pairwise similarities between regions naturally constructs the multiple edges between each vertex pair. The unified tag analysis is then generally described as the tag propagation between a vertex and its edges, as well as between all edges cross the entire image repository. We develop a core vertex-vs-edge tag equation unique for multi-edge graph to unify the image/vertex tag(s) and region-pair/edge tag(s). Finally, unified tag analysis is formulated as a constrained optimization problem, where the objective function characterizing the cross-patch tag consistency is constrained by the core equations for all vertex pairs, and the cutting plane method is used for efficient optimization. Extensive experiments on various tag analysis tasks over three widely used benchmark datasets validate the effectiveness of our proposed unified solution. Dong Liu 0001, Shuicheng Yan, Yong Rui, HongJiang Zhang |
ACM Multimedia | 3 |
| 2010 | Image Classification With Kernelized Spatial-ContextabstractThe goal of image classification is to classify a collection of unlabeled images into a set of semantic classes. Many methods have been proposed to approach this goal by leveraging visual appearances of local patches in images. However, the spatial context between these local patches also provides significant information to improve the classification accuracy. Traditional spatial contextual models, such as two-dimensional hidden Markov model, attempt to construct one common model for each image category to depict the spatial structures of the images in this class. However due to large intra-class variances in an image category, one single model has difficulties in representing various spatial contexts in different images. In contrast, we propose to construct a prototype set of spatial contextual models by leveraging the kernel methods rather than only one model. Such an algorithm combines the advantages of rich representation ability of spatial contextual models as well as the powerful classification ability of kernel method. In particular, we propose a new distance measure between different spatial contextual models by integrating joint appearance-spatial image features. Such a distance measure can be efficiently computed in a recursive formulation that scales well to image size. Extensive experiments demonstrate that the proposed approach significantly outperforms the state-of-the-art approaches. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
IEEE Trans. Multim. | 3 |
| 2009 | Two-Dimensional Multilabel Active Learning with an Efficient Online Adaptation Model for Image ClassificationabstractConventional active learning dynamically constructs the training set only along the sample dimension. While this is the right strategy in binary classification, it is suboptimal for multilabel image classification. We argue that for each selected sample, only some effective labels need to be annotated while others can be inferred by exploring the label correlations. The reason is that the contributions of different labels to minimizing the classification error are different due to the inherent label correlations. To this end, we propose to select sample-label pairs, rather than only samples, to minimize a multilabel Bayesian classification error bound. We call it two-dimensional active learning because it considers both the sample dimension and the label dimension. Furthermore, as the number of training samples increases rapidly over time due to active learning, it becomes intractable for the offline learner to retrain a new model on the whole training set. So we develop an efficient online learner to adapt the existing model with the new one by minimizing their model distance under a set of multilabel constraints. The effectiveness and efficiency of the proposed method are evaluated on two benchmark data sets and a realistic image collection from a real-world image sharing Web site-Corbis. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Event Tactic Analysis Based on Broadcast Sports VideoabstractMost existing approaches on sports video analysis have concentrated on semantic event detection. Sports professionals, however, are more interested in tactic analysis to help improve their performance. In this paper, we propose a novel approach to extract tactic information from the attack events in broadcast soccer video and present the events in a tactic mode to the coaches and sports professionals. We extract the attack events with far-view shots using the analysis and alignment of web-casting text and broadcast video. For a detected event, two tactic representations, aggregate trajectory and play region sequence, are constructed based on multi-object trajectories and field locations in the event shots. Based on the multi-object trajectories tracked in the shot, a weighted graph is constructed via the analysis of temporal-spatial interaction among the players and the ball. Using the Viterbi algorithm, the aggregate trajectory is computed based on the weighted graph. The play region sequence is obtained using the identification of the active field locations in the event based on line detection and competition network. The interactive relationship of aggregate trajectory with the information of play region and the hypothesis testing for trajectory temporal-spatial distribution are employed to discover the tactic patterns in a hierarchical coarse-to-fine framework. Extensive experiments on FIFA World Cup 2006 show that the proposed approach is highly effective. Guangyu Zhu 0002, Changsheng Xu, Qingming Huang, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
IEEE Trans. Multim. | 4 |
| 2008 | Two-Dimensional Active Learning for image classificationabstractIn this paper, we propose a two-dimensional active learning scheme and show its application in image classification. Traditional active learning methods select samples only along the sample dimension. While this is the right strategy in binary classification, it is sub-optimal for multi-label classification. In multi-label classification, we argue that, for each selected sample, only a part of more effective labels are necessary to be annotated while others can be inferred by exploring the correlations among the labels. The reason is that the contributions of different labels to minimizing the classification error are different due to the inherent label correlations. To this end, we propose to select sample-label pairs, rather than only samples, to minimize a multi-label Bayesian classification error bound. This new active learning strategy not only considers the sample dimension but also the label dimension, and we call it Two-Dimensional Active Learning (2DAL). We also show that the traditional active learning formulation is a special case of 2DAL when there is only one label. Extensive experiments conducted on two real-world applications show that the 2DAL significantly outperforms the best existing approaches which did not take label correlation into account. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
CVPR | 3 |
| 2008 | A joint appearance-spatial distance for kernel-based image categorizationabstractThe goal of image categorization is to classify a collection of unlabeled images into a set of predefined classes to support semantic-level image retrieval. The distance measures used in most existing approaches either ignored the spatial structures or used them in a separate step. As a result, these distance measures achieved only limited success. To address these difficulties, in this paper, we propose a new distance measure that integrates joint appearance-spatial image features. Such a distance measure is computed as an upper bound of an information-theoretic discrimination, and can be computed efficiently in a recursive formulation that scales well to image size. In addition, the upper bound approximation can be further tightened via adaption learning from a universal reference model. Extensive experiments on two widely-used data sets show that the proposed approach significantly outperforms the state-of-the-art approaches. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Zhengjun Zha, HongJiang Zhang |
CVPR | 3 |
| 2008 | Topic mining on web-shared videosabstractInternet videos have grown exponentially with the help from video sharing Websites. Automatic topic mining is therefore increasingly important for organizing and navigating such large video databases. Most of current solutions of topic detection and mining were done on news videos and cannot be directly applied on Web videos, because of their limited and noisy semantic information. In this paper, we will try to address this problem and propose an automatic topic mining framework on Web videos. We develop an iterative weight-updated co-clustering scheme to filter "noisy" tags and mine the "hot" topics. We then propose a visual-based clustering approach to further group the videos with similar content, and rank the visual-similar groups by their similarity to the topic center. Experiments on a large Web video database demonstrate the superior performance of our weight-updated co-clustering to both of the traditional co-clustering and k-means. The experiments also demonstrate significant improvement of users' experience by our visual-based clustering and ranking. Lu Liu 0005, Yong Rui, Lifeng Sun, Bo Yang 0008, Jianwei Zhang 0001, Shiqiang Yang |
ICASSP | 2 |
| 2008 | Web video topic discovery and tracking via bipartite graph reinforcement modelabstractAutomatic topic discovery and tracking on web-shared videos can greatly benefit both web service providers and end users. Most of current solutions of topic detection and tracking were done on news and cannot be directly applied on web videos, because the semantic information of web videos is much less than that of news videos. In this paper, we propose a bipartite graph model to address this issue. The bipartite graph represents the correlation between web videos and their keywords, and automatic topic discovery is achieved through two steps - coarse topic filtering and fine topic re-ranking. First, a weight-updating co-clustering algorithm is employed to filter out topic candidates at a coarse level. Then the videos on each topic are re-ranked by analyzing the link structures of the corresponding bipartite graph. After the topics are discovered, the interesting ones can also be tracked over a period of time using the same bipartite graph model. The key is to propagate the relevant scores and keywords from the videos of interests to other relevant ones through the bipartite graph links. Experimental results on real web videos from YouKu, a YouTube counterpart in China, demonstrate the effectiveness of the proposed methods. We report very promising results. Lu Liu 0005, Lifeng Sun, Yong Rui, Shiqiang Yang |
WWW | 3 |
| 2008 | Application Potential of Multimedia Information RetrievalabstractThis paper will first briefly survey the existing impact of multimedia information retrieval (MIR) in applications. It will then analyze the current trends of MIR research which can have an influence on future applications. It will then detail the future possibilities and bottlenecks in applying the MIR research results in the main target application areas, such as the consumer (e.g., personal video recorders, web information retrieval), public safety (e.g., automated smart surveillance systems), and professional world (e.g., automated meeting capture and summarization). In particular, recommendations will be made to the research community regarding the challenges that need to be met to make the knowledge transfer towards the applications more efficient and effective. It will also attempt to study the trends in the applications which can inform the MIR community on directing intellectual resources towards MIR problems which can have a maximal real-world impact. Mohan Kankanhalli, Yong Rui |
Proc. IEEE | 2 |
| 2008 | Using Webcast Text for Semantic Event Detection in Broadcast Sports VideoabstractSports video semantic event detection is essential for sports video summarization and retrieval. Extensive research efforts have been devoted to this area in recent years. However, the existing sports video event detection approaches heavily rely on either video content itself, which face the difficulty of high-level semantic information extraction from video content using computer vision and image processing techniques, or manually generated video ontology, which is domain specific and difficult to be automatically aligned with the video content. In this paper, we present a novel approach for sports video semantic event detection based on analysis and alignment of Webcast text and broadcast video. Webcast text is a text broadcast channel for sports game which is co-produced with the broadcast video and is easily obtained from the Web. We first analyze Webcast text to cluster and detect text events in an unsupervised way using probabilistic latent semantic analysis (pLSA). Based on the detected text event and video structure analysis, we employ a conditional random field model (CRFM) to align text event and video event by detecting event moment and event boundary in the video. Incorporation of Webcast text into sports video analysis significantly facilitates sports video semantic event detection. We conducted experiments on 33 hours of soccer and basketball games for Webcast analysis, broadcast video analysis and text/video semantic alignment. The results are encouraging and compared with the manually labeled ground truth. Changsheng Xu, Yifan Zhang 0001, Guangyu Zhu 0002, Yong Rui, Hanqing Lu, Qingming Huang |
IEEE Trans. Multim. | 4 |
| 2008 | Boosting-Based Multimodal Speaker Detection for Distributed Meeting VideosabstractIdentifying the active speaker in a video of a distributed meeting can be very helpful for remote participants to understand the dynamics of the meeting. A straightforward application of such analysis is to stream a high resolution video of the speaker to the remote participants. In this paper, we present the challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and propose a novel boosting-based multimodal speaker detection (BMSD) algorithm. Instead of separately performing sound source localization (SSL) and multiperson detection (MPD) and subsequently fusing their individual results, the proposed algorithm fuses audio and visual information at feature level by using boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. In experiments that includes hundreds of real-world meetings, the proposed BMSD algorithm reduces the error rate of SSL-only approach by 24.6%, and the SSL and MPD fusion approach by 20.9%. To the best of our knowledge, this is the first real-time multimodal speaker detection algorithm that is deployed in commercial products. Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola, Xinding Sun, Nelson Pinto, Zhengyou Zhang |
IEEE Trans. Multim. | 3 |
| 2008 | Correlative multilabel video annotation with temporal kernelsabstractAutomatic video annotation is an important ingredient for semantic-level video browsing, search and navigation. Much attention has been paid to this topic in recent years. These researches have evolved through two paradigms. In the first paradigm, each concept is individually annotated by a pre-trained binary classifier. However, this method ignores the rich information between the video concepts and only achieves limited success. Evolved from the first paradigm, the methods in the second paradigm add an extra step on the top of the first individual classifiers to fuse the multiple detections of the concepts. However, the performance of these methods can be degraded by the error propagation incurred in the first step to the second fusion one. In this article, another paradigm of the video annotation method is proposed to address these problems. It simultaneously annotates the concepts as well as model correlations between them in one step by the proposed Correlative Multilabel (CML) method, which benefits from the compensation of complementary information between different labels. Furthermore, since the video clips are composed by temporally ordered frame sequences, we extend the proposed method to exploit the rich temporal information in the videos. Specifically, a temporal-kernel is incorporated into the CML method based on the discriminative information between Hidden Markov Models (HMMs) that are learned from the videos. We compare the performance between the proposed approach and the state-of-the-art approaches in the first and second paradigms on the widely used TRECVID data set. As to be shown, superior performance of the proposed method is gained. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Tao Mei 0001, Meng Wang 0001, HongJiang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2008 | An automated end-to-end lecture capture and broadcasting systemabstractRemote viewing of lectures presented to a live audience is becoming increasingly popular. At the same time, the lectures can be recorded for subsequent on-demand viewing over the Internet. Providing such services, however, is often prohibitive due to the labor-intensive cost of capturing and pre/post-processing. This article presents a complete automated end-to-end system that supports capturing, broadcasting, viewing, archiving and searching of presentations. Specifically, we describe a system architecture that minimizes the pre- and post-production time, and a fully automated lecture capture system called iCam2 that synchronously captures all contents of the lecture, including audio, video, and presentation material. No staff is needed during lecture capture and broadcasting, so the operational cost of the system is negligible. The system has been used on a daily basis for more than 4 years, during which 522 lectures have been captured. These lectures have been viewed over 20,000 times. Cha Zhang, Yong Rui, Jim Crawford, Li-wei He |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2007 | Read, write, and navigation awareness in realistic multi-view collaborations
Sasa Junuzovic, Prasun Dewan, Yong Rui |
CollaborateCom | 3 |
| 2007 | Concurrent Multiple Instance Learning for Image CategorizationabstractWe propose a new multiple instance learning (MIL) algorithm to learn image categories. Unlike existing MIL algorithms, in which the individual instances in a bag are assumed to be independent with each other, we develop concurrent tensors to explicitly model the inter-dependency between the instances to better capture image's inherent semantics. Rank-1 tensor factorization is then applied to obtain the label of each instance. Furthermore, we formulate the classification problem in the reproducing kernel Hilbert space (RKHS) to extend instance label prediction to the whole feature space. Finally, a regularizer is introduced, which avoids overfitting and significantly improves learning machine's generalization capability, similar to that in SVMs. We report superior categorization performances compared with key existing approaches on both the COREL and the Caltech datasets. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Tao Mei 0001, Jinhui Tang 0001, HongJiang Zhang |
CVPR | 3 |
| 2007 | Semantic Event Extraction from Basketball Games using Multi-Modal AnalysisabstractIn this paper, we present a novel multi-modal framework for semantic event extraction from basketball games based on Webcasting text and broadcast video. We propose novel approaches to text analysis for event detection and semantics extraction, video analysis for event structure modeling and event moment detection, and text/video alignment for event boundary detection in the video. Compared with existing approaches to event detection in sports video which rely heavily on low-level features directly extracted from video itself, our approach aims to bridge the semantic gap between low-level features and high-level events and facilitates personalization of the sports video. Promising results are reported on real-world video clips by using text analysis, video analysis and text/video alignment. Yifan Zhang 0001, Changsheng Xu, Yong Rui, Jinqiao Wang, Hanqing Lu |
ICME | 3 |
| 2007 | Correlative multi-label video annotationabstractAutomatically annotating concepts for video is a key to semantic-level video browsing, search and navigation. The research on this topic evolved through two paradigms. The first paradigm used binary classification to detect each individual concept in a concept set. It achieved only limited success, as it did not model the inherent correlation between concepts, e.g., urban and building. The second paradigm added a second step on top of the individual concept detectors to fuse multiple concepts. However, its performance varies because the errors incurred in the first detection step can propagate to the second fusion step and therefore degrade the overall performance. To address the above issues, we propose a third paradigm which simultaneously classifies concepts and models correlations between them in a single step by using a novel Correlative Multi-Label (CML) framework. We compare the performance between our proposed approach and the state-of-the-art approaches in the first and second paradigms on the widely used TRECVID data set. We report superior performance from the proposed approach. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Tao Mei 0001, HongJiang Zhang |
ACM Multimedia | 3 |
| 2007 | Trajectory based event tactics analysis in broadcast sports videoabstractMost of existing approaches on event detection in sports video are general audience oriented. The extracted events are then presented to the audience without further analysis. However, professionals, such as soccer coaches, are more interested in the tactics used in the events. In this paper, we present a novel approach to extract tactic information from the goal event in broadcast soccer video and present the goal event in a tactic mode to the coaches and sports professionals. We first extract goal events with far-view shots based on analysis and alignment of web-casting text and broadcast video. For a detected goal event, we employ a multi-object detection and tracking algorithm to obtain the players and ball trajectories in the shot. Compared with existing work, we proposed an effective tactic representation called aggregate trajectory which is constructed based on multiple trajectories using a novel analysis of temporal-spatial interaction among the players and the ball. The interactive relationship with play region information and hypothesis testing for trajectory temporal-spatial distribution are exploited to analyze the tactic patterns in a hierarchical coarse-to-fine framework. The experimental results on the data of FIFA World Cup 2006 are promising and demonstrate our approach is effective. Guangyu Zhu 0002, Qingming Huang, Changsheng Xu, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
ACM Multimedia | 4 |
| 2006 | Light Weight Background Blurring for Video Conferencing ApplicationsabstractBackground blurring is an effective way to both preserving privacy and keeping communication effective during video conferencing. This paper proposes a light weight real-time algorithm to perform background blurring using a fast background modeling algorithm combined with a face detector/tracker. A soft decision is made at each pixel whether it belongs to the foreground or the background based on multiple vision features. The classification results are mapped to a per-pixel blurring radius image to blur the background. The algorithm produces satisfactory results under a wide range of conditions, and occupies less than 30% of the CPU cycles on a 3 GHz Pentium 4 machine without further optimization. Cha Zhang, Yong Rui, Li-wei He |
ICIP | 2 |
| 2006 | PING: a Group-to-Individual Distributed Meeting SystemabstractGroup-to-individual (G2I) distributed meeting is an important but understudied area. Because of the asymmetry between different parties in G2I meetings, it has two unique challenges: l)the remote participant tends to be ignored by the local participants; and 2) the remote participant has inferior audio, video, and data experience than the local participants. To address these issues, in this paper we present PING, a system explicitly designed for G2I distributed meetings that combines recent advances in both hardware, e.g., microphone arrays, remote person stand-in devices, and software, e.g., audio-video processing, to improve users' G2I meeting experience. We report how PING addresses the above two challenges and its system design and implementation Yong Rui, Eric Rudolph, Li-wei He, Rico Malvar, Michael F. Cohen, Ivan Tashev |
ICME | 1 |
| 2006 | PASS: Peer-Aware Silence Suppression for Internet Voice ConferencesabstractA novel tandem-free solution for multiparty VoIP conferences called PASS (peer-aware silence suppression) is presented. Similar to traditional tandem-free solutions, PASS introduces a limit on the number of concurrent speakers in a conference. But in contrast to traditional solutions, PASS silence suppression and speaker selection are completely distributed, running on each client. No speaker selection is performed at the bridge at all. This configuration leads to better scalability, lower bandwidth occupation and jitter buffer delay, and higher compatibility with a wide variety of network topologies. The key component of PASS, distributed silence suppression and speaker selection, is realized through a robust approach proposed in this paper. Based on a voice activity measure derived using machine learning techniques, this approach is able to reliably suppress silence in complex environments, and perform accurate and transparent speaker selection as well Xun Xu 0003, Li-wei He, Dinei A. F. Florêncio, Yong Rui |
ICME | 4 |
| 2006 | Recognizing Faces in Recorded Meetings via MRC-BoostingabstractPerson-based indices and timelines can enable fast and non-linear access to recorded meetings. This paper focuses on how to automatically construct those indices and timelines by using face recognition techniques. While there exist extensive research in generic face recognition, recognizing faces in recorded meetings is still an understudied area. Real-world meeting videos impose several interesting and unique challenges including complex lighting, low imaging quality, and large variations in head pose and size. In this paper, a promising approach based on MRC-Boosting is presented to address these challenges, which achieves encouraging performance on real-world meeting videos and shows superior accuracy and robustness compared to two popular existing approaches Xun Xu 0003, Yong Rui, Thomas S. Huang |
ICME | 2 |
| 2006 | A Three-Layer Virtual Director Model for Supporting Automated Multi-Site Distributed EducationabstractIn multi-site distributed education (MSDE), video streams from multiple sites are available. To best utilize the limited screen space at each site, we develop a customizable, automated display management system in this paper, i.e., only user-preferred streams will be shown as triggered by events and timers. The configuration of such user preference, however, is challenging because it has to be both human-friendly and machine-friendly. To address this challenge, we propose a three-layer virtual director model. In the user layer, we identify three categories of parameters that can represent a wide range of user preferences yet are easy to use. These preferences are then automatically translated into a machine-friendly timed automaton in the execution layer. The automaton is simulated dynamically, which selects a subset of streams to show on the screen through a display layer. Evaluation results demonstrate the correctness and efficiency of the proposed framework Bin Yu 0010, Cha Zhang, Yong Rui, Klara Nahrstedt |
ICME | 3 |
| 2006 | Boosting-Based Multimodal Speaker Detection for Distributed MeetingsabstractSpeaker detection is a very important task in distributed meeting applications. This paper discusses a number of challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and proposes a boosting-based multimodal speaker detection (BMSD) algorithm. Instead of performing sound source localization (SSL) and multi-person detection (MPD) separately and subsequently fusing their individual results, the proposed algorithm uses boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. The algorithm reduces the error rate of SSL-only approach by 47%, and the SSL and MPD fusion approach by 27% Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola |
MMSP | 3 |
| 2006 | Multicue HMM-UKF for Real-Time Contour TrackingabstractWe propose an HMM model for contour detection based on multiple visual cues in spatial domain and improve it by joint probabilistic matching to reduce background clutter. It is further integrated with unscented Kalman filter to exploit object dynamics in nonlinear systems for robust contour tracking. Yunqiang Chen, Yong Rui, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Direct Kernel Biased Discriminant Analysis: A New Content-Based Image Retrieval Relevance Feedback AlgorithmabstractIn recent years, a variety of relevance feedback (RF) schemes have been developed to improve the performance of content-based image retrieval (CBIR). Given user feedback information, the key to a RF scheme is how to select a subset of image features to construct a suitable dissimilarity measure. Among various RF schemes, biased discriminant analysis (BDA) based RF is one of the most promising. It is based on the observation that all positive samples are alike, while in general each negative sample is negative in its own way. However, to use BDA, the small sample size (SSS) problem is a big challenge, as users tend to give a small number of feedback samples. To explore solutions to this issue, this paper proposes a direct kernel BDA (DKBDA), which is less sensitive to SSS. An incremental DKBDA (IDKBDA) is also developed to speed up the analysis. Experimental results are reported on a real-world image collection to demonstrate that the proposed methods outperform the traditional kernel BDA (KBDA) and the support vector machine (SVM) based RF algorithms. Dacheng Tao, Xiaoou Tang, Xuelong Li 0001, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2005 | Sound source localization for circular arrays of directional microphonesabstractPrevious research in sound source localization has helped increase the robustness of estimates to noise and reverberation. Circular arrays are of particular interest for a number of scenarios, particularly because they can be placed in the center of the sources. First, that improves the sound capture due to the reduced distance. Second, it helps on the direction estimation, not only because of the reduced distance, but also because it increases the angle differences. Nevertheless, most research on circular arrays focused on the case of omni-directional microphones. In this paper, we present a new algorithm for sound source localization developed specifically for directional microphones. Results obtained from real meeting room setups show a typical error of less than 3 degrees. Yong Rui, Dinei A. F. Florêncio, Warren Lam, Jinyan Su |
ICASSP (3) | 1 |
| 2005 | Hybrid speaker tracking in an automated lecture roomabstractWe present a hybrid speaker tracking scheme based on a single pan/tilt/zoom (PTZ) camera in an automated lecture capturing system. Given that the camera's video resolution is higher than the required output resolution, we frame the output video as a sub-region of the camera's input video. This allows us to track the speaker both digitally and mechanically. Digital tracking has the advantage of being smooth, and mechanical tracking can cover a wide area. The hybrid tracking achieves the benefits of both worlds. In addition to hybrid tracking, we present an intelligent pan/zoom selection scheme to improve the aestheticity of the lecture scene. Cha Zhang, Yong Rui, Li-wei He, Michael N. Wallick |
ICME | 2 |
| 2005 | What is the state of our community?abstract10.1145/1101149.1101297 Yong Rui, Ramesh Jain 0001, Nicolas D. Georganas, HongJiang Zhang, Klara Nahrstedt, John R. Smith, Mohan Kankanhalli |
ACM Multimedia | 1 |
| 2005 | An automated end-to-end lecture capturing and broadcasting systemabstractWe present a complete end-to-end system that is fully automated and supports capturing, broadcasting, viewing, archiving and search. Specifically, we describe a system architecture that minimizes the pre- and post-production time, and a fully automated lecture capturing system called iCam2, which synchronously captures all the contents of the lecture, including audio, video and visual aids. As no staff is needed during the capturing and broadcasting process, the operation cost of our system is negligible. The system has been used on a daily basis for more than 4 years, during which 467 lectures were captured with 17,000+ online viewers. Cha Zhang, Jim Crawford, Yong Rui, Li-wei He |
ACM Multimedia | 3 |
| 2004 | Time delay estimation in the presence of correlated noise and reverberationabstractWe propose a new two-stage framework for time delay estimation in the presence of correlated noise and reverberation. The new framework allows us to develop a set of new approaches as well as to unify existing ones. We further develop the maximum likelihood estimation when reverberation is present. The corresponding weighting function is a more accurate form of the weighting function proposed by H. Wang and P. Chu (Proc. ICASSP, 1997), one of the best existing techniques. We compare our new algorithms with the existing ones and report superior performance. Yong Rui, Dinei A. F. Florêncio |
ICASSP (2) | 1 |
| 2004 | A portable solution for automatic lecture room camera managementabstractRapid advances in technology and decreasing costs have made it possible to attach high resolution video cameras to just about any computer and record the interactions in a lecture room. Additionally, lecture rooms may be outfitted with several cameras for this purpose. However, recording the interactions alone does not create effective video. We present a method for not only recording, but also editing in real-time, lectures. Unlike previous work, the system is highly portable, allowing quick set-up in various types of lecture rooms. This portability is achieved by using the abstraction of virtual cameramen and physical cameras, and a scriptable interface to the editing rules. Michael N. Wallick, Yong Rui, Li-wei He |
ICME | 2 |
| 2004 | Breaking the clock face HIPabstractWeb services designed for human users are being abused by computer programs (bots). The bots steal thousands of free email accounts in a minute, participate in online polls to skew results, and irritate people in chat rooms. These real-world issues have recently generated a new research area called human interactive proofs (HIP), whose goal is to defend Web services from malicious attacks by differentiating bots from human users. For a HIP challenge to be effective, it needs to be both easy for human and robust against bots attacks. Recently, there is a new HIP design which is based on telling time from a clock face. This is a very innovative idea and quite easy for human to pass. However, its robustness to attacks needs verification. We present an algorithm that can break the clock face HIP at 87.4%. ZhenQiu Zhang, Yong Rui, Thomas S. Huang, Cem Paya |
ICME | 2 |
| 2004 | Constraint Based Region Matching for Image Retrieval
Yong Rui, Jia-Guang Sun 0001 |
Int. J. Comput. Vis. | 2 |
| 2004 | Improving Retrieval Performance by Region Constraints and Relevance Feedback
Yong Rui, Jia-Guang Sun 0001 |
J. Comput. Sci. Technol. | 2 |
| 2004 | Automating lecture capture and broadcast: technology and videography
Yong Rui, Anoop Gupta, Jonathan Grudin, Li-wei He |
Multim. Syst. | 1 |
| 2004 | ARTiFACIAL: Automated Reverse Turing test using FACIAL features
Yong Rui, Zicheng Liu 0001 |
Multim. Syst. | 1 |
| 2004 | Real-time speaker tracking using particle filter sensor fusionabstractSensor fusion for object tracking has become an active research direction during the past few years. But how to do it in a robust and principled way is still an open problem. In this paper, we propose a new fusion framework that combines both the bottom-up and top-down approaches to probabilistically fuse multiple sensing modalities. At the lower level, individual vision and audio trackers are designed to generate effective proposals for the fuser. At the higher level, the fuser performs reliable tracking by verifying hypotheses over multiple likelihood models from multiple cues. Unlike traditional fusion algorithms, the proposed framework is a closed-loop system where the fuser and trackers coordinate their tracking information. Furthermore, to handle nonstationary situations, the proposed framework evaluates the performance of the individual trackers and dynamically updates their object states. We present a real-time speaker tracking system based on the proposed framework by fusing object contour, color and sound source location. We report robust tracking results. Yunqiang Chen, Yong Rui |
Proc. IEEE | 2 |
| 2003 | Videography for telepresentationsabstractOur goal is to help automate the capture and broadcast of lectures to remote audiences. There are two inter-related components to the design of such systems. The technology component includes the hardware (e.g., video cameras) and associated software (e.g., speaker-tracking). The aesthetic component embodies the rules and idioms that human videographers follow to make a video visually engaging. We present a lecture room automation system and a substantial number of new video-production rules obtained from professional videographers who critiqued it. We also describe rules for a variety of lecture room environments differing in the numbers and types of cameras. We further discuss gaps between what professional videographers do and what is technologically feasible today. Yong Rui, Anoop Gupta, Jonathan Grudin |
CHI | 1 |
| 2003 | New direct approaches to robust sound source localizationabstractWhen more than two microphones are used, the traditional time-delay-of-arrival (TDOA) based sound source localization (SSL) approach involves two steps. The first step computes TDOA for each microphone pair, and the second step combines these estimates. This two-step process discards relevant information in the first step, thus degrading the SSL accuracy and robustness. Although less used, one-step processes do exist. In this paper, we review these processes, create a unified framework, and introduce two new one-step algorithms. We compare our proposed approaches against existing 1and 2-step approaches and demonstrate significantly better SSL performance. Yong Rui, Dinei A. F. Florêncio |
ICME | 1 |
| 2003 | ARTiFACIAL: automated reverse turing test using FACIAL featuresabstractWeb services designed for human users are being abused by computer programs (bots). The bots steal thousands of free email accounts in a minute; participate in online polls to skew results; and irritate people by joining online chat rooms. These real-world issues have recently generated a new research area called Human Interactive Proofs (HIP), whose goal is to defend services from malicious attacks by differentiating bots from human users. In this paper, we propose a new HIP algorithm based on detecting human face and facial features. Human faces are the most familiar object to humans, rendering it possibly the best candidate for HIP. We conducted user studies and showed the ease of use of our system to human users. We designed attacks using the best existing face detectors and demonstrated the difficulty to bots. Yong Rui, Zicheng Liu 0001 |
ACM Multimedia | 1 |
| 2003 | Excuse me, but are you human?abstractWeb services designed for human users are being abused by computer programs (bots). The bots steal thousands of free email accounts in a minute; participate in online polls to skew results; and irritate people by joining online chat rooms. These real-world issues have recently generated a new research area called Human Interactive Proofs (HIP), whose goal is to defend services from malicious attacks by differentiating bots from human users. We propose a new HIP algorithm based on detecting human face and facial features. Human faces are the most familiar object to humans, rendering it possibly the best candidate for HIP. We conducted user studies and showed the ease of use of our system to human users. We designed attacks using the best existing face detectors and demonstrated the difficulty to bots. Yong Rui, Zicheng Liu 0001 |
ACM Multimedia | 1 |
| 2003 | Adaptive tree similarity learning for image retrieval
Yong Rui, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
Multim. Syst. | 2 |
| 2002 | Parametric contour tracking using unscented Kalman filterabstractThis paper presents an efficient method to integrate various spatial-temporal constraints to regularize the contour tracking. Specifically, the global shape prior, contour smoothness and object dynamics are addressed. First, the contour is represented as a parametric shape, based on which a causal smoothness constraint can be developed to exploit the local spatial constraint. The causality nature of the constraint allows us to do efficient probabilistic contour detection using the powerful hidden Markov model (HMM). Finally, a unscented Kalman filter (UKF) is applied to estimate object parameters based on the nonlinear observation model (i.e. the relationship between the detected contour points and the contour parameters) and the object dynamics. Better than other variants of the recursive least mean square estimators (e.g., extended Kalman filter), the UKF approximates nonlinear systems up to the second order (third for Gaussian prior) with similar computational cost. This novel tracking algorithm is running in real-time and robust to severe distractions due to the comprehensive spatial-temporal constraints. It is applied to track people in bad illumination and cluttered environments. Promising results are reported. Yunqiang Chen, Thomas S. Huang, Yong Rui |
ICIP (3) | 3 |
| 2002 | Distributed meetings: a meeting capture and broadcasting systemabstractThe common meeting is an integral part of everyday life for most workgroups. However, due to travel, time, or other constraints, people are often not able to attend all the meetings they need to. Teleconferencing and recording of meetings can address this problem. In this paper we describe a system that provides these features, as well as a user study evaluation of the system. The system uses a variety of capture devices (a novel 360° camera, a whiteboard camera, an overview camera, and a microphone array) to provide a rich experience for people who want to participate in a meeting from a distance. The system is also combined with speaker clustering, spatial indexing, and time compression to provide a rich experience for people who miss a meeting and want to watch it afterward. Ross Cutler, Yong Rui, Anoop Gupta, Jonathan J. Cadiz, Ivan Tashev, Li-wei He, Alex Colburn, Zhengyou Zhang, Zicheng Liu 0001, Steve Silverberg |
ACM Multimedia | 2 |
| 2001 | Automating camera management for lecture room environmentsabstractGiven rapid improvements in network infrastructure and streaming-media technologies, a large number of corporations and universities are recording lectures and making them available online for anytime, anywhere access. However, producing high-quality lecture videos is still labor intensive and expensive. Fortunately, recent technology advances are making it feasible to build automated camera management systems to capture lectures. In this paper we report on our design, implementation and study of such a system. Compared to previous work-which has tended to be technology centric-we started with interviews with professional video producers and used their knowledge and expertise to create video production rules. We then targeted technology components that allowed us to implement a substantial portion of these rules, including the design of a virtual video director. The system's performance was compared to that of a human operator via a user study. Results suggest that our system's quality in close to that of a human-controlled system. In fact most remote audience members could not tell if the video was produced by a computer or a person. Qiong Liu 0003, Yong Rui, Anoop Gupta, Jonathan J. Cadiz |
CHI | 2 |
| 2001 | Viewing meeting captured by an omni-directional cameraabstractOne vision of future technology is the ability to easily and inexpensively capture any group meeting that occurs, store it, and make it available for people to view anytime and anywhere on the network. One barrier to achieving this vision has been the design of low-cost camera systems that can capture important aspects of the meeting without needing a human camera operator. A promising solution that has emerged recently is omni-directional cameras that can capture a 360-degree video of the entire meeting. Yong Rui, Anoop Gupta, Jonathan J. Cadiz |
CHI | 1 |
| 2001 | JPDAF Based HMM or Real-Time Contour TrackingabstractTracking objects using multiple cues yields more robust results. The well-known hidden Markov model (HMM) provides a powerful framework to incorporate multiple cues by expanding its observation. However, a plain HMM does not capture the inter-correlation between measurements of neighboring states when computing the transition probabilities. This can seriously damage the tracking performance. To overcome this difficulty, we propose a novel HMM framework targeted at contour-based object tracking. A joint probability data association filter (JPDAF) is used to compute the HMM's transition probabilities, taking into account the intercorrelated neighboring measurements. To ensure real-time performance, we have further developed an efficient method to calculate the data association probability via dynamic programming, which allows the proposed JPDAF-HMM to run comfortably at 30 frames/sec. This new tracking framework can easily incorporate various image cues (e.g., edge intensity, foreground region color and background region color), and also offers an online learning process to adapt to changes in the scene. To evaluate its tracking performance, we have applied the proposed JPDAF-HMM in various real-world video sequences. We report promising tracking results in complex environments. Yunqiang Chen, Yong Rui, Thomas S. Huang |
CVPR (1) | 2 |
| 2001 | Better Proposal Distributions: Object Tracking Using Unscented Particle FilterabstractTracking objects involves the modeling of non-linear non-Gaussian systems. On one hand, variants of Kalman filters are limited by their Gaussian assumptions. On the other hand, conventional particle filter, e.g., CONDENSATION, uses transition prior as the proposal distribution. The transition prior does not take into account current observation data, and many particles can therefore be wasted in low likelihood area. To overcome these difficulties, unscented particle filter (UPF) has recently been proposed in the field of filtering theory. In this paper, we introduce the UPF framework into audio and visual tracking. The UPF uses the unscented Kalman filter to generate sophisticated proposal distributions that seamlessly integrate the current observation, thus greatly improving the tracking performance. To evaluate the efficacy of the UPF framework, we apply it in two real-world tracking applications. One is the audio-based speaker localization, and the other is the vision-based human tracking. The experimental results are compared against those of the widely used CONDENSATION approach and have demonstrated superior tracking performance. Yong Rui, Yunqiang Chen |
CVPR (2) | 1 |
| 2001 | Optimal Adaptive Learning for Image RetrievalabstractLearning-enhanced relevance feedback is one of the most promising and active research directions in content-based image retrieval. However, the existing approaches either require prior knowledge of the data or entail high computation costs, making them less practical. To overcome these difficulties and motivated by the successful history of optimal adaptive filters, we present a new approach to interactive image retrieval. Specifically, we cast the image retrieval problem in the optimal filtering framework, which does not require prior knowledge of the data, supports incremental learning, is simple to implement and achieves better performance than state-of-the-art approaches. To evaluate the effectiveness and robustness of the proposed approach, extensive experiments have been carried out on a large heterogeneous image collection with 17,000 images. We report promising results on a wide variety of queries. Yong Rui, Shi-Min Hu 0001 |
CVPR (1) | 2 |
| 2001 | Optimal radial contour tracking by dynamic programmingabstractA common problem in most active contour methods is that the recursive searching scheme can only return a local optimal solution. Furthermore, the internal energy of the snake is not strong enough to control the shape of the contour. To overcome these difficulties, we develop a causal internal energy term based on a radial contour representation to encode the smooth constraint of the contour, and develop a global shape priori to control contour's shape and position based on object's dynamics. The causality nature of the representation allows us to efficiently find global optimal solution using dynamic programming. To validate the efficacy and robustness of the proposed approach, we apply this approach to track people in bad illumination and cluttered environments. We report promising results in the paper. Yunqiang Chen, Thomas S. Huang, Yong Rui |
ICIP (1) | 3 |
| 2001 | Building an intelligent camera management systemabstractGiven rapid improvements in storage devices, network infrastructure and streaming-media technologies, a large number of corporations and universities are recording lectures and making them available online for anytime, anywhere access. However, producing high-quality lecture videos is still labor intensive and expensive. Fortunately, recent technology advances are making it feasible to build automated camera management systems to capture lectures. In this paper we report our design of such a system, including system configuration, audio-visual tracking techniques, software architecture, and user study. Motivated by different roles in a professional video production team, we have developed a multi-cinematographer single-director camera management system. The system performs lecturer tracking, audience tracking, and video editing all fully automatically, and offers quality close to that of human-operated systems. Yong Rui, Li-wei He, Anoop Gupta, Qiong Liu 0003 |
ACM Multimedia | 1 |
| 2000 | Browsing digital videoabstractVideo in digital format played on programmable devices presents opportunities for significantly enhancing the user's viewing experience. For example, time compression and pause removal can shorten the viewing time for a video, textual and visual indices can allow personalized navigation through the content, and random-access digital storage allows instantaneous seeks into the content. To understand user behavior when such capabilities are available, we built a software video browsing application that combines many such features. We present results from a user study where users browsed video in six different categories: classroom lectures, conference presentations, entertainment shows, news, sports, and travel. Our results show that the most frequently used features were time compression, pause removal, and navigation using shot boundaries. Also, the behavior was different depending on the content type, and we present a classification. Finally, the users found the browser to be very useful. Two main reasons were: i) the ability to save time and ii) the feeling of control over what content they watched. Francis C. Li, Anoop Gupta, Elizabeth Sanocki, Li-wei He, Yong Rui |
CHI | 5 |
| 2000 | Segmenting Visual Actions Based on Spatio-Temporal Motion PatternsabstractThe analysis of human action captured in video sequences has been a topic of considerable interest in computer vision. Much of the previous work has focused on the problem of action or activity recognition, but ignored the problem of detecting action boundaries in a video sequence containing unfamiliar and arbitrary visual actions. This paper presents an approach to this problem based on detecting temporal discontinuities of the spatial pattern of image motion that captures the action. We represent frame to frame optical-flow in terms of the coefficients of the most significant principal components computed from all the flow-fields within a given video sequence. We then detect the discontinuities in the temporal trajectories of these coefficients based on three different measures. We compare our segment boundaries against those detected by human observers on the same sequences in a recent independent psychological study of human perception of visual events. We show experimental results on the two sequences that were used in this study. Our experimental results are promising both from visual evaluation and when compared against the results of the psychological study. Yong Rui, P. Anandan 0001 |
CVPR | 1 |
| 2000 | Optimizing Learning in Image RetrievalabstractCombining learning with vision techniques in interactive image retrieval has been an active research topic during the past few years. However, existing learning techniques either are based on heuristics or fail to analyze the working conditions. Furthermore, there is almost no in depth study on how to effectively learn from the users when there are multiple visual features in the retrieval system. To address these limitations, in this paper we present a vigorous optimization formulation of the learning process and solve the problem in a principled way. By using Lagrange multipliers, we have derived explicit solutions, which are both optimal and fast to compute. Extensive comparisons against state-of-the-art techniques have been performed. Experiments were carried out on a large-size heterogeneous image collection consisting of 17,000 images. Retrieval performance was tested under a wide range of conditions. Various evaluation criteria, including precision-recall curve and rank measure, have demonstrated the effectiveness and robustness of the proposed technique. Yong Rui, Thomas S. Huang |
CVPR | 1 |
| 2000 | Automatically extracting highlights for TV Baseball programsabstractIn today's fast-paced world, while the number of channels of television programming available is increasing rapidly, the time available to watch them remains the same or is decreasing. Users desire the capability to watch the programs time-shifted (on-demand) and/or to watch just the highlights to save time. In this paper we explore how to provide for the latter capability, that is the ability to extract highlights automatically, so that viewing time can be reduced. Yong Rui, Anoop Gupta, Alex Acero |
ACM Multimedia | 1 |
| 1999 | Video Sequence Learning and Recognition Via Dynamic SomabstractInformation contained in video sequences is crucial for an autonomous robot or a computer to learn and respond to its surrounding environment. In the past, robot vision mainly concentrated on still image processing and small "image cube" processing. Continuous video sequence learning and recognition is rarely addressed in the literature due to its high requirement of dynamic processing. In this paper, we propose a novel neural network structure called dynamic self-organizing map (DSOM) for video sequence processing. The proposed technique has been tested on simulation data sets, and the results validate its learning/recognition ability. Qiong Liu 0003, Yong Rui, Thomas S. Huang, Stephen E. Levinson |
ICIP (4) | 2 |
| 1999 | Water-Filling: A Novel Way for Image Structural Feature ExtractionabstractThe performance of a content based image retrieval (CBIR) system is inherently constrained by the features adopted to represent the images in the database. In this paper, a new approach is proposed for image feature extraction based on edge maps. The feature vector with multiple feature components is computed through a “water-filling algorithm” applied on the edge map of the original image. The idea of this algorithm is to obtain measures of the edge length and complexity by graph traverse. The new feature is move generally applicable than texture or shape features. We call this structure feature. Experiments show that the new feature is capable of catching salient edge/structure information in the images. An experimental retrieval system utilizing the proposed new features yields better results in retrieving city/building images than some global texture features (wavelet moments). The new feature is ideal for images with clear edge structure. After combining the new features with other features in a relevance feedback framework, satisfactory retrieval results are observed. Xiang Sean Zhou, Yong Rui, Thomas S. Huang |
ICIP (2) | 2 |
| 1999 | A novel relevance feedback technique in image retrievalabstractThe relevance feedback based approach to image retrieval has been an active research direction in the past few years.Many parameter estimation techniques have been proposed for relevance feedback.However, most of them are either based on ad-hoc heuristics or only partial solutions.In this paper, we introduce the first technique that not only has a solid theoretical framework but also takes into account the multi-level image content model.This technique formulates a vigorous optimization problem.By using Lagrange multipliers, we have derived the explicit optimal solutions for both the query vectors and the weights associated with the two-level image model.Experimental results on realworld image collections have shown the effectiveness and robustness of our proposed algorithm. Yong Rui, Thomas S. Huang |
ACM Multimedia (2) | 1 |
| 1999 | Video key frame extraction by unsupervised clustering and feedback adjustment
Yueting Zhuang, Yong Rui, Thomas S. Huang |
J. Comput. Sci. Technol. | 2 |
| 1999 | Image Retrieval: Current Techniques, Promising Directions, and Open Issues
Yong Rui, Thomas S. Huang, Shih-Fu Chang |
J. Vis. Commun. Image Represent. | 1 |
| 1999 | Constructing Table-of-Content for Videos
Yong Rui, Thomas S. Huang, Sharad Mehrotra |
Multim. Syst. | 1 |
| 1998 | Digital image/video library and MPEG-7: standardization and research issuesabstractMuch research activity and interest has emerged in two closely related areas: the digital image/video library (DIVL) and MPEG-7. We review the critical research issues in DIVL from a signal processing viewpoint, the objectives and scope of MPEG-7, and the relationships between these two. Yong Rui, Thomas S. Huang, Shih-Fu Chang |
ICASSP | 1 |
| 1998 | Adaptive Key Frame Extraction using Unsupervised ClusteringabstractKey frame extraction has been recognized as one of the important research issues in video information retrieval. Although progress has been made in key frame extraction, the existing approaches are either computationally expensive or ineffective in capturing salient visual content. We first discuss the importance of key frame selection; and then review and evaluate the existing approaches. To overcome the shortcomings of the existing approaches, we introduce a new algorithm for key frame extraction based on unsupervised clustering. The proposed algorithm is both computationally simple and able to adapt to the visual content. The efficiency and effectiveness are validated by large amount of real-world videos. Yueting Zhuang, Yong Rui, Thomas S. Huang, Sharad Mehrotra |
ICIP (1) | 2 |
| 1998 | Browsing and retrieving video content in a unified frameworkabstractIn this paper, we first review the recent research progress in video analysis, representation, browsing, and retrieval. Motivated by the mechanism used to access a book's content, we then present novel techniques for constructing video table-of-contents and index to facilitate accessing video's content. We further explore the relationship between video browsing and retrieval and propose a unified framework to incorporate both entities in a seamless way. Preliminary research results justify our proposed framework for providing access to videos based on their content. Yong Rui, Thomas S. Huang, Sharad Mehrotra |
MMSP | 1 |
| 1998 | Relevance feedback: a power tool for interactive content-based image retrievalabstractContent-based image retrieval (CBIR) has become one of the most active research areas in the past few years. Many visual feature representations have been explored and many systems built. While these research efforts establish the basis of CBIR, the usefulness of the proposed approaches is limited. Specifically, these efforts have relatively ignored two distinct characteristics of CBIR systems: (1) the gap between high-level concepts and low-level features, and (2) the subjectivity of human perception of visual content. This paper proposes a relevance feedback based interactive retrieval approach, which effectively takes into account the above two characteristics in CBIR. During the retrieval process, the user's high-level query and perception subjectivity are captured by dynamically updated weights based on the user's feedback. The experimental results over more than 70000 images show that the proposed approach greatly reduces the user's effort of composing a query, and captures the user's information need more precisely. Yong Rui, Thomas S. Huang, Michael Ortega-Binderberger, Sharad Mehrotra |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Supporting Ranked Boolean Similarity Queries in MARSabstractTo address the emerging needs of applications that require access to and retrieval of multimedia objects, we are developing the Multimedia Analysis and Retrieval System (MARS). In this paper, we concentrate on the retrieval subsystem of MARS and its support for content-based queries over image databases. Content-based retrieval techniques have been extensively studied for textual documents in the area of automatic information retrieval. This paper describes how these techniques can be adapted for ranked retrieval over image databases. Specifically, we discuss the ranking and retrieval algorithms developed in MARS based on the Boolean retrieval model and describe the results of our experiments that demonstrate the effectiveness of the developed model for image retrieval. Michael Ortega-Binderberger, Yong Rui, Kaushik Chakrabarti, Kriengkrai Porkaew, Sharad Mehrotra, Thomas S. Huang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1997 | Content-Based Image Retrieval with Relevance Feedback in MARSabstractTechnology advances in the areas of image processing (IP) and information retrieval (IR) have evolved separately for a long time. However, successful content-based image retrieval systems require the integration of the two. There is an urgent need to develop integration mechanisms to link the image retrieval model to text retrieval model, such that the well established text retrieval techniques can be utilized. Approaches of converting image feature vectors (IF domain) to weighted-term vectors (IR domain) are proposed in this paper. Furthermore, the relevance feedback technique from the IR domain is used in content-based image retrieval to demonstrate the effectiveness of this conversion. Experimental results show that the image retrieval precision increases considerably by using the proposed integration approach. Yong Rui, Thomas S. Huang, Sharad Mehrotra |
ICIP (2) | 1 |
| 1997 | Supporting Similarity Queries in MARSabstractTo address the emerging needs of applications that require access to and retrieval of multimedia objects, we are developing the Multimedia Analysis and Retrieval System (MARS) in our group at the University of Illinois [13].In this paper, we concentrate on the retrieval subsystem of MARS and its support for content-based queries over image databases.Content-based retrieval techniques have been extensively studied for textual documents in the area of automatic information retrieval [24, 21.This paper describes how these techniques can be adapted for ranked retried over image databases.Specifically, we discuss the ranking and retrieval algorithms developed in MARS based on the Boolean re-trievaI model and describe the results of our experiments that demonstrate the effectiveness of the developed model for image retrieval. Michael Ortega-Binderberger, Yong Rui, Kaushik Chakrabarti, Sharad Mehrotra, Thomas S. Huang |
ACM Multimedia | 2 |
| 1996 | Automated region segmentation using attraction-based grouping in spatial-color-texture spaceabstractA great deal of attention has been paid to image content-based retrieval systems (CBRS). One important goal in CBRS is to extract local low level image features such as color, texture and shape, to allow queries based on these features. A large CBRS containing tens of thousands of images requires an automatic feature-extraction method since human aided segmentation is impractical. We address this problem in a particular application setting by using an attraction-based grouping method in spatial-color-texture space. The attraction concept makes this approach similar to human aided segmentation. Experimental results show that the method is reasonably better than existing methods, and has the potential to be used in other CBRS-related applications. Yong Rui, Alfred C. She, Thomas S. Huang |
ICIP (1) | 1 |