VLDB 2026 Research / reviewers in the wild / expert
Yueting Zhuang
dblp:218/7793
· DBLP profile ↗
389ranked-venue papers
28as first author
134since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 224 · 19 first-author · 61 since 2021Artificial intelligence and machine learning · 172 · 3 first-author · 89 since 2021Databases, data management, data science and information retrieval · 37 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 6 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GUI-G²: Gaussian Reward Modeling for GUI GroundingabstractGraphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks. Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang |
AAAI | 12 |
| 2026 | MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language ModelsabstractJie Cao, Tianwei Lin, Bo Yuan, Rolan Yan, Hongyang He, Wenqiao Zhang, Juncheng Li, Dongping Zhang, Siliang Tang, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianwei Lin 0001, Rolan Yan, Hongyang He, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang |
ACL (1) | 10 |
| 2026 | UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy OptimizationabstractZhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001 |
ACL (1) | 10 |
| 2026 | Experience-driven Multi-turn Reinforcement Learning for GUI AgentsabstractZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang |
ACL (1) | 11 |
| 2026 | Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsabstractHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou, Yang Zhang, Longtao Huang, Hui Xue, Yongliang Shen, Weiming Lu, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haolei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 10 |
| 2026 | Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive TasksabstractWenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenqi Zhang 0001, Mengna Wang, Gangao Liu, Huixin Xu, Yongliang Shen 0001, Guiyang Hou, Xin Li 0056, Weiming Lu 0001, Peng Li 0031, Yueting Zhuang |
ACL (1) | 14 |
| 2026 | MSR-Net: Multi-model Synergistic Residual Network for Efficient Large-Small Model Collaboration
Guangyu Dai, Siliang Tang, Yueting Zhuang |
ICIC (9) | 3 |
| 2026 | Towards Meta-Cognitive Knowledge Editing for Multimodal LLMsabstractKnowledge editing enables multimodal large language models (MLLMs) to efficiently update outdated or incorrect information. However, existing benchmarks primarily emphasize cognitive-level modifications while lacking a focus on deeper meta-cognitive processes. To bridge this gap, we introduce CogEdit, a novel benchmark designed to evaluate MLLMs' meta-cognitive knowledge editing abilities across three levels: (1) Counterfactual-Driven Editing, assessing self-awareness of knowledge correctness changes; (2) Boundary Constraint Editing, ensuring appropriate generalization without unintended interference; and (3) Noise-Robust Editing, promoting reflective evaluation of uncertain information. To advance meta-cognitive editing, we propose MIND (Meta-cognitive INtegrated Dynamic Knowledge Editing), a framework that constructs a meta-knowledge memory for self-awareness, employs game-theoretic interactions to monitor knowledge activation, and incorporates label refinement for noise-robust updates. Extensive experiments show that MIND significantly outperforms existing cognitive editing approaches, achieving strong performance on both traditional and meta-cognitive knowledge editing benchmarks. Zhaoyu Fan 0002, Kaihang Pan, Mingze Zhou, Bosheng Qin, Juncheng Li 0006, Shengyu Zhang 0001, Wenqiao Zhang, Siliang Tang, Fei Wu 0001, Yueting Zhuang |
WWW | 10 |
| 2026 | Structural-temporal coupling anomaly detection with dynamic graph transformer
Chang Zong, Yueting Zhuang, Jian Shao 0001, Weiming Lu 0001 |
Data Min. Knowl. Discov. | 2 |
| 2026 | MAKIMA: Tuning-free multi-attribute open-domain video editing via mask-guided attention modulation
Wenqiao Zhang, Hongyang He, Juncheng Li 0006, Zheqi Lv, Siliang Tang, Yueting Zhuang |
Expert Syst. Appl. | 9 |
| 2026 | Improving large models with small models: Lower costs and better performance
Dong Chen 0017, Shuo Zhang 0014, Yueting Zhuang, Siliang Tang, Qidong Liu 0001, Xin Yang 0011, Mingliang Xu 0001 |
Neural Networks | 4 |
| 2026 | Exploring financial sentiment analysis via fine-tuning large language model and attributed graph neural network
Zongshen Mu, Yueting Zhuang, Jie Tan 0001, Hong Cheng 0001 |
Neural Networks | 3 |
| 2026 | Momentor++: Advancing Video Large Language Models With Fine-Grained Long Video ReasoningabstractLarge Language Models (LLMs) exhibit remarkable proficiency in understanding and managing text-based tasks.Many works try to transfer these capabilities to the video domain, which are referred to as Video-LLMs. However, current Video-LLMs can only grasp the coarse-grained semantics and are unable to efficiently handle tasks involving the comprehension or localization of specific video segments. To address these challenges, we propose Momentor, a Video-LLM designed to perform fine-grained temporal understanding tasks. To facilitate the training of Momentor, we develop an automatic data generation engine to build Moment-10M, a large-scale video instruction dataset with segment-level instruction data. Building upon the foundation of the previously published Momentor and the Moment-10M dataset, we further extend this work by introducing a Spatio-Temporal Token Consolidation (STTC) method, which can merge redundant visual tokens spatio-temporally in a parameter-free manner, thereby significantly promoting computational efficiency while preserving fine-grained visual details. We integrate STTC with Momentor to develop Momentor++ and validate its performance on various benchmarks. Momentor demonstrates robust capabilities in fine-grained temporal understanding and localization. Further, Momentor++ excels in efficiently processing and analyzing extended videos with complex events, showcasing marked advancements in handling extensive temporal contexts. Juncheng Li 0006, Minghe Gao, Xiangnan He 0001, Siliang Tang, Wei-Shi Zheng 0001, Jun Xiao 0001, Meng Wang 0001, Tat-Seng Chua, Yueting Zhuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | Generalized Visual Relation Detection With Diffusion ModelsabstractVisual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance, they are all restricted to pre-defined relation categories, while failing to consider thesemantic ambiguitycharacteristic of visual relations. Unlike objects, the appearance of visual relations is always subtle and can be described by multiple predicate words from different perspectives, e.g., “ride” can be depicted as “race” and “sit on”, from the sports and spatial position views, respectively. To this end, we propose to model visual relations as continuous embeddings, and design diffusion models to achieve generalized VRD in a conditional generative manner, termed Diff-VRD. We model the diffusion process in a latent space and generate all possible relations in the image as an embedding sequence. During the generation, the visual and text embeddings of subject-object pairs serve as conditional signals and are injected via cross-attention. After the generation, we design a subsequent matching stage to assign the relation words to subject-object pairs by considering their semantic similarities. Benefiting from the diffusion-based generative process, our Diff-VRD is able to generate visual relations beyond the pre-defined category labels of datasets. To properly evaluate this generalized VRD task, we introduce two evaluation metrics, i.e., text-to-image retrieval and SPICE PR Curve inspired by image captioning. Extensive experiments in both human-object interaction (HOI) detection and scene graph generation (SGG) benchmarks attest to the superiority and effectiveness of Diff-VRD. Kaifeng Gao, Hanwang Zhang, Jun Xiao 0001, Yueting Zhuang, Qianru Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Physically Plausible Human-Object Rendering From Sparse Views via 3D Gaussian SplattingabstractRendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applications. Existing methods often struggle to simultaneously achieve high rendering quality, physical plausibility, and computational efficiency. To address these limitations, we propose HOGS (Human-Object Rendering via 3D Gaussian Splatting), a novel framework for efficient HOI rendering with physically plausible geometric constraints from sparse views. HOGS represents both humans and objects as dynamic 3D Gaussians. Central to HOGS is a novel optimization process that operates directly on these Gaussians to enforce geometric consistency (i.e., preventing inter-penetration or floating contacts) to achieve physical plausibility. To support this core optimization under sparse-view ambiguity, our framework incorporates two pre-trained modules: an optimization-guided Human Pose Refiner for robust estimation under sparse-view occlusions, and a Human-Object Contact Predictor that efficiently identifies interaction regions to guide our novel contact and separation losses. Extensive experiments on both human-object and hand-object interaction datasets demonstrate that HOGS achieves state-of-the-art rendering quality and maintains high computational efficiency. Jun Xiao 0001, Yi Yang 0001, Yueting Zhuang, Long Chen 0016 |
IEEE Trans. Image Process. | 4 |
| 2025 | Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language ModelsabstractDiffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly matching the input text with the generated image, but due to cross-modal information asymmetry, this leads to unreliable or incomplete assessment results. Motivated by this, we introduce the Image Regeneration task in this study to assess text-to-image models by tasking the T2I model with generating an image according to the reference image. We use GPT4V to bridge the gap between the reference image and the text input for the T2I model, allowing T2I models to understand image content. This evaluation process is simplified as comparisons between the generated image and the reference image are straightforward. Two regeneration datasets spanning content-diverse and style-diverse evaluation dataset are introduced to evaluate the leading diffusion models currently available. Additionally, we present ImageRepainter framework to enhance the quality of generated images by improving content comprehension via MLLM guided iterative generation and revision. Our comprehensive experiments have showcased the effectiveness of this framework in assessing the generative capabilities of models. By leveraging MLLM, we have demonstrated that a robust T2M can produce images more closely resembling the reference image. Chutian Meng, Fan Ma, Jiaxu Miao, Yi Yang 0001, Yueting Zhuang |
AAAI | 6 |
| 2025 | TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and CompetitionabstractWhile Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) effectively address resource constraints during fine-tuning, their performance often falls short, especially in multidimensional task scenarios. To address this issue, one straightforward solution is to introduce task-specific LoRA as domain experts, leveraging the modeling of multiple capabilities of experts and thus enhancing the general capability of multi-task learning.Although promising, these additional components often add complexity to the training and inference process, contravening the efficiency that PEFT is designed to deliver. Considering this, we introduce an innovative PEFT method, TeamLoRA, consisting of a collaboration and competition module for LoRA experts, thus achieving the right balance of effectiveness and efficiency:(i) For collaboration, we introduce a novel knowledge sharing and organization mechanism designed to optimize hierarchical learning while enhancing the efficiency of model training and inference.(ii) For competition, we propose leveraging a game-theoretic interaction mechanism for experts, encouraging experts to transfer their domain-specific knowledge while facing diverse downstream tasks, thus enhancing the performance.By doing so, TeamLoRA elegantly connects the experts as a “Team” with internal collaboration and competition, enabling a faster and more accurate PEFT paradigm. Meanwhile, we curate a Comprehensive Multi-Task Evaluation (CME) benchmark to thoroughly assess the capability of multi-task learning. Experiments conducted on our CME and other benchmarks indicate the effectiveness and efficiency of TeamLoRA. Our project is available at https://github.com/DCDmllm/TeamLoRA. Tianwei Lin 0001, Wenqiao Zhang, Haoyuan Li 0002, Zhelun Yu, Wanggui He, Juncheng Li 0006, Jiannan Guo 0003, Hao Jiang 0014, Siliang Tang, Yueting Zhuang |
ACL (1) | 12 |
| 2025 | Meta-Reflection: A Feedback-Free Reflection Learning FrameworkabstractYaoke Wang, Yun Zhu, XintongBao XintongBao, Wenqiao Zhang, Suyang Dai, Kehan Chen, Wenqiang Li, Gang Huang, Siliang Tang, Yueting Zhuang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yaoke Wang, Yun Zhu 0007, XintongBao XintongBao, Wenqiao Zhang, Suyang Dai, Siliang Tang, Yueting Zhuang |
ACL (1) | 10 |
| 2025 | STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-TrainingabstractVideo Large Language Models (Video-LLMs) have recently shown strong performance in basic video understanding tasks, such as captioning and coarse-grained question answering, but struggle with compositional reasoning that requires multi-step spatio-temporal inference across object relations, interactions, and events. The hurdles to enhancing this capability include extensive manual labor, the lack of spatio-temporal compositionality in existing training data and the absence of explicit reasoning supervision. In this paper, we propose STEP, a novel graph-guided self-training method that enables Video-LLMs to generate reasoning-rich fine-tuning data from any raw videos to improve itself. Specifically, we first induce Spatio-Temporal Scene Graph (STSG) representation of diverse videos to capture fine-grained, multi-granular video semantics. Then, the STSGs guide the derivation of multi-step reasoning Question-Answer (QA) data with Chain-of-Thought (CoT) rationales. Both answers and rationales are integrated as training objective, aiming to enhance model’s reasoning abilities by supervision over explicit reasoning steps. Experimental results demonstrate the effectiveness of STEP across models of varying scales, with a significant 21.3% improvement in tasks requiring three or more reasoning steps. Furthermore, it achieves superior performance with a minimal amount of self-generated rationale-enriched training samples in both compositional reasoning and comprehensive understanding benchmarks, highlighting the broad applicability and vast potential. Haiyi Qiu, Minghe Gao, Kaihang Pan, Juncheng Li 0006, Wenjie Wang 0007, Siliang Tang, Yueting Zhuang, Tat-Seng Chua |
CVPR | 9 |
| 2025 | AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaabstractInstruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to execute complex user instructions accurately, as they are trained on low-quality data with limited editing types. We present AnyEdit, a comprehensive multi-modal instruction editing dataset, comprising 2.5 million high-quality editing pairs spanning over 20 editing types and five domains. We ensure the diversity and quality of the AnyEdit collection through three aspects: initial data diversity, adaptive editing process, and automated selection of editing results. Using the dataset, we further train a novel AnyEdit Stable Diffusion with task-aware routing and learnable task embedding for unified image editing. Comprehensive experiments on three benchmark datasets show that AnyEdit consistently boosts the performance of diffusion-based editing models. This presents prospects for developing instruction-driven image editing models that support human creativity. Wei Chow, Zhongqi Yue, Kaihang Pan, Xiaoyang Wan, Juncheng Li 0006, Siliang Tang, Hanwang Zhang, Yueting Zhuang |
CVPR | 10 |
| 2025 | VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMabstractVideo Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a largescale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities. Yuqian Yuan, Wentong Li 0001, Zesen Cheng, Boqiang Zhang, Xin Li 0056, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing |
CVPR | 10 |
| 2025 | Benchmarking Multimodal CoT Reward Model Stepwise by Visual ProgramabstractRecent advancements in reward signal usage for Large Language Models (LLMs) are remarkable. However, significant challenges exist when transitioning reward signal to the multimodal domain, including labor-intensive annotations, over-reliance on one-step rewards, and inadequate evaluation. To address these issues, we propose SVIP, a novel approach to train a step-level multi-dimensional Chain-of-Thought (CoT) reward model automatically. It generates code for solving visual tasks and transforms the analysis of code blocks into the evaluation of CoT step as training samples. Then, we train SVIP-Reward model using a multi-head attention mechanism called TriAtt-CoT. The advantages of SVIP-Reward are evident throughout the entire process of MLLM. We also introduce a benchmark for CoT reward model training and testing. Experimental results demonstrate that SVIP-Reward improves MLLM performance across training and inference-time scaling, yielding better results on benchmarks while reducing hallucinations and enhancing reasoning ability. Minghe Gao, Xuqi Liu, Zhongqi Yue, Juncheng Li 0006, Siliang Tang, Fei Wu 0001, Tat-Seng Chua, Yueting Zhuang |
ICCV | 10 |
| 2025 | Iris: Breaking GUI Complexity with Adaptive Focus and Self-RefiningabstractDigital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based agents built on Large Language Models (LLMs) often require frequent updates due to platform-specific APIs, visual agents leveraging Multimodal Large Language Models (MLLMs) offer enhanced adaptability by interacting directly with Graphical User Interfaces (GUIs). However, these agents face significant challenges in visual perception, particularly when handling high-resolution, visually complex digital environments. This paper introduces Iris, a foundational visual agent that addresses these challenges through two key innovations: Information-Sensitive Cropping (ISC) and Self-Refining Dual Learning (SRDL). ISC dynamically identifies and prioritizes visually dense regions using a edge detection algorithm, enabling efficient processing by allocating more computational resources to areas with higher information density. SRDL enhances the agent's ability to handle complex tasks by leveraging a dual-learning loop, where improvements in referring (describing UI elements) reinforce grounding (locating elements) and vice versa, all without requiring additional annotated data. Empirical evaluations demonstrate that Iris achieves state-of-the-art performance across multiple benchmarks with only 850K GUI annotations, outperforming methods using 10x more training data. These improvements further translate to significant gains in both web and OS agent downstream tasks. Zhiqi Ge, Juncheng Li 0006, Xinglei Pang, Minghe Gao, Kaihang Pan, Hao Fei 0001, Wenqiao Zhang, Siliang Tang, Yueting Zhuang |
ICCV | 10 |
| 2025 | Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessabstractInstruction tuning fine-tunes pre-trained Multi-modal Large Language Models (MLLMs) to handle real-world tasks. However, the rapid expansion of visual instruction datasets introduces data redundancy, leading to excessive computational costs. We propose a collaborative framework, DataTailor, which leverages three key principles--informativeness, uniqueness, and representativeness--for effective data selection. We argue that a valuable sample should be informative of the task, non-redundant, and represent the sample distribution (i.e., not an outlier). We further propose practical ways to score against each principle, which automatically adapts to a given dataset without tedious hyperparameter tuning. Comprehensive experiments on various benchmarks demonstrate that DataTailor achieves 101.3% of the performance of full-data fine-tuning with only 15% of the data, significantly reducing computational costs while maintaining superior results. This exemplifies the "Less is More" philosophy in MLLM development. The code and data is available in this \href{https://github.com/Yuqifan1117/DataTailor}{URL}. Zhebei Shen, Zhongqi Yue, Bosheng Qin, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang |
ICCV | 10 |
| 2025 | 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingabstractCompared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook. Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing |
ICCV | 8 |
| 2025 | What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent CapabilitiesabstractAs multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. Training on our graph-structured data shows that it improves generalization across environments. We conduct multidimensional evaluations for virtual agents, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io. Wendong Bu, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Liyunfei, Mengze Li 0001, Wei Ji 0008, Juncheng Li 0006, Siliang Tang, Yueting Zhuang |
ICML | 13 |
| 2025 | HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge AdaptationabstractWe present **HealthGPT**, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation knowledge to pre-trained Large Language Models (LLMs). This is achieved through a novel heterogeneous low-rank adaptation **(H-LoRA)** technique, which is complemented by a tailored hierarchical visual perception **(HVP)** approach and a three-stage learning strategy **(TLS)**. To effectively learn the HealthGPT, we devise a comprehensive medical domain-specific comprehension and generation dataset called **VL-Health**. Experimental results demonstrate exceptional performance and scalability
of HealthGPT in medical visual unified tasks. Our project can be accessed at https://github.com/DCDmllm/HealthGPT. Tianwei Lin 0001, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li 0002, Wanggui He, Hao Jiang 0014, Mengze Li 0001, Siliang Tang, Jun Xiao 0001, Yueting Zhuang, Beng Chin Ooi |
ICML | 14 |
| 2025 | Logic Distillation: Learning from Code Function by Function for Decision-making TasksabstractLarge language models (LLMs) have garnered increasing attention owing to their powerful comprehension and generation capabilities. Generally, larger LLMs (L-LLMs) that require paid interfaces exhibit significantly superior performance compared to smaller LLMs (S-LLMs) that can be deployed on a variety of devices. Knowledge distillation (KD) aims to empower S-LLMs with the capabilities of L-LLMs, while S-LLMs merely mimic the outputs of L-LLMs, failing to get the powerful decision-making capability for new situations. Consequently, S-LLMs are helpless when it comes to continuous decision-making tasks that require logical reasoning. To tackle the identified challenges, we propose a novel framework called Logic Distillation (LD). Initially, LD employs L-LLMs to instantiate complex instructions into discrete functions and illustrates their usage to establish a function base. Subsequently, LD fine-tunes S-LLMs based on the function base to learn the logic employed by L-LLMs in decision-making. During testing, S-LLMs will yield decision-making outcomes, function by function, based on current states. Experiments demonstrate that with the assistance of LD, S-LLMs can achieve outstanding results in continuous decision-making tasks, comparable to, or even surpassing, those of L-LLMs. The code and data for the proposed method are provided for research purposes https://github.com/Anfeather/Logic-Distillation. Dong Chen 0017, Yueting Zhuang, Siliang Tang, Qidong Liu 0001, Mingliang Xu 0001 |
IJCAI | 4 |
| 2025 | SVGenius: Benchmarking LLMs in SVG Understanding, Editing and GenerationabstractLarge Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius. Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang |
ACM Multimedia | 13 |
| 2025 | Chart-HQA: A Benchmark for Hypothetical Question Answering in ChartsabstractMultimodal Large Language Models (MLLMs) have garnered significant attention for their strong visual-semantic understanding. Most existing chart benchmarks evaluate MLLMs' ability to parse information from charts to answer questions. However, they overlook the inherent output biases of MLLMs, where models rely on their parametric memory to answer questions rather than genuinely understanding the chart content. To address this limitation, we introduce a novel Chart Hypothetical Question Answering (HQA) task, which imposes assumptions on the same question to compel models to engage in counterfactual reasoning based on the chart content. Furthermore, we introduce HAI, a human-AI interactive data synthesis approach that leverages the efficient text-editing capabilities of LLMs alongside human expert knowledge to generate diverse and high-quality HQA data at a low cost. Using HAI, we construct Chart-HQA, a challenging benchmark synthesized from publicly available data sources. Evaluation results on 18 MLLMs of varying model sizes reveal that current models face significant generalization challenges and exhibit imbalanced reasoning performance on the HQA task. Our codebase and newly generated datasets are available at https://github.com/chenxn2020/Chart-HQA. Xiangnan Chen, Yuancheng Fang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 7 |
| 2025 | EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
Sijing Li, Tianwei Lin 0001, Lingshuai Lin, Wenqiao Zhang, Xiaoda Yang, Juncheng Li 0006, Jun Xiao 0001, Yueting Zhuang, Beng Chin Ooi |
ACM Multimedia | 11 |
| 2025 | Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with BenchmarkabstractMultimodal Industrial Anomaly Detection (MIAD)-fusing 3D point clouds and 2D RGB for product defect detection-is critical to quality inspection. However, existing MIAD methods assume all modalities are available and paired, overlooking real-scenario modality-missing and risking overfitting to incomplete data. To address these, we conduct the first comprehensive study on Modality-Incomplete Industrial Anomaly Detection (MIIAD) and establish MIIAD Bench , a benchmark covering diverse missing settings. Meanwhile, we propose RADAR, a robust two-stage Robust modAlity-instructive fusing & Detecting frAmewoRk. RADAR integrates i) a Modality-Incomplete Instruction mechanism-guiding the multimodal Transformer to focus more on available modal info, and ii) a Double-Pseudo Hybrid Module to highlight unique modality combinations and reduce overfitting. Our results show RADAR outperforms prior methods markedly on MIIAD Bench. Bingchen Miao, Wenqiao Zhang, Juncheng Li 0006, Wangyu Wu, Siliang Tang, Zhaocheng Li, Jun Xiao 0001, Yueting Zhuang |
ACM Multimedia | 9 |
| 2025 | Counterfactual Evolution of Multimodal Datasets via Visual ProgrammingabstractThe rapid development of Multimodal Large Language Models (MLLMs) poses increasing demands on the diversity and complexity of multimodal datasets. Yet manual annotation pipelines can no longer keep pace. Existing augmentation methods often follow fixed rules and lack verifiable control over sample diversity and reasoning complexity. To address this, we introduce Scalable COunterfactual Program Evolution (SCOPE), a framework that uses symbolic Visual Programming to guide program evolution via counterfactual reasoning. SCOPE performs the three steps of counterfactual inference: (1) Abduction, by generating verifiable programs to model reasoning associations; (2) Action, by intervening on program structure along three axes—reasoning path, visual context, and cross-instance composition; and (3) Prediction, by categorizing evolved instances by difficulty, structure, and input multiplicity. Based on this process, we build SCOPE-Train and SCOPE-Test, evolving benchmarks with expert validation. To support training, we propose MAP, a curriculum learning strategy that aligns model capacity with sample difficulty. Experiments show that SCOPE improves reasoning performance, exposes model blind spots, and enhances visual dialog capabilities. Minghe Gao, Zhongqi Yue, Wei Ji 0008, Siliang Tang, Jun Xiao 0001, Tat-Seng Chua, Yueting Zhuang, Juncheng Li 0006 |
NeurIPS | 9 |
| 2025 | Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement LearningabstractRecent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same model. Consequently, visual comprehension does not enhance visual generation, and the reasoning mechanisms of LLMs have not been fully integrated to revolutionize image generation. In this paper, we propose to enable the collaborative co-evolution of visual comprehension and generation, advancing image generation into an iterative introspective process. We introduce a two-stage training approach: supervised fine-tuning teaches the MLLM with the foundational ability to generate genuine CoT for visual generation, while reinforcement learning activates its full potential via an exploration-exploitation trade-off. Ultimately, we unlock the Aha moment in visual generation, advancing MLLMs from text-to-image tasks to unified image generation. Extensive experiments demonstrate that our model not only excels in text-to-image generation and image editing, but also functions as a superior image semantic evaluator with enhanced visual comprehension capabilities. Project Page: \url{https://janus-pro-r1.github.io}. Kaihang Pan, Wendong Bu, Juncheng Li 0006, Yingting Wang, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
NeurIPS | 12 |
| 2025 | EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction EvolutionabstractRecent advances in reinforcement learning (RL) methods such as Grouped Relative Policy Optimization (GRPO) have strengthened the reasoning capabilities of Large Vision-Language Models (LVLMs). However, due to the inherent entanglement between visual and textual modalities, applying GRPO to LVLMs often leads to reward convergence across different responses to the same sample as training progresses, hindering effective gradient updates and causing the enhancement of chain-of-thought reasoning to stagnate or even collapse.
To address this issue, we propose a progressive instruction evolution framework, EvolvedGRPO, to gradually generate more complex questions via editing instructions in an adversarial way, progressively aligned with the model’s evolving capabilities. Specifically, we design two instruction editing strategies across modalities, incorporating incrementally increasing editing instructions and RL-based adversarial data augmentation to improve the effectiveness of model training. To address GRPO's limitations on overly difficult problems, we first train on basic subproblem versions of complex multi-modal questions in both the visual and textual modalities, progressively increasing difficulty to enable prefix-style process rewards, effectively combining the strengths of both process rewards and group-wise relative rewards. Finally, EvolvedGRPO achieves state-of-the-art performance among open-source RL models on multi-modal reasoning tasks, even approaching the closed-source GPT-4o in reasoning capabilities, and demonstrates better performance on unseen LVLM general benchmarks. The Code for EvolvedGRPO is available at https://github.com/SHENZHEBEI/EvolvedGRPO. Zhebei Shen, Juncheng Li 0006, Wei Ji 0008, Siliang Tang, Yueting Zhuang |
NeurIPS | 7 |
| 2025 | Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought TuningabstractLarge language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from **Thought Leaps** due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called **ScaleQM+**, based on the structured ScaleQuestMath dataset, and trained **CoT-Bridge** to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87\% on NuminaMath. Our approach effectively enhances distilled data (+3.02\%) and provides better starting points for reinforcement learning (+3.1\%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrates improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits. Haolei Xu, Yongliang Shen 0001, Wenqi Zhang 0001, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 10 |
| 2025 | EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?abstractThe emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions.capabilities in object-level spatiotemporal reasoning required for real-world interactions.To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios.Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types.To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation frameworkBased on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems. Yuqian Yuan, Ronghao Dang, Wentong Li 0001, Xin Li 0056, Deli Zhao, Fan Wang 0019, Wenqiao Zhang, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 11 |
| 2025 | Let LRMs Break Free from Overthinking via Self-Braking TuningabstractLarge reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking.
Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions.
In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point.
Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models. Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 10 |
| 2025 | From Easy to Hard: Learning Curricular Shape-Aware Features for Robust Panoptic Scene Graph Generation
Hanrong Shi, Lin Li 0065, Jun Xiao 0001, Yueting Zhuang, Long Chen 0016 |
Int. J. Comput. Vis. | 4 |
| 2025 | Learning Combinatorial Prompts for Universal Controllable Image CaptioningabstractAbstract Controllable Image Captioning (CIC)—generating natural language descriptions about images under the guidance of given control signals—is one of the most promising directions toward next-generation captioning systems. Till now, various kinds of control signals for CIC have been proposed, ranging from content-related control to structure-related control. However, due to the format and target gaps of different control signals, all existing CIC works (or architectures) only focus on one certain control signal, and overlook the human-like combinatorial ability. By “combinatorial", we mean that our humans can easily meet multiple needs (or constraints) simultaneously when generating descriptions. To this end, we propose a novel prompt-based framework for CIC by learning Com binatorial Pro mpts, dubbed as ComPro . Specifically, we directly utilize a pretrained language model GPT-2 Radford et al. (OpenAI blog 1:9, 2019) as our language model, which can help to bridge the gap between different signal-specific CIC architectures. Then, we reformulate the CIC as a prompt-guide sentence generation problem, and propose a new lightweight prompt generation network to generate the combinatorial prompts for different kinds of control signals. For different control signals, we further design a new mask attention mechanism to realize the prompt-based CIC. Due to its simplicity, our ComPro can be further extended to more kinds of combined control signals by concatenating these prompts. Extensive experiments on two prevalent CIC benchmarks have verified the effectiveness and efficiency of our ComPro on both single and combined control signals. Zhen Wang 0004, Jun Xiao 0001, Yueting Zhuang, Fei Gao 0014, Jian Shao 0001, Long Chen 0016 |
Int. J. Comput. Vis. | 3 |
| 2025 | Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion ModelsabstractWe study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world. Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang |
IEEE Trans. Big Data | 9 |
| 2025 | Improving Vision Anomaly Detection With the Guidance of Language ModalityabstractRecent years have seen a surge of interest in anomaly detection. However, existing unsupervised anomaly detectors, particularly those for the vision modality, face significant challenges due to redundant information and sparse latent space. In contrast, anomaly detectors demonstrate superior performance in the language modality due to the unimodal nature of the data. This paper tackles the aforementioned challenges for vision modality from a multimodal point of view. Specifically, we propose Cross-modal Guidance (CMG), comprising of Cross-modal Entropy Reduction (CMER) and Cross-modal Linear Embedding (CMLE), to address the issues of redundant information and sparse latent space, respectively. CMER involves masking portions of the raw image and computing the matching score with the corresponding text. Essentially, CMER eliminates irrelevant pixels to direct the detector's focus towards critical content. To learn a more compact latent space for the vision anomaly detection, CMLE learns a correlation structure matrix from the language modality. Then, the acquired matrix compels the distribution of images to resemble that of texts in the latent space. Extensive experiments demonstrate the effectiveness of the proposed methods. Particularly, compared to the baseline that only utilizes images, the performance of CMG has been improved by 16.81%. Ablation experiments further confirm the synergy among the proposed CMER and CMLE, as each component depends on the other to achieve optimal performance. Dong Chen 0017, Kaihang Pan, Guangyu Dai, Guoming Wang, Yueting Zhuang, Siliang Tang, Mingliang Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Truncate Diffusion: Efficient Video Editing With Low-Rank TruncateabstractDiffusion models have demonstrated remarkable capabilities for text-to-video (T2V) editing tasks, relying on fine-tuning for pretrained text-to-image (T2I) diffusion models with only one video-prompt pair. However, conventional fine-tuning approaches require tuning and storing numerous parameters for each video, leading to substantial parameter and memory costs. To mitigate these issues, we propose Truncating Diffusion, an efficient fine-tuning method for video editing that optimizes both parameter and memory usage. Specifically, we propose the Truncating Diffusion module, which is designed with a focus on module architecture and initialization, specifically targeting optimization with a small training set, such as a single video-prompt pair. Theoretical analysis using the Johnson-Lindenstrauss lemma and the Eckart-Young-Mirsky theorem shows that Truncating Diffusion can achieve a minimal Frobenius norm distance to the original attention algorithm with appropriate initialization, which enhances the ease of optimization and improves video editing performance. During fine-tuning, the Truncating Diffusion module integrates seamlessly with the original diffusion model. We freeze the weights of the denoising network within the original pretrained diffusion model, updating only the introduced low-rank alternatives to ensure parameter and memory efficiency. Additionally, we propose Latent Flow Loss and Bidirectional Inter-Frame Attention (BIFA) to improve temporal consistency in synthesized videos. The Latent Flow Loss leverages global temporal information from the input video during training, while BIFA utilizes local temporal information from adjacent frames during inference. These enhancements do not incur additional memory or parameter costs during fine-tuning. Comparisons with state-of-the-art approaches demonstrate that Truncating Diffusion provides superior text alignment, video quality, and inter-frame temporal consistency in video editing. Importantly, Truncating Diffusion requires fine-tuning only 3.2% of the parameters and uses just 62% of the memory compared to the baseline model. Bosheng Qin, Wentao Ye, Wenqiao Zhang, Siliang Tang, Yueting Zhuang |
IEEE Trans. Multim. | 7 |
| 2025 | FADngs: Federated Learning for Anomaly DetectionabstractWith the increasing demand for data privacy, federated learning (FL) has gained popularity for various applications. Most existing FL works focus on the classification task, overlooking those scenarios where anomaly detection may also require privacy-preserving. Traditional anomaly detection algorithms cannot be directly applied to the FL setting due to false and missing detection issues. Moreover, with common aggregation methods used in FL (e.g., averaging model parameters), the global model cannot keep the capacities of local models in discriminating anomalies deviating from local distributions, which further degrades the performance. For the aforementioned challenges, we propose Federated Anomaly Detection with Noisy Global Density Estimation, and Self-supervised Ensemble Distillation (FADngs). Specifically, FADngs aligns the knowledge of data distributions from each client by sharing processed density functions. Besides, FADngs trains local models in an improved contrastive learning way that learns more discriminative representations specific for anomaly detection based on the shared density functions. Furthermore, FADngs aggregates capacities by ensemble distillation, which distills the knowledge learned from different distributions to the global model. Our experiments demonstrate that the proposed method significantly outperforms state-of-the-art federated anomaly detection methods. We also empirically show that the shared density function is privacy-preserving. The code for the proposed method is provided for research purposes https://github.com/kanade00/Federated_Anomaly_detection. Boyu Dong, Dong Chen 0017, Yu Wu 0011, Siliang Tang, Yueting Zhuang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | DBA: Efficient Transformer With Dynamic Bilinear Low-Rank AttentionabstractMany studies have aimed to improve Transformer model efficiency using low-rank-based methods that compress sequence length with predetermined or learned compression matrices. However, these methods fix compression coefficients for tokens in the same position during inference, ignoring sequence-specific variations. They also overlook the impact of hidden state dimensions on efficiency gains. To address these limitations, we propose dynamic bilinear low-rank attention (DBA), an efficient and effective attention mechanism that compresses sequence length using input-sensitive dynamic compression matrices. DBA achieves linear time and space complexity by jointly optimizing sequence length and hidden state dimension while maintaining state-of-the-art performance. Specifically, we demonstrate through experiments and the properties of low-rank matrices that sequence length can be compressed with compression coefficients dynamically determined by the input sequence. In addition, we illustrate that the hidden state dimension can be approximated by extending the Johnson-Lindenstrauss lemma, thereby introducing only a small amount of error. DBA optimizes the attention mechanism through bilinear forms that consider both the sequence length and hidden state dimension. Moreover, the theoretical analysis substantiates that DBA excels at capturing high-order relationships in cross-attention problems. Experimental results across different tasks with varied sequence length conditions demonstrate that DBA achieves state-of-the-art performance compared to several robust baselines. DBA also maintains higher processing speed and lower memory usage, highlighting its efficiency and effectiveness across diverse applications. Bosheng Qin, Juncheng Li 0006, Siliang Tang, Yueting Zhuang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better PerformanceabstractPretrained large models, particularly large language models, have garnered increasing attention, as they have demonstrated remarkable abilities through contextual learning. Pretrained large models are increasingly recognized as fundamental tools for solving various tasks. However, the substantial computational demands of large models have dissuaded most product teams and individuals from running them. In such scenarios, to leverage the exceptional performance of large models, one must solely depend on costly APIs, further burdening product teams and individuals. On the other hand, despite the overall inferior performance of small models compared to large models, there are certain distributions where small models can achieve comparable or even superior results. For instance, during training, small models may become trapped in a local optimum that is unique to certain distributions, leading to superior performance. Hence, we propose Data Shunt (DS), a general paradigm for collaboration of small and large models. DS not only substantially reduces the cost associated with deploying large models but also effectively enhances overall performance. Specifically, DS determines the shunting direction by evaluating the confidence level of small models. When the confidence level falls below a specific threshold, the input data is forwarded to large models. To further leverage the advantages of the small and large models, we introduce Prompt Pruning (PP) and 2-Stage Confidence Distillation (2CD), which facilitate mutual collaboration, leading to better results and less cost. The remarkable performance across diverse modalities and tasks demonstrates the superiority of the proposed DS over large models. For instance, ChatGPT achieves an accuracy of 94.43% on Amazon Product sentiment analysis, and DS achieves an accuracy of 95.64%, while the cost has been reduced to only 31.18%. The code for the proposed method are provided for research purposes https://github.com/Anfeather/Data-Shunt. Dong Chen 0017, Yueting Zhuang, Shuo Zhang 0014, Jinfeng Liu 0007, Su Dong 0002, Siliang Tang |
AAAI | 2 |
| 2024 | Learning Global Controller in Latent Space for Parameter-Efficient Fine-TuningabstractZeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong, Wenqi Zhang, Jian Shao, Weiming Lu, Yueting Zhuang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zeqi Tan, Yongliang Shen 0001, Xiaoxia Cheng, Chang Zong, Wenqi Zhang 0001, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 8 |
| 2024 | T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from TextabstractIn this work, we propose a two-stage sign language production (SLP) paradigm that first encodes sign language sequences into discrete codes and then autoregressively generates sign language from text based on the learned codebook. However, existing vector quantization (VQ) methods are fixed-length encodings, overlooking the uneven information density in sign language, which leads to under-encoding of important regions and over-encoding of unimportant regions. To address this issue, we propose a novel dynamic vector quantization (DVA-VAE) model that can dynamically adjust the encoding length based on the information density in sign language to achieve accurate and compact encoding. Then, a GPT-like model learns to generate code sequences and their corresponding durations from spoken language text. Extensive experiments conducted on the PHOENIX14T dataset demonstrate the effectiveness of our proposed method. To promote sign language research, we propose a new large German sign language dataset, PHOENIX-News, which contains 486 hours of sign language videos, audio, and transcription texts. Experimental analysis on PHOENIX-News shows that the performance of our model can be further improved by increasing the size of the training data. Our project homepage is https://t2sgpt-demo.yinaoxiong.cn. Aoxiong Yin, Haoyuan Li 0002, Siliang Tang, Yueting Zhuang |
ACL (1) | 5 |
| 2024 | Self-Contrast: Better Reflection Through Inconsistent Solving PerspectivesabstractWenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001 |
ACL (1) | 6 |
| 2024 | Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationabstractWenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenqi Zhang 0001, Mengna Wang, Yongliang Shen 0001, Guiyang Hou, Zeqi Tan, Peng Li 0031, Yueting Zhuang, Weiming Lu 0001 |
ACL (1) | 9 |
| 2024 | HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction DataabstractMulti-modal Large Language Models (MLLMs) tuned on machine-generated instruction-following data have demonstrated remarkable performance in various multi-modal understanding and generation tasks. However, the hallucinations inherent in machine-generated data, which could lead to hallucinatory outputs in MLLMs, remain under-explored. This work aims to investigate various hallucinations (i.e., object, relation, attribute hallucinations) and mitigate those hallucinatory toxicities in large-scale machine-generated visual instruction datasets. Drawing on the human ability to identify factual errors, we present a novel hallucination detection and elimination framework, HalluciDoctor, based on the cross-checking paradigm. We use our framework to identify and eliminate hallucinations in the training data automatically. Interestingly, HalluciDoctor also indicates that spurious correlations arising from long-tail object cooccurrences contribute to hallucinations. Based on that, we execute counterfactual visual instruction expansion to balance data distribution, thereby enhancing MLLMs' resistance to hallucinations. Comprehensive experiments on hallucination evaluation benchmarks show that our method successfully mitigates 44.6% hallucinations relatively and maintains competitive performance compared to LLaVA. The data and code for this paper are publicly available.11https://github.com/Yuqifan1117/HalluciDoctor Juncheng Li 0006, Longhui Wei, Liang Pang 0001, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian 0001, Yueting Zhuang |
CVPR | 9 |
| 2024 | Bridging Local Details and Global Context in Text-Attributed GraphsabstractRepresentation learning on text-attributed graphs (TAGs) is vital for real-world applications, as they combine semantic textual and contextual structural information.Research in this field generally consist of two main perspectives: local-level encoding and global-level aggregating, respectively refer to textual node information unification (e.g., using Language Models) and structure-augmented modeling (e.g., using Graph Neural Networks).Most existing works focus on combining different information levels but overlook the interconnections, i.e., the contextual textual information among nodes, which provides semantic insights to bridge local and global levels.In this paper, we propose GraphBridge, a multi-granularity integration framework that bridges local and global perspectives by leveraging contextual textual information, enhancing fine-grained understanding of TAGs.Besides, to tackle scalability and efficiency challenges, we introduce a graph-aware token reduction module.Extensive experiments across various models and datasets show that our method achieves state-of-the-art performance, while our graph-aware token reduction module significantly enhances efficiency and solves scalability issues.Codes are available at https://github.com/wykk00/GraphBridge Yaoke Wang, Yun Zhu 0007, Wenqiao Zhang, Yueting Zhuang, Liyunfei, Siliang Tang |
EMNLP | 4 |
| 2024 | Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language ModelabstractWenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wenqi Zhang 0001, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen 0001, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu 0001, Yueting Zhuang |
EMNLP | 11 |
| 2024 | Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question AnsweringabstractRecent progress with LLM-based agents has shown promising results across various tasks.However, their use in answering questions from knowledge bases remains largely unexplored.Implementing a KBQA system using traditional methods is challenging due to the shortage of task-specific training data and the complexity of creating task-focused model structures.In this paper, we present Triad, a unified framework that utilizes an LLM-based agent with multiple roles for KBQA tasks.The agent is assigned three roles to tackle different KBQA subtasks: agent as a generalist for mastering various subtasks, as a decision maker for the selection of candidates, and as an advisor for answering questions with knowledge.Our KBQA framework is executed in four phases, involving the collaboration of the agent's multiple roles.We evaluated the performance of our framework using three benchmark datasets, and the results show that our framework outperforms state-of-the-art systems on the LC-QuAD and YAGO-QA benchmarks, yielding F1 scores of 11.8% and 20.7%, respectively. Chang Zong, Weiming Lu 0001, Jian Shao 0001, Heng Chang, Yueting Zhuang |
EMNLP | 7 |
| 2024 | An Optimal Transport-Based Method For Medical Image GenerationabstractOptimal transport (OT) is an effective technique for mapping between two distributions. Introducing a definition of distance between the distributions, OT provides the best calculable mapping plan, especially from an irregular distribution to a regular one. For this reason, OT has found wide applications, including deep learning, to mitigate some inherent risks in existing techniques. For example, generative adversarial networks (GANs) are widely used to generate medical images but sometimes experience mode collapse, a phenomenon that can be avoided by using OT. We use OT to map latent space distribution in a GAN model to a reasonable Gaussian distribution and generate more convincing and complete medical images to help train medical AI models. Our method consists of three steps. First, we train an auto-encoder as a basic structure of our model. Second, we introduce OT to the trained auto-encoder, which aligns the latent space distribution with a Gaussian distribution. This mapping helps to enforce a more realistic distribution of the generated images, enhancing their clinical relevance. Third, we construct a GAN-type model by training the decoder part of the auto-encoder and discriminator networks. We conducted a comprehensive experimental evaluation on a diverse medical dataset by comparing the performance of our proposed method to other state-of-the-art models. Results show that our method performs relatively better than other models by FID scores and provides more detailed information. Bohan Lei, Yueting Zhuang, Xiaoyin Xu |
ICIP | 2 |
| 2024 | Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsabstractRecent advancements in Multimodal Large Language Models (MLLMs) have been utilizing Visual Prompt Generators (VPGs) to convert visual features into tokens that LLMs can recognize. This is achieved by training the VPGs on millions of image-caption pairs, where the VPG-generated tokens of images are fed into a frozen LLM to generate the corresponding captions. However, this image-captioning based training objective inherently biases the VPG to concentrate solely on the primary visual contents sufficient for caption generation, often neglecting other visual details. This shortcoming results in MLLMs’ underperformance in comprehending demonstrative instructions consisting of multiple, interleaved, and multimodal instructions that demonstrate the required context to complete a task. To address this issue, we introduce a generic and lightweight Visual Prompt Generator Complete module (VPG-C), which can infer and complete the missing details essential for comprehending demonstrative instructions. Further, we propose a synthetic discriminative training strategy to fine-tune VPG-C, eliminating the need for supervised demonstrative instructions. As for evaluation, we build DEMON, a comprehensive benchmark for demonstrative instruction understanding. Synthetically trained with the proposed strategy, VPG-C achieves significantly stronger zero-shot performance across all tasks of DEMON. Further evaluation on the MME and OwlEval benchmarks also demonstrate the superiority of VPG-C. The code and models are available at https://github.com/DCDmllm/Cheetah. Juncheng Li 0006, Kaihang Pan, Zhiqi Ge, Minghe Gao, Wei Ji 0008, Wenqiao Zhang, Tat-Seng Chua, Siliang Tang, Hanwang Zhang, Yueting Zhuang |
ICLR | 10 |
| 2024 | InstructVid2Vid: Controllable Video Editing with Natural Language InstructionsabstractWe introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for per-example fine-tuning or inversion. The proposed InstructVid2Vid model modifies a pretrained image generation model, Stable Diffusion, to generate a time-dependent sequence of video frames. By harnessing the collective intelligence of disparate models, we engineer a training dataset rich in video-instruction triplets, which is a more cost-efficient alternative to collecting data in real-world scenarios. To enhance the coherence between successive frames within the generated videos, we propose the Inter-Frames Consistency Loss and incorporate it during the training process. With multimodal classifier-free guidance during the inference stage, the generated videos is able to resonate with both the input video and the accompanying instructions. Experimental results demonstrate that InstructVid2Vid is capable of generating high-quality, temporally coherent videos and performing diverse edits, including attribute editing, background changes, and style transfer. These results underscore the versatility and effectiveness of our proposed method. Bosheng Qin, Juncheng Li 0006, Siliang Tang, Tat-Seng Chua, Yueting Zhuang |
ICME | 5 |
| 2024 | Auto-Encoding Morph-Tokens for Multimodal LLMabstractFor multimodal LLMs, the synergy of visual comprehension (textual output) and generation (visual output) presents an ongoing challenge. This is due to a conflicting objective: for comprehension, an MLLM needs to abstract the visuals; for generation, it needs to preserve the visuals as much as possible. Thus, the objective is a dilemma for visual-tokens. To resolve the conflict, we propose encoding images into morph-tokens to serve a dual purpose: for comprehension, they act as visual prompts instructing MLLM to generate texts; for generation, they take on a different, non-conflicting role as complete visual-tokens for image reconstruction, where the missing visual cues are recovered by the MLLM. Extensive experiments show that morph-tokens can achieve a new SOTA for multimodal comprehension and generation simultaneously. Our project is available at https://github.com/DCDmllm/MorphTokens. Kaihang Pan, Siliang Tang, Juncheng Li 0006, Zhaoyu Fan 0002, Wei Chow, Shuicheng Yan, Tat-Seng Chua, Yueting Zhuang, Hanwang Zhang |
ICML | 8 |
| 2024 | Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningabstractLarge Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. However, existing Video-LLMs can only capture the coarse-grained semantics and are unable to effectively handle tasks related to comprehension or localization of specific video segments. In light of these challenges, we propose Momentor, a Video-LLM capable of accomplishing fine-grained temporal understanding tasks. To support the training of Momentor, we design an automatic data generation engine to construct Moment-10M, a large-scale video instruction dataset with segment-level instruction data. We train Momentor on Moment-10M, enabling it to perform segment-level reasoning and localization. Zero-shot evaluations on several tasks demonstrate that Momentor excels in fine-grained temporally grounded comprehension and localization. Juncheng Li 0006, Yu Wu 0011, Yaobo Ye, Hao Fei 0001, Tat-Seng Chua, Yueting Zhuang, Siliang Tang |
ICML | 7 |
| 2024 | De-fine: Decomposing and Refining Visual Programs with Auto-FeedbackabstractVisual programming, a modular paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that need task-specific data, it performs visual processing and inference in an unsupervised manner. Current visual programming methods generate programs in a single pass where the ability to evaluate and optimize based on feedback, unfortunately, is lacking, which consequentially limits their effectiveness for complex, multi-step problems. Inspired by benders decomposition, we introduce De-fine, a training-free framework that automatically decomposes complex tasks into simpler subtasks and refines programs through auto-feedback. This model-agnostic approach can improve logical reasoning performance by integrating the strengths of multiple models. Our experiments across various visual tasks show that De-fine creates more robust programs. Moreover, viewing each feedback module as an independent agent will yield fresh prospects for the field of agent research. Minghe Gao, Juncheng Li 0006, Hao Fei 0001, Liang Pang 0001, Wei Ji 0008, Guoming Wang, Zheqi Lv, Wenqiao Zhang, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 10 |
| 2024 | Fact : Teaching MLLMs with Faithful, Concise and Transferable RationalesabstractThe remarkable performance of Multimodal Large Language Models (MLLMs) has demonstrated their proficient understanding capabilities in handling various visual tasks. Nevertheless, the opaque nature of black-box reasoning processes persists as an enigma, rendering them uninterpretable and struggling with hallucination. Their ability to execute intricate reasoning tasks is also constrained, culminating in stagnation of progression. In this work, we introduce Fact, a novel paradigm designed to generate multimodal rationales that are faithful, concise, and transferable for teaching MLLMs. This paradigm utilizes verifiable visual programming to generate executable code guaranteeing faithfulness. Through a series of operations including pruning, merging, and bridging, the rationale enhances its conciseness. Furthermore, we filter rationales that can be transferred to end-to-end paradigms from programming paradigms to guarantee transferability. Empirical evidence from experiments demonstrates the superiority of Fact across models of varying parameter sizes, significantly enhancing their compositional reasoning and generalization ability and reducing hallucinations owing to its high correlation between images and text. Minghe Gao, Liang Pang 0001, Yuan Yao 0013, Jisheng Dang, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang, Tat-Seng Chua |
ACM Multimedia | 9 |
| 2024 | DEMON24: ACM MM24 Demonstrative Instruction Following ChallengeabstractWe introduce the DEMON Challenge, defined as a benchmark for demonstrative instruction following, to ACM Multimedia 2024. The DEMON Challenge aims to assess the ability of models and systems to comprehend demonstrative instructions consisting of multiple, interleaved, and multimodal context that demonstrate the required information to complete a task. These instructions are curated from a diverse range of multi-modal datasets, spanning various fields and scenarios, to ensure comprehensive coverage and challenge diversity. The challenge details and participation information are available on the https://dcdmllm.github.io/DEMON-challenge/. Zhiqi Ge, Juncheng Li 0006, Wei Zhou 0101, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 6 |
| 2024 | WorldGPT: Empowering LLM as Multimodal World ModelabstractWorld models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and confined to single-modality state representations. In this paper, We introduce WorldGPT, a generalist world model built upon Multimodal Large Language Model (MLLM). WorldGPT acquires an understanding of world dynamics through analyzing millions of videos across various domains. To further enhance WorldGPT's capability in specialized scenarios and long-term tasks, we have integrated it with a novel cognitive architecture that combines memory offloading, knowledge retrieval, and context reflection. As for evaluation, we build WorldNet, a multimodal state transition prediction benchmark encompassing varied real-life scenarios. Conducting evaluations on WorldNet directly demonstrates WorldGPT's capability to accurately model state transition patterns, affirming its effectiveness in understanding and predicting the dynamics of complex scenarios. We further explore WorldGPT's emerging potential in serving as a world simulator, helping multimodal agents generalize to unfamiliar domains through efficiently synthesising multimodal instruction instances which are proved to be as reliable as authentic data for fine-tuning purposes. The code and dataset are available on the https://github.com/DCDmllm/WorldGPT Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li 0006, Guoming Wang, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 7 |
| 2024 | TaskBench: Benchmarking Large Language Models for Task AutomationabstractIn recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents. Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Kan Ren, Weiming Lu 0001, Dongsheng Li 0002, Yueting Zhuang |
NeurIPS | 9 |
| 2024 | Contrastive Hawkes graph neural networks with dynamic sampling for event prediction
Zongshen Mu, Yueting Zhuang, Siliang Tang |
Neurocomputing | 2 |
| 2024 | Position-aware compositional embeddings for compressed recommendation systems
Zongshen Mu, Yueting Zhuang, Siliang Tang |
Neurocomputing | 2 |
| 2024 | Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
Knowl. Based Syst. | 6 |
| 2024 | Ask Questions With Double Hints: Visual Question Generation With Answer-Awareness and Region-ReferenceabstractThe visual question generation (VQG) task aims to generate human-like questions from an image and potentially other side information (e.g., answer type). Previous works on VQG fall in two aspects: i) They suffer from one image to many questions mapping problem, which leads to the failure of generating referential and meaningful questions from an image. ii) They fail to model complex implicit relations among the visual objects in an image and also overlook potential interactions between the side information and image. To address these limitations, we first propose a novel learning paradigm to generate visual questions with answer-awareness and region-reference. Concretely, we aim to ask the right visual questions with Double Hints - textual answers and visual regions of interests, which could effectively mitigate the existing one-to-many mapping issue. Particularly, we develop a simple methodology to self-learn the visual hints without introducing any additional human annotations. Furthermore, to capture these sophisticated relationships, we propose a new double-hints guided Graph-to-Sequence learning framework, which first models them as a dynamic graph and learns the implicit topology end-to-end, and then utilizes a graph-to-sequence model to generate the questions with double hints. Experimental results demonstrate the priority of our proposed method. Lingfei Wu 0001, Siliang Tang, Fangli Xu, Bo Long, Yueting Zhuang, Jian Pei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | IDPro: Flexible Interactive Video Object Segmentation by ID-Queried Concurrent PropagationabstractInteractive Video Object Segmentation (iVOS) is inherently demanding, requiring real-time interaction between humans and computers. Enhancing user experience involves considerations such as user input habits, segmentation quality, running time, and memory consumption. However, existing methods compromise user experience by employing a single input mode and exhibiting slow running speeds. Specifically, these approaches restrict user interaction to a single frame, limiting the expression of user intent. To overcome these limitations and better align with user habits, we introduce a framework that facilitates flexible input modes by ID-queried concurrent propagation (IDPro). In particular, we have devised the Across-Frame Interaction Module (AFI), allowing users to freely annotate various objects across multiple frames. The AFI module transfers scribble information across interactive frames, generating multi-frame masks. Additionally, we leverage an id-queried mechanism to process multiple objects. To achieve more efficient propagation and a lightweight model, we propose a truncated re-propagation strategy, replacing the previous multi-round fusion module, which employs an across-round memory that stores crucial interaction information. Our SwinB-IDPro attains a new state-of-the-art performance on DAVIS 2017 (89.6%,${\mathcal {J}}\& {\mathcal {F}}\text{@}60$). Furthermore, our R50-IDPro exhibits over${3 \times }$faster performance than the leading competitor in challenging multi-object scenarios. Tao Jiang 0042, Zongxin Yang, Yi Yang 0001, Yueting Zhuang, Jun Xiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Improving Reference-Based Distinctive Image Captioning with Contrastive RewardsabstractDistinctive Image Captioning (DIC)—generating distinctive captions that describe the unique details of a target image—has received considerable attention over the last few years. A recent DIC method proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-Based DIC (Ref-DIC). It aims to force the generated captions to distinguish between the target image and the reference image. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects “ towel ” and “ toilet ” while all reference images are without them, then a simple caption “ A bathroom with a towel and a toilet ” is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at the object-/attribute-level (vs. scene-level). Second, to generate distinctive captions, we develop a Transformer-based Ref-DIC baseline TransDIC . It not only extracts visual features from the target image but also encodes the differences between objects in the target and reference images. Taking one step further, we propose a stronger TransDIC \({++}\) , which consists of an extra contrastive learning module to make full use of the reference images. This new module is model-agnostic, which can be easily incorporated into various Ref-DIC architectures. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC \({++}\) can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics. Yangjun Mao, Jun Xiao 0001, Meng Cao 0002, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | DiffusionNER: Boundary Diffusion for Named Entity RecognitionabstractIn this paper, we propose DIFFUSIONNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans.During training, DIFFUSIONNER gradually adds noises to the golden entity boundaries by a fixed forward diffusion process and learns a reverse diffusion process to recover the entity boundaries.In inference, DIFFU-SIONNER first randomly samples some noisy spans from a standard Gaussian distribution and then generates the named entities by denoising them with the learned reverse diffusion process.The proposed boundary-denoising diffusion process allows progressive refinement and dynamic sampling of entities, empowering DIFFUSIONNER with efficient and flexible entity generation capability.Experiments on multiple flat and nested NER datasets demonstrate that DIFFUSIONNER achieves comparable or even better performance than previous state-of-the-art models 1 . Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 6 |
| 2023 | PromptNER: Prompt Locating and Typing for Named Entity RecognitionabstractYongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang, Rongsheng Zhang, Yadong Xi, Weiming Lu, Yueting Zhuang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yongliang Shen 0001, Zeqi Tan, Shuhui Wu, Wenqi Zhang 0001, Yadong Xi, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 8 |
| 2023 | Unsupervised Prompt Tuning for Text-Driven Object DetectionabstractGrounded language-image pre-trained models have shown strong zero-shot generalization to various downstream object detection tasks. Despite their promising performance, the models rely heavily on the laborious prompt engineering. Existing works typically address this problem by tuning text prompts using downstream training data in a few-shot or fully supervised manner. However, a rarely studied problem is to optimize text prompts without using any annotations. In this paper, we delve into this problem and propose an Unsupervised Prompt Tuning framework for text-driven object detection, which is composed of two novel mean teaching mechanisms. In conventional mean teaching, the quality of pseudo boxes is expected to optimize better as the training goes on, but there is still a risk of overfitting noisy pseudo boxes. To mitigate this problem, 1) we propose Nested Mean Teaching, which adopts nested-annotation to supervise teacher-student mutual learning in a bi-level optimization manner; 2) we propose Dual Complementary Teaching, which employs an offline pre-trained teacher and an online mean teacher via data-augmentation-based complementary labeling so as to ensure learning without accumulating confirmation bias. By integrating these two mechanisms, the proposed unsupervised prompt tuning framework achieves significant performance improvement on extensive object detection datasets. Weizhen He, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Donglian Qi, Yueting Zhuang |
ICCV | 8 |
| 2023 | Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsabstractPrompt tuning, a recently emerging paradigm, enables the powerful vision-language pre-training models to adapt to downstream tasks in a parameter- and data- efficient way, by learning the "soft prompts" to condition frozen pretraining models. Though effective, it is particularly problematic in the few-shot scenario, where prompt tuning performance is sensitive to the initialization and requires a time-consuming process to find a good initialization, thus restricting the fast adaptation ability of the pre-training models. In addition, prompt tuning could undermine the generalizability of the pre-training models, because the learnable prompt tokens are easy to overfit to the limited training samples. To address these issues, we introduce a novel Gradient-RegulAted Meta-prompt learning (GRAM) framework that jointly meta-learns an efficient soft prompt initialization for better adaptation and a lightweight gradient regulating function for strong cross-domain generalizability in a meta-learning paradigm using only the unlabeled image-text pre-training data. Rather than designing a specific prompt tuning method, our GRAM can be easily incorporated into various prompt tuning methods in a model-agnostic way, and comprehensive experiments show that GRAM brings about consistent improvement for them in several settings (i.e., few-shot learning, cross-domain generalization, cross-dataset generalization, etc.) over 11 datasets. Further, experiments show that GRAM enables the orthogonal methods of textual and visual prompt tuning to work in a mutually-enhanced way, offering better generalizability beyond the uni-modal prompt tuning methods. Juncheng Li 0006, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li 0001, Wei Ji 0008, Qi Tian 0001, Tat-Seng Chua, Yueting Zhuang |
ICCV | 10 |
| 2023 | Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldabstractScene Graph Generation (SGG) aims to extractrelationships in images for vision understanding. Although recent works have made steady progress on SGG, they still suffer long-tail distribution issues that tail-predicates are more costly to train and hard to distinguish due to a small amount of annotated data compared to frequent predicates. Existing re-balancing strategies try to handle it via prior rules but are still confined to pre-defined conditions, which are not scalable for various models and datasets. In this paper, we propose a Cross-modal prediCate boosting (CaCao) framework, where a visually-prompted language model is learned to generate diverse fine-grained predicates in a low-resource way. The proposed CaCao can be applied in a plug-and-play fashion and automatically strengthen existing SGG to tackle the long-tailed problem. Based on that, we further introduce a novel Entangled cross-modal prompt approach for open-world predicate scene graph generation (Epic), where models can generalize to unseen predicates in a zero-shot manner. Comprehensive experiments on three benchmark datasets show that CaCao consistently boosts the performance of multiple scene graph generation models in a model-agnostic way. Moreover, our Epic achieves competitive performance on open-world predicate prediction. The data and code for this paper are publicly available.1 Juncheng Li 0006, Yu Wu 0011, Siliang Tang, Wei Ji 0008, Yueting Zhuang |
ICCV | 6 |
| 2023 | Learning in Imperfect Environment: Multi-Label Classification with Long-Tailed Distribution and Partial LabelsabstractConventional multi-label classification (MLC) methods assume that all samples are fully labeled and identically distributed. Unfortunately, this assumption is unrealistic in large-scale MLC data that has long-tailed (LT) distribution and partial labels (PL). To address the problem, we introduce a novel task, Partial labeling and Long-Tailed Multi-Label Classification (PLT-MLC), to jointly consider the above two imperfect learning environments. Not surprisingly, we find that most LT-MLC and PL-MLC approaches fail to solve the PLT-MLC, resulting in significant performance degradation on the two proposed PLT-MLC benchmarks. Therefore, we propose an end-to-end learning framework: COrrection → ModificatIon → balanCe, abbreviated as COMIC. Our bootstrapping philosophy is to simultaneously correct the missing labels (Correction) with convinced prediction confidence over a class-aware threshold and to learn from these recall labels during training. We next propose a novel multi-focal modifier loss that simultaneously addresses head-tail imbalance and positive-negative imbalance to adaptively modify the attention to different samples (Modification) under the LT class distribution. In addition, we develop a balanced training strategy by distilling the model’s learning effect from head and tail samples, and thus design a balanced classifier (Balance) conditioned on the head and tail learning effect to maintain stable performance for all samples. Our experimental study shows that the proposed COMIC significantly outperforms general MLC, LT-MLC and PL-MLC methods in terms of effectiveness and robustness on our newly created PLT-MLC datasets. Codes and benchmarks are available on the link https://https://github.com/wannature/COMIC Wenqiao Zhang, Changshuo Liu, Lingze Zeng, Beng Chin Ooi, Siliang Tang, Yueting Zhuang |
ICCV | 6 |
| 2023 | Continual Vision-Language Representation Learning with Off-Diagonal InformationabstractLarge-scale multi-modal contrastive learning frameworks like CLIP typically require a large amount of image-text samples for training. However, these samples are always collected continuously in real scenarios. This paper discusses the feasibility of continual CLIP training using streaming data. Unlike continual learning based on self-supervised learning methods for pure images, which is empirically robust against catastrophic forgetting, CLIP’s performance degeneration in the continual setting is significant and non-neglectable. By analyzing the changes in the model’s representation space during continual CLIP training from a spatial geometry perspective, we explore and summarize these spatial variations as Spatial Disorder (SD), which can be divided into Intra-modal Rotation and Inter-modal Deviation. Moreover, we empirically and theoretically demonstrate how SD leads to a performance decline for CLIP on cross-modal retrieval tasks. To alleviate SD, we propose a new continual vision-language representation learning framework Mod-X: Maintain off-diagonal information-matriX. By selectively aligning the off-diagonal information distribution of contrastive matrices, the Mod-X improves the capability of the multi-modal model by maintaining the multi-modal representation space alignment on the old data domain during continuously fitting the new training data domain. Experiments on commonly used datasets with different scales and scopes have demonstrated the effectiveness of our method. Zixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang, Qi Tian 0001 |
ICML | 4 |
| 2023 | FedAA: Using Non-sensitive Modalities to Improve Federated Learning while Preserving Image PrivacyabstractFederated learning aims to train a better global model without sharing the sensitive training samples (usually images) of local clients. Since the sample distributions in local clients tend to be different from each other (i.e., non-IID), one of the major challenges for federated learning is to alleviate model degradation when aggregating local models. The degradation can be attributed to the weight divergence that quantifies the difference of local models from different training processes. Furthermore, non-IID also results in feature space heterogeneity during local training, making neurons of local models in the same location have different functions and further exacerbating weight divergence. In this paper, we demonstrate that the problem can be solved by sharing information from the non-sensitive modality (e.g., metadata, non-sensitive descriptions, etc.) while keeping the sensitive information of images protected. In particular, we propose Federated Learning with Adversarial Example and Adversarial Identifier (FedAA) that trains adversarial examples based on the shared non-sensitive modality to fine-tune local models before global aggregation. The training of local models is enhanced by client identifiers that discriminate the source of inputs to force different local models to get similar outputs and be more homogeneous during the local training. Experiments show that FedAA significantly outperforms recent non-IID federated learning algorithms while preserving image privac, by sharing information from non-sensitive modalities. Dong Chen 0017, Siliang Tang, Zijin Shen, Guoming Wang, Jun Xiao 0001, Yueting Zhuang, Carl Yang 0001 |
ACM Multimedia | 6 |
| 2023 | Unsupervised Domain Adaptation for Video Object Grounding with Cascaded Debiasing LearningabstractThis paper addresses the Unsupervised Domain Adaptation (UDA) for the dense frame prediction task - Video Object Grounding (VOG). This investigation springs from the recognition of the limited generalization capabilities of data-driven approaches when confronted with unseen test scenarios. We set the goal of enhancing the adaptability of the source-dominated model from a labeled domain to the unlabeled target domain through re-training on pseudo-labels (i.e., predicted boxes of language-described objects). Given the potential for source-domain biases in the pseudo-label generation, we decompose the labeling refinement as two cascaded debiasing subroutines: (1) we develop a discarded training strategy to correct the Biased Proposal Selection by filtering out the examples with uncertain proposals selected from the proposal (candidate box) set. The identifier of these uncertain examples is the discordance between the predictions of the source-dominated model and those of a target-domain clustered classifier, which remains free from the source-domain bias. (2) With the refined proposals as a foundation, we measure Grounding Coordinate Offset based on the semantic distance of the model's prediction across domains, based on which we alleviate source-domain bias in the target model through adversarial learning. To verify the superiority of the proposed method, we collected two UDA-VOG datasets called I2O-VOG and R2M-VOG by manually dividing and combining the well-known VOG datasets. The extensive experiments on them show our model significantly outperforms SOTA methods by a large margin. Mengze Li 0001, Juncheng Li 0006, Zhou Zhao 0001, Wenqiao Zhang, Shengyu Zhang 0001, Shiliang Pu, Yueting Zhuang, Fei Wu 0001 |
ACM Multimedia | 8 |
| 2023 | Degeneration-Tuning: Using Scrambled Grid shield Unwanted Concepts from Stable DiffusionabstractOwing to the unrestricted nature of the content in the training data, large text-to-image diffusion models, such as Stable Diffusion (SD), are capable of generating images with potentially copyrighted or dangerous content based on corresponding textual concepts information. This includes specific intellectual property (IP), human faces, and various artistic styles. However, Negative Prompt, a widely used method for content removal, frequently fails to conceal this content due to inherent limitations in its inference logic. In this work, we propose a novel strategy named Degeneration-Tuning (DT) to shield contents of unwanted concepts from SD weights. By utilizing Scrambled Grid to reconstruct the correlation between undesired concepts and their corresponding image domain, we guide SD to generate meaningless content when such textual concepts are provided as input. As this adaptation occurs at the level of the model's weights, the SD, after DT, can be grafted onto other conditional diffusion frameworks like ControlNet to shield unwanted concepts. In addition to qualitatively showcasing the effectiveness of our DT method in protecting various types of concepts, a quantitative comparison of the SD before and after DT indicates that the DT method does not significantly impact the generative quality of other contents. The FID and IS scores of the model on COCO-30K exhibit only minor changes after DT, shifting from 12.61 and 39.20 to 13.04 and 38.25, respectively, which clearly outperforms the previous methods. Zixuan Ni, Longhui Wei, Jiacheng Li 0002, Siliang Tang, Yueting Zhuang, Qi Tian 0001 |
ACM Multimedia | 5 |
| 2023 | Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationabstractData-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best. Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
ACM Multimedia | 6 |
| 2023 | HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceabstractSolving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI tasks autonomously. Considering large language models (LLMs) have exhibited exceptional abilities in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks, with language serving as a generic interface to empower this. Based on this philosophy, we present HuggingGPT, an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities (e.g., Hugging Face) to solve AI tasks. Specifically, we use ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. By leveraging the strong language capability of ChatGPT and abundant AI models in Hugging Face, HuggingGPT can tackle a wide range of sophisticated AI tasks spanning different modalities and domains and achieve impressive results in language, vision, speech, and other challenging tasks, which paves a new way towards the realization of artificial general intelligence. Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang |
NeurIPS | 6 |
| 2023 | Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsabstractPretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs within an image. However, naively utilizing CLIP with prevalent class-based prompts for zero-shot VRD has several weaknesses, e.g., it struggles to distinguish between different fine-grained relation types and it neglects essential spatial information of two objects. To this end, we propose a novel method for zero-shot VRD: RECODE, which solves RElation detection via COmposite DEscription prompts. Specifically, RECODE first decomposes each predicate category into subject, object, and spatial components. Then, it leverages large language models (LLMs) to generate description-based prompts (or visual cues) for each component. Different visual cues enhance the discriminability of similar relation categories from different perspectives, which significantly boosts performance in VRD. To dynamically fuse different cues, we further introduce a chain-of-thought method that prompts LLMs to generate reasonable weights for different visual cues. Extensive experiments on four VRD benchmarks have demonstrated the effectiveness and interpretability of RECODE. Lin Li 0065, Jun Xiao 0001, Guikun Chen, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
NeurIPS | 5 |
| 2023 | Graph neural networks meet with distributed graph partitioners and reconciliations
Zongshen Mu, Siliang Tang, Chang Zong, Dianhai Yu, Yueting Zhuang |
Neurocomputing | 5 |
| 2023 | A knowledge-guided and traditional Chinese medicine informed approach for herb recommendationabstractTraditional Chinese medicine (TCM) is an interesting research topic in China’s thousands of years of history. With the recent advances in artificial intelligence technology, some researchers have started to focus on learning the TCM prescriptions in a data-driven manner. This involves appropriately recommending a set of herbs based on patients’ symptoms. Most existing herb recommendation models disregard TCM domain knowledge, for example, the interactions between symptoms and herbs and the TCM-informed observations (i.e., TCM formulation of prescriptions). In this paper, we propose a knowledge-guided and TCM-informed approach for herb recommendation. The knowledge used includes path interactions and co-occurrence relationships among symptoms and herbs from a knowledge graph generated from TCM literature and prescriptions. The aforementioned knowledge is used to obtain the discriminative feature vectors of symptoms and herbs via a graph attention network. To increase the ability of herb prediction for the given symptoms, we introduce TCM-informed observations in the prediction layer. We apply our proposed model on a TCM prescription dataset, demonstrating significant improvements over state-of-the-art herb recommendation methods. Zhe Jin 0003, Yin Zhang 0006, Jiaxu Miao, Yi Yang 0001, Yueting Zhuang, Yunhe Pan |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Federated unsupervised representation learningabstractTo leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins. Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 10 |
| 2023 | Attribute-driven streaming edge partitioning with reconciliations for distributed graph neural network training
Zongshen Mu, Siliang Tang, Yueting Zhuang, Dianhai Yu |
Neural Networks | 3 |
| 2023 | Variational Cross-Graph Reasoning and Adaptive Structured Semantics Learning for Compositional Temporal GroundingabstractTemporal grounding is the task of locating a specific segment from an untrimmed video according to a query sentence. This task has achieved significant momentum in the computer vision community as it enables activity grounding beyond pre-defined activity classes by utilizing the semantic diversity of natural language descriptions. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, existing temporal grounding datasets are not carefully designed to evaluate the compositional generalizability. To systematically benchmark the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. We empirically find that they fail to generalize to queries with novel combinations of seen words. We argue that the inherent compositional structure (i.e., composition constituents and their relationships) inside the videos and language is the crucial factor to achieve compositional generalization. Based on this insight, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into hierarchical semantic graphs, respectively, and learns fine-grained semantic correspondence between the two graphs. Meanwhile, we introduce a novel adaptive structured semantics learning approach to derive the structure-informed and domain-generalizable graph representations, which facilitate the fine-grained semantic correspondence reasoning between the two graphs. To further evaluate the understanding of the compositional structure, we also introduce a more challenging setting, where one of the components in the novel composition is unseen. This requires more sophisticated understanding of the compositional structure to infer the potential semantics of the unseen word based on the other learned composition constituents appearing in both the video and language context, and their relationships. Extensive experiments validate the superior compositional generalizability of our approach, demonstrating its ability to handle queries with novel combinations of seen words as well as novel words in the testing composition. Juncheng Li 0006, Siliang Tang, Linchao Zhu, Wenqiao Zhang, Yi Yang 0001, Tat-Seng Chua, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Single image super-resolution based on progressive fusion of orientation-aware features
Zewei He, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, Xin Li 0003, Siliang Tang, Yueting Zhuang, Zheming Lu 0001 |
Pattern Recognit. | 8 |
| 2023 | Stable Prediction With Leveraging Seed VariableabstractIn this paper, we focus on the problem of stable prediction across unknown test data, where the test distribution might be different from the training one and is always agnostic when model training. In such a case, previous machine learning methods might exploit subtly spurious correlations induced by non-causal variables in training data for prediction. Those spurious correlations are changeable across data, leading to instability of prediction across unknown test data. To address this problem, we propose a conditional independence test based algorithm to screen out part of non-causal features and reduce those spurious correlations for a more stable prediction by leveraging a seed variable. We show, both theoretically and with empirical experiments, that our algorithm can precisely screen out the isolated non-causal variables, which have no causal relationship with other variables, and remove the spurious correlations induced by them, increasing the stability of prediction across unknown test data. Extensive experiments on both synthetic and real-world datasets demonstrate that our algorithm outperforms state-of-the-art methods for stable prediction across unknown test data. Kun Kuang 0001, Haotian Wang 0001, Ruoxuan Xiong, Runze Wu 0001, Weiming Lu 0001, Yueting Zhuang, Fei Wu 0001, Peng Cui 0001, Bo Li 0064 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Learning Decomposed Representations for Treatment Effect EstimationabstractIn observational studies, confounder separation and balancing are the fundamental problems of treatment effect estimation. Most of the previous methods focused on addressing the problem of confounder balancing by treating all observed pre-treatment variables as confounders, ignoring confounder separation. In general, not all the observed pre-treatment variables are confounders that refer to the common causes of the treatment and the outcome, some variables only contribute to the treatment (i.e., instrumental variables) and some only contribute to the outcome (i.e., adjustment variables). Balancing those non-confounders, including instrumental variables and adjustment variables, would generate additional bias for treatment effect estimation. By modeling the different causal relations among observed pre-treatment variables, treatment variables and outcome variables, we propose a synergistic learning framework to i) separate confounders by learning decomposed representations of both confounders and non-confounders, ii) balance confounder with sample re-weighting technique, and simultaneously iii) estimate the treatment effect in observational studies via counterfactual inference. Empirical results on synthetic and real-world datasets demonstrate that the proposed method can precisely decompose confounders and achieve a more precise estimation of treatment effect than baselines. Anpeng Wu, Junkun Yuan, Kun Kuang 0001, Bo Li 0064, Runze Wu 0001, Yueting Zhuang, Fei Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Cross-Modal Data Augmentation for Tasks of Different ModalitiesabstractData augmentation has become one of the keys to alleviating the over-fitting of models on training data and improving the generalization capabilities on testing data. Most existing data augmentation methods only focus on one modality, which is incapable when facing multiple data modalities. Some prior works try to interpolate with random coefficients in the latent space to generate new samples, which can generically work for any data modality. However, these works ignore the extra information conveyed by multimodality data. In fact, the extra information in one modality can provide semantic directions to generate more meaningful samples in another modality. This paper proposes Cross-modal Data Augmentation (CMDA), a simple yet effective data augmentation method to alleviate the over-fitting issue and improve the generalization performance. We evaluate CMDA on unsupervised and supervised tasks of different modalities, on which CMDA consistently and significantly outperforms baselines. For instance, CMDA improves the unsupervised anomaly detection baseline in vision modality from the AUROC$76.46\%, 73.07\%$and 64.36% to$83.25\%, 76.22\%$and 70.57% on three different datasets, respectively. Besides, extensive experiments demonstrate that CMDA is applicable to various neural network architectures. Furthermore, prior methods that interpolate in the latent space need to work with downstream tasks to construct the latent space. In contrast, CMDA can work with or without downstream tasks, which makes the applicability of CMDA more extensive. The source code is publicly available for non-commercial or research use athttps://github.com/Anfeather/CMDA Dong Chen 0017, Yueting Zhuang, Zijin Shen, Carl Yang 0001, Guoming Wang, Siliang Tang, Yi Yang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Deep Residual Weight-Sharing Attention Network With Low-Rank Attention for Visual Question AnsweringabstractThe attention-based networks have become prevailing recently in visual question answering (VQA) due to their high performances. However, the extensive memory consumption of attention-based models poses excessive-high demand for the implementation equipment, raising concerns about their future application scenarios. Therefore, designing an efficient and lightweight VQA model is central to expanding possible application areas. Our work presents a novel lightweight attention-based VQA model, namely residual weight-sharing attention network (RWSAN), consisting of residual weight-sharing attention (RWSA) layers cascaded in depth. Each RWSA layer models the textual representation with self residual weight-sharing attention (SRWSA) and captures question features and question-image interactions with self-guided residual weight-sharing attention (SGRWSA). Inside each RWSA layer, the proposed low-rank attention units perform residual learning with learned connection patterns and shared parameters, and every stacked RWSA layer also uses the same parameters. Extensive ablation experiments with quantitative and qualitative analysis are conducted to illustrate the effectiveness and generality of RWSA. Experiments on VQA-v2, GQA, and CLEVR datasets show that the RWSAN achieves competitive performance with much fewer parameters over the state-of-the-art methods. Bosheng Qin, Haoji Hu, Yueting Zhuang |
IEEE Trans. Multim. | 3 |
| 2023 | Elastic Knowledge Distillation by Learning From RecollectionabstractModel performance can be further improved with the extra guidance apart from the one-hot ground truth. To achieve it, recently proposed recollection-based methods utilize the valuable information contained in the past training history and derive a "recollection" from it to provide data-driven prior to guide the training. In this article, we focus on two fundamental aspects of this method, i.e., recollection construction and recollection utilization. Specifically, to meet the various demands of models with different capacities and at different training periods, we propose to construct a set of recollections with diverse distributions from the same training history. After that, all the recollections collaborate together to provide guidance, which is adaptive to different model capacities, as well as different training periods, according to our similarity-based elastic knowledge distillation (KD) algorithm. Without any external prior to guide the training, our method achieves a significant performance gain and outperforms the methods of the same category, even as well as KD with well-trained teacher. Extensive experiments and further analysis are conducted to demonstrate the effectiveness of our method. Yongjian Fu 0002, Hanbin Zhao, Wenfu Wang, Weihao Fang, Yueting Zhuang, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language MatchingabstractThe prevailing framework for matching multimodal inputs is based on a two-stage process: (1) detecting proposals with an object detector and (2) matching text queries with proposals. Existing two-stage solutions mostly focus on the matching step. In this article, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., query-agnostic), hoping that the proposals contain all instances mentioned in the text query (i.e., query-aware). Due to this mismatch, chances are that proposals relevant to the text query are suppressed during the filtering process, which in turn bounds the matching performance. To this end, we propose VL-NMS, which is the first method to yield query-aware proposals at the first stage. VL-NMS regards all mentioned instances as critical objects and introduces a lightweight module to predict a score for aligning each proposal with a critical object. These scores can guide the NMS operation to filter out proposals irrelevant to the text query, increasing the recall of critical objects, and resulting in a significantly improved matching performance. Since VL-NMS is agnostic to the matching step, it can be easily integrated into any state-of-the-art two-stage matching method. We validate the effectiveness of VL-NMS on three multimodal matching tasks, namely referring expression grounding, phrase grounding, and image-text matching. Extensive ablation studies on several baselines and benchmarks consistently demonstrate the superiority of VL-NMS. Chenchi Zhang, Jun Xiao 0001, Hanwang Zhang, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | On the Efficacy of Small Self-Supervised Contrastive Models without Distillation SignalsabstractIt is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be suitable for some resource-restricted scenarios due to the huge computational expenses of deploying a large model. In this paper, we study the issue of training self-supervised small models without distillation signals. We first evaluate the representation spaces of the small models and make two non-negligible observations: (i) the small models can complete the pretext task without overfitting despite their limited capacity and (ii) they universally suffer the problem of over clustering. Then we verify multiple assumptions that are considered to alleviate the over-clustering phenomenon. Finally, we combine the validated techniques and improve the baseline performances of five small architectures with considerable margins, which indicates that training small self-supervised contrastive models is feasible even without distillation signals. The code is available at https://github.com/WOWNICE/ssl-small. Haizhou Shi, Youcai Zhang, Siliang Tang, Wenjie Zhu 0003, Yandong Guo, Yueting Zhuang |
AAAI | 7 |
| 2022 | MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningabstractText-based image captioning (TextCap) requires simultaneous comprehension of visual content and reading the text of images to generate a natural language description. Although a task can teach machines to understand the complex human environment further given that text is omnipresent in our daily surroundings, it poses additional challenges in normal captioning. A text-based image intuitively contains abundant and complex multimodal relational content, that is, image details can be described diversely from multiview rather than a single caption. Certainly, we can introduce additional paired training data to show the diversity of images' descriptions, this process is labor-intensive and time-consuming for TextCap pair annotations with extra texts. Based on the insight mentioned above, we investigate how to generate diverse captions that focus on different image parts using an unpaired training paradigm. We propose the Multimodal relAtional Graph adversarIal InferenCe (MAGIC) framework for diverse and unpaired TextCap. This framework can adaptively construct multiple multimodal relational graphs of images and model complex relationships among graphs to represent descriptive diversity. Moreover, a cascaded generative adversarial network is developed from modeled graphs to infer the unpaired caption generation in image–sentence feature alignment and linguistic coherence levels. We validate the effectiveness of MAGIC in generating diverse captions from different relational information items of an image. Experimental results show that MAGIC can generate very promising outcomes without using any image–caption training pairs. Wenqiao Zhang, Jiannan Guo 0003, Shengyu Zhang 0001, Qingpeng Cai 0002, Juncheng Li 0006, Sihui Luo 0003, Yueting Zhuang |
AAAI | 8 |
| 2022 | Parallel Instance Query Network for Named Entity RecognitionabstractYongliang Shen, Xiaobin Wang, Zeqi Tan, Guangwei Xu, Pengjun Xie, Fei Huang, Weiming Lu, Yueting Zhuang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yongliang Shen 0001, Xiaobin Wang, Zeqi Tan, Pengjun Xie, Fei Huang 0002, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 8 |
| 2022 | Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningabstractTemporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at ht tps: / / gi thub. com/YYJMJC/ Composi tional- Temporal-Grounding. Juncheng Li 0006, Junlin Xie, Linchao Zhu, Siliang Tang, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang, Xin Wang 0061 |
CVPR | 8 |
| 2022 | Label Matching Semi-Supervised Object DetectionabstractSemi-supervised object detection has made significant progress with the development of mean teacher driven self-training. Despite the promising results, the label mismatch problem is not yet fully explored in the previous works, leading to severe confirmation bias during self-training. In this paper, we delve into this problem and propose a simple yet effective LabelMatch framework from two different yet complementary perspectives, i.e., distribution-level and instance-level. For the former one, it is reasonable to approximate the class distribution of the unlabeled data from that of the labeled data according to Monte Carlo Sampling. Guided by this weakly supervision cue, we introduce a re-distribution mean teacher, which leverages adaptive label-distribution-aware confidence thresholds to generate unbiased pseudo labels to drive student learning. For the latter one, there exists an overlooked label assignment ambiguity problem across teacher-student models. To remedy this issue, we present a novel label assignment mechanism for self-training framework, namely proposal self-assignment, which injects the proposals from student into teacher and generates accurate pseudo labels to match each proposal in the student model accordingly. Experiments on both MS-COCO and PASCAL-VOC datasets demonstrate the considerable superiority of our proposed framework to other state-of-the-arts. Code will be available at https://github.com/HIK-LAB/SSOD. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Jie Song 0011, Di Xie, Shiliang Pu, Mingli Song, Yueting Zhuang |
CVPR | 9 |
| 2022 | Learning to Learn by Jointly Optimizing Neural Architecture and WeightsabstractMeta-learning enables models to adapt to new environments rapidly with a few training examples. Current gradient-based meta-learning methods concentrate on finding good model-agnostic initialization (meta-weights) for learners. In this paper, we aim to obtain better meta-learners by co-optimizing the architecture and meta-weights simultaneously. Existing NAS-based meta-learning methods apply a two-stage strategy, i.e., first searching architectures and then re-training meta-weights on the searched architecture. However, this two-stage strategy would break the mutual impact of the architecture and meta-weights since they are optimized separately. Differently, we propose progressive connection consolidation, fixing the architecture layer by layer, in which the layer with the largest weight value would be fixed first. In this way, we can jointly search architectures and train the meta-weights on fixed layers. Besides, to improve the generalization performance of the searched meta-learner on all tasks, we propose a more effective rule for co-optimization, namely Connection-Adaptive Meta-learning (CAML). By searching only once, we can obtain both adaptive architecture and meta-weights for meta-learning. Extensive experiments show that our method achieves state-of-the-art performance with 3x less computational cost, revealing our method's effectiveness and efficiency. Yadong Ding, Yu Wu 0011, Chengyue Huang, Siliang Tang, Yi Yang 0001, Longhui Wei, Yueting Zhuang, Qi Tian 0001 |
CVPR | 7 |
| 2022 | Slimmable Domain AdaptationabstractVanilla unsupervised domain adaptation methods tend to optimize the model with fixed neural architecture, which is not very practical in real-world scenarios since the target data is usually processed by different resource-limited devices. It is therefore of great necessity to facilitate architecture adaptation across various devices. In this paper, we introduce a simple framework, Slimmable Domain Adaptation, to improve cross-domain generalization with a weight-sharing model bank, from which models of different capacities can be sampled to accommodate different accuracy-efficiency trade-offs. The main challenge in this frame-work lies in simultaneously boosting the adaptation performance of numerous models in the model bank. To tackle this problem, we develop a Stochastic EnsEmble Distillation method to fully exploit the complementary knowledge in the model bank for inter-model interaction. Nevertheless, considering the optimization conflict between inter-model interaction and intra-model adaptation, we augment the existing bi-classifier domain confusion architecture into an Optimization-Separated Tri-Classifier counterpart. After optimizing the model bank, architecture adaptation is leveraged via our proposed Unsupervised Performance Evaluation Metric. Under various resource constraints, our framework surpasses other competing approaches by a very large margin on multiple benchmarks. It is also worth emphasizing that our framework can preserve the performance improvement against the source-only model even when the computing complexity is reduced to 1/64. Code will be available at https://github.com/HIK-LAB/SlimDA. Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Luojun Lin, Di Xie, Shiliang Pu, Xinchao Wang, Mingli Song, Yueting Zhuang |
CVPR | 10 |
| 2022 | Query-based Instance Discrimination Network for Relational Triple ExtractionabstractJoint entity and relation extraction has been a core task in the field of information extraction.Recent approaches usually consider the extraction of relational triples from a stereoscopic perspective, either learning a relation-specific tagger or separate classifiers for each relation type.However, they still suffer from error propagation, relation redundancy and lack of highlevel connections between triples.To address these issues, we propose a novel query-based approach to construct instance-level representations for relational triples.By metric-based comparison between query embeddings and token embeddings, we can extract all types of triples in one step, thus eliminating the error propagation problem.In addition, we learn the instance-level representation of relational triples via contrastive learning.In this way, relational triples can not only enclose rich classlevel semantics but also access to high-order global connections.Experimental results show that our proposed method achieves the state of the art on five widely used benchmarks. Zeqi Tan, Yongliang Shen 0001, Xuming Hu, Wenqi Zhang 0001, Xiaoxia Cheng, Weiming Lu 0001, Yueting Zhuang |
EMNLP | 7 |
| 2022 | Transductive Clip with Class-Conditional Contrastive LearningabstractInspired by the remarkable zero-shot generalization capacity of vision-language pre-trained model, we seek to leverage the supervision from CLIP model to alleviate the burden of data labeling. However, such supervision inevitably contains the label noise, which significantly degrades the discriminative power of the classification model. In this work, we propose Transductive CLIP, a novel framework for learning a classification network with noisy labels from scratch. Firstly, a class-conditional contrastive learning mechanism is proposed to mitigate the reliance on pseudo labels and boost the tolerance to noisy labels. Secondly, ensemble labels is adopted as a pseudo label updating strategy to stabilize the training of deep neural networks with noisy labels. This framework can reduce the impact of noisy labels from CLIP model effectively by combining both techniques. Experiments on multiple benchmark datasets demonstrate the substantial improvements over other state-of-the-art methods. Junchu Huang, Weijie Chen 0006, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang |
ICASSP | 6 |
| 2022 | Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object DetectionabstractVanilla unsupervised domain adaptive (UDA) object detection typically requires the labeled source data for joint-training with the unlabeled target data, which is usually unavailable in real-world scenarios due to data privacy, leading to source data-free UDA object detection. Herein, we first analyze the phenomenon of cross-domain detection degradation varying from easy to hard samples (e.g. the objects with different scales or occlusion degrees), termed as domain generalization differentiation. In detail, the ability to detect easy samples is well transferred while the one to detect hard samples is dramatically degraded. To this end, we then revisit the existing self-training method, which is of great challenge to deal with the abundant false negatives (hard samples). Assumed that true positives (easy samples) labeled by the source model can be exploited as supervision cues. UDA is finally modeled into an unsupervised false negatives mining problem. Thus, we propose a Simulation-and-Mining (S&M) framework, which simulates false negatives by augmenting true positives and mines back false negatives alternatively and iteratively. Experimental results show the effectiveness. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Di Xie, Yueting Zhuang, Shiliang Pu |
ICASSP | 6 |
| 2022 | Learning Domain Adaptive Object Detection with Probabilistic TeacherabstractSelf-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts. Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu |
ICML | 9 |
| 2022 | Robust Meta-learning with Sampling Noise and Label Noise via Eigen-ReptileabstractRecent years have seen a surge of interest in meta-learning techniques for tackling the few-shot learning (FSL) problem. However, the meta-learner is prone to overfitting since there are only a few available samples, which can be identified as sampling noise on a clean dataset. Besides, when handling the data with noisy labels, the meta-learner could be extremely sensitive to label noise on a corrupted dataset. To address these two challenges, we present Eigen-Reptile (ER) that updates the meta-parameters with the main direction of historical task-specific parameters. Specifically, the main direction is computed in a fast way, where the scale of the calculated matrix is related to the number of gradient steps for the specific task instead of the number of parameters. Furthermore, to obtain a more accurate main direction for Eigen-Reptile in the presence of many noisy labels, we further propose Introspective Self-paced Learning (ISPL). We have theoretically and experimentally demonstrated the soundness and effectiveness of the proposed Eigen-Reptile and ISPL. Particularly, our experiments on different tasks show that the proposed method is able to outperform or achieve highly competitive performance compared with other gradient-based methods with or without noisy labels. The code and data for the proposed method are provided for research purposes https://github.com/Anfeather/Eigen-Reptile. Dong Chen 0017, Lingfei Wu 0001, Siliang Tang, Xiao Yun, Bo Long, Yueting Zhuang |
ICML | 6 |
| 2022 | Self-Supervised Noisy Label Learning for Source-Free Unsupervised Domain AdaptationabstractDomain adaptation is an important property in robot vision, which enables the neural networks pre-trained on source domains to adapt target domains automatically without any annotation efforts. During this process, source data is not always accessible due to the constraints of expensive storage overhead and data privacy protection. Therefore, the source domain pre-trained model is expected to optimize with only unlabeled target data, termed as source-free unsupervised domain adaptation. In this paper, we view this problem as a special case of noisy label learning, since the given pre-trained model can generate noisy labels for unlabeled target data via network inference. The potential semantic cues for unsupervised domain adaptation exactly lie on these noisy labels. Inspired by this problem modeling, we propose a simple yet effective Self-Supervised Noisy Label Learning method, which injects self-supervised learning to impose the intrinsic data structure and facilitate label-denoising. Extensive experiments have been conducted on diverse benchmarks to validate the effectiveness. Our method achieves state-of-the-art performance. Weijie Chen 0006, Luojun Lin, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang |
IROS | 6 |
| 2022 | Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosabstractUnderstanding human emotions is a crucial ability for intelligent robots to provide better human-robot interactions. The existing works are limited to trimmed video-level emotion classification, failing to locate the temporal window corresponding to the emotion. In this paper, we introduce a new task, named Temporal Emotion Localization in videos (TEL), which aims to detect human emotions and localize their corresponding temporal boundaries in untrimmed videos with aligned subtitles. TEL presents three unique challenges compared to temporal action localization: 1) The emotions have extremely varied temporal dynamics; 2) The emotion cues are embedded in both appearances and complex plots; 3) The fine-grained temporal annotations are complicated and labor-intensive. To address the first two challenges, we propose a novel dilated context integrated network with a coarse-fine two-stream architecture. The coarse stream captures varied temporal dynamics by modeling multi-granularity temporal contexts. The fine stream achieves complex plots understanding by reasoning the dependency between the multi-granularity temporal contexts from the coarse stream and adaptively integrates them into fine-grained video segment features. To address the third challenge, we introduce a cross-modal consensus learning paradigm, which leverages the inherent semantic consensus between the aligned video and subtitle to achieve weakly-supervised learning. We contribute a new testing set with 3,000 manually-annotated temporal boundaries so that future research on the TEL problem can be quantitatively evaluated. Extensive experiments show the effectiveness of our approach on temporal emotion localization. The repository of this work is at https://github.com/YYJMJC/TemporalEmotion-Localization-in-Videos. Juncheng Li 0006, Junlin Xie, Linchao Zhu, Siliang Tang, Wenqiao Zhang, Shengyu Zhang 0001, Longhui Wei, Qi Tian 0001, Yueting Zhuang |
ACM Multimedia | 11 |
| 2022 | Learning Hybrid Behavior Patterns for Multimedia RecommendationabstractMultimedia recommendation aims to predict user preferences where users interact with multimodal items. Collaborative filtering based on graph convolutional networks manifests impressive performance gains in multimedia recommendation. This is attributed to the capability of learning good user and item embeddings by aggregating the collaborative signals from high-order neighbors. However, previous researches [37,38] fail to explicitly mine different behavior patterns (i.e., item categories, common user interests) by exploiting user-item and item-item graphs simultaneously, which plays an important role in modeling user preferences. And it is the lack of different behavior pattern constraints and multimodal feature reconciliations that results in performance degradation. Towards this end, We propose a Hybrid Clustering Graph Convolutional Network (HCGCN) for multimedia recommendation. We perform high-order graph convolutions inside user-item clusters and item-item clusters to capture various user behavior patterns. Meanwhile, we design corresponding clustering losses to enhance user-item preference feedback and multimodal representation learning constraint to adjust the modality importance, making more accurate recommendations. Experimental results on three real-world multimedia datasets not only demonstrate the significant improvement of our model over the state-of-the-art methods, but also validate the effectiveness of integrating hybrid user behavior patterns for multimedia recommendation. Zongshen Mu, Yueting Zhuang, Jun Xiao 0001, Siliang Tang |
ACM Multimedia | 2 |
| 2022 | Fine-Grained Semantically Aligned Vision-Language Pre-TrainingabstractLarge-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarity of the global representations of images and text, or advanced cross-modal attention upon image and text features. However, they fail to explicitly learn the fine-grained semantic alignment between visual regions and textual phrases, as only global image-text alignment information is available. In this paper, we introduce LOUPE, a fine-grained semantically aLigned visiOn-langUage PrE-training framework, which learns fine-grained semantic alignment from the novel perspective of game-theoretic interactions. To efficiently estimate the game-theoretic interactions, we further propose an uncertainty-aware neural Shapley interaction learning module. Experiments show that LOUPE achieves state-of-the-art performance on a variety of vision-language tasks. Without any object-level human annotations and fine-tuning, LOUPE achieves competitive performance on object detection and visual grounding. More importantly, LOUPE opens a new promising direction of learning fine-grained semantics from large-scale raw image-text pairs. Juncheng Li 0006, Longhui Wei, Linchao Zhu, Lingxi Xie, Yueting Zhuang, Qi Tian 0001, Siliang Tang |
NeurIPS | 7 |
| 2022 | NAP: Neural architecture search with pruning
Yadong Ding, Yu Wu 0011, Chengyue Huang, Siliang Tang, Fei Wu 0001, Yi Yang 0001, Wenwu Zhu 0001, Yueting Zhuang |
Neurocomputing | 8 |
| 2022 | Deep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Feifei Shao, Long Chen 0016, Jian Shao 0001, Wei Ji 0008, Shaoning Xiao, Lu Ye, Yueting Zhuang, Jun Xiao 0001 |
Neurocomputing | 7 |
| 2022 | Boosting RGB-D Saliency Detection by Leveraging Unlabeled RGB ImagesabstractTraining deep models for RGB-D salient object detection (SOD) often requires a large number of labeled RGB-D images. However, RGB-D data is not easily acquired, which limits the development of RGB-D SOD techniques. To alleviate this issue, we present a Dual-Semi RGB-D Salient Object Detection Network (DS-Net) to leverage unlabeled RGB images for boosting RGB-D saliency detection. We first devise a depth decoupling convolutional neural network (DDCNN), which contains a depth estimation branch and a saliency detection branch. The depth estimation branch is trained with RGB-D images and then used to estimate the pseudo depth maps for all unlabeled RGB images to form the paired data. The saliency detection branch is used to fuse the RGB feature and depth feature to predict the RGB-D saliency. Then, the whole DDCNN is assigned as the backbone in a teacher-student framework for semi-supervised learning. Moreover, we also introduce a consistency loss on the intermediate attention and saliency maps for the unlabeled data, as well as a supervised depth and saliency loss for labeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our DDCNN outperforms state-of-the-art methods both quantitatively and qualitatively. We also demonstrate that our semi-supervised DS-Net can further improve the performance, even when using an RGB image with the pseudo depth map. Xiaoqiang Wang 0007, Lei Zhu 0003, Siliang Tang, Huazhu Fu, Ping Li 0016, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 8 |
| 2022 | Balance-Subsampled Stable Prediction Across Unknown Test DataabstractIn data mining and machine learning, it is commonly assumed that training and test data share the same population distribution. However, this assumption is often violated in practice because of the sample selection bias, which might induce the distribution shift from training data to test data. Such a model-agnostic distribution shift usually leads to prediction instability across unknown test data. This article proposes a novel balance-subsampled stable prediction (BSSP) algorithm based on the theory of fractional factorial design. It isolates the clear effect of each predictor from the confounding variables. A design-theoretic analysis shows that the proposed method can reduce the confounding effects among predictors induced by the distribution shift, improving both the accuracy of parameter estimation and the stability of prediction across unknown test data. Numerical experiments on synthetic and real-world datasets demonstrate that our BSSP algorithm can significantly outperform the baseline methods for stable prediction across unknown test data. Kun Kuang 0001, Hengtao Zhang, Runze Wu 0001, Fei Wu 0001, Yueting Zhuang, Aijun Zhang |
ACM Trans. Knowl. Discov. Data | 5 |
| 2021 | Empower Distantly Supervised Relation Extraction with Collaborative Adversarial TrainingabstractWith recent advances in distantly supervised (DS) relation extraction (RE), considerable attention is attracted to leverage multi-instance learning (MIL) to distill high-quality supervision from the noisy DS. Here, we go beyond label noise and identify the key bottleneck of DS-MIL to be its low data utilization: as high-quality supervision being refined by MIL, MIL abandons a large amount of training instances, which leads to a low data utilization and hinders model training from having abundant supervision. In this paper, we propose collaborative adversarial training to improve the data utilization, which coordinates virtual adversarial training (VAT) and adversarial training (AT) at different levels. Specifically, since VAT is label-free, we employ the instance-level VAT to recycle instances abandoned by MIL. Besides, we deploy AT at the bag-level to unleash the full potential of the high-quality supervision got by MIL. Our proposed method brings consistent improvements (∼ 5 absolute AUC score) to the previous state of the art, which verifies the importance of the data utilization issue and the effectiveness of our method. Siliang Tang, Jian Shao 0001, Zhigang Chen 0003, Yueting Zhuang |
AAAI | 7 |
| 2021 | A Free Lunch for Unsupervised Domain Adaptive Object Detection without Source DataabstractUnsupervised domain adaptation (UDA) assumes that source and target domain data are freely available and usually trained together to reduce the domain gap. However, considering the data privacy and the inefficiency of data transmission, it is impractical in real scenarios. Hence, it draws our eyes to optimize the network in the target domain without accessing labeled source data. To explore this direction in object detection, for the first time, we propose a source data-free domain adaptive object detection (SFOD) framework via modeling it into a problem of learning with noisy labels. Generally, a straightforward method is to leverage the pre-trained network from the source domain to generate the pseudo labels for target domain optimization. However, it is difficult to evaluate the quality of pseudo labels since no labels are available in target domain. In this paper, self-entropy descent (SED) is a metric proposed to search an appropriate confidence threshold for reliable pseudo label generation without using any handcrafted labels. Nonetheless, completely clean labels are still unattainable. After a thorough experimental analysis, false negatives are found to dominate in the generated noisy labels. Undoubtedly, false negatives mining is helpful for performance improvement, and we ease it to false negatives simulation through data augmentation like Mosaic. Extensive experiments conducted in four representative adaptation tasks have demonstrated that the proposed framework can easily achieve state-of-the-art performance. From another view, it also reminds the UDA community that the labeled source data are not fully exploited in the existing methods. Weijie Chen 0006, Di Xie, Shicai Yang, Shiliang Pu, Yueting Zhuang |
AAAI | 7 |
| 2021 | Disentangled Motif-aware Graph Learning for Phrase GroundingabstractIn this paper, we propose a novel graph learning framework for phrase grounding in the image. Developing from the sequential to the dense graph model, existing works capture coarse-grained context but fail to distinguish the diversity of context among phrases and image regions. In contrast, we pay special attention to different motifs implied in the context of the scene graph and devise the disentangled graph network to integrate the motif-aware contextual information into representations. Besides, we adopt interventional strategies at the feature and the structure levels to consolidate and generalize representations. Finally, the cross-modal attention network is utilized to fuse intra-modal features, where each phrase can be computed similarity with regions to select the best-grounded one. We validate the efficiency of disentangled and interventional graph network (DIGN) through a series of ablation studies, and our model achieves state-of-the-art performance on Flickr30K Entities and ReferIt Game benchmarks. Zongshen Mu, Siliang Tang, Yueting Zhuang |
AAAI | 5 |
| 2021 | Consensus Graph Representation Learning for Better Grounded Image CaptioningabstractThe contemporary visual captioning models frequently hallucinate objects that are not actually in a scene, due to the visual misclassification or over-reliance on priors that resulting in the semantic inconsistency between the visual information and the target lexical words. The most common way is to encourage the captioning model to dynamically link generated object words or phrases to appropriate regions of the image, i.e., the grounded image captioning (GIC). However, GIC utilizes an auxiliary task (grounding objects) that has not solved the key issue of object hallucination, i.e., the semantic inconsistency. In this paper, we take a novel perspective on the issue above: exploiting the semantic coherency between the visual and language modalities. Specifically, we propose the Consensus Rraph Representation Learning framework (CGRL) for GIC that incorporates a consensus representation into the grounded captioning pipeline. The consensus is learned by aligning the visual graph (e.g., scene graph) to the language graph that consider both the nodes and edges in a graph. With the aligned consensus, the captioning model can capture both the correct linguistic characteristics and visual relevance, and then grounding appropriate image regions further. We validate the effectiveness of our model, with a significant decline in object hallucination (-9% CHAIRi) on the Flickr30k Entities dataset. Besides, our CGRL also evaluated by several automatic metrics and human evaluation, the results indicate that the proposed approach can simultaneously improve the performance of image captioning (+2.9 Cider) and grounding (+2.3 F1LOC}). Wenqiao Zhang, Siliang Tang, Jun Xiao 0001, Yueting Zhuang |
AAAI | 6 |
| 2021 | CIL: Contrastive Instance Learning Framework for Distantly Supervised Relation ExtractionabstractTao Chen, Haizhou Shi, Siliang Tang, Zhigang Chen, Fei Wu, Yueting Zhuang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haizhou Shi, Siliang Tang, Zhigang Chen 0003, Fei Wu 0001, Yueting Zhuang |
ACL/IJCNLP (1) | 6 |
| 2021 | Natural Language Video Localization with Learnable Moment ProposalsabstractGiven an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query.To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment candidates and then find out the best-matching one.2) proposal-free models directly predict two temporal boundaries of the referential moment from frames.Currently, almost all the propose-and-rank methods have inferior performance than proposal-free counterparts.In this paper, we argue that propose-and-rank approach is underestimated due to the predefined manners: 1) Hand-designed rules are hard to guarantee the complete coverage of targeted segments.2) Densely sampled candidate moments cause redundant computation and degrade the performance of ranking process.To this end, we propose a novel model termed LP-Net (Learnable Proposal Network for NLVL) with a fixed set of learnable moment proposals.The position and length of these proposals are dynamically adjusted during training process.Moreover, a boundary-aware loss has been proposed to leverage frame-level information and further improve the performance.Extensive ablations on two challenging NLVL benchmarks have demonstrated the effectiveness of LPNet over existing state-of-the-art methods 1 . Shaoning Xiao, Long Chen 0016, Jian Shao 0001, Yueting Zhuang, Jun Xiao 0001 |
EMNLP (1) | 4 |
| 2021 | Semi-supervised Active Learning for Semi-supervised Models: Exploit Adversarial Examples with Graph-based Virtual LabelsabstractThe performance of computer vision models significantly improves with more labeled data. However, the acquisition of labeled data is limited by the high cost. To mitigate the reliance on large labeled datasets, active learning (AL) and semi-supervised learning (SSL) are frequently adopted. Although current mainstream methods begin to combine SSL and AL (SSL-AL) to excavate the diverse expressions of unlabeled samples, these methods’ fully supervised task models are still trained only with labeled data. Besides, these method’s SSL-AL frameworks suffer from mismatch problems. Here, we propose a graph-based SSL-AL framework to unleash the SSL task models’ power and make an effective SSL-AL interaction. In the framework, SSL leverages graph-based label propagation to deliver virtual labels to unlabeled samples, rendering AL samples’ structural distribution and boosting AL. AL finds samples near the clusters’ boundary to help SSL perform better label propagation by exploiting adversarial examples. The information exchange in the closed-loop realizes mutual enhancement of SSL and AL. Experimental results show that our method outperforms the state-of-the-art methods against classification and segmentation benchmarks. Jiannan Guo 0003, Yangyang Kang, Kun Kuang 0001, Siliang Tang, Zhuoren Jiang, Changlong Sun, Fei Wu 0001, Yueting Zhuang |
ICCV | 9 |
| 2021 | Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceabstractVideo-and-Language Inference is a recently proposed task for joint video-and-language understanding. This new task requires a model to draw inference on whether a natural language statement entails or contradicts a given video clip. In this paper, we study how to address three critical challenges for this task: judging the global correctness of the statement involved multiple semantic meanings, joint reasoning over video and subtitles, and modeling long-range relationships and complex social interactions. First, we propose an adaptive hierarchical graph network that achieves in-depth understanding of the video over complex interactions. Specifically, it performs joint reasoning over video and subtitles in three hierarchies, where the graph structure is adaptively adjusted according to the semantic structures of the statement. Secondly, we introduce semantic coherence learning to explicitly encourage the semantic coherence of the adaptive hierarchical graph network from three hierarchies. The semantic coherence learning can further improve the alignment between vision and linguistics, and the coherence across a sequence of video segments. Experimental results show that our method significantly outperforms the baseline by a large margin. Juncheng Li 0006, Siliang Tang, Linchao Zhu, Xuanwen Huang, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang |
ICCV | 8 |
| 2021 | A Sequence-to-Set Network for Nested Named Entity RecognitionabstractNamed entity recognition (NER) is a widely studied task in natural language processing. Recently, a growing number of studies have focused on the nested NER. The span-based methods, considering the entity recognition as a span classification task, can deal with nested entities naturally. But they suffer from the huge search space and the lack of interactions between entities. To address these issues, we propose a novel sequence-to-set neural network for nested NER. Instead of specifying candidate spans in advance, we provide a fixed set of learnable vectors to learn the patterns of the valuable spans. We utilize a non-autoregressive decoder to predict the final set of entities in one pass, in which we are able to capture dependencies between entities. Compared with the sequence-to-sequence method, our model is more suitable for such unordered recognition task as it is insensitive to the label order. In addition, we utilize the loss function based on bipartite matching to compute the overall training loss. Experimental results show that our proposed model achieves state-of-the-art on three nested NER corpora: ACE 2004, ACE 2005 and KBP 2017. The code is available at https://github.com/zqtan1024/sequence-to-set. Zeqi Tan, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang |
IJCAI | 5 |
| 2021 | WAB'21: 1st Workshop on Multimodal Product Identification in Livestreaming and WAB ChallengeabstractProduct identification has become a very important component in the modern E-commerce shopping system. Consumers could enjoy watching livingstreaming and buying products that livestream hosts recommended. However, with hundreds of products presented in a livingstreaming video, finding the specific product could be laboursome for consumers. Hence, automatic product identification is desired in livingstreaming based E-commerce system. Compared with the image-based visual searching system, the complicated contents in the livestreaming videos make the identification even more challenging. To promote the research on product identification in livestreaming, we present the largest multimodal product retrieval dataset named "Watch and Buy" (WAB) and launch the multimodal product retrieval challenge. We hope this workshop could help researchers further advance the performance and applicability of livestreaming product identification in real-world systems. Yueting Zhuang, Guilin Wu, Yahong Han, Haihong Tang, Baoming Yan, Yi Yang 0001 |
ACM Multimedia | 1 |
| 2021 | Learning to Generate Visual Questions with Noisy SupervisionabstractThe task of visual question generation (VQG) aims to generate human-like neural questions from an image and potentially other side information (e.g., answer type or the answer itself). Existing works often suffer from the severe one image to many questions mapping problem, which generates uninformative and non-referential questions. Recent work has demonstrated that by leveraging double visual and answer hints, a model can faithfully generate much better quality questions. However, visual hints are not available naturally. Despite they proposed a simple rule-based similarity matching method to obtain candidate visual hints, they could be very noisy practically and thus restrict the quality of generated questions. In this paper, we present a novel learning approach for double-hints based VQG, which can be cast as a weakly supervised learning problem with noises. The key rationale is that the salient visual regions of interest can be viewed as a constraint to improve the generation procedure for producing high-quality questions. As a result, given the predicted salient visual regions of interest, we can focus on estimating the probability of being ground-truth questions, which in turn implicitly measures the quality of predicted visual hints. Experimental results on two benchmark datasets show that our proposed method outperforms the state-of-the-art approaches by a large margin on a variety of metrics, including both automatic machine metrics and human evaluation. Lingfei Wu 0001, Siliang Tang, Yueting Zhuang, Zhuoye Ding, Bo Long |
NeurIPS | 4 |
| 2021 | Hierarchical Cross-Modal Graph Consistency Learning for Video-Text RetrievalabstractDue to the popularity of video contents on the Internet, the information retrieval between videos and texts has attracted broad interest from researchers, which is a challenging cross-modal retrieval task. A common solution is to learn a joint embedding space to measure the cross-modal similarity. However, many existing approaches either pay more attention to textual information, video information, or cross-modal matching methods, but less to all three. We believe that a good video-text retrieval system should take into account all three points, fully exploiting the semantic information of both modalities and considering a comprehensive match. In this paper, we propose a Hierarchical Cross-Modal Graph Consistency Learning Network (HCGC) for video-text retrieval task, which considers multi-level graph consistency for video-text matching. Specifically, we first construct a hierarchical graph representation for the video, which includes three levels from global to local: video, clips and objects. Similarly, the corresponding text graph is constructed according to the semantic relationships among sentence, actions and entities. Then, in order to learn a better match between the video and text graph, we design three types of graph consistency (both direct and indirect): inter-graph parallel consistency, inter-graph cross consistency and intra-graph cross consistency. Extensive experimental results on different video-text datasets demonstrate the effectiveness of our approach on both text-to-video and video-to-text retrieval. Weike Jin, Zhou Zhao 0001, Jieming Zhu, Xiuqiang He 0001, Yueting Zhuang |
SIGIR | 6 |
| 2021 | Multiple knowledge representation for big data artificial intelligence: framework, applications, and case studiesabstract提出一种多重知识表示框架, 探讨了其对推动大数据人工智能技术在各个领域中发展的重要意义及深远影响. 传统知识表达和现代基于深度学习的知识表达通常着眼于利用特定变换方式, 将输入转换为符号编码或者向量. 例如, 知识图谱关注于描述各个概念之间的语义联系, 而深度神经网络更像是感知原始信号输入的工具. 多重知识表达是一种更为先进的人工智能表征框架, 具备更完整的智能功能, 比如原始信号感知、 特征提取及向量化、 知识符号化和逻辑推断. 多重知识表达有如下两点优势: (1) 与现有以深度学习为主导的人工智能技术相比, 具有更强的解释性以及更好的泛化能力; (2) 将多重知识表达集成于现有人工智能技术, 有利于各种表征 (例如原始信号感知以及符号化编码) 发挥互补优势. 我们希望多重知识表达相关研究以及应用能够驱动新一代人工智能蓬勃发展. Yi Yang 0001, Yueting Zhuang, Yunhe Pan |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Visual knowledge: an attempt to explore machine creativityabstract长期以来困扰人工智能领域的一个问题是: 人工智能是否具有创造力, 或者说, 算法的推理过程是否可以具有创造性. 本文从思维科学的角度探讨人工智能创造力的问题. 首先, 列举形象思维推理的相关研究; 然后, 重点介绍一种特殊的视觉知识表示形式, 即视觉场景图; 最后, 详细介绍视觉场景图构造问题与潜在应用. 所有证据表明, 视觉知识和视觉思维不仅可以改善当前人工智能任务的性能, 而且可以用于机器创造力的实践. Yueting Zhuang, Siliang Tang |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2021 | Tell and guess: cooperative learning for natural image caption generation with hierarchical refined attention
Wenqiao Zhang, Siliang Tang, Jiajie Su, Jun Xiao 0001, Yueting Zhuang |
Multim. Tools Appl. | 5 |
| 2021 | Adaptive Spatio-Temporal Graph Enhanced Vision-Language Representation for Video QAabstractVision-language research has become very popular, which focuses on understanding of visual contents, language semantics and relationships between them. Video question answering (Video QA) is one of the typical tasks. Recently, several BERT style pre-training methods have been proposed and shown effectiveness on various vision-language tasks. In this work, we leverage the successful vision-language transformer structure to solve the Video QA problem. However, we do not pre-train it with any video data, because video pre-training requires massive computing resources and is hard to perform with only a few GPUs. Instead, our work aims to leverage image-language pre-training to help with video-language modeling, by sharing a common module design. We further introduce an adaptive spatio-temporal graph to enhance the vision-language representation learning. That is, we adaptively refine the spatio-temporal tubes of salient objects according to their spatio-temporal relations learned through a hierarchical graph convolution process. Finally, we can obtain a number of fine-grained tube-level video object representations, as the visual inputs of the vision-language transformer module. Experiments on three widely used Video QA datasets show that our model achieves the new state-of-the-art results. Weike Jin, Zhou Zhao 0001, Xiaochun Cao, Jieming Zhu, Xiuqiang He 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 6 |
| 2021 | Mining Fraudsters and Fraudulent Strategies in Large-Scale Mobile Social NetworksabstractThe rapid development of modern communication technologies-in particular, (mobile) phone communications-has largely facilitated human social interactions and information exchange. However, the emergence of telemarketing frauds can significantly dissipate individual fortune and social wealth, resulting in a potential slow down or damage to economics. In this work, we propose to spot telemarketing frauds, with an emphasis on unveiling the “precise fraud” phenomenon and the strategies that are used by fraudsters to precisely select targets. To study this problem, we employ a one-month complete dataset of telecommunication metadata in Shanghai with 54 million users and 698 million call logs. Through our study, we find that user's information might have been seriously leaked, and fraudsters have a preference over the target user's age and activity in mobile network. We further propose a novel semi-supervised learning framework to distinguish fraudsters from non-fraudsters. Experimental results on a real-world data show that our approach outperforms several state-of-the-art algorithms in accuracy of detecting fraudsters (e.g., +0.278 in terms of F1 on average). We believe that our study can potentially inform policymaking for government and mobile service providers. Yang Yang 0009, Yuhong Xu, Yizhou Sun, Yuxiao Dong, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2021 | Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video ApplicationsabstractWe present a self-supervised spatiotemporal learning approach by exploring the temporal coherence of videos. The chronological order of shuffled clips from the video is used as the supervisory signal to guide the 3D Convolutional Neural Networks (CNNs) to learn meaningful visual knowledge. Unlike the existing approaches which use frames, we utilize dynamic video clips to reduce the uncertainty of order. We test three types of representative 3D CNNs, all of which benefit from the proposed approach. The learned 3D CNNs can be used either as a feature extractor or a pre-trained model for further fine-tuning on downstream tasks. We also propose two curriculum learning strategies to make the 3D CNNs easier to train and get the state-of-the-art results in nearest neighbor retrieval and action recognition tasks compared with other self-supervised learning methods. Meanwhile, it is further extended to the field of visual question answering application and has achieved promising results. Besides, comprehensive and extensive experimental results and analyses are provided for readers to better understand the video clip order we explore with self-supervised and curriculum learning for video application. Jun Xiao 0001, Lin Li 0065, Dejing Xu, Chengjiang Long, Jian Shao 0001, Shiliang Pu, Yueting Zhuang |
IEEE Trans. Multim. | 8 |
| 2021 | End-to-End Video Saliency Detection via a Deep Contextual Spatiotemporal NetworkabstractAs an interesting and important problem in computer vision, learning-based video saliency detection aims to discover the visually interesting regions in a video sequence. Capturing the information within frame and between frame at different aspects (such as spatial contexts, motion information, temporal consistency across frames, and multiscale representation) is important for this task. A key issue is how to jointly model all these factors within a unified data-driven scheme in an end-to-end fashion. In this article, we propose an end-to-end spatiotemporal deep video saliency detection approach, which captures the information on spatial contexts and motion characteristics. Furthermore, it encodes the temporal consistency information across the consecutive frames by implementing a convolutional long short-term memory (Conv-LSTM) model. In addition, the multiscale saliency properties for each frame are adaptively integrated for final saliency prediction in a collaborative feature-pyramid way. Finally, the proposed deep learning approach unifies all the aforementioned parts into an end-to-end joint deep learning scheme. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches. Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang, Junwei Han 0001, Mingliang Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Time2Graph: Revisiting Time Series Modeling with Dynamic ShapeletsabstractTime series modeling has attracted extensive research efforts; however, achieving both reliable efficiency and interpretability from a unified model still remains a challenging problem. Among the literature, shapelets offer interpretable and explanatory insights in the classification tasks, while most existing works ignore the differing representative power at different time slices, as well as (more importantly) the evolution pattern of shapelets. In this paper, we propose to extract time-aware shapelets by designing a two-level timing factor. Moreover, we define and construct the shapelet evolution graph, which captures how shapelets evolve over time and can be incorporated into the time series embeddings by graph embedding algorithms. To validate whether the representations obtained in this way can be applied effectively in various scenarios, we conduct experiments based on three public time series datasets, and two real-world datasets from different domains. Experimental results clearly show the improvements achieved by our approach compared with 16 state-of-the-art baselines. Ziqiang Cheng, Yang Yang 0009, Wei Wang 0059, Wenjie Hu 0003, Yueting Zhuang, Guojie Song |
AAAI | 5 |
| 2020 | Neural-DINF: A Neural Network based Framework for Measuring Document InfluenceabstractMeasuring the scholarly impact of a document without citations is an important and challenging problem. Existing approaches such as Document Influence Model (DIM) are based on dynamic topic models, which only consider the word frequency change. In this paper, we use both frequency changes and word semantic shifts to measure document influence by developing a neural network framework. Our model has three steps. Firstly, we train the word embeddings for different time periods. Subsequently, we propose an unsupervised method to align vectors for different time periods. Finally, we compute the influence value of documents. Our experimental results show that our model outperforms DIM. Changlin Yang, Siliang Tang, Yueting Zhuang |
ACL | 6 |
| 2020 | Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationabstractVisual navigation is a task of training an embodied agent by intelligently navigating to a target object (e.g., television) using only visual observations. A key challenge for current deep reinforcement learning models lies in the requirements for a large amount of training data. It is exceedingly expensive to construct sufficient 3D synthetic environments annotated with the target object information. In this paper, we focus on visual navigation in the low-resource setting, where we have only a few training environments annotated with object information. We propose a novel unsupervised reinforcement learning approach to learn transferable meta-skills (e.g., bypass obstacles, go straight) from unannotated environments without any supervisory signals. The agent can then fast adapt to visual navigation through learning a high-level master policy to combine these meta-skills, when the visual-navigation-specified reward is provided. Experimental results show that our method significantly outperforms the baseline by 53.34% relatively on SPL, and further qualitative analysis demonstrates that our method learns transferable motor primitives for visual navigation. Juncheng Li 0006, Xin Wang 0061, Siliang Tang, Haizhou Shi, Fei Wu 0001, Yueting Zhuang, William Yang Wang |
CVPR | 6 |
| 2020 | Counterfactual Samples Synthesizing for Robust Visual Question AnsweringabstractDespite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the train set and fail to generalize to the test set with different QA distributions. To reduce the language biases, several recent works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on VQA-CP. However, since the complexity of design, current methods are unable to equip the ensemble-based models with two indispensable characteristics of an ideal VQA model: 1) visual-explainable: the model should rely on the right visual regions when making decisions. 2) question-sensitive: the model should be sensitive to the linguistic variations in question. To this end, we propose a model-agnostic Counterfactual Samples Synthesizing (CSS) training scheme. The CSS generates numerous counterfactual training samples by masking critical objects in images or words in questions, and assigning different ground-truth answers. After training with the complementary samples (ie, the original and generated samples), the VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. In return, the performance of these models is further boosted. Extensive ablations have shown the effectiveness of CSS. Particularly, by building on top of the model LMH, we achieve a record-breaking performance of 58.95% on VQA-CP v2, with 6.5% gains. Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shiliang Pu, Yueting Zhuang |
CVPR | 6 |
| 2020 | De-Biased Court's View Generation with CausalityabstractYiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu, Changlong Sun, Jun Xiao, Yueting Zhuang, Luo Si, Fei Wu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yiquan Wu 0001, Kun Kuang 0001, Xiaozhong Liu 0001, Changlong Sun, Jun Xiao 0001, Yueting Zhuang, Luo Si, Fei Wu 0001 |
EMNLP (1) | 7 |
| 2020 | Hierarchical Attention Based Spatial-Temporal Graph-to-Sequence Learning for Grounded Video DescriptionabstractThe task of Grounded Video Description~(GVD) is to generate sentences whose objects can be grounded with the bounding boxes in the video frames. Existing works often fail to exploit structural information both in modeling the relationships among the region proposals and in attending them for text generation. To address these issues, we cast the GVD task as a spatial-temporal Graph-to-Sequence learning problem, where we model video frames as spatial-temporal sequence graph in order to better capture implicit structural relationships. In particular, we exploit two ways to construct a sequence graph that captures spatial-temporal correlations among different objects in each frame and further present a novel graph topology refinement technique to discover optimal underlying graph structure. In addition, we also present hierarchical attention mechanism to attend sequence graph in different resolution levels for better generating the sentences. Our extensive experiments demonstrate the effectiveness of our proposed method compared to state-of-the-art methods. Lingfei Wu 0001, Fangli Xu, Siliang Tang, Jun Xiao 0001, Yueting Zhuang |
IJCAI | 6 |
| 2020 | Topic Adaptation and Prototype Encoding for Few-Shot Visual StorytellingabstractVisual Storytelling~(VIST) is a task to tell a narrative story about a certain topic according to the given photo stream. The existing studies focus on designing complex models, which rely on a huge amount of human-annotated data. However, the annotation of VIST is extremely costly and many topics cannot be covered in the training dataset due to the long-tail topic distribution. In this paper, we focus on enhancing the generalization ability of the VIST model by considering the few-shot setting. Inspired by the way humans tell a story, we propose a topic adaptive storyteller to model the ability of inter-topic generalization. In practice, we apply the gradient-based meta-learning algorithm on multi-modal seq2seq models to endow the model the ability to adapt quickly from topic to topic. Besides, We further propose a prototype encoding structure to model the ability of intra-topic derivation. Specifically, we encode and restore the few training story text to serve as a reference to guide the generation at inference time. Experimental results show that topic adaptation and prototype encoding structure mutually bring benefit to the few-shot model on BLEU and METEOR metric. The further case study shows that the stories generated after few-shot adaptation are more relative and expressive. Jiacheng Li 0002, Siliang Tang, Juncheng Li 0006, Jun Xiao 0001, Fei Wu 0001, Shiliang Pu, Yueting Zhuang |
ACM Multimedia | 7 |
| 2020 | Photo Stream Question AnswerabstractUnderstanding and reasoning over partially observed visual clues are often regarded as a challenging real-world problem even for human beings. In this paper, we present a new visual question answering (VQA) task -- Photo Stream QA, which aims to answer the open-ended questions about a narrative photo stream. Photo Stream QA is more challenging and interesting than the existing VQA tasks, since the temporal and visual variance among photos in the stream is huge and hard to observe. Therefore, instead of learning simple vision-text mappings, the AI algorithms must fill these variance gaps with more recollection, reasoning, even the knowledge from our daily experiences. To tackle the problems in Photo Stream QA, we propose an end-to-end baseline (E-TAA) with a novel Experienced Unit (E-unit) and Three-stage Alternating Attention (TAA). E-unit yields a better visual representation which captures the temporal semantic relation among visual clues in the photo stream, while TAA creates three levels of attention that gradually refines visual features by using the textual representation from the question as the guidance. Experimental results on our developed dataset demonstrate that, as the first attempt at the Photo Stream QA task, E-TAA provides promising results outperforming all the other baseline methods. Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Jun Xiao 0001, Shiliang Pu, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 7 |
| 2020 | Relational Graph Learning for Grounded Video Description GenerationabstractGrounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the model from hallucinating object words in its description. However, such design mainly focuses on object word generation and thus may ignore fine-grained information and suffer from missing visual concepts. Moreover, relational words (e.g., 'jump left or right') are usual spatio-temporal inference results, i.e., these words cannot be grounded on certain spatial regions. To tackle the above limitations, we design a novel relational graph learning framework for GVD, in which a language-refined scene graph representation is designed to explore fine-grained visual concepts. Furthermore, the refined graph can be regarded as relational inductive knowledge to assist captioning models in selecting the relevant information it needs to generate correct words. We validate the effectiveness of our model through automatic metrics and human evaluation, and the results indicate that our approach can generate more fine-grained and accurate description, and it solves the problem of object hallucination to some extent. Wenqiao Zhang, Xin Wang 0061, Siliang Tang, Haizhou Shi, Jun Xiao 0001, Yueting Zhuang, William Yang Wang |
ACM Multimedia | 7 |
| 2020 | Pixel-Level Cycle Association: A New Perspective for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to train a model performing satisfactory pixel-level predictions on the target with only out-of-domain (source) annotations. The conventional solution to this task is to minimize the discrepancy between source and target to enable effective knowledge transfer. Previous domain discrepancy minimization methods are mainly based on the adversarial training. They tend to consider the domain discrepancy globally, which ignore the pixel-wise relationships and are less discriminative. In this paper, we propose to build the pixel-level cycle association between source and target pixel pairs and contrastively strengthen their connections to diminish the domain gap and make the features more discriminative. To the best of our knowledge, this is a new perspective for tackling such a challenging task. Experiment results on two representative domain adaptation benchmarks, i.e. GTAV $\rightarrow$ Cityscapes and SYNTHIA $\rightarrow$ Cityscapes, verify the effectiveness of our proposed method and demonstrate that our method performs favorably against previous state-of-the-arts. Our method can be trained end-to-end in one stage and introduce no additional parameters, which is expected to serve as a general framework and help ease future research in domain adaptive semantic segmentation. Code is available at https://github.com/kgl-prml/Pixel-Level-Cycle-Association. Guoliang Kang, Yunchao Wei, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
NeurIPS | 4 |
| 2020 | Bi-Decoder Augmented Network for Neural Machine Translation
Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001 |
Neurocomputing | 4 |
| 2020 | Learning embeddings of a heterogeneous behavior network for potential behavior predictionabstractPotential behavior prediction involves understanding the latent human behavior of specific groups, and can assist organizations in making strategic decisions. Progress in information technology has made it possible to acquire more and more data about human behavior. In this paper, we examine behavior data obtained in real-world scenarios as an information network composed of two types of objects (humans and actions) associated with various attributes and three types of relationships (human-human, human-action, and action-action), which we call the heterogeneous behavior network (HBN). To exploit the abundance and heterogeneity of the HBN, we propose a novel network embedding method, human-action-attribute-aware heterogeneous network embedding (a 4 HNE), which jointly considers structural proximity, attribute resemblance, and heterogeneity fusion. Experiments on two real-world datasets show that this approach outperforms other similar methods on various heterogeneous information network mining tasks for potential behavior prediction. Shiliang Pu, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2020 | Human-Centric Clothing Segmentation via Deformable Semantic Locality-Preserving NetworkabstractIn the fields of computer vision and graphics, clothing segmentation is a challenging and practical task which is typically implemented in a fine-grained semantic segmentation framework. Unlike the generic semantic segmentation task, clothing segmentation has some domain-specific properties such as diverse appearance variations, non-rigid geometry deformations, and small sample learning. To deal with these points, we propose a semantic locality-preserving segmentation model, which adaptively attaches an original clothing image with a semantically similar (e.g., appearance or pose) auxiliary exemplar by search. Through considering the interactions of the clothing image and its exemplar, more intrinsic knowledge about the locality manifold structures of clothing images is discovered to make the learning process of small sample problem more stable and tractable. Besides, we present a CNN model based on the deformable convolutions to extract the non-rigid geometry-aware features for clothing images. Furthermore, we apply our semantic locality-preserving segmentation model in both image and video cases, resulting in favorable clothing segmentation performance. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art approaches. Wei Ji 0008, Xi Li 0001, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Context-Aware Graph Label Propagation Network for Saliency DetectionabstractRecently, a large number of existing methods for saliency detection have mainly focused on designing complex network architectures to aggregate powerful features from backbone networks. However, contextual information is not well utilized, which often causes false background regions and blurred object boundaries. Motivated by these issues, we propose an easyto-implement module that utilizes the edge-preserving ability of superpixels and the graph neural network to interact the context of superpixel nodes. In more detail, we first extract the features from the backbone network and obtain the superpixel information of images. This step is followed by superpixel pooling in which we transfer the irregular superpixel information to a structured feature representation. To propagate the information among the foreground and background regions, we use a graph neural network and self-attention layer to better evaluate the degree of saliency degree. Additionally, an affinity loss is proposed to regularize the affinity matrix to constrain the propagation path. Moreover, we extend our module to a multiscale structure with different numbers of superpixels. Experiments on five challenging datasets show that our approach can improve the performance of three baseline methods in terms of some popular evaluation metrics. Wei Ji 0008, Xi Li 0001, Lina Wei, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 5 |
| 2020 | Open-Ended Video Question Answering via Multi-Modal Conditional Adversarial NetworksabstractAs a challenging task in visual information retrieval, open-ended long-form video question answering automatically generates the natural language answer from the referenced video content according to the given question. However, the existing video question answering works mainly focus on the short-form video, which may be ineffectively applied for long-form video question answering directly, due to the insufficiency of modeling the semantic representation of long-form video content. In this paper, we study the problem of open-ended long-form video question answering from the viewpoint of hierarchical multimodal conditional adversarial network learning. We propose the hierarchical attentional encoder network to learn the joint representation of long-form video content and given question with adaptive video segmentation. We then devise the reinforced decoder network to generate the natural language answer for openended video question answering with multi-modal conditional adversarial network learning. We construct three large-scale open-ended video question answering datasets. The extensive experiments validate the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zehan Song, Chujie Lu, Jun Xiao 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 6 |
| 2020 | MRFN: Multi-Receptive-Field Network for Fast and Accurate Single Image Super-ResolutionabstractRecently, convolutional neural network (CNN) based models have shown great potential in the task of single image superresolution (SISR). However, many state-of-the-art SISR solutions are reproducing some tricks proven effective in other vision tasks, such as pursuing a deeper model. In this paper, we propose a new solution (named as Multi-Receptive-Field Network - MRFN), which outperforms existing SISR solutions in three different aspects. First, from receptive field: a novel multi-receptive-field (MRF) module is proposed to extract and fuse features in different receptive fields from local to global. Integrating these hierarchical features can generate better mappings on recovering high-fidelity details at different scales. Second, from network architectures: both dense skip connections and deep supervision are utilized to combine features from the current MRF module and preceding ones for training more representative features. Moreover, a deconvolution layer is embedded at the end of the network to avoid artificial priors induced by numerical data pre-processing (e.g., bicubic stretching), and speed up the restoration process. Finally, from error modeling: different from L1 and L2 loss functions, we proposed a novel two-parameter training loss called Weighted Huber loss function which can adaptively adjust the value of back-propagated derivative according to the residual value, thus fit the reconstruction error more effectively. Extensive qualitative and quantitative evaluation results on benchmark datasets demonstrate that our proposed MRFN can achieve more accurate recovering results than most state-of-the-art methods with significantly less complexity. Zewei He, Yanpeng Cao, Baobei Xu, Jiangxin Yang, Yanlong Cao, Siliang Tang, Yueting Zhuang |
IEEE Trans. Multim. | 8 |
| 2020 | Frame Augmented Alternating Attention Network for Video Question AnsweringabstractVision and language understanding is one of the most fundamental and challenging problems in Multimedia Intelligence. Simultaneously understanding video actions with a related natural language question, and further produces accurate answer is even more challenging since it requires joint modeling information across modality. In the past few years, some studies begin to attack this problem by utilizing attention enhanced deep neural networks. However, simple attention mechanisms such as unidirectional attention fail to yield a better mapping between different modalities. Moreover, none of these Video QA models explore high-level semantics in augmented video-frame level. In this paper, we augmented each frame representation with its context information by a novel feature extractor that combines the advantages of Resnet and a variant of C3D. In addition, we proposed a novel alternating attention network which can alternately attend frame regions, video frames and words in the question in multi-turns. This yields better joint representations of video and question, further help the deep model to discover the deeper relationship between two modalities. Our method outperforms the state-of-the-art Video QA models on two existing video question answering datasets. Further ablation studies proved that our feature extractor and the alternating attention mechanism can improve the performance jointly. Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Shiliang Pu, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Multim. | 6 |
| 2020 | Multichannel Attention Refinement for Video Question AnsweringabstractVideo Question Answering (VideoQA) is the extension of image question answering (ImageQA) in the video domain. Methods are required to give the correct answer after analyzing the provided video and question in this task. Comparing to ImageQA, the most distinctive part is the media type. Both tasks require the understanding of visual media, but VideoQA is much more challenging, mainly because of the complexity and diversity of videos. Particularly, working with the video needs to model its inherent temporal structure and analyze the diverse information it contains. In this article, we propose to tackle the task from a multichannel perspective. Appearance, motion, and audio features are extracted from the video, and question-guided attentions are refined to generate the expressive clues that support the correct answer. We also incorporate the relevant text information acquired from Wikipedia as an attempt to extend the capability of the method. Experiments on TGIF-QA and ActivityNet-QA datasets show the advantages of our method compared to existing methods. We also demonstrate the effectiveness and interpretability of our method by analyzing the refined attention weights during the question-answering procedure. Yueting Zhuang, Dejing Xu, Wenzhuo Cheng, Zhou Zhao 0001, Shiliang Pu, Jun Xiao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Heterogeneous Attributed Network Embedding with Graph Convolutional NetworksabstractNetwork embedding which assigns nodes in networks to lowdimensional representations has received increasing attention in recent years. However, most existing approaches, especially the spectral-based methods, only consider the attributes in homogeneous networks. They are weak for heterogeneous attributed networks that involve different node types as well as rich node attributes and are common in real-world scenarios. In this paper, we propose HANE, a novel network embedding method based on Graph Convolutional Networks, that leverages both the heterogeneity and the node attributes to generate high-quality embeddings. The experiments on the real-world dataset show the effectiveness of our method. Ziheng Duan, Binbing Liao, Fei Wu 0001, Yueting Zhuang |
AAAI | 5 |
| 2019 | ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question AnsweringabstractRecent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos. Zhou Yu 0001, Dejing Xu, Jun Yu 0002, Ting Yu 0016, Zhou Zhao 0001, Yueting Zhuang, Dacheng Tao |
AAAI | 6 |
| 2019 | Cross-Relation Cross-Bag Attention for Distantly-Supervised Relation ExtractionabstractDistant supervision leverages knowledge bases to automatically label instances, thus allowing us to train relation extractor without human annotations. However, the generated training data typically contain massive noise, and may result in poor performances with the vanilla supervised learning. In this paper, we propose to conduct multi-instance learning with a novel Cross-relation Cross-bag Selective Attention (C2SA), which leads to noise-robust training for distant supervised relation extractor. Specifically, we employ the sentence-level selective attention to reduce the effect of noisy or mismatched sentences, while the correlation among relations were captured to improve the quality of attention weights. Moreover, instead of treating all entity-pairs equally, we try to pay more attention to entity-pairs with a higher quality. Similarly, we adopt the selective attention mechanism to achieve this goal. Experiments with two types of relation extractor demonstrate the superiority of the proposed approach over the state-of-the-art, while further ablation studies verify our intuitions and demonstrate the effectiveness of our proposed two techniques. Yujin Yuan, Siliang Tang, Zhongfei Zhang, Yueting Zhuang, Shiliang Pu, Fei Wu 0001, Xiang Ren 0001 |
AAAI | 5 |
| 2019 | Understanding Default Behavior in Online LendingabstractMicrocredit, very small loans given out without any collaterals, is a new form of financial instrument that serves the segment of population that are typically underserved by traditional financial services. When microcredit takes the form of lending over the internet, it has the advantage of easy online application process and fast funding for borrowers, as well as attractive rate of return for individual lenders. For platforms that facilitate such activities, the key challenge lies in risk management, i.e. adequately pricing each loan's risk so as to balance borrowers' lending cost and lenders' risk-adjusted return. In fact, identifying default borrowers is of critical importance for the ecosystem. Traditionally, credit risk depends heavily on borrowers' historical loan records. However, most borrowers do not have any bureau history, and therefore cannot provide sufficient loan records. In this paper, we study default prediction in online lending by using social behavior. Specifically, we based our work on a dataset provided by PPDai, one of the leading platforms in China. Our dataset consists of over 11 million users and more than 1.5 billion call logs between them. We establish a mobile network and explore social factors that predict borrowers' default. Based on this, we focused on cheating agents, who recruit and teach borrowers to cheat by providing false information and faking application materials. Cheating agents represent a type of default, especially detrimental to the system. We propose a novel probabilistic framework to identify default borrowers and cheating agents simultaneously. Experimental results on production dataset demonstrate significant improvement over several baseline methods. Moreover, our model can effectively identify cheating agents without any labels. Yang Yang 0009, Yuhong Xu, Chunping Wang 0001, Yizhou Sun, Fei Wu 0001, Yueting Zhuang |
CIKM | 6 |
| 2019 | Self-Supervised Spatiotemporal Learning via Video Clip Order PredictionabstractWe propose a self-supervised spatiotemporal learning technique which leverages the chronological order of videos. Our method can learn the spatiotemporal representation of the video by predicting the order of shuffled clips from the video. The category of the video is not required, which gives our technique the potential to take advantage of infinite unannotated videos. There exist related works which use frames, while compared to frames, clips are more consistent with the video dynamics. Clips can help to reduce the uncertainty of orders and are more appropriate to learn a video representation. The 3D convolutional neural networks are utilized to extract features for clips, and these features are processed to predict the actual order. The learned representations are evaluated via nearest neighbor retrieval experiments. We also use the learned networks as the pre-trained models and finetune them on the action recognition task. Three types of 3D convolutional neural networks are tested in experiments, and we gain large improvements compared to existing self-supervised methods. Dejing Xu, Jun Xiao 0001, Zhou Zhao 0001, Jian Shao 0001, Di Xie, Yueting Zhuang |
CVPR | 6 |
| 2019 | Video Dialog via Progressive Inference and Cross-TransformerabstractWeike Jin, Zhou Zhao, Mao Gu, Jun Xiao, Furu Wei, Yueting Zhuang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Xiao 0001, Furu Wei, Yueting Zhuang |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Learning Dynamic Context Augmentation for Global Entity LinkingabstractXiyuan Yang, Xiaotao Gu, Sheng Lin, Siliang Tang, Yueting Zhuang, Fei Wu, Zhigang Chen, Guoping Hu, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiyuan Yang, Xiaotao Gu, Siliang Tang, Yueting Zhuang, Fei Wu 0001, Zhigang Chen 0003, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Weak Supervision Enhanced Generative Network for Question GenerationabstractAutomatic question generation according to an answer within the given passage is useful for many applications, such as question answering system, dialogue system, etc. Current neural-based methods mostly take two steps which extract several important sentences based on the candidate answer through manual rules or supervised neural networks and then use an encoder-decoder framework to generate questions about these sentences. These approaches still acquire two steps and neglect the semantic relations between the answer and the context of the whole passage which is sometimes necessary for answering the question. To address this problem, we propose the Weakly Supervision Enhanced Generative Network (WeGen) which automatically discovers relevant features of the passage given the answer span in a weakly supervised manner to improve the quality of generated questions. More specifically, we devise a discriminator, Relation Guider, to capture the relations between the passage and the associated answer and then the Multi-Interaction mechanism is deployed to transfer the knowledge dynamically for our question generation system. Experiments show the effectiveness of our method in both automatic evaluations and human evaluations. Jiyuan Zheng, Qijiong Liu, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang |
IJCAI | 6 |
| 2019 | Multi-interaction Network with Object Relation for Video Question AnsweringabstractVideo question answering is an important task for testing machine's ability of video understanding. The existing methods normally focus on the combination of recurrent and convolutional neural networks to capture spatial and temporal information of the video. Recently, some work has also shown that using attention mechanism can achieve better performance. In this paper, we propose a new model called Multi-interaction network for video question answering. There are two types of interactions in our model. The first type is the multi-modal interaction between the visual and textual information. The second type is the multi-level interaction inside the multi-modal interaction. Specifically, instead of using original self-attention, we propose a new attention mechanism called multi-interaction, which can capture both element-wise and segment-wise sequence interactions, simultaneously. And in addition to the normal frame-level interaction, we also take the object relations into consideration, in order to obtain more fine-grained information, such as motions and other potential relations among these objects. We evaluate our method on TGIF-QA and other two video QA datasets. The qualitative and quantitative experimental results show the effectiveness of our model, which achieves the new state-of-the-art performance. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang |
ACM Multimedia | 6 |
| 2019 | Informative Visual Storytelling with Cross-modal RulesabstractExisting methods in the Visual Storytelling field often suffer from the problem of generating general descriptions, while the image contains a lot of meaningful contents remaining unnoticed. The failure of informative story generation can be concluded to the model's incompetence of capturing enough meaningful concepts. The categories of these concepts include entities, attributes, actions, and events, which are in some cases crucial to grounded storytelling. To solve this problem, we propose a method to mine the cross-modal rules to help the model infer these informative concepts given certain visual input. We first build the multimodal transactions by concatenating the CNN activations and the word indices. Then we use the association rule mining algorithm to mine the cross-modal rules, which will be used for the concept inference. With the help of the cross-modal rules, the generated stories are more grounded and informative. Besides, our proposed method holds the advantages of interpretation, expandability, and transferability, indicating potential for wider application. Finally, we leverage these concepts in our encoder-decoder framework with the attention mechanism. We conduct several experiments on the VIsual StoryTelling~(VIST) dataset, the results of which demonstrate the effectiveness of our approach in terms of both automatic metrics and human evaluation. Additional experiments are also conducted showing that our mined cross-modal rules as additional knowledge helps the model gain better performance when trained on a small dataset. Jiacheng Li 0002, Haizhou Shi, Siliang Tang, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2019 | Walking with MIND: Mental Imagery eNhanceD Embodied QAabstractThe EmbodiedQA is a task of training an embodied agent by intelligently navigating in a simulated environment and gathering visual information to answer questions. Existing approaches fail to explicitly model the mental imagery function of the agent, while the mental imagery is crucial to embodied cognition, and has a close relation to many high-level meta-skills such as generalization and interpretation. In this paper, we propose a novel Mental Imagery eNhanceD (MIND) module for the embodied agent, as well as a relevant deep reinforcement framework for training. The MIND module can not only model the dynamics of the environment (e.g. 'what might happen if the agent passes through a door') but also help the agent to create a better understanding of the environment (e.g. 'The refrigerator is usually in the kitchen'). Such knowledge makes the agent a faster and better learner in locating a feasible policy with only a few trails. Furthermore, the MIND module can generate mental images that are treated as short-term subgoals by our proposed deep reinforcement framework. These mental images facilitate policy learning since short-term subgoals are easy to achieve and reusable. This yields better planning efficiency than other algorithms that learn a policy directly from primitive actions. Finally, the mental images visualize the agent's intentions in a way that human can understand, and this endows our agent's actions with more interpretability. The experimental results and further analysis prove that the agent with the MIND module is superior to its counterparts not only in EQA performance but in many other aspects such as route planning, behavioral interpretation, and the ability to generalize from a few examples. Juncheng Li 0006, Siliang Tang, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2019 | Video Relation Detection with Spatio-Temporal GraphabstractWhat we perceive from visual content are not only collections of objects but the interactions between them. Visual relations, denoted by the triplet , could convey a wealth of information for visual understanding. Different from static images and because of the additional temporal channel, dynamic relations in videos are often correlated in both spatial and temporal dimensions, which make the relation detection in videos a more complex and challenging task. In this paper, we abstract videos into fully-connected spatial-temporal graphs. We pass message and conduct reasoning in these 3D graphs with a novel VidVRD model using graph convolution network. Our model can take advantage of spatial-temporal contextual cues to make better predictions on objects as well as their dynamic relationships. Furthermore, an online association method with a siamese network is proposed for accurate relation instances association. By combining our model (VRD-GCN) and the proposed association method, our framework for video relation detection achieves the best performance in the latest benchmarks. We validate our approach on benchmark ImageNet-VidVRD dataset. The experimental results show that our framework outperforms the state-of-the-art by a large margin and a series of ablation studies demonstrate our method's effectiveness. Xufeng Qian, Yueting Zhuang, Shaoning Xiao, Shiliang Pu, Jun Xiao 0001 |
ACM Multimedia | 2 |
| 2019 | Video Dialog via Multi-Grained Convolutional Self-Attention Context NetworksabstractVideo dialog is a new and challenging task, which requires an AI agent to maintain a meaningful dialog with humans in natural language about video contents. Specifically, given a video, a dialog history and a new question about the video, the agent has to combine video information with dialog history to infer the answer. And due to the complexity of video information, the methods of image dialog might be ineffectively applied directly to video dialog. In this paper, we propose a novel approach for video dialog called multi-grained convolutional self-attention context network, which combines video information with dialog history. Instead of using RNN to encode the sequence information, we design a multi-grained convolutional self-attention mechanism to capture both element and segment level interactions which contain multi-grained sequence information. Then, we design a hierarchical dialog history encoder to learn the context-aware question representation and a two-stream video encoder to learn the context-aware video representation. We evaluate our method on two large-scale datasets. Due to the flexibility and parallelism of the new attention mechanism, our method can achieve higher time efficiency, and the extensive experiments also show the effectiveness of our method. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang |
SIGIR | 6 |
| 2019 | What Makes a Good Team? A Large-scale Study on the Effect of Team Composition in Honor of KingsabstractTeam composition is a central factor in determining the effectiveness of a team. In this paper, we present a large-scale study on the effect of team composition on multiple measures of team effectiveness. We use a dataset from the largest multiplayer online battle arena (MOBA) game, Honor of Kings, with 96 million matches involving 100 million players. We measure team effectiveness based on team performance (whether a team is going to win), team tenacity (whether a team is going to surrender), and team rapport (whether a team uses abusive language). Our results confirm the importance of team diversity with respect to player roles, and show that diversity has varying effects on team effectiveness: although diverse teams perform well and show tenacity in adversity, they are more likely to abuse when losing than less diverse teams. Our study also contributes to the situation vs. personality debate and show that abusive players tend to choose the leading role and players do not become more abusive when taking such roles. Ziqiang Cheng, Yang Yang 0009, Chenhao Tan, Denny Cheng, Yueting Zhuang |
WWW | 6 |
| 2019 | A Bilinear Ranking SVM for Knowledge Based Relation Prediction and ClassificationabstractAs an important and challenging problem, knowledge representation and inference are typically carried out in a knowledge embedding framework over a multi-relational knowledge graph, and thus have a wide range of applications such as semantic retrieval and question answering. In this paper, we propose a bilinear learning framework which performs cross-entity knowledge relation analysis in the continuous vector space (derived from knowledge embedding). In the framework, we effectively model the intrinsic correlations among different types of knowledge relations within a max-margin multi-relational ranking scheme, which jointly optimizes the tasks of entity embedding and cross-entity relation prediction in terms of multi-relational structures of the knowledge graph. Specifically, we devise a bilinear scoring function that aims to evaluate the confidence degree of semantic relation prediction for entity pairs through a multi-relational learning-to-rank pipeline. In essence, the pipeline formulates the problem of relation prediction for entity pairs as that of learning relation-specific ranking functions by max-margin optimization. Experimental results demonstrate the effectiveness of the proposed framework on two common benchmark datasets. Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Jingdong Wang 0001, Yueting Zhuang, Xuelong Li 0001 |
IEEE Trans. Big Data | 7 |
| 2019 | Deep Group-Wise Fully Convolutional Network for Co-Saliency Detection With Graph PropagationabstractA key problem in co-saliency detection is how to effectively model the interactive relationship of a whole image group and the individual perspective of each image in a united data-driven manner. In this paper, we propose a group-wise deep co-saliency detection approach to address the co-saliency object discovery problem based on the fully convolutional network (FCN). The proposed approach captures the group-wise interaction information for group images by learning a semantics-aware image representation based on a convolutional neural network, which adaptively learns the group-wise features for co-saliency detection. Furthermore, the proposed approach discovers the collaborative and interactive relationships between group-wise feature representation and single image individual feature representation, and model this in a collaborative learning framework. Then, we set up a unified deep learning scheme to jointly optimize the process of group-wise feature representation learning and the collaborative learning, leading to more reliable and robust co-saliency detection results. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches. Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 6 |
| 2019 | Video Question Answering via Knowledge-based Progressive Spatial-Temporal Attention NetworkabstractVisual Question Answering (VQA) is a challenging task that has gained increasing attention from both the computer vision and the natural language processing communities in recent years. Given a question in natural language, a VQA system is designed to automatically generate the answer according to the referenced visual content. Though there recently has been much intereset in this topic, the existing work of visual question answering mainly focuses on a single static image, which is only a small part of the dynamic and sequential visual data in the real world. As a natural extension, video question answering (VideoQA) is less explored. Because of the inherent temporal structure in the video, the approaches of ImageQA may be ineffectively applied to video question answering. In this article, we not only take the spatial and temporal dimension of video content into account but also employ an external knowledge base to improve the answering ability of the network. More specifically, we propose a knowledge-based progressive spatial-temporal attention network to tackle this problem. We obtain both objects and region features of the video frames from a region proposal network. The knowledge representation is generated by a word-level attention mechanism using the comment information of each object that is extracted from DBpedia. Then, we develop a question-knowledge-guided progressive spatial-temporal attention network to learn the joint video representation for video question answering task. We construct a large-scale video question answering dataset. The extensive experiments based on two different datasets validate the effectiveness of our method. Weike Jin, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | VPModel: High-Fidelity Product Simulation in a Virtual-Physical EnvironmentabstractIn the development of a new product, the design team must describe the expected effects of the final products to potential users and stakeholders. However, existing prototyping tools can only present a product imperfectly, due to limitations at different levels. Specifically, the physical product model, which may be the product of 3D printing, could lack a visual interface; the presentation of the product through modeling software such as Rhinoceros 3D does not provide good realistic tactile perception; or the interface platforms, such as Axure RP, used to display the interactive effects differ from those to be used in the actual operation. Thus, we present the VPModel, a high-fidelity prototyping tool, able to integrate multiple prototyping methods simultaneously. It combines a touchable 3D-printed product model (3DPM) and a corresponding visualized virtual model, and the interactive interfaces are rendered synchronously in a mixed-reality device. Through the tangible, visual, and interactive demonstration, designers and normal users can each obtain a similar experience to the experience of the finished product. Furthermore, the VPModel also enhances design practices by enabling comparisons between modular models. However, the implementation of this system is a challenging task, which subsumes several fundamental problems as sub-tasks: object detection, real-time matching, hand-gesture detection and action recognition. To achieve the expected goals of the VPModel, this system uses physical hardware (a Microsoft MR HoloLens headset, a Leap Motion Controller, and a 3D printer) and existing machine learning algorithms. To evaluate our VPModel, we report the user experience of 16 participants, evaluated using a closed-ended questionnaire survey, a quantitative analysis of task performance, and a qualitative analysis of open-ended interviews. The results show a significant improvement in realism and enjoyment using the VPModel over the two traditional camera prototype approaches. In summary, the VPModel can be used to support design strategy and to convey design concepts fully and efficiently, which indicates a potential use for the VPModel in shortening product development cycles and reducing communication costs. Xin Min, Wenqiao Zhang, Shouqian Sun, Siliang Tang, Yueting Zhuang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Multi-Label Community-Based Question Classification via Personalized Sequence Memory Network LearningabstractMulti-label community-based question classification is a challenging problem in Community-based Question Answering (CQA) services, arising in many real applications such as question navigation and expert finding. Most of the existing approaches consider the problem as content-based tag suggestion task, which suffers from the textual sparsity issue. Unlike the previous studies, we consider the problem of multi-label community-based question classification from the viewpoint of personalized sequence learning. We introduce the personalized sequence memory network that leverages not only the semantics of questions but also the personalized information of askers to provide the sequence tag learning function to capture the high-order tag dependency. The experiment on real-world dataset shows the effectiveness of our method. Xinyu Duan, Shengyu Zhang 0001, Zhou Zhao 0001, Fei Wu 0001, Yueting Zhuang |
AAAI | 5 |
| 2018 | Urban Dreams of Migrants: A Case Study of Migrant Integration in ShanghaiabstractUnprecedented human mobility has driven the rapid urbanization around the world. In China, the fraction of population dwelling in cities increased from 17.9% to 52.6% between 1978 and 2012. Such large-scale migration poses challenges for policymakers and important questions for researchers. To investigate the process of migrant integration, we employ a one-month complete dataset of telecommunication metadata in Shanghai with 54 million users and 698 million call logs. We find systematic differences between locals and migrants in their mobile communication networks and geographical locations. For instance, migrants have more diverse contacts and move around the city with a larger radius than locals after they settle down. By distinguishing new migrants (who recently moved to Shanghai) from settled migrants (who have been in Shanghai for a while), we demonstrate the integration process of new migrants in their first three weeks. Moreover, we formulate classification problems to predict whether a person is a migrant. Our classifier is able to achieve an F1-score of 0.82 when distinguishing settled migrants from locals, but it remains challenging to identify new migrants because of class imbalance. This classification setup holds promise for identifying new migrants who will successfully integrate into locals (new migrants that misclassified as locals). Yang Yang 0009, Chenhao Tan, Zongtao Liu, Fei Wu 0001, Yueting Zhuang |
AAAI | 5 |
| 2018 | Dynamic Network Embedding by Modeling Triadic Closure ProcessabstractNetwork embedding, which aims to learn the low-dimensional representations of vertices, is an important task and has attracted considerable research efforts recently. In real world, networks, like social network and biological networks, are dynamic and evolving over time. However, almost all the existing network embedding methods focus on static networks while ignore network dynamics. In this paper, we present a novel representation learning approach, DynamicTriad, to preserve both structural information and evolution patterns of a given network. The general idea of our approach is to impose triad, which is a group of three vertices and is one of the basic units of networks. In particular, we model how a closed triad, which consists of three vertices connected with each other, develops from an open triad that has two of three vertices not connected with each other. This triadic closure process is a fundamental mechanism in the formation and evolution of networks, thereby makes our model being able to capture the network dynamics and to learn representation vectors for each vertex at different time steps. Experimental results on three real-world networks demonstrate that, compared with several state-of-the-art techniques, DynamicTriad achieves substantial gains in several application scenarios. For example, our approach can effectively be applied and help to identify telephone frauds in a mobile network, and to predict whether a user will repay her loans or not in a loan network. Le-kui Zhou, Yang Yang 0009, Xiang Ren 0001, Fei Wu 0001, Yueting Zhuang |
AAAI | 5 |
| 2018 | Discourse Marker Augmented Network with Reinforcement Learning for Natural Language InferenceabstractNatural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is one of the most important problems in natural language processing.It requires to infer the logical relationship between two given sentences.While current approaches mostly focus on the interaction architectures of the sentences, in this paper, we propose to transfer knowledge from some important discourse markers to augment the quality of the NLI model.We observe that people usually use some discourse markers such as "so" or "but" to represent the logical relationship between two sentences.These words potentially have deep connections with the meanings of the sentences, thus can be utilized to help improve the representations of them.Moreover, we use reinforcement learning to optimize a new objective function with a reward defined by the property of the NLI datasets to make full use of the labels information.Experiments show that our method achieves the state-of-the-art performance on several large-scale datasets. Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001 |
ACL (1) | 4 |
| 2018 | Semantic Locality-Aware Deformable Network for Clothing SegmentationabstractClothing segmentation is a challenging vision problem typically implemented within a fine-grained semantic segmentation framework. Different from conventional segmentation, clothing segmentation has some domain-specific properties such as texture richness, diverse appearance variations, non-rigid geometry deformations, and small sample learning. To deal with these points, we propose a semantic locality-aware segmentation model, which adaptively attaches an original clothing image with a semantically similar (e.g., appearance or pose) auxiliary exemplar by search. Through considering the interactions of the clothing image and its exemplar, more intrinsic knowledge about the locality manifold structures of clothing images is discovered to make the learning process of small sample problem more stable and tractable. Furthermore, we present a CNN model based on the deformable convolutions to extract the non-rigid geometry-aware features for clothing images. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art approaches. Wei Ji 0008, Xi Li 0001, Yueting Zhuang, Omar El Farouk Bourahla, Yixin Ji, Jiabao Cui |
IJCAI | 3 |
| 2018 | Feature Enhancement in Attention for Visual Question AnsweringabstractAttention mechanism has been an indispensable part of Visual Question Answering (VQA) models, due to the importance of its selective ability on image regions and/or question words. However, attention mechanism in almost all the VQA models takes as input the image visual and question textual features, which stem from different sources and between which there exists essential semantic gap. In order to further improve the accuracy of correlation between region and question in attention, we focus on region representation and propose the idea of feature enhancement, which includes three aspects. (1) We propose to leverage region semantic representation which is more consistent with the question representation. (2) We enrich the region representation using features from multiple hierarchies and (3) we refine the semantic representation for richer information. With these three incremental feature enhancement mechanisms, we improve the region representation and achieve better attentive effect and VQA performance. We conduct extensive experiments on the largest VQA v2.0 benchmark dataset and achieve competitive results without additional training data, and prove the effectiveness of our proposed feature-enhanced attention by visual demonstrations. Yuetan Lin, Zhangyang Pang, Yueting Zhuang |
IJCAI | 4 |
| 2018 | Deep Convolutional Neural Networks with Merge-and-Run MappingsabstractA deep residual network, built by stacking a sequence of residual blocks, is easy to train, because identity mappings skip residual branches and thus improve information flow. To further reduce the training difficulty, we present a simple network architecture, deep merge-and-run neural networks. The novelty lies in a modularized building block, merge-and-run block, which assembles residual branches in parallel through a merge-and-run mapping: average the inputs of these residual branches (Merge), and add the average to the output of each residual branch as the input of the subsequent residual branch (Run), respectively. We show that the merge-and-run mapping is a linear idempotent function in which the transformation matrix is idempotent, and thus improves information flow, making training easy. In comparison with residual networks, our networks enjoy compelling advantages: they contain much shorter paths and the width, i.e., the number of channels, is increased, and the time complexity remains unchanged. We evaluate the performance on the standard recognition tasks. Our approach demonstrates consistent improvements over ResNets with the comparable setup, and achieves competitive results (e.g., 3.06% testing error on CIFAR-10, 17.55% on CIFAR-100, 1.51% on SVHN). Mingjie Li 0007, Depu Meng, Xi Li 0001, Zhaoxiang Zhang 0001, Yueting Zhuang, Zhuowen Tu, Jingdong Wang 0001 |
IJCAI | 6 |
| 2018 | Attentional Image Retweet Modeling via Multi-Faceted Ranking Network LearningabstractRetweet prediction is a challenging problem in social media sites (SMS). In this paper, we study the problem of image retweet prediction in social media, which predicts the image sharing behavior that the user reposts the image tweets from their followees. Unlike previous studies, we learn user preference ranking model from their past retweeted image tweets in SMS. We first propose heterogeneous image retweet modeling network (IRM) that exploits users' past retweeted image tweets with associated contexts, their following relations in SMS and preference of their followees. We then develop a novel attentional multi-faceted ranking network learning framework with multi-modal neural networks for the proposed heterogenous IRM network to learn the joint image tweet representations and user preference representations for prediction task. The extensive experiments on a large-scale dataset from Twitter site shows that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Lingtao Meng, Jun Xiao 0001, Min Yang 0007, Fei Wu 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 8 |
| 2018 | Open-Ended Long-form Video Question Answering via Adaptive Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced long-form video content according to the question. However, the existing video question answering works mainly focus on the short-form video question answering, due to the lack of modeling the semantic representation of long-form video contents. In this paper, we consider the problem of long-form video question answering from the viewpoint of adaptive hierarchical reinforced encoder-decoder network learning. We propose the adaptive hierarchical encoder network to learn the joint representation of the long-form video contents according to the question with adaptive video segmentation. we then develop the reinforced decoder network to generate the natural language answer for open-ended video question answering. We construct a large-scale long-form video question answering dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhou Yu 0001, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001, Yueting Zhuang |
IJCAI | 8 |
| 2018 | MacNet: Transferring Knowledge from Machine Comprehension to Sequence-to-Sequence ModelsabstractMachine Comprehension (MC) is one of the core problems in natural language processing, requiring both understanding of the natural language and knowledge about the world. Rapid progress has been made since the release of several benchmark datasets, and recently the state-of-the-art models even surpass human performance on the well-known SQuAD evaluation. In this paper, we transfer knowledge learned from machine comprehension to the sequence-to-sequence tasks to deepen the understanding of the text. We propose MacNet: a novel encoder-decoder supplementary architecture to the widely used attention-based sequence-to-sequence models. Experiments on neural machine translation (NMT) and abstractive text summarization show that our proposed framework can significantly improve the performance of the baseline models, and our method for the abstractive text summarization achieves the state-of-the-art results on the Gigaword dataset. Boyuan Pan, Yazheng Yang, Hao Li 0009, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001 |
NeurIPS | 5 |
| 2018 | To Stay or to Leave: Churn Prediction for Urban Migrants in the Initial PeriodabstractIn China, 260 million people migrate to cities to realize their urban dreams. Despite that these migrants play an important role in the rapid urbanization process, many of them fail to settle down and eventually leave the city. The integration process of migrants thus raises an important issue for scholars and policymakers. In this paper, we use Shanghai as an example to investigate migrants' behavior in their first weeks and in particular, how their behavior relates to early departure. Our dataset consists of a one-month complete dataset of 698 telecommunication logs between 54 million users, plus a novel and publicly available housing price data for 18K real estates in Shanghai. We find that migrants who end up leaving early tend to neither develop diverse connections in their first weeks nor move around the city. Their active areas also have higher housing prices than that of staying migrants. We formulate a churn prediction problem to determine whether a migrant is going to leave based on her behavior in the first few days. The prediction performance improves as we include data from more days. Interestingly, when using the same features, the classifier trained from only the first few days is already as good as the classifier trained using full data, suggesting that the performance difference mainly lies in the difference between features. Yang Yang 0009, Zongtao Liu, Chenhao Tan, Fei Wu 0001, Yueting Zhuang |
WWW | 5 |
| 2018 | Entity mention aware document representation
Siliang Tang, Fei Wu 0001, Yueting Zhuang |
Inf. Sci. | 4 |
| 2018 | Temporality-enhanced knowledgememory network for factoid question answeringabstractQuestion answering is an important problem that aims to deliver specific answers to questions posed by humans in natural language. How to efficiently identify the exact answer with respect to a given question has become an active line of research. Previous approaches in factoid question answering tasks typically focus on modeling the semantic relevance or syntactic relationship between a given question and its corresponding answer. Most of these models suffer when a question contains very little content that is indicative of the answer. In this paper, we devise an architecture named the temporality-enhanced knowledge memory network (TE-KMN) and apply the model to a factoid question answering dataset from a trivia competition called quiz bowl. Unlike most of the existing approaches, our model encodes not only the content of questions and answers, but also the temporal cues in a sequence of ordered sentences which gradually remark the answer. Moreover, our model collaboratively uses external knowledge for a better understanding of a given question. The experimental results demonstrate that our method achieves better performance than several state-of-the-art methods. Xinyu Duan, Siliang Tang, Shengyu Zhang 0001, Yin Zhang 0006, Zhou Zhao 0001, Jianru Xue, Yueting Zhuang, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2018 | Active instance matching with pairwise constraints and its application to Chinese knowledge base construction
Weiming Lu 0001, Zhenyu Zhang 0008, Yueting Zhuang |
Knowl. Inf. Syst. | 5 |
| 2018 | Multimodal Deep Embedding via Hierarchical Grounded Compositional SemanticsabstractFor a number of important problems, isolated semantic representations of individual syntactic words or visual objects do not suffice, but instead a compositional semantic representation is required; for example, a literal phrase or a set of spatially concurrent objects. In this paper, we aim to harness the existing image-sentence databases to exploit the compositional nature of image-sentence data for multimodal deep embedding. In particular, we propose an approach called hierarchical-alike (bottom-up two layers) multimodal grounded compositional semantics (hiMoCS) learning. The proposed hiMoCS systemically captures the compositional semantic connotation of multimodal data in the setting of hierarchical-alike deep learning by modeling the inherent correlations between two modalities of collaboratively grounded semantics, such as the textual entity (with its describing attribute) and visual object, the phrase (e.g., subject-verb-object triplet), and spatially concurrent objects. We argue that hiMoCS is more appropriate to reflect the multimodal compositional semantics of the image and its narrative textual sentence, which are strongly coupled. We evaluate hiMoCS on the several benchmark data sets and show that the utilization of the hiMoCS (textual entities and visual objects, textual phrase, and spatially concurrent objects) achieves a much better performance than only using the flat grounded compositional semantics. Yueting Zhuang, Jun Song 0004, Fei Wu 0001, Xi Li 0001, Zhongfei Zhang, Yong Rui |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Fusing Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM NetworksabstractRecent skeleton-based action recognition approaches achieve great improvement by using recurrent neural network (RNN) models. Currently, these approaches build an end-to-end network from coordinates of joints to class categories and improve accuracy by extending RNN to spatial domains. First, while such well-designed models and optimization strategies explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, according to the evolution of previous work, we select a set of simple geometric features, and then separately feed each type of features to a three-layer LSTM framework. Second, we propose a multistream LSTM architecture with a new smoothed score fusion technique to learn classification from different geometric feature streams. Furthermore, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and the fusion results achieve the state-of-the-art performance on four datasets. We also show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training. Songyang Zhang 0004, Yang Yang 0009, Jun Xiao 0001, Xiaoming Liu 0002, Yi Yang 0001, Di Xie, Yueting Zhuang |
IEEE Trans. Multim. | 7 |
| 2018 | Social-Aware Movie Recommendation via Multimodal Network LearningabstractWith the rapid development of Internet movie industry social-aware movie recommendation systems (SMRs) have become a popular online web service that provide relevant movie recommendations to users. In this effort many existing movie recommendation approaches learn a user ranking model from user feedback with respect to the movie's content. Unfortunately this approach suffers from the sparsity problem inherent in SMR data. In the present work we address the sparsity problem by learning a multimodal network representation for ranking movie recommendations. We develop a heterogeneous SMR network for movie recommendation that exploits the textual description and movie-poster image of each movie as well as user ratings and social relationships. With this multimodal data we then present a heterogeneous information network learning framework called SMR-multimodal network representation learning (MNRL) for movie recommendation. To learn a ranking metric from the heterogeneous information network we also developed a multimodal neural network model. We evaluated this model on a large-scale dataset from a real world SMR Web site and we find that SMR-MNRL achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Tim Weninger, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Multim. | 7 |
| 2018 | Identifying Objective and Subjective Words via Topic ModelingabstractIt is observed that distinct words in a given document have either strong or weak ability in delivering facts (i.e., the objective sense) or expressing opinions (i.e., the subjective sense) depending on the topics they associate with. Motivated by the intuitive assumption that different words have varying degree of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as dentified bjective- ubjective latent Dirichlet allocation (LDA) ( osLDA) is proposed in this paper. In the osLDA model, the simple Pólya urn model adopted in traditional topic models is modified by incorporating it with a probabilistic generative process, in which the novel "Bag-of-Discriminative-Words" (BoDW) representation for the documents is obtained; each document has two different BoDW representations with regard to objective and subjective senses, respectively, which are employed in the joint objective and subjective classification instead of the traditional Bag-of-Topics representation. The experiments reported on documents and images demonstrate that: 1) the BoDW representation is more predictive than the traditional ones; 2) osLDA boosts the performance of topic modeling via the joint discovery of latent topics and the different objective and subjective power hidden in every word; and 3) osLDA has lower computational complexity than supervised LDA, especially under an increasing number of topics. Hanqi Wang, Fei Wu 0001, Weiming Lu 0001, Yi Yang 0001, Xi Li 0001, Xuelong Li 0001, Yueting Zhuang |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2017 | Community-Based Question Answering via Asymmetric Multi-Faceted Ranking Network LearningabstractNowadays the community-based question answering (CQA) sites become the popular Internet-based web service, which have accumulated millions of questions and their posted answers over time. Thus, question answering becomes an essential problem in CQA sites, which ranks the high-quality answers to the given question. Currently, most of the existing works study the problem of question answering based on the deep semantic matching model to rank the answers based on their semantic relevance, while ignoring the authority of answerers to the given question. In this paper, we consider the problem of community-based question answering from the viewpoint of asymmetric multi-faceted ranking network embedding. We propose a novel asymmetric multi-faceted ranking network learning framework for community-based question answering by jointly exploiting the deep semantic relevance between question-answer pairs and the answerers' authority to the given question. We then develop an asymmetric ranking network learning method with deep recurrent neural networks by integrating both answers' relative quality rank to the given question and the answerers' following relations in CQA sites. The extensive experiments on a large-scale dataset from a real world CQA site show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
AAAI | 6 |
| 2017 | Integrating Side Information for Boosting Machine ComprehensionabstractMachine Reading and Comprehension recently has drawn a fair amount of attention in the field of natural language processing. In this paper, we consider integrating side information to improve machine comprehension on answering cloze-style questions more precisely. To leverage the external information, we present a novel attention-based architecture which could feed the side information representations into word level embeddings to explore the comprehension performance. Our experiments show consistent improvements of our model over various baselines. Min Yang 0007, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang |
CIKM | 6 |
| 2017 | Zero-Shot Recognition Using Dual Visual-Semantic Mapping PathsabstractZero-shot recognition aims to accurately recognize objects of unseen classes by using a shared visual-semantic mapping between the image feature space and the semantic embedding space. This mapping is learned on training data of seen classes and is expected to have transfer ability to unseen classes. In this paper, we tackle this problem by exploiting the intrinsic relationship between the semantic space manifold and the transfer ability of visual-semantic mapping. We formalize their connection and cast zero-shot recognition as a joint optimization problem. Motivated by this, we propose a novel framework for zero-shot recognition, which contains dual visual-semantic mapping paths. Our analysis shows this framework can not only apply prior semantic knowledge to infer underlying semantic manifold in the image feature space, but also generate optimized semantic embedding space, which can enhance the transfer ability of the visual-semantic mapping to unseen classes. The proposed method is evaluated for zero-shot recognition on four benchmark datasets, achieving outstanding results. Yanan Li 0002, Huanhang Hu, Yuetan Lin, Yueting Zhuang |
CVPR | 5 |
| 2017 | NITE: A Neural Inductive Teaching Framework for Domain Specific NERabstractIn domain-specific NER, due to insufficient labeled training data, deep models usually fail to behave normally.In this paper, we proposed a novel Neural Inductive TEaching framework (NITE) to transfer knowledge from existing domain-specific NER models into an arbitrary deep neural network in a teacher-student training manner.NITE is a general framework that builds upon transfer learning and multiple instance learning, which collaboratively not only transfers knowledge to a deep student network but also reduces the noise from teachers.NITE can help deep learning methods to effectively utilize existing resources (i.e., models, labeled and unlabeled data) in a small domain.The experiment resulted on Disease NER proved that without using any labeled data, NITE can significantly boost the performance of a CNN-bidirectional LSTM-CRF NER neural network nearly over 30% in terms of F1-score. Siliang Tang, Jinjian Zhang, Fei Wu 0001, Yueting Zhuang |
EMNLP | 5 |
| 2017 | Deeply-Learned Part-Aligned Representations for Person Re-identificationabstractIn this paper, we address the problem of person re-identification, which refers to associating the persons captured from different cameras. We propose a simple yet effective human part-aligned representation for handling the body part misalignment problem. Our approach decomposes the human body into regions (parts) which are discriminative for person matching, accordingly computes the representations over the regions, and aggregates the similarities computed between the corresponding regions of a pair of probe and gallery images as the overall matching score. Our formulation, inspired by attention models, is a deep neural network modeling the three steps together, which is learnt through minimizing the triplet loss function without requiring body part labeling information. Unlike most existing deep learning algorithms that learn a global or spatial partition-based local representation, our approach performs human body partition, and thus is more robust to pose changes and various human spatial distributions in the person bounding box. Our approach shows state-of-the-art results over standard datasets, Market-1501, CUHK03, CUHK01 and VIPeR. Xi Li 0001, Yueting Zhuang, Jingdong Wang 0001 |
ICCV | 3 |
| 2017 | Link Prediction via Ranking Metric Dual-Level Attention Network LearningabstractLink prediction is a challenging problem for complex network analysis, arising in many disciplines such as social networks and telecommunication networks. Currently, many existing approaches estimate the proximity of the link endpoints for link prediction from their feature or the local neighborhood around them, which suffer from the localized view of network connections and insufficiency of discriminative feature representation. In this paper, we consider the problem of link prediction from the viewpoint of learning discriminative path-based proximity ranking metric embedding. We propose a novel ranking metric network learning framework by jointly exploiting both node-level and path-level attentional proximity of the endpoints for link prediction. We then develop the path-based dual-level reasoning attentional learning method with recurrent neural network for proximity ranking metric embedding. The extensive experiments on two large-scale datasets show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Ben Gao, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 6 |
| 2017 | Microblog Sentiment Classification via Recurrent Random Walk Network LearningabstractMicroblog Sentiment Classification (MSC) is a challenging task in microblog mining, arising in many applications such as stock price prediction and crisis management. Currently, most of the existing approaches learn the user sentiment model from their posted tweets in microblogs, which suffer from the insufficiency of discriminative tweet representation. In this paper, we consider the problem of microblog sentiment classification from the viewpoint of heterogeneous MSC network embedding. We propose a novel recurrent random walk network learning framework for the problem by exploiting both users’ posted tweets and their social relations in microblogs. We then introduce the deep recurrent neural networks with random-walk layer for heterogeneous MSC network embedding, which can be trained end-to-end from the scratch. Weemploytheback-propagationmethodfortraining the proposed recurrent random walk network model. The extensive experiments on the large-scale public datasets from Twitter show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 5 |
| 2017 | Video Question Answering via Hierarchical Spatio-Temporal Attention NetworksabstractOpen-ended video question answering is a challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced video content according to the question. However, the existing visual question answering works only focus on the static image, which may be ineffectively applied to video question answering due to the temporal dynamics of video contents. In this paper, we consider the problem of open-ended video question answering from the viewpoint of spatio-temporal attentional encoder-decoder learning framework. We propose the hierarchical spatio-temporal attention network for learning the joint representation of the dynamic video contents according to the given question. We then develop the encoder-decoder learning method with reasoning recurrent neural networks for open-ended video question answering. We construct a large-scale video question answering dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 5 |
| 2017 | Detecting Temporal Proposal for Action Localization with Tree-structured Search PolicyabstractUnderstanding the semantics in videos is a complex but crucial task in video analysis. This paper focuses on localizing category-independent events, actions or other semantics in an untrimmed video, referred as salient temporal proposal localization. Traditional methods like sliding window have a high computational cost due to the densely sampling of different video segments. We propose a reinforcement learning based method, which trains a localizer that learns a search policy that, instead of exploring every video segment, finds an optimal search path to locate a salient proposal based on the currently observing video segment in a tree structure, therefore reduces the number of video segments fed into the proposal detector. In each search step, a localizer is trained to iteratively select the next sub-region containing salient proposals to continue the search, and a proposal detector is trained to recognize salient proposal from the sub-regions. The experiments demonstrate that our method is able to precisely detect salient proposals with a comparable recall and with much fewer candidate windows. Xinyang Jiang, Siliang Tang, Yang Yang 0009, Zhou Zhao 0001, Yin Zhang 0006, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 7 |
| 2017 | Video Question Answering via Gradually Refined Attention over Appearance and MotionabstractRecently image question answering (ImageQA) has gained lots of attention in the research community. However, as its natural extension, video question answering (VideoQA) is less explored. Although both tasks look similar, VideoQA is more challenging mainly because of the complexity and diversity of videos. As such, simply extending the ImageQA methods to videos is insufficient and suboptimal. Particularly, working with the video needs to model its inherent temporal structure and analyze the diverse information it contains. In this paper, we consider exploiting the appearance and motion information resided in the video with a novel attention mechanism. More specifically, we propose an end-to-end model which gradually refines its attention over the appearance and motion features of the video using the question as guidance. The question is processed word by word until the model generates the final optimized attention. The weighted representation of the video, as well as other contextual information, are used to generate the answer. Extensive experiments show the advantages of our model compared to other baseline models. We also demonstrate the effectiveness of our model by analyzing the refined attention weights during the question answering procedure. Dejing Xu, Zhou Zhao 0001, Jun Xiao 0001, Fei Wu 0001, Hanwang Zhang, Xiangnan He 0001, Yueting Zhuang |
ACM Multimedia | 7 |
| 2017 | Video Question Answering via Hierarchical Dual-Level Attention Network LearningabstractVideo question answering is a challenging task in visual information retrieval, which provides the accurate answer from the referenced video contents according to the given question. However, the existing visual question answering approaches mainly tackle the problem of static image question answering, which may be ineffectively applied for video question answering directly, due to the insufficiency of modeling the video temporal dynamics. In this paper, we study the problem of video question answering from the viewpoint of hierarchical dual-level attention network learning. We obtain the object appearance and movement information in the video based on both frame-level and segment-level feature representation methods. We then develop the hierarchical duallevel attention networks to learn the question-aware video representations with word-level and question-level attention mechanisms. We next devise the question-level fusion attention mechanism for our proposed networks to learn the questionaware joint video representation for video question answering. We construct two large-scale video question answering datasets. The extensive experiments validate the effectiveness of our method. Zhou Zhao 0001, Jinghao Lin, Xinghua Jiang, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
ACM Multimedia | 6 |
| 2017 | Panel: Cross-media IntelligenceabstractIn this panel, we attempt to review and discuss the recent emerging theoretical and technological advances and trends of cross-media. Integrating data-driven machine learning with human knowledge can effectively lead to explainable, robust, and general models. Thus, the effective employment of the interaction between cross-media data during inference and reasoning becomes a challenge to populate the cross-media knowledge graph. Some other fundamental and controversial issues such as leveraging the auxiliary information to boost the cross-media understanding, the existence of unified framework to bridge the gap between multi-modality will also be discussed in this panel. Yueting Zhuang, Ramesh Jain 0001, Wen Gao 0001, Kiyoharu Aizawa |
ACM Multimedia | 1 |
| 2017 | ENCORE: External Neural Constraints Regularized Distant Supervision for Relation ExtractionabstractDistant Supervision is a widely used approach for training relation extraction models. It generates noisy training samples by heuristically labeling a corpus using an existing knowledge base. Previous noise reduction methods for distant supervision fail to utilize information such as data credibility and sample confidence. In this paper, we proposed a novel neural framework, named ENCORE (External Neural COnstraints REgularized distant supervision), which allows an integration of other information for standard DS through regularizations under multiple external neural networks. In ENCORE, a teacher-student co-training mechanism is used to iterative distilling information from external neural networks to an existing relation extraction model. The experiment results demonstrated that without increasing any data or reshaping its original structure, ENCORE enhanced a CNN based relation extraction model for over 12%. The enhanced model also outperforms the state-of-the-art relation extraction method on the same dataset. Siliang Tang, Jinjian Zhang, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang |
SIGIR | 6 |
| 2017 | Video Question Answering via Attribute-Augmented Attention Network LearningabstractVideo Question Answering is a challenging problem in visual information retrieval, which provides the answer to the referenced video content according to the question. However, the existing visual question answering approaches mainly tackle the problem of static image question, which may be ineffectively for video question answering due to the insufficiency of modeling the temporal dynamics of video contents. In this paper, we study the problem of video question answering by modeling its temporal dynamics with frame-level attention mechanism. We propose the attribute-augmented attention network learning framework that enables the joint frame-level attribute detection and unified video representation learning for video question answering. We then incorporate the multi-step reasoning process for our proposed attention network to further improve the performance. We construct a large-scale video question answering dataset. We conduct the experiments on both multiple-choice and open-ended video question answering tasks to show the effectiveness of the proposed method. Yunan Ye, Zhou Zhao 0001, Long Chen 0016, Jun Xiao 0001, Yueting Zhuang |
SIGIR | 6 |
| 2017 | Learning Max-Margin GeoSocial Multimedia Network Representations for Point-of-Interest SuggestionabstractWith the rapid development of mobile devices, point-of-interest (POI) suggestion has become a popular online web service, which provides attractive and interesting locations to users. In order to provide interesting POIs, many existing POI recommendation works learn the latent representations of users and POIs from users' past visiting POIs, which suffers from the sparsity problem of POI data. In this paper, we consider the problem of POI suggestion from the viewpoint of learning geosocial multimedia network representations. We propose a novel max-margin metric geosocial multimedia network representation learning framework by exploiting users' check-in behavior and their social relations. We then develop a random-walk based learning method with max-margin metric network embedding. We evaluate the performance of our method on a large-scale geosocial multimedia network dataset and show that our method achieves the best performance than other state-of-the-art solutions. Zhou Zhao 0001, Hanqing Lu, Min Yang 0007, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
SIGIR | 7 |
| 2017 | Disambiguating named entities with deep supervised learning via crowd labelsabstractNamed entity disambiguation (NED) is the task of linking mentions of ambiguous entities to their referenced entities in a knowledge base such as Wikipedia. We propose an approach to effectively disentangle the discriminative features in the manner of collaborative utilization of collective wisdom (via human-labeled crowd labels) and deep learning (via human-generated data) for the NED task. In particular, we devise a crowd model to elicit the underlying features (crowd features) from crowd labels that indicate a matching candidate for each mention, and then use the crowd features to fine-tune a dynamic convolutional neural network (DCNN). The learned DCNN is employed to obtain deep crowd features to enhance traditional hand-crafted features for the NED task. The proposed method substantially benefits from the utilization of crowd knowledge (via crowd labels) into a generic deep learning for the NED task. Experimental analysis demonstrates that the proposed approach is superior to the traditional hand-crafted features when enough crowd labels are gathered. Le-kui Zhou, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2017 | Challenges and opportunities: from big data to knowledge in AI 2.0abstractIn this paper, we review recent emerging theoretical and technological advances of artificial intelligence (AI) in the big data settings. We conclude that integrating data-driven machine learning with human knowledge (common priors or implicit intuitions) can effectively lead to explainable, robust, and general AI, as follows: from shallow computation to deep neural reasoning; from merely data-driven model to data-driven with structured logic rules models; from task-oriented (domain-specific) intelligence (adherence to explicit instructions) to artificial general intelligence in a general context (the capability to learn from experience). Motivated by such endeavors, the next generation of AI, namely AI 2.0, is positioned to reinvent computing itself, to transform big data into structured knowledge, and to enable better decision-making for our society. Yueting Zhuang, Fei Wu 0001, Chun Chen 0001, Yunhe Pan |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | A human motion feature based on semi-supervised learning of GMM
Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Hanzhi Zhang, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Multim. Syst. | 5 |
| 2017 | Flickr group recommendation with auxiliary information in heterogeneous information networks
Yuanfang Xia, Siliang Tang, Fei Wu 0001, Yueting Zhuang |
Multim. Syst. | 5 |
| 2017 | Regularized Deep Belief Network for Image Attribute DetectionabstractIn general, an image attribute is a human-nameable visual property that has a semantic connotation. Appropriate modeling of the intrinsic contextual correlations among attributes plays a fundamental role in attribute detection. In this paper, we consider image attribute detection from the perspective of regularized deep learning. In particular, we propose a regularized deep belief network (rDBN) to perform the image attribute detection task, which is composed of two parts: 1) a detection DBN (dDBN) that models the joint distribution of images and their corresponding attributes, which acts as an attribute detector and 2) a contextual restricted Boltzmann machine that explicitly models the correlations among attributes acting as a regularizer that restraints the output detection result given by the dDBN to meet the contextual prior of attributes. Furthermore, we propose an efficient fine-tuning scheme that can further optimize the performance of the dDBN by backpropagation. Experimental results show that the proposed rDBN obtains improvements over the state-of-the-art methods for attribute detection on the benchmark data sets. Fei Wu 0001, Zhuhao Wang, Weiming Lu 0001, Xi Li 0001, Yi Yang 0001, Jiebo Luo 0001, Yueting Zhuang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2017 | Data-Dependent Label Distribution Learning for Age EstimationabstractAs an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples with respect to several ordinal age labels, which have intrinsically cross-age correlations across adjacent age dimensions. As a result, such correlations usually lead to the age label ambiguities of the face samples. Namely, each face sample is associated with a latent label distribution that encodes the cross-age correlation information on label ambiguities. Motivated by this observation, we propose a totally data-driven label distribution learning approach to adaptively learn the latent label distributions. The proposed approach is capable of effectively discovering the intrinsic age distribution patterns for cross-age correlation analysis on the basis of the local context structures of face samples. Without any prior assumptions on the forms of label distribution learning, our approach is able to flexibly model the sample-specific context aware label distribution properties by solving a multi-task problem, which jointly optimizes the tasks of age-label distribution learning and age prediction for individuals. Experimental results demonstrate the effectiveness of our approach. Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Xin Geng 0001, Ming-Hsuan Yang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 8 |
| 2017 | Temporal Interaction and Causal Influence in Community-Based Question AnsweringabstractDuring the last decade, community-based question answering (CQA) sites have accumulated a vast amount of questions and their crowdsourced answers over time. How to efficiently identify the quality of answers that are relevant to a given question has become an active line of research in CQA. The major challenge of CQA is the accurate selection of high-quality answers w.r.t given questions. Previous approaches tend to model the semantic matching between individual pair of one question and its corresponding answer (how fitting an answer is to a posted question). However, these works ignore the temporal interactions between answers (how previous answers influence the late posted answers). For example, a rational user likely adapts others' opinions, revises his inclinations, and posts a more appropriate answer after understanding the given question and previously posted answers. As a result, this paper devises an architecture named Temporal Interaction and Causal Influence LSTM (TC-LSTM) to effectively leverage not only the causal influence between question-answer (how appropriate an answer is for a given question) but also the temporal interactions between answers-answer (how a high-quality answer gradually forms). In particular, long short-term memory (LSTM) is used to capture the explicit question-answer influence and the implicit answers-answer interactions. Experiments are conducted on SemEval 2015 CQA dataset for answer classification task and Baidu Zhidao Dataset for answer ranking task. The experimental results show the advantage of our model comparing with other state-of-the-art methods. Fei Wu 0001, Xinyu Duan, Jun Xiao 0001, Zhou Zhao 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2017 | Bag-of-Discriminative-Words (BoDW) Representation via Topic ModelingabstractMany of the words in a given document either deliver facts (objective) or express opinions (subjective), respectively, depending on the topics they are involved in. For example, given a bunch of documents, the word “bug” assigned to the topic “order Hemiptera” apparently remarks one object (i.e., one kind of insects), while the same word assigned to the topic “software” probably conveys a negative opinion. Motivated by the intuitive assumption that different words have varying degrees of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as discriminatively objective-subjective LDA (dosLDA) is proposed in this paper. The essential idea underlying the proposed dosLDA is that a pair of objective and subjective selection variables are explicitly employed to encode the interplay between topics and discriminative power for the words in documents in a supervised manner. As a result, each document is appropriately represented as “bag-of-discriminativewords” (BoDW). The experiments reported on documents and images demonstrate that dosLDA not only performs competitively over traditional approaches in terms of topic modeling and document classification, but also has the ability to discern the discriminative power of each word in terms of its objective or subjective sense with respect to its assigned topic. Yueting Zhuang, Hanqi Wang, Jun Xiao 0001, Fei Wu 0001, Yi Yang 0001, Weiming Lu 0001, Zhongfei Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Community-Based Question Answering via Heterogeneous Social Network LearningabstractCommunity-based question answering (cQA) sites have accumulated vast amount of questions and corresponding crowdsourced answers over time. How to efficiently share the underlying information and knowledge from reliable (usually highly-reputable) answerers has become an increasingly popular research topic. A major challenge in cQA tasks is the accurate matching of high-quality answers w.r.t given questions. Many of traditional approaches likely recommend corresponding answers merely depending on the content similarity between questions and answers, therefore suffer from the sparsity bottleneck of cQA data. In this paper, we propose a novel framework which encodes not only the contents of question-answer(Q-A) but also the social interaction cues in the community to boost the cQA tasks. More specifically, our framework collaboratively utilizes the rich interaction among questions, answers and answerers to learn the relative quality rank of different answers w.r.t a same question. Moreover, the information in heterogeneous social networks is comprehensively employed to enhance the quality of question-answering (QA) matching by our deep random walk learning framework. Extensive experiments on a large-scale dataset from a real world cQA site show that leveraging the heterogeneous social information indeed achieves better performance than other state-of-the-art cQA methods. Hanyin Fang, Fei Wu 0001, Zhou Zhao 0001, Xinyu Duan, Yueting Zhuang, Martin Ester |
AAAI | 5 |
| 2016 | Relational Knowledge Transfer for Zero-Shot LearningabstractGeneral zero-shot learning (ZSL) approaches exploit transfer learning via semantic knowledge space. In this paper, we reveal a novel relational knowledge transfer (RKT) mechanism for ZSL, which is simple, generic and effective. RKT resolves the inherent semantic shift problem existing in ZSL through restoring the missing manifold structure of unseen categories via optimizing semantic mapping. It extracts the relational knowledge from data manifold structure in semantic knowledge space based on sparse coding theory. The extracted knowledge is then transferred backwards to generate virtual data for unseen categories in the feature space. On the one hand, the generalizing ability of the semantic mapping function can be enhanced with the added data. On the other hand, the mapping function for unseen categories can be learned directly from only these generated data, achieving inspiring performance. Incorporated with RKT, even simple baseline methods can achieve good results. Extensive experiments on three challenging datasets show prominent performance obtained by RKT, and we obtain 82.43% accuracy on the Animals with Attributes dataset. Yanan Li 0002, Yuetan Lin, Yueting Zhuang |
AAAI | 4 |
| 2016 | Hierarchical Recurrent Neural Encoder for Video Representation with Application to CaptioningabstractRecently, deep learning approach, especially deep Convolutional Neural Networks (ConvNets), have achieved overwhelming accuracy with fast processing speed for image classification. Incorporating temporal structure with deep ConvNets for video representation becomes a fundamental problem for video content analysis. In this paper, we propose a new approach, namely Hierarchical Recurrent Neural Encoder (HRNE), to exploit temporal information of videos. Compared to recent video representation inference approaches, this paper makes the following three contributions. First, our HRNE is able to efficiently exploit video temporal structure in a longer range by reducing the length of input information flow, and compositing multiple consecutive inputs at a higher level. Second, computation operations are significantly lessened while attaining more non-linearity. Third, HRNE is able to uncover temporal tran-sitions between frame chunks with different granularities, i.e. it can model the temporal transitions between frames as well as the transitions between segments. We apply the new method to video captioning where temporal information plays a crucial role. Experiments demonstrate that our method outperforms the state-of-the-art on video captioning benchmarks. Pingbo Pan, Zhongwen Xu, Yi Yang 0001, Fei Wu 0001, Yueting Zhuang |
CVPR | 5 |
| 2016 | Self-Paced Boost Learning for Classification
Te Pi, Xi Li 0001, Zhongfei Zhang, Deyu Meng, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang |
IJCAI | 7 |
| 2016 | Diverse Image Captioning via GroupTalk
Zhuhao Wang, Fei Wu 0001, Weiming Lu 0001, Jun Xiao 0001, Xi Li 0001, Yueting Zhuang |
IJCAI | 7 |
| 2016 | Expert Finding for Community-Based Question Answering via Ranking Metric Network Learning
Zhou Zhao 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 5 |
| 2016 | Ad Recommendation for Sponsored Search Engine via Composite Long-Short Term MemoryabstractSearch engine logs contain a large amount of users' click-through data that can be leveraged as implicit indicators of relevance. In this paper we address ad recommendation problem that finding and ranking the most relevant ads with respect to users' search queries. Due to the click sparsity, the conventional methods can hardly model the both inter- and intra-relations among users, queries and ads. We utilize the long-short term memory(LSTM) network to effectively encode two kinds of sequences: the (user, query) sequence and the query word sequence to represent users' query intention in a continuous vector space and decode them as distributions over ads respectively. Further more, we combine these two LSTM networks in an appropriate way to build up a more robust model referred as composite LSTM model(cLSTM) for ad recommendation. We evaluate the proposed cLSTM on real world click-through data set comparing with two baseline methods, the results demonstrate that our proposed model outperforms two baselines and mitigate the click sparsity problem to a certain degree. Dejiang Kong, Fei Wu 0001, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 4 |
| 2016 | Partial Multi-Modal Sparse Coding via Adaptive Similarity Structure RegularizationabstractMulti-modal sparse coding has played an important role in many multimedia applications, where data are usually with multiple modalities. Recently, various multi-modal sparse coding approaches have been proposed to learn sparse codes of multi-modal data, which assume that data appear in all modalities, or at least there is one modality containing all data. However, in real applications, it is often the case that some modalities of the data may suffer from missing information and thus result in partial multi-modality data. In this paper, we propose to solve the partial multi-modal sparse coding problem via multi-modal similarity structure regularization. Specifically, we propose a partial multi-modal sparse coding framework termed Adaptive Partial Multi-Modal Similarity Structure Regularization for Sparse Coding (AdaPM2SC), which preserves the similarity structure within the same modality and between different modalities. Experimental results conducted on two real-world datasets demonstrate that AdaPM2SC significantly outperforms the state-of-the-art methods under partial multi-modality scenario. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2016 | D-Ocean: an unstructured data management system for data ocean environment
Yueting Zhuang, Yaoguang Wang, Jian Shao 0001, Ling Chen 0001, Weiming Lu 0001, Jianling Sun, Baogang Wei, Jiangqin Wu |
Frontiers Comput. Sci. | 1 |
| 2016 | Recognizing an Action Using Its Name: A Knowledge-Based Approach
Chuang Gan 0001, Yi Yang 0001, Linchao Zhu, Deli Zhao, Yueting Zhuang |
Int. J. Comput. Vis. | 5 |
| 2016 | Online Metric-Weighted Linear Representations for Robust Visual TrackingabstractIn this paper, we propose a visual tracker based on a metric-weighted linear representation of appearance. In order to capture the interdependence of different feature dimensions, we develop two online distance metric learning methods using proximity comparison information and structured output learning. The learned metric is then incorporated into a linear representation of appearance. We show that online distance metric learning significantly improves the robustness of the tracker, especially on those sequences exhibiting drastic appearance changes. In order to bound growth in the number of training samples, we design a time-weighted reservoir sampling method. Moreover, we enable our tracker to automatically perform object identification during the process of object tracking, by introducing a collection of static template samples belonging to several object classes of interest. Object identification results for an entire video sequence are achieved by systematically combining the tracking information and visual recognition at each frame. Experimental results on challenging video sequences demonstrate the effectiveness of the method for both inter-frame tracking and object identification. Xi Li 0001, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang, Yueting Zhuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Fast view-based 3D model retrieval via unsupervised multiple feature fusion and online projection learning
Jun Xiao 0001, Yinfu Feng, Mingming Ji, Yueting Zhuang |
Signal Process. | 4 |
| 2016 | Aspect Learning for Multimedia Summarization via Nonparametric BayesianabstractSummarization is desirable for efficient comprehension of an increasingly vast amount of data. A summary of multiple documents is a concise description of the main topic. Generally speaking, a topic delivers various aspects. For example, the natural disaster topic is likely to imply the aspects of casualties and rescue. Therefore, a good summary is expected to cover all the informative aspects of a topic in order to enhance both diversity and coverage of the topic. However, for the real-world data, the profile of aspects in a given topic (e.g., the number of the aspects as well as their appropriate describing sentences or images) is hardly specified in advance. To address this problem, this paper proposes an approach to learn the hidden aspects in the topics via a nonparametric Bayesian model for multimedia summarization, namely, aspect learning for multimedia summarization via nonparametric Bayesian (ALSNB). More specifically, we introduce the priors of beta-Bernoulli process and Dirichlet process into the traditional dictionary learning. As a result, the proposed approach is able to adaptively identify the particular aspects of an individual topic. The experimental results on several datasets for text summarization and image summarization show the superiority of the proposed ALSNB over other methods. Fei Wu 0001, Hanyin Fang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yi Yang 0001, Wenwu Zhu 0001, Yueting Zhuang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2016 | Deep Learning Driven Visual Path Prediction From a Single ImageabstractCapabilities of inference and prediction are the significant components of visual systems. Visual path prediction is an important and challenging task among them, with the goal to infer the future path of a visual object in a static scene. This task is complicated as it needs high-level semantic understandings of both the scenes and underlying motion patterns in video sequences. In practice, cluttered situations have also raised higher demands on the effectiveness and robustness of models. Motivated by these observations, we propose a deep learning framework, which simultaneously performs deep feature learning for visual representation in conjunction with spatiotemporal context modeling. After that, a unified path-planning scheme is proposed to make accurate path prediction based on the analytic results returned by the deep context models. The highly effective visual representation and deep context models ensure that our framework makes a deep semantic understanding of the scenes and motion patterns, consequently improving the performance on visual path prediction task. In experiments, we extensively evaluate the model's performance by constructing two large benchmark datasets from the adaptation of video tracking datasets. The qualitative and quantitative experimental results show that our approach outperforms the state-of-the-art approaches and owns a better generalization capability. Siyu Huang, Xi Li 0001, Zhongfei Zhang, Zhouzhou He, Fei Wu 0001, Wei Liu 0005, Jinhui Tang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 8 |
| 2016 | DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object DetectionabstractA key problem in salient object detection is how to effectively model the semantic properties of salient objects in a data-driven manner. In this paper, we propose a multi-task deep saliency model based on a fully convolutional neural network with global input (whole raw images) and global output (whole saliency maps). In principle, the proposed saliency model takes a data-driven strategy for encoding the underlying saliency prior information, and then sets up a multi-task learning scheme for exploring the intrinsic correlations between saliency detection and semantic image segmentation. Through collaborative feature learning from such two correlated tasks, the shared fully convolutional layers produce effective features for object perception. Moreover, it is capable of capturing the semantic information on salient objects across different levels using the fully convolutional layers, which investigate the feature-sharing properties of salient object detection with a great reduction of feature redundancy. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches. Xi Li 0001, Lina Wei, Ming-Hsuan Yang 0001, Fei Wu 0001, Yueting Zhuang, Haibin Ling, Jingdong Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2016 | Joint Multilabel Classification With Community-Aware Label Graph LearningabstractAs an important and challenging problem in machine learning and computer vision, multilabel classification is typically implemented in a max-margin multilabel learning framework, where the inter-label separability is characterized by the sample-specific classification margins between labels. However, the conventional multilabel classification approaches are usually incapable of effectively exploring the intrinsic inter-label correlations as well as jointly modeling the interactions between inter-label correlations and multilabel classification. To address this issue, we propose a multilabel classification framework based on a joint learning approach called label graph learning (LGL) driven weighted Support Vector Machine (SVM). In principle, the joint learning approach explicitly models the inter-label correlations by LGL, which is jointly optimized with multilabel classification in a unified learning scheme. As a result, the learned label correlation graph well fits the multilabel classification task while effectively reflecting the underlying topological structures among labels. Moreover, the inter-label interactions are also influenced by label-specific sample communities (each community for the samples sharing a common label). Namely, if two labels have similar label-specific sample communities, they are likely to be correlated. Based on this observation, LGL is further regularized by the label Hypergraph Laplacian. Experimental results have demonstrated the effectiveness of our approach over several benchmark data sets. Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Yueting Zhuang, Jingdong Wang 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2016 | Learning of Multimodal Representations With Random Walks on the Click GraphabstractIn multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. With the click data collected from the users' searching behavior, existing approaches take either one-to-one paired data (text-image pairs) or ranking examples (text-query-image and/or image-query-text ranking lists) as training examples, which do not make full use of the click data, particularly the implicit connections among the data objects. In this paper, we treat the click data as a large click graph, in which vertices are images/text queries and edges indicate the clicks between an image and a query. We consider learning a multimodal representation from the perspective of encoding the explicit/implicit relevance relationship between the vertices in the click graph. By minimizing both the truncated random walk loss as well as the distance between the learned representation of vertices and their corresponding deep neural network output, the proposed model which is named multimodal random walk neural network (MRW-NN) can be applied to not only learn robust representation of the existing multimodal data in the click graph, but also deal with the unseen queries and images to support cross-modal retrieval. We evaluate the latent representation learned by MRW-NN on a public large-scale click log data set Clickture and further show that MRW-NN achieves much better cross-modal retrieval performance on the unseen queries/images than the other state-of-the-art methods. Fei Wu 0001, Jun Song 0004, Shuicheng Yan, Zhongfei Zhang, Yong Rui, Yueting Zhuang |
IEEE Trans. Image Process. | 7 |
| 2016 | Graph Regularized Feature Selection with Data ReconstructionabstractFeature selection is a challenging problem for high dimensional data processing, which arises in many real applications such as data mining, information retrieval, and pattern recognition. In this paper, we study the problem of unsupervised feature selection. The problem is challenging due to the lack of label information to guide feature selection. We formulate the problem of unsupervised feature selection from the viewpoint of graph regularized data reconstruction. The underlying idea is that the selected features not only preserve the local structure of the original data space via graph regularization, but also approximately reconstruct each data point via linear combination. Therefore, the graph regularized data reconstruction error becomes a natural criterion for measuring the quality of the selected features. By minimizing the reconstruction error, we are able to select the features that best preserve both the similarity and discriminant information in the original data. We then develop an efficient gradient algorithm to solve the corresponding optimization problem. We evaluate the performance of our proposed algorithm on text clustering. The extensive experiments demonstrate the effectiveness of our proposed approach. Zhou Zhao 0001, Xiaofei He 0001, Deng Cai 0001, Lijun Zhang 0005, Wilfred Ng, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | User Preference Learning for Online Social RecommendationabstractA social recommendation system has attracted a lot of attention recently in the research communities of information retrieval, machine learning, and data mining. Traditional social recommendation algorithms are often based on batch machine learning methods which suffer from several critical limitations, e.g., extremely expensive model retraining cost whenever new user ratings arrive, unable to capture the change of user preferences over time. Therefore, it is important to make social recommendation system suitable for real-world online applications where data often arrives sequentially and user preferences may change dynamically and rapidly. In this paper, we present a new framework of online social recommendation from the viewpoint of online graph regularized user preference learning (OGRPL), which incorporates both collaborative user-item relationship as well as item content features into an unified preference learning process. We further develop an efficient iterative procedure, OGRPL-FW which utilizes the Frank-Wolfe algorithm, to solve the proposed online optimization problem. We conduct extensive experiments on several large-scale datasets, in which the encouraging results demonstrate that the proposed algorithms obtain significantly lower errors (in terms of both RMSE and MAE) than the state-of-the-art online recommendation methods when receiving the same amount of training data in the online learning process. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Scalable Linear Visual Feature Learning via Online Parallel Nonnegative Matrix FactorizationabstractVisual feature learning, which aims to construct an effective feature representation for visual data, has a wide range of applications in computer vision. It is often posed as a problem of nonnegative matrix factorization (NMF), which constructs a linear representation for the data. Although NMF is typically parallelized for efficiency, traditional parallelization methods suffer from either an expensive computation or a high runtime memory usage. To alleviate this problem, we propose a parallel NMF method called alternating least square block decomposition (ALSD), which efficiently solves a set of conditionally independent optimization subproblems based on a highly parallelized fine-grained grid-based blockwise matrix decomposition. By assigning each block optimization subproblem to an individual computing node, ALSD can be effectively implemented in a MapReduce-based Hadoop framework. In order to cope with dynamically varying visual data, we further present an incremental version of ALSD, which is able to incrementally update the NMF solution with a low computational cost. Experimental results demonstrate the efficiency and scalability of the proposed methods as well as their applications to image clustering and image retrieval. Xueyi Zhao, Xi Li 0001, Zhongfei Zhang, Chunhua Shen, Yueting Zhuang, Lixin Gao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2016 | Effective deep learning-based multi-modal retrieval
Wei Wang 0059, Beng Chin Ooi, Dongxiang Zhang, Yueting Zhuang |
VLDB J. | 5 |
| 2015 | Exploring Semantic Inter-Class Relationships (SIR) for Zero-Shot Action RecognitionabstractAutomatically recognizing a large number of action categories from videos is of significant importance for video understanding. Most existing works focused on the design of more discriminative feature representation, and have achieved promising results when the positive samples are enough. However, very limited efforts were spent on recognizing a novel action without any positive exemplars, which is often the case in the real settings due to the large amount of action classes and the users' queries dramatic variations. To address this issue, we propose to perform action recognition when no positive exemplars of that class are provided, which is often known as the zero-shot learning. Different from other zero-shot learning approaches, which exploit attributes as the intermediate layer for the knowledge transfer, our main contribution is SIR, which directly leverages the semantic inter-class relationships between the known and unknown actions followed by label transfer learning. The inter-class semantic relationships are automatically measured by continuous word vectors, which learned by the skip-gram model using the large-scale text corpus. Extensive experiments on the UCF101 dataset validate the superiority of our method over fully-supervised approaches using few positive exemplars. Chuang Gan 0001, Ming Lin 0002, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
AAAI | 4 |
| 2015 | Structured Embedding via Pairwise Relations and Long-Range Interactions in Knowledge BaseabstractWe consider the problem of embedding entities and relations of knowledge bases into low-dimensional continuous vector spaces (distributed representations). Unlike most existing approaches, which are primarily efficient for modelling pairwise relations between entities, we attempt to explicitly model both pairwise relations and long-range interactions between entities, by interpreting them as linear operators on the low-dimensional embeddings of the entities. Therefore, in this paper we introduces Path-Ranking to capture the long-range interactions of knowledge graph and at the same time preserve the pairwise relations of knowledge graph; we call it 'structured embedding via pairwise relation and long-range interactions' (referred to as SePLi). Comparing with the-state-of-the-art models, SePLi achieves better performances of embeddings. Fei Wu 0001, Jun Song 0004, Yi Yang 0001, Xi Li 0001, Zhongfei Zhang, Yueting Zhuang |
AAAI | 6 |
| 2015 | Metric Learning Driven Multi-Task Structured Output Optimization for Robust Keypoint TrackingabstractAs an important and challenging problem in computer vision and graphics, keypoint-based object tracking is typically formulated in a spatio-temporal statistical learning framework. However, most existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this issue, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames, while spatial model consistency is modeled by solving a geometric verification based structured learning problem. Discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. Finally, the above three modules are simultaneously optimized in a joint learning scheme. Experimental results have demonstrated the effectiveness of our tracker. Xi Li 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
AAAI | 5 |
| 2015 | Sketch the Storyline with CHARCOAL: A Non-Parametric Approach
Siliang Tang, Fei Wu 0001, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang |
IJCAI | 6 |
| 2015 | Mobile Query Recommendation via Tensor Function Learning
Zhou Zhao 0001, Ruihua Song, Xing Xie 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 5 |
| 2015 | Deep Compositional Cross-modal Learning to Rank via Local-Global AlignmentabstractCross-modal retrieval is a very hot research topic that is imperative to many applications involving multi-modal data. Discovering an appropriate representation for multi-modal data and learning a ranking function are essential to boost the cross-media retrieval. Motivated by the assumption that a compositional cross-modal semantic representation (pairs of images and text) is more attractive for cross-modal ranking, this paper exploits the existing image-text databases to optimize a ranking function for cross-modal retrieval, called deep compositional cross-modal learning to rank (C2MLR). In this paper, C2MLR considers learning a multi-modal embedding from the perspective of optimizing a pairwise ranking problem while enhancing both local alignment and global alignment. In particular, the local alignment (i.e., the alignment of visual objects and textual words) and the global alignment (i.e., the image-level and sentence-level alignment) are collaboratively utilized to learn the multi-modal embedding common space in a max-margin learning to rank manner. The experiments demonstrate the superiority of our proposed C2MLR due to its nature of multi-modal compositional embedding. Xinyang Jiang, Fei Wu 0001, Xi Li 0001, Zhou Zhao 0001, Weiming Lu 0001, Siliang Tang, Yueting Zhuang |
ACM Multimedia | 7 |
| 2015 | Topic aspect-oriented summarization via group selection
Hanyin Fang, Weiming Lu 0001, Fei Wu 0001, Yin Zhang 0006, Xindi Shang, Jian Shao 0001, Yueting Zhuang |
Neurocomputing | 7 |
| 2015 | A locally weighted sparse graph regularized Non-Negative Matrix Factorization method
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang |
Neurocomputing | 4 |
| 2015 | Efficient semi-supervised multiple feature fusion with out-of-sample extension for 3D model retrieval
Mingming Ji, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Neurocomputing | 4 |
| 2015 | The classification of multi-modal data with hidden conditional random field
Xinyang Jiang, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Weiming Lu 0001, Yueting Zhuang |
Pattern Recognit. Lett. | 6 |
| 2015 | Sparse motion bases selection for human motion denoisingabstractHuman motion denoising is an indispensable step of data preprocessing for many motion data based applications. In this paper, we propose a data-driven based human motion denoising method that sparsely selects the most correlated subset of motion bases for clean motion reconstruction. Meanwhile, it takes the statistic property of two common noises, i.e., Gaussian noise and outliers, into account in deriving the objective functions. In particular, our method firstly divides each human pose into five partitions termed as poselets to gain a much fine-grained pose representation. Then, these poselets are reorganized into multiple overlapped poselet groups using a lagged window moving across the entire motion sequence to preserve the embedded spatial–temporal motion patterns. Afterward, five compacted and representative motion dictionaries are constructed in parallel by means of fast K-SVD in the training phase; they are used to remove the noise and outliers from noisy motion sequences in the testing phase by solving ℓ 1 -minimization problems. Extensive experiments show that our method outperforms its competitors. More importantly, compared with other data-driven based method, our method does not need to specifically choose the training data , it can be more easily applied to real-world applications. Jun Xiao 0001, Yinfu Feng, Mingming Ji, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang |
Signal Process. | 6 |
| 2015 | Weakly Semi-Supervised Deep Learning for Multi-Label Image AnnotationabstractIn this paper, we study leveraging both weakly labeled images and unlabeled images for multi-label image annotation. Motivated by the recent advance in deep learning, we propose an approach called weakly semi-supervised deep learning for multi-label image annotation (WeSed). In WeSed, a novel weakly weighted pairwise ranking loss is effectively utilized to handle weakly labeled images, while a triplet similarity loss is employed to harness unlabeled images. WeSed enables us to train deep convolutional neural network (CNN) with images from social networks where images are either only weakly labeled with several labels or unlabeled. We also design an efficient algorithm to sample high-quality image triplets from large image datasets to fine-tune the CNN. WeSed is evaluated on benchmark datasets for multi-label annotation. The experiments demonstrate the effectiveness of our proposed approach and show that the leverage of the weakly labeled images and unlabeled images leads to a significantly better performance. Fei Wu 0001, Zhuhao Wang, Zhongfei Zhang, Yi Yang 0001, Jiebo Luo 0001, Wenwu Zhu 0001, Yueting Zhuang |
IEEE Trans. Big Data | 7 |
| 2015 | Mining Spatial-Temporal Patterns and Structural Sparsity for Human Motion Data DenoisingabstractMotion capture is an important technique with a wide range of applications in areas such as computer vision, computer animation, film production, and medical rehabilitation. Even with the professional motion capture systems, the acquired raw data mostly contain inevitable noises and outliers. To denoise the data, numerous methods have been developed, while this problem still remains a challenge due to the high complexity of human motion and the diversity of real-life situations. In this paper, we propose a data-driven-based robust human motion denoising approach by mining the spatial-temporal patterns and the structural sparsity embedded in motion data. We first replace the regularly used entire pose model with a much fine-grained partlet model as feature representation to exploit the abundant local body part posture and movement similarities. Then, a robust dictionary learning algorithm is proposed to learn multiple compact and representative motion dictionaries from the training data in parallel. Finally, we reformulate the human motion denoising problem as a robust structured sparse coding problem in which both the noise distribution information and the temporal smoothness property of human motion have been jointly taken into account. Compared with several state-of-the-art motion denoising methods on both the synthetic and real noisy motion data, our method consistently yields better performance than its counterparts. The outputs of our approach are much more stable than that of the others. In addition, it is much easier to setup the training dataset of our method than that of the other data-driven-based methods. Yinfu Feng, Mingming Ji, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 6 |
| 2015 | Cross-Modal Learning to Rank via Latent Joint RepresentationabstractCross-modal ranking is a research topic that is imperative to many applications involving multimodal data. Discovering a joint representation for multimodal data and learning a ranking function are essential in order to boost the cross-media retrieval (i.e., image-query-text or text-query-image). In this paper, we propose an approach to discover the latent joint representation of pairs of multimodal data (e.g., pairs of an image query and a text document) via a conditional random field and structural learning in a listwise ranking manner. We call this approach cross-modal learning to rank via latent joint representation (CML²R). In CML²R, the correlations between multimodal data are captured in terms of their sharing hidden variables (e.g., topics), and a hidden-topic-driven discriminative ranking function is learned in a listwise ranking manner. The experiments show that the proposed approach achieves a good performance in cross-media retrieval and meanwhile has the capability to learn the discriminative representation of multimodal data. Fei Wu 0001, Xinyang Jiang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang |
IEEE Trans. Image Process. | 7 |
| 2015 | Probabilistic Word Selection via Topic ModelingabstractWe propose selective supervised Latent Dirichlet Allocation (ssLDA) to boost the prediction performance of the widely studied supervised probabilistic topic models. We introduce a Bernoulli distribution for each word in one given document to selectthis word as a strongly or weakly discriminative one with respect to its assigned topic. The Bernoulli distribution is parameterized by the discrimination power of the word for its assigned topic. As a result, the document is represented as a “bag-of-selective-words” instead of the probabilistic “bag-of-topics” in the topic modeling domain or the flat “bag-of-words” in the traditional natural language processing domain to form a new perspective. Inheriting the general framework of supervised LDA (sLDA), ssLDA can also predict many types of response specified by a Gaussian Linear Model (GLM). Focusing on the utilization of this word selection mechanism for singe-label document classification in this paper, we conduct the variational inference for approximating the intractable posterior and derive a maximum-likelihood estimation of parameters in ssLDA. The experiments reported on textual documents show that ssLDA not only performs competitively over “state-of-the-art” classification approaches based on both the flat “bag-of-words” and probabilistic “bag-of-topics” representation in terms of classification performance, but also has the ability to discover the discrimination power of the words specified in the topics (compatible with our rational knowledge). Yueting Zhuang, Haidong Gao, Fei Wu 0001, Siliang Tang, Yin Zhang 0006, Zhongfei Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Structured Visual Feature Learning for Classification via Supervised Probabilistic Tensor FactorizationabstractIn this paper, structured visual feature learning aims at exploiting the intrinsic structural properties of mutually correlated multimedia collections (e.g., video frames or facial images) to learn a more effective feature representation for multimedia data classification. We pose structured visual feature learning as a problem of supervised tensor factorization (STF), which is capable of effectively learning multi-view visual features from structural tensorial multimedia data. In mathematics , STF is formulated as a joint optimization framework of probabilistic inference and$\epsilon $-insensitive support vector regression. As a result, the feature representation obtained by STF not only preserves the intrinsic multi-view structural information on tensorial multimedia data, but also includes the discriminative information derived from the max-margin learning process. Using the learned discriminative visual features, we conduct a set of multimedia classification experiments on several challenging datasets, including images and videos, which demonstrate the effectiveness of our method. Xu Tan 0003, Fei Wu 0001, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yueting Zhuang |
IEEE Trans. Multim. | 6 |
| 2014 | Attribute prediction with long-range interactions via path codingabstractDue to the describable or human-nameable nature of visual attributes, the appropriate utilization of attributes has been receiving much attention in recent years in many applications. Motivated by the assumption that the long-range interactions between attributes can boost image understanding and classification, path coding is utilized in this paper to model the long-range interactions between attributes for the attribute prediction, we call it attribute prediction via a path coding penalty (abbreviated as AP2CP). AP2CP not only introduces structured sparsity penalties over paths on a directed acyclic graph, but also captures the intrinsical long-range dependent interactions between attributes. The proposed AP2CP can be efficiently solved by leveraging network flow optimization. The experiments show that the proposed AP2CP achieves a better performance in attribute prediction. Zhuhao Wang, Fei Wu 0001, Yahong Han, Jiebo Luo 0001, Qi Tian 0001, Yueting Zhuang |
ICIP | 6 |
| 2014 | Geo-informative discriminative image representation by semi-supervised hierarchical topic modelingabstractNowadays, the prevalence of sharing tourist photos to online communities has created an increasing demand for mining discriminative architecture aspects from historic landmarks. Some previous researches have demonstrated that topic models could discover discriminative features represented by meaningful visual-topics. However, they seldom exploited the indicative function of geo-tags and the hierarchy in architecture characteristics. In order to utilize this information, we proposed a semi-supervised hierarchical topic modeling approach (namely, shTM). In our approach, every image could be represented by a probability distribution over selected geo-related visual-topics from a partly randomized topic tree. We evaluated our approach on a real-world dataset with over 26 thousand geo-informative photos from Flickr. Experiments show that shTM topics could reveal more discriminative aspects of a specific architecture than other well-known image features, such as HOG and SIFT, on the tasks of automatic photo categorization and geographical information retrieval. Zijian Li 0002, Siliang Tang, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
ICME | 5 |
| 2014 | Learning Multimodal Neural Network with Ranking ExamplesabstractTo support cross-modal information retrieval, cross-modal learning to rank approaches utilize ranking examples (e.g., an example may be a text query and its corresponding ranked images) to learn appropriate ranking (similarity) function. However, the fact that each modality is represented with intrinsically different low-level features hinders these approaches from better reducing the heterogeneity-gap between the modalities and thus giving satisfactory retrieval results. In this paper, we consider learning with neural networks, from the perspective of optimizing the listwise ranking loss of the cross-modal ranking examples. The proposed model, named Cross-Modal Ranking Neural Network (CMRNN), benefits from the advance of both neural networks on learning high-level semantics and learning to rank techniques on learning ranking function, such that the learned cross-modal ranking function is implicitly embedded in the learned high-level representation for data objects with different modalities (e.g., text and imagery) to perform cross-modal retrieval directly. We compare CMRNN to existing state-of-the-art cross-modal ranking methods on two datasets and show that it achieves a better performance. Fei Wu 0001, Xi Li 0001, Yin Zhang 0006, Weiming Lu 0001, Yueting Zhuang |
ACM Multimedia | 7 |
| 2014 | Jointly Discovering Fine-grained and Coarse-grained Sentiments via Topic ModelingabstractThe ever-increasing user-generated contents in social media and other web services make it highly desirable to discover opinions of users on all kinds of topics. Motivated by the assumption that individual word and paragraph in documents will deliver fine-grained (e.g., "laudatory", "annoyed" or "boring") and coarse-grained (e.g., positive, negative or neutral) sentiments about certain topics respectively, this paper focuses on a deeper thematic level to jointly disentangle fine-grained and coarse-grained opinions towards topics in terms of sentiment analysis, named as LDA with multi-grained sentiments (MgS-LDA). As a result, the proposed MgS-LDA not only discovers the topics in social media, but also identifies opinions about a given topic in terms of fine-grained and coarse-grained sentiment. Results of several experiments show that our proposed MgS-LDA achieves better performance on both sentimental classification and topic modeling than related methods. Hanqi Wang, Fei Wu 0001, Xi Li 0001, Siliang Tang, Jian Shao 0001, Yueting Zhuang |
ACM Multimedia | 6 |
| 2014 | Multi-modal Mutual Topic Reinforce Modeling for Cross-media RetrievalabstractAs an important and challenging problem in the multimedia area, multi-modal data understanding aims to explore the intrinsic semantic information across different modalities in a collaborative manner. To address this problem, a possible solution is to effectively and adaptively capture the common cross-modal semantic information by modeling the inherent correlations between the latent topics from different modalities. Motivated by this task, we propose a supervised multi-modal mutual topic reinforce modeling (M$^3$R) approach, which seeks to build a joint cross-modal probabilistic graphical model for discovering the mutually consistent semantic topics via appropriate interactions between model factors (e.g., categories, latent topics and observed multi-modal data). In principle, M$^3$R is capable of simultaneously accomplishing the following two learning tasks: 1) modality-specific (e.g., image-specific or text-specific ) latent topic learning; and 2) cross-modal mutual topic consistency learning. By investigating the cross-modal topic-related distribution information, M$^3$R encourages to disentangle the semantically consistent cross-modal topics (containing some common semantic information across different modalities). In other words, the semantically co-occurring cross-modal topics are reinforced by M$^3$R through adaptively passing the mutually reinforced messages to each other in the model-learning process. To further enhance the discriminative power of the learned latent topic representations, M$^3$R incorporates the auxiliary information (i.e., categories or labels) into the process of Bayesian modeling, which boosts the modeling capability of capturing the inter-class discriminative information. Experimental results over two benchmark datasets demonstrate the effectiveness of the proposed M$^3$R in cross-modal retrieval. Fei Wu 0001, Jun Song 0004, Xi Li 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2014 | Cross-Media Hashing with Neural NetworksabstractCross-media hashing, which conducts cross-media retrieval by embedding data from different modalities into a common low-dimensional hamming space, has attracted intensive attention in recent years. This is motivated by the facts a) the multi-modal data is widespread, e.g., the web images on Flickr are associated with tags, and b) hashing is an effective technique towards large-scale high-dimensional data processing, which is exactly the situation of cross-media retrieval. Inspired by recent advances in deep learning, we propose a cross-media hashing approach based on multi-modal neural networks. By restricting in the learning objective a) the hash codes for relevant cross-media data being similar, and b) the hash codes being discriminative for predicting the class labels, the learned Hamming space is expected to well capture the cross-media semantic relationships and to be semantically discriminative. The experiments on two real-world data sets show that our approach achieves superior cross-media retrieval performance compared with the state-of-the-art methods. Yueting Zhuang, Zhou Yu 0001, Wei Wang 0059, Fei Wu 0001, Siliang Tang, Jian Shao 0001 |
ACM Multimedia | 1 |
| 2014 | Discriminative coupled dictionary hashing for fast cross-media retrievalabstractCross-media hashing, which conducts cross-media retrieval by embedding data from different modalities into a common low-dimensional Hamming space, has attracted intensive attention in recent years. The existing cross-media hashing approaches only aim at learning hash functions to preserve the intra-modality and inter-modality correlations, but do not directly capture the underlying semantic information of the multi-modal data. We propose a discriminative coupled dictionary hashing (DCDH) method in this paper. In DCDH, the coupled dictionary for each modality is learned with side information (e.g., categories). As a result, the coupled dictionaries not only preserve the intra-similarity and inter-correlation among multi-modal data, but also contain dictionary atoms that are semantically discriminative (i.e., the data from the same category is reconstructed by the similar dictionary atoms). To perform fast cross-media retrieval, we learn hash functions which map data from the dictionary space to a low-dimensional Hamming space. Besides, we conjecture that a balanced representation is crucial in cross-media retrieval. We introduce multi-view features on the relatively ``weak'' modalities into DCDH and extend it to multi-view DCDH (MV-DCDH) in order to enhance their representation capability. The experiments on two real-world data sets show that our DCDH and MV-DCDH outperform the state-of-the-art methods significantly on cross-media retrieval. Zhou Yu 0001, Fei Wu 0001, Yi Yang 0001, Qi Tian 0001, Jiebo Luo 0001, Yueting Zhuang |
SIGIR | 6 |
| 2014 | Hashing with List-Wise learning to rankabstractHashing techniques have been extensively investigated to boost similarity search for large-scale high-dimensional data. Most of the existing approaches formulate the their objective as a pair-wise similarity-preserving problem. In this paper, we consider the hashing problem from the perspective of optimizing a list-wise learning to rank problem and propose an approach called List-Wise supervised Hashing (LWH). In LWH, the hash functions are optimized by employing structural SVM in order to explicitly minimize the ranking loss of the whole list-wise permutations instead of merely the point-wise or pair-wise supervision. We evaluate the performance of LWH on two real-world data sets. Experimental results demonstrate that our method obtains a significant improvement over the state-of-the-art hashing approaches due to both structural large margin and list-wise ranking pursuing in a supervised manner. Zhou Yu 0001, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Jian Shao 0001, Yueting Zhuang |
SIGIR | 6 |
| 2014 | Special section on learning from multiple evidences for large scale multimedia analysis
Yi Yang 0001, Nicu Sebe, Cees Snoek, Xian-Sheng Hua 0001, Yueting Zhuang |
Comput. Vis. Image Underst. | 5 |
| 2014 | Exploiting temporal stability and low-rank structure for motion capture data refinementabstractInspired by the development of the matrix completion theories and algorithms, a low-rank based motion capture (mocap) data refinement method has been developed, which has achieved encouraging results. However, it does not guarantee a stable outcome if we only consider the low-rank property of the motion data. To solve this problem, we propose to exploit the temporal stability of human motion and convert the mocap data refinement problem into a robust matrix completion problem, where both the low-rank structure and temporal stability properties of the mocap data as well as the noise effect are considered. An efficient optimization method derived from the augmented Lagrange multiplier algorithm is presented to solve the proposed model. Besides, a trust data detection method is also introduced to improve the degree of automation for processing the entire set of the data and boost the performance. Extensive experiments and comparisons with other methods demonstrate the effectiveness of our approaches on both predicting missing data and de-noising. Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001, Rong Song |
Inf. Sci. | 3 |
| 2014 | A GPU-accelerated non-negative sparse latent semantic analysis algorithm for social tagging data
Yin Zhang 0006, Deng Yi, Baogang Wei, Yueting Zhuang |
Inf. Sci. | 4 |
| 2014 | Real-time motion data annotation via action stringabstractABSTRACT Even though there is an explosive growth of motion capture data, there is still a lack of efficient and reliable methods to automatically annotate all the motions in a database. Moreover, because of the popularity of mocap devices in home entertainment systems, real‐time human motion annotation or recognition becomes more and more imperative. This paper presents a new motion annotation method that achieves both the aforementioned two targets at the same time. It uses a probabilistic pose feature based on the Gaussian Mixture Model to represent each pose. After training a clustered pose feature model, a motion clip could be represented as an action string. Then, a dynamic programming‐based string matching method is introduced to compare the differences between action strings. Finally, in order to achieve the real‐time target, we construct a hierarchical action string structure to quickly label each given action string. The experimental results demonstrate the efficacy and efficiency of our method. Copyright © 2014 John Wiley & Sons, Ltd. Qi Tian 0001, Jun Xiao 0001, Yueting Zhuang, Hanzhi Zhang, Xiaosong Yang, Jian J. Zhang 0001, Yinfu Feng |
Comput. Animat. Virtual Worlds | 3 |
| 2014 | Multiple kernel learning with NOn-conVex group spArsity
Weiming Lu 0001, Fei Wu 0001, Yueting Zhuang |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Effective Multi-Modal Retrieval based on Stacked Auto-EncodersabstractMulti-modal retrieval is emerging as a new search paradigm that enables seamless information retrieval from various types of media. For example, users can simply snap a movie poster to search relevant reviews and trailers. To solve the problem, a set of mapping functions are learned to project high-dimensional features extracted from data of different media types into a common low-dimensional space so that metric distance measures can be applied. In this paper, we propose an effective mapping mechanism based on deep learning (i.e., stacked auto-encoders) for multi-modal retrieval. Mapping functions are learned by optimizing a new objective function, which captures both intra-modal and inter-modal semantic relationships of data from heterogeneous sources effectively. Compared with previous works which require a substantial amount of prior knowledge such as similarity matrices of intra-modal data and ranking examples, our method requires little prior knowledge. Given a large training dataset, we split it into mini-batches and continually adjust the mapping functions for each batch of input. Hence, our method is memory efficient with respect to the data volume. Experiments on three real datasets illustrate that our proposed method achieves significant improvement in search accuracy over the state-of-the-art methods. Wei Wang 0059, Beng Chin Ooi, Dongxiang Zhang, Yueting Zhuang |
Proc. VLDB Endow. | 5 |
| 2014 | Sparse Multi-Modal HashingabstractLearning hash functions across heterogenous high-dimensional features is very desirable for many applications involving multi-modal data objects. In this paper, we propose an approach to obtain the sparse codesets for the data objects across different modalities via joint multi-modal dictionary learning, which we call sparse multi-modal hashing (abbreviated as${\rm SM}^{2}{\rm H}$). In${\rm SM}^{2}{\rm H}$, both intra-modality similarity and inter-modality similarity are first modeled by a hypergraph, then multi-modal dictionaries are jointly learned by Hypergraph Laplacian sparse coding. Based on the learned dictionaries, the sparse codeset of each data object is acquired and conducted for multi-modal approximate nearest neighbor retrieval using a sensitive Jaccard metric. The experimental results show that${\rm SM}^{2}{\rm H}$outperforms other methods in terms of mAP and Percentage on two real-world data sets. Fei Wu 0001, Zhou Yu 0001, Yi Yang 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang |
IEEE Trans. Multim. | 6 |
| 2013 | Digital Library Engine: Adapting Digital Library for Cloud ComputingabstractWith the rapid growth of digital libraries, more data and smart services are involved. People come to recognize the importance of digital libraries and the convenience they might bring to the society. However, the cost of owning a digital library is quite high, and many institutions do not have the ability to run and maintain a digital library by themselves, especially for massive data and complex services which require lots of storage and computing resources. In this paper, we proposed the Digital Library Engine, which aims to provide a new Platform as a Service for fast developing and deploying digital libraries in cloud. To the best of our knowledge, it is the first work to create a PaaS system for digital libraries. With the help of Digital Library Engine, institutions only need to develop some service bundles, which can be deployed in the engine, and then their own digital libraries could be running well with features of scalability, reliability, security, extensibility, availability and manageability. The practice in CADAL and the experiments demonstrate the feasibility and efficiency of our engine. Weiming Lu 0001, Liangju Zheng, Jian Shao 0001, Baogang Wei, Yueting Zhuang |
IEEE CLOUD | 5 |
| 2013 | Supervised Nonnegative Tensor Factorization with Maximum-Margin ConstraintabstractNon-negative tensor factorization (NTF) has attracted great attention in the machine learning community. In this paper, we extend traditional non-negative tensor factorization into a supervised discriminative decomposition, referred as Supervised Non-negative Tensor Factorization with Maximum-Margin Constraint(SNTFM2). SNTFM2 formulates the optimal discriminative factorization of non-negative tensorial data as a coupled least-squares optimization problem via a maximum-margin method. As a result, SNTFM2 not only faithfully approximates the tensorial data by additive combinations of the basis, but also obtains a strong generalization power to discriminative analysis (in particularfor classification in this paper). The experimental results show the superiority of our proposed model over state-of-the-art techniques on both toy and real world data sets. Fei Wu 0001, Xu Tan 0003, Yi Yang 0001, Dacheng Tao, Siliang Tang, Yueting Zhuang |
AAAI | 6 |
| 2013 | Supervised Coupled Dictionary Learning with Group Structures for Multi-modal RetrievalabstractA better similarity mapping function across heterogeneous high-dimensional features is very desirable for many applications involving multi-modal data. In this paper, we introduce coupled dictionary learning (DL) into supervised sparse coding for multi-modal (cross-media) retrieval. We call this Supervised coupled dictionary learning with group structures for Multi-Modal retrieval (SliM2). SliM2 formulates the multi-modal mapping as a constrained dictionary learning problem. By utilizing the intrinsic power of DL to deal with the heterogeneous features, SliM2 extends unimodal DL to multi-modal DL. Moreover, the label information is employed in SliM2 to discover the shared structure inside intra-modality within the same class by a mixed norm (i.e., `l1/l2`-norm). As a result, the multimodal retrieval is conducted via a set of jointly learned mapping functions across multi-modal data. The experimental results show the effectiveness of our proposed model when applied to cross-media retrieval. Yueting Zhuang, Fei Wu 0001, Yin Zhang 0006, Weiming Lu 0001 |
AAAI | 1 |
| 2013 | πLDA: document clustering with selective structural constraintsabstractSegments, such as sentence boundaries in texts or annotated regions in images, can be considered as useful structural constraints (i.e., priors) for unsupervised topic modeling. However, some segment units (e.g., words in texts or visual words in images) inside a given segment may be irrelevant to the topic of this segment due to their characteristics. This paper proposes a model called πLDA, which introduces a latent variable π into LDA, a traditional topic model, to capture the characteristic of each segment unit. That is to say, the πLDA model is conducted to determine whether a segment unit is assigned (or selected) to the topic embedded in its corresponding segment. Compared with other approaches that assume all the segment units in one segment to share a common topic, our proposed πLDA has the selective ability to discover the discriminative segment units (e.g., informative words or visual words). Experimental results and interpretations of them are presented for demonstrating the promising performance of our method. Siliang Tang, Hanqi Wang, Jian Shao 0001, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 6 |
| 2013 | Cross-media semantic representation via bi-directional learning to rankabstractIn multimedia information retrieval, most classic approaches tend to represent different modalities of media in the same feature space. Existing approaches take either one-to-one paired data or uni-directional ranking examples (i.e., utilizing only text-query-image ranking examples or image-query-text ranking examples) as training examples, which do not make full use of bi-directional ranking examples (bi-directional ranking means that both text-query-image and image-query-text ranking examples are utilized in the training period) to achieve a better performance. In this paper, we consider learning a cross-media representation model from the perspective of optimizing a listwise ranking problem while taking advantage of bi-directional ranking examples. We propose a general cross-media ranking algorithm to optimize the bi-directional listwise ranking loss with a latent space embedding, which we call Bi-directional Cross-Media Semantic Representation Model (Bi-CMSRM). The latent space embedding is discriminatively learned by the structural large margin learning for optimization with certain ranking criteria (mean average precision in this paper) directly. We evaluate Bi-CMSRM on the Wikipedia and NUS-WIDE datasets and show that the utilization of the bi-directional ranking examples achieves a much better performance than only using the uni-directional ranking examples. Fei Wu 0001, Zhongfei Zhang, Shuicheng Yan, Yong Rui, Yueting Zhuang |
ACM Multimedia | 6 |
| 2013 | A low rank structural large margin method for cross-modal rankingabstractCross-modal retrieval is a classic research topic in multimedia information retrieval. The traditional approaches study the problem as a pairwise similarity function problem. In this paper, we consider this problem from a new perspective as a listwise ranking problem and propose a general cross-modal ranking algorithm to optimize the listwise ranking loss with a low rank embedding, which we call Latent Semantic Cross-Modal Ranking (LSCMR). The latent low-rank embedding space is discriminatively learned by structural large margin learning to optimize for certain ranking criteria directly. We evaluate LSCMR on the Wikipedia and NUS-WIDE dataset. Experimental results show that this method obtains significant improvements over the state-of-the-art methods. Fei Wu 0001, Siliang Tang, Zhongfei Zhang, Xiaofei He 0001, Yueting Zhuang |
SIGIR | 6 |
| 2013 | Hypergraph Spectral Hashing for image retrieval with heterogeneous social contexts
Yang Liu 0098, Jian Shao 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
Neurocomputing | 5 |
| 2013 | Learning for scalable multimedia representation
Qingshan Liu 0001, Yueting Zhuang |
Neurocomputing | 2 |
| 2013 | A semantic feature for human motion retrievalabstractABSTRACT With the explosive growth of motion capture data, it becomes very imperative in animation production to have an efficient search engine to retrieve motions from large motion repository. However, because of the high dimension of data space and complexity of matching methods, most of the existing approaches cannot return the result in real time. This paper proposes a high level semantic feature in a low dimensional space to represent the essential characteristic of different motion classes. On the basis of the statistic training of Gauss Mixture Model, this feature can effectively achieve motion matching on both global clip level and local frame level. Experiment results show that our approach can retrieve similar motions with rankings from large motion database in real‐time and also can make motion annotation automatically on the fly. Copyright © 2013 John Wiley & Sons, Ltd. Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2013 | Image annotation by semi-supervised cross-domain learning with group sparsity
Fei Wu 0001, Jian Shao 0001, Yueting Zhuang |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Retrieval-based cartoon gesture recognition and applications via semi-supervised heterogeneous classifiers learning
Zhang Liang, Yueting Zhuang, Yi Yang 0001, Jun Xiao 0001 |
Pattern Recognit. | 2 |
| 2012 | Adaptive Unsupervised Multi-view Feature Selection for Visual Concept Recognition
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaoming Liu 0002 |
ACCV (1) | 3 |
| 2012 | Graph-guided sparse reconstruction for region taggingabstractMany of contextual correlations co-exist within the segmented regions among images, like the visual context and semantic context. The appropriate integration and utilization of such contexts are very important to boost the performance of region tagging. Inspired by the recent advances of sparse reconstruction methods, this paper proposes an approach, called Graph-Guided Sparse Reconstruction for Region Tagging (G2SRRT). The G2SRRT consists of two steps: sparse reconstruction for testing regions and tag propagation from training regions to testing regions. In G2SRRT, graph is conducted to flexibly model the contextual correlations among regions. To integrate the graph structure learned from training regions into the sparse reconstruction, we define a Graph-Guided Fusion (G2F) penalty over the graph to encourage the sparsity of differences between two reconstruction coefficients, which corresponds to the linked regions in the graph. Guided by this G2F penalty, the highly correlated regions tend to be jointly selected for the reconstruction, which results in a better performance of region tagging. Experiments on three open benchmark image datasets demonstrate the effectiveness of the proposed algorithm. Yahong Han, Fei Wu 0001, Jian Shao 0001, Qi Tian 0001, Yueting Zhuang |
CVPR | 5 |
| 2012 | Supervised cross-collection topic modelingabstractNowadays, vast amounts of multimedia data can be obtained across different collections (or domains). Therefore, it poses significant challenges for the utilization of those cross-collection data, for examples, the summarization of similarities and differences of data across different domains (e.g., CNN and NYT), as well as finding visually similar images across different visual domains (e.g., photos, paintings and hand-drawn sketches). In this paper, a supervised cross-collection Latent Dirichlet Allocation (scLDA) approach is proposed to utilize the data across different collections. As a natural extension of traditional Latent Dirichlet Allocation (LDA), scLDA not only takes the structural priors of different collections into consideration, but also exploits the category information. The strength of this work lies in integrating topic modeling, cross-domain learning and supervised learning together. We conduct scLDA for comparative text mining as well as classification of news articles and images from different collections. The results suggest that our proposed scLDA can generate meaningful collection-specific topics and achieves better retrieval accuracy than other related topic models. Haidong Gao, Siliang Tang, Yin Zhang 0006, Dapeng Jiang, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 6 |
| 2012 | Correlated attribute transfer with multi-task graph-guided fusionabstractDue to the describable or human-nameable nature of visual attributes, the attribute-based methods have been receiving much attentions in recent years in many applications. The advantages of the utilization of visual attributes are that they can be composed to create descriptions at various levels of specificity or they can be learned once and then applied to recognize new objects or categories. Therefore, attribute prediction becomes an essential problem to boost image understanding. This paper proposes an approach for correlated attribute transfer from a well-defined source image set to an uncontrolled target image set for attribute prediction. We call it correlated attribute transfer with multi-task graph-guided fusion (CAT-MtG2F). The novelty of CAT-MtG2F is to encourage highly correlated attributes to share a common set of relevant low-level features and transfer the learned common structure from the source image set to the target image set. The experiments show that the proposed CAT-MtG2F achieves better performance in attribute prediction. Yahong Han, Fei Wu 0001, Qi Tian 0001, Yueting Zhuang, Jiebo Luo 0001 |
ACM Multimedia | 5 |
| 2012 | Annotating web images using NOVA: NOn-conVex group spArsityabstractAs image feature vector is large, selecting the right features plays a fundamental role in Web image annotation. Most existing approaches are either based on individual feature selection, which leads to local optima, or using a convex penalty, which leads to inconsistency. To address these difficulties, in this paper we propose a new sparsity-based approach NOVA (NOn-conVex group spArsity). To the best of our knowledge, NOVA is the first to introduce non-convex penalty for group selection in high-dimensional heterogeneous features space. Because it is a group-sparsity approach, it approximately reaches global optima. Because it uses non-convex penalty, it achieves the consistency. We demonstrate the superior performance of NOVA via three means. First, we present theoretical proof that NOVA is consistent, satisfying un-biasness, sparsity and continuity. Second, we show NOVA converges to the true underlying model by using a ground-truth-available generative-model simulation. Third, we report extensive experimental results on three diverse and widely-used data sets Kodak, MSRA-MM 2.0, and NUS-WIDE. We also compare NOVA against the state-of-the-art approaches, and report superior experimental results. Fei Wu 0001, Yong Rui, Shuicheng Yan, Yueting Zhuang |
ACM Multimedia | 5 |
| 2012 | Synthesizing style-preserving cartoons via non-negative style factorizationabstractWe present a complete framework for synthesizing style-preserving 2D cartoons by learning from traditional Chinese cartoons. In contrast to reusing-based approaches which rely on rearranging or retrieving existing cartoon sequences, we aim to generate stylized cartoons with the idea of style factorization. Specifically, starting with 2D skeleton features of cartoon characters extracted by an improved rotoscoping system, we present a non-negative style factorization (NNSF) algorithm to obtain style basis and weights and simultaneously preserve class separability. Thus, factorized style basis can be combined with heterogeneous weights to re-synthesize style-preserving features, and then these features are used as the driving source in the character reshaping process via our proposed subkey-driving strategy. Extensive experiments and examples demonstrate the effectiveness of the proposed framework. Zhang Liang, Jun Xiao 0001, Yueting Zhuang |
J. Zhejiang Univ. Sci. C | 3 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 4 |
| 2012 | A Multimedia Retrieval Framework Based on Semi-Supervised Ranking and Relevance FeedbackabstractWe present a new framework for multimedia content analysis and retrieval which consists of two independent algorithms. First, we propose a new semi-supervised algorithm called ranking with Local Regression and Global Alignment (LRGA) to learn a robust Laplacian matrix for data ranking. In LRGA, for each data point, a local linear regression model is used to predict the ranking scores of its neighboring points. A unified objective function is then proposed to globally align the local models from all the data points so that an optimal ranking score can be assigned to each data point. Second, we propose a semi-supervised long-term Relevance Feedback (RF) algorithm to refine the multimedia data representation. The proposed long-term RF algorithm utilizes both the multimedia data distribution in multimedia feature space and the history RF information provided by users. A trace ratio optimization problem is then formulated and solved by an efficient algorithm. The algorithms have been applied to several content-based multimedia retrieval applications, including cross-media retrieval, image retrieval, and 3D motion/pose data retrieval. Comprehensive experiments on four data sets have demonstrated its advantages in precision, robustness, scalability, and computational efficiency. Yi Yang 0001, Feiping Nie 0001, Dong Xu 0001, Jiebo Luo 0001, Yueting Zhuang, Yunhe Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | A unified framework for web video topic discovery and visualization
Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
Pattern Recognit. Lett. | 4 |
| 2012 | Dynamic Time Warping for Chinese calligraphic character matching and recognizing
Xiafen Zhang, Yueting Zhuang |
Pattern Recognit. Lett. | 2 |
| 2012 | Sparse Unsupervised Dimensionality Reduction for Multiple View DataabstractDifferent kinds of high-dimensional visual features can be extracted from a single image. Images can thus be treated as multiple view data when taking each type of extracted high-dimensional visual feature as a particular understanding of images. In this paper, we propose a framework of sparse unsupervised dimensionality reduction for multiple view data. The goal of our framework is to find a low-dimensional optimal consensus representation from multiple heterogeneous features by multiview learning. In this framework, we first learn low-dimensional patterns individually from each view, considering the specific statistical property of each view. We construct a low-dimensional optimal consensus representation from those learned patterns, the goal of which is to leverage the complementary nature of the multiple views. We formulate the construction of the low-dimensional consensus representation to approximate the matrix of patterns by means of a low-dimensional consensus base matrix and a loading matrix. To select the most discriminative features for the spectral embedding of multiple views, we propose to add anl1-norm into the loading matrix's columns and impose orthogonal constraints on the base matrix. We develop a new alternating algorithm, i.e., spectral sparse multiview embedding, to efficiently obtain the solution. Each row of the loading matrix encodes structured information corresponding to multiple patterns. In order to gain flexibility in sharing information across subsets of the views, we impose a novel structured sparsity-inducing norm penalty on the loading matrix's rows. This penalty makes the loading coefficients adaptively load shared information across subsets of the learned patterns. We call this method structured sparse multiview dimensionality reduction. Experiments on a toy benchmark image data set and two real-world Web image data sets demonstrate the effectiveness of the proposed algorithms. Yahong Han, Fei Wu 0001, Dacheng Tao, Jian Shao 0001, Yueting Zhuang, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2012 | Image Annotation by Input-Output Structural Grouping SparsityabstractAutomatic image annotation (AIA) is very important to image retrieval and image understanding. Two key issues in AIA are explored in detail in this paper, i.e., structured visual feature selection and the implementation of hierarchical correlated structures among multiple tags to boost the performance of image annotation. This paper simultaneously introduces an input and output structural grouping sparsity into a regularized regression model for image annotation. For input high-dimensional heterogeneous features such as color, texture, and shape, different kinds (groups) of features have different intrinsic discriminative power for the recognition of certain concepts. The proposed structured feature selection by structural grouping sparsity can be used not only to select group-of-features but also to conduct within-group selection. Hierarchical correlations among output labels are well represented by a tree structure, and therefore, the proposed tree-structured grouping sparsity can be used to boost the performance of multitag image annotation. In order to efficiently solve the proposed regression model, we relax the solving process as a framework of the bilayer regression model for multilabel boosting by the selection of heterogeneous features with structural grouping sparsity (Bi-MtBGS). The first-layer regression is to select the discriminative features for each label. The aim of the second-layer regression is to refine the feature selection model learned from the first layer, which can be taken as a multilabel boosting process. Extensive experiments on public benchmark image data sets and real-world image data sets demonstrate that the proposed approach has better performance of multitag image annotation and leads to a quite interpretable model for image understanding. Yahong Han, Fei Wu 0001, Qi Tian 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 4 |
| 2012 | Spline Regression Hashing for Fast Image SearchabstractTechniques for fast image retrieval over large databases have attracted considerable attention due to the rapid growth of web images. One promising way to accelerate image search is to use hashing technologies, which represent images by compact binary codewords. In this way, the similarity between images can be efficiently measured in terms of the Hamming distance between their corresponding binary codes. Although plenty of methods on generating hash codes have been proposed in recent years, there are still two key points that needed to be improved: 1) how to precisely preserve the similarity structure of the original data and 2) how to obtain the hash codes of the previously unseen data. In this paper, we propose our spline regression hashing method, in which both the local and global data similarity structures are exploited. To better capture the local manifold structure, we introduce splines developed in Sobolev space to find the local data mapping function. Furthermore, our framework simultaneously learns the hash codes of the training data and the hash function for the unseen data, which solves the out-of-sample problem. Extensive experiments conducted on real image datasets consisting of over one million images show that our proposed method outperforms the state-of-the-art techniques. Yang Liu 0098, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 4 |
| 2012 | Web and Personal Image Annotation by Mining Label Correlation With Relaxed Visual Graph EmbeddingabstractThe number of digital images rapidly increases, and it becomes an important challenge to organize these resources effectively. As a way to facilitate image categorization and retrieval, automatic image annotation has received much research attention. Considering that there are a great number of unlabeled images available, it is beneficial to develop an effective mechanism to leverage unlabeled images for large-scale image annotation. Meanwhile, a single image is usually associated with multiple labels, which are inherently correlated to each other. A straightforward method of image annotation is to decompose the problem into multiple independent single-label problems, but this ignores the underlying correlations among different labels. In this paper, we propose a new inductive algorithm for image annotation by integrating label correlation mining and visual similarity mining into a joint framework. We first construct a graph model according to image visual features. A multilabel classifier is then trained by simultaneously uncovering the shared structure common to different labels and the visual graph embedded label prediction matrix for image annotation. We show that the globally optimal solution of the proposed framework can be obtained by performing generalized eigen-decomposition. We apply the proposed framework to both web image annotation and personal album labeling using the NUS-WIDE, MSRA MM 2.0, and Kodak image data sets, and the AUC evaluation metric. Extensive experiments on large-scale image databases collected from the web and personal album show that the proposed algorithm is capable of utilizing both labeled and unlabeled data for image annotation and outperforms other algorithms. Yi Yang 0001, Fei Wu 0001, Feiping Nie 0001, Heng Tao Shen, Yueting Zhuang, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 5 |
| 2011 | Tag Clustering and Refinement on Semantic Unity GraphabstractRecently, there has been extensive research towards the user-provided tags on photo sharing websites which can greatly facilitate image retrieval and management. However, due to the arbitrariness of the tagging activities, these tags are often imprecise and incomplete. As a result, quite a few technologies has been proposed to improve the user experience on these photo sharing systems, including tag clustering and refinement, etc. In this work, we propose a novel framework to model the relationships among tags and images which can be applied to many tag based applications. Different from previous approaches which model images and tags as heterogeneous objects, images and their tags are uniformly viewed as compositions of Semantic Unities in our framework. Then Semantic Unity Graph (SUG) is introduced to represent the complex and high-order relationships among these Semantic Unities. Based on the representation of Semantic Unity Graph, the relevance of images and tags can be naturally measured in terms of the similarity of their Semantic Unities. Then Tag clustering and refinement can then be performed on SUG and the polysemy of images and tags is explicitly considered in this framework. The experiment results conducted on NUS-WIDE and MIR-Flickr datasets demonstrate the effectiveness and efficiency of the proposed approach. Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001, Yueting Zhuang |
ICDM | 5 |
| 2011 | Inverse-degree Sampling for Spectral ClusteringabstractAmong those classical clustering algorithms, spectral clustering performs much better than K-means in most cases. However, for the sake of cubic time complexity, spectral clustering is hardly used for clustering large-scale data sets. Therefore, sampling-based methods such as Nystrom method and Column sampling are respectively conducted as potential approaches to tackle this challenge. As we know, current sampling-based methods often utilize the uniform or other random sampling policies to select representative data and tend to disregard the data in small size clusters. This paper proposes an unbiased sampling framework, derives a new sampling method called inverse-degree sampling and then introduces an entropy criterion to prove it in theory simply. According to the selection of representative data by inverse-degree sampling in spectral clustering, the time complexity of spectral clustering becomes quadratic. Experiments on both toy data and real-world data demonstrate both the good sampling performance and the comparable clustering quality. Haidong Gao, Yueting Zhuang, Fei Wu 0001, Jian Shao 0001 |
ICIG | 2 |
| 2011 | Image annotation by composite kernel learning with group structureabstractWe can obtain more and more kinds of heterogeneous features (such as color, shape and texture) in images which can be extracted to describe various aspects of visual characteristics. Those high-dimensional heterogeneous visual features are intrinsically embedded in a non-linear space. In order to effectively utilize these heterogeneous features, this paper proposes an approach, called Composite Kernel Learning with Group Structure (CKLGS), to select groups of discriminative features for image annotation. For each image label, the CKLGS method embeds the nonlinear image data with discriminative features into different Reproducing Kernel Hilbert Spaces (RKHS), and then composes these kernels to select groups of discriminative features. Thus a classification model can be trained for image annotation. By the comparisons with other image annotation algorithms, experiments show that the proposed CKLGS for image annotation achieves a better performance. Fei Wu 0001, Yueting Zhuang, Jian Shao 0001 |
ACM Multimedia | 3 |
| 2011 | Hypergraph spectral hashing for similarity search of social imageabstractThe development of social media brings great challenges to image retrieval on both efficiency and accuracy. In addition to achieving fast similarity search over large scale data, it is very crucial to represent the complex and high-order relationships among the social contents to improve the semantic understanding of social images.In this paper, unified hypergraph is implemented to model the various relationships among images and other contexts in social media. Moreover, we extend traditional spectral hashing to hypergraph to accelerate similarity search of social images by mapping semantically related vertices into similar binary codes within a short Hamming distance. Furthermore, the proposed HSH approach is extended to out-of-sample data in a supervised manner. We evaluated our approach on the dataset crawled from Flickr and the experiment results indicate that our proposed HSH approach is both efficient and effective. Yueting Zhuang, Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001 |
ACM Multimedia | 1 |
| 2011 | Group sparse representation for image categorization and semantic video retrieval
Fei Wu 0001, Yueting Zhuang |
Sci. China Inf. Sci. | 3 |
| 2011 | Stable multi-label boosting for image annotation with structural feature selection
Yueting Zhuang, Yahong Han, Fei Wu 0001, JiaCheng Yang |
Sci. China Inf. Sci. | 1 |
| 2011 | Efficient shape matching for Chinese calligraphic character retrievalabstractAn efficient search method is desired for calligraphic characters due to the explosive growth of calligraphy works in digital libraries. However, traditional optical character recognition (OCR) and handwritten character recognition (HCR) technologies are not suitable for calligraphic character retrieval. In this paper, a novel shape descriptor called SC-HoG is proposed by integrating global and local features for more discriminability, where a gradient descent algorithm is used to learn the optimal combining parameter. Then two efficient methods, keypoint-based method and locality sensitive hashing (LSH) based method, are proposed to accelerate the retrieval by reducing the feature set and converting the feature set to a feature vector. Finally, a re-ranking method is described for practicability. The approach filters query-dissimilar characters using the LSH-based method to obtain candidates first, and then re-ranks the candidates using the keypoint- or sample-based method. Experimental results demonstrate that our approaches are effective and efficient for calligraphic character retrieval. Weiming Lu 0001, Jiangqin Wu, Baogang Wei, Yueting Zhuang |
J. Zhejiang Univ. Sci. C | 4 |
| 2011 | A hybrid brain-computer interface control strategy in a virtual environmentabstractThis paper presents a hybrid brain-computer interface (BCI) control strategy, the goal of which is to expand control functions of a conventional motor imagery or a P300 potential based BCI in a virtual environment. The hybrid control strategy utilizes P300 potential to control virtual devices and motor imagery related sensorimotor rhythms to navigate in the virtual world. The two electroencephalography (EEG) patterns serve as source signals for different control functions in their corresponding system states, and state switch is achieved in a sequential manner. In the current system, imagination of left/right hand movement was translated into turning left/right in the virtual apartment continuously, while P300 potentials were mapped to discrete virtual device control commands using a five-oddball paradigm. The combination of motor imagery and P300 patterns in one BCI system for virtual environment control was tested and the results were compared with those of a single motor imagery or P300-based BCI. Subjects obtained similar performances in the hybrid and single control tasks, which indicates the hybrid control strategy works well in the virtual environment. Jian-xun Luo, Yueting Zhuang, Xiaoxiang Zheng, Weidong Chen 0002 |
J. Zhejiang Univ. Sci. C | 7 |
| 2011 | Cartoon synthesis using constrained spreading activation network
Jun Yu 0002, Seah Hock Soon, Yueting Zhuang |
Multim. Tools Appl. | 3 |
| 2011 | Learning a 3D Human Pose Distance Metric from Geometric Pose DescriptorabstractEstimating 3D pose similarity is a fundamental problem on 3D motion data. Most previous work calculates L2-like distance of joint orientations or coordinates, which does not sufficiently reflect the pose similarity of human perception. In this paper, we present a new pose distance metric. First, we propose a new rich pose feature set called Geometric Pose Descriptor (GPD). GPD is more effective in encoding pose similarity by utilizing features on geometric relations among body parts, as well as temporal information such as velocities and accelerations. Based on GPD, we propose a semisupervised distance metric learning algorithm called Regularized Distance Metric Learning with Sparse Representation (RDSR), which integrates information from both unsupervised data relationship and labels. We apply the proposed pose distance metric to applications of motion transition decision and content-based pose retrieval. Quantitative evaluations demonstrate that our method achieves better results with only a small amount of human labels, showing that the proposed pose distance metric is a promising building block for various 3D-motion related applications. Cheng Chen 0023, Yueting Zhuang, Feiping Nie 0001, Yi Yang 0001, Fei Wu 0001, Jun Xiao 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Multi-Task Sparse Discriminant Analysis (MtSDA) with Overlapping CategoriesabstractMulti-task learning aims at combining information across tasks to boost prediction performance, especially when the number of training samples is small and the number of predictors is very large. In this paper, we first extend the Sparse Discriminate Analysis (SDA) of Clemmensen et al.. We call this Multi-task Sparse Discriminate Analysis (MtSDA). MtSDA formulates multi-label prediction as a quadratic optimization problem whereas SDA obtains single labels via a nearest class mean rule. Second, we propose a class of equicorrelation matrices to use in MtSDA which includes the identity matrix. MtSDA with both matrices are compared with singletask learning (SVM and LDA+SVM) and multi-task learning (HSML). The comparisons are made on real data sets in terms of AUC and F-measure. The data results show that MtSDA outperforms other methods substantially almost all the time and in some cases MtSDA with the equicorrelation matrix substantially outperforms MtSDA with identity matrix. Yahong Han, Fei Wu 0001, Jinzhu Jia, Yueting Zhuang, Bin Yu 0001 |
AAAI | 4 |
| 2010 | Local and Global Regressive Mapping for Manifold Learning with Out-of-Sample ExtrapolationabstractOver the past few years, a large family of manifold learning algorithms have been proposed, and applied to various applications. While designing new manifold learning algorithms has attracted much research attention, fewer research efforts have been focused on out-of-sample extrapolation of learned manifold. In this paper, we propose a novel algorithm of manifold learning. The proposed algorithm, namely Local and Global Regressive Mapping (LGRM), employs local regression models to grasp the manifold structure. We additionally impose a global regression term as regularization to learn a model for out-of-sample data extrapolation. Based on the algorithm, we propose a new manifold learning framework. Our framework can be applied to any manifold learning algorithms to simultaneously learn the low dimensional embedding of the training data and a model which provides explicit mapping of the out-of-sample data to the learned manifold. Experiments demonstrate that the proposed framework uncover the manifold structure precisely and can be freely applied to unseen data. Yi Yang 0001, Feiping Nie 0001, Shiming Xiang, Yueting Zhuang |
AAAI | 4 |
| 2010 | Sparse representation using nonnegative curds and wheyabstractIt has been of great interest to find sparse and/or nonnegative representations in computer vision literature. In this paper we propose a novel method to such a purpose and refer to it as nonnegative curds and whey (NNCW). The NNCW procedure consists of two stages. In the first stage we consider a set of sparse and nonnegative representations of a test image, each of which is a linear combination of the images within a certain class, by solving a set of regression-type nonnegative matrix factorization problems. In the second stage we incorporate these representations into a new sparse and nonnegative representation by using the group nonnegative garrote. This procedure is particularly appropriate for discriminant analysis owing to its supervised and nonnegativity nature in sparsity pursuing. Experiments on several benchmark face databases and Caltech 101 image dataset demonstrate the efficiency and effectiveness of our nonnegative curds and whey method. Fei Wu 0001, Yueting Zhuang, Shuicheng Yan |
CVPR | 4 |
| 2010 | Automatic annotation of geo-information in panoramic street view by image retrievalabstractPanoramic street view is now becoming a popular service in digital map due to its expedient of virtual walking through. Recently, some geo-tagged photos have been added into street view as additional illustration images for the same scenes. However, these images come directly from certain photo sharing websites where users manually tag the locations of their uploaded images; the quality of annotated location is very poor. In this paper, we propose a system to annotate the location tags of users' images in panoramic street view only with visual features. By searching the exemplar region which mostly matches the query image, the system can annotate the query image with geo-information provided by the matched exemplar image. In order to boost matching performance, various techniques such as the motion estimation, visual clustering and feature matching are implemented in our system. The experiments show satisfactory performance and promising results. Yueting Zhuang, Fei Wu 0001 |
ICIP | 2 |
| 2010 | Topic discovery of web video using star-structured K-partite graphabstractAs the explosive growth of web videos on video-shared sites like YouTube, the discovery of video topics has become a hot research area. In order to utilize all kinds of characteristics in web video such as visual features (SIFT, shape or color) and contextual cues (such as title or tags) effectively, this paper proposes an approach to represent the explicit and implicit correlations hidden in web videos by a star-structured K-partite graph model, and then a co-clustering process is conducted to discover video topics. The experimental results demonstrate the feasibility and effectiveness of the proposed approach. Jian Shao 0001, Wentao Yin, Yueting Zhuang |
ACM Multimedia | 4 |
| 2010 | Multi-label boosting for image annotation by structural grouping sparsityabstractWe can obtain high-dimensional heterogenous features from real-world images to describe their various aspects of visual characteristics, such as color, texture and shape etc.Different kinds of heterogenous features have different intrinsic discriminative power for image understanding. The selection of groups of discriminative features for certain semantics is hence crucial to make the image understanding more interpretable. This paper formulates the multi-label image annotation as a regression model with a regularized penalty. We call it Multi-label Boosting by the selection of heterogeneous features with structural Grouping Sparsity (MtBGS). MtBGS induces a (structural ) sparse selection model to identify subgroups of homogenous features for predicting a certain label. Moreover, the correlations among multiple tags are utilized in MtBGS to boost the performance of multi-label annotation. Extensive experiments on public image datasets show that the proposed approach has better multi-label image annotation performance and leads to a quite interpretable model for image understanding. Fei Wu 0001, Yahong Han, Qi Tian 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2010 | Heterogeneous feature selection by group lasso with logistic regressionabstractThe selection of groups of discriminative features is critical for image understanding since the irrelevant features could deteriorate the performance of image understanding. This paper formulates the selection of groups of discriminative features by the extension of group lasso with logistic regression for high-dimensional feature setting, we call it as the heterogeneous feature selection by Group Lasso with Logistic Regression (GLLR). GLLR encodes a sparse grouping prior to seek after a more interpretable model for feature selection and can identify most of discriminative groups of homogeneous features. The utilization of GLLR for image annotation shows the proposed GLLR achieves a better performance. Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 3 |
| 2010 | Overview of ACM international workshop on connected multimediaabstractFollowing the very first international workshop on connected multimedia held in Hangzhou, China, in October of 2009 jointly sponsored by US National Science Foundation and Zhejiang University of China, this is the very first ACM International Workshop on Connected Multimedia in conjunction with ACM International Conference on Multimedia held in Florence, Italy, in October of 2010. In this workshop overview, we first define what we mean by connected multimedia, and then briefly overview the program of this workshop. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2010 | Classification by semi-supervised discriminative regularization
Fei Wu 0001, Yi Yang 0001, Yueting Zhuang, Feiping Nie 0001 |
Neurocomputing | 4 |
| 2010 | Silhouette representation and matching for 3D pose discrimination - A comparative study
Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001 |
Image Vis. Comput. | 2 |
| 2010 | Multiple Hypergraph Clustering of Web Images by MiningWord2Image Correlations
Fei Wu 0001, Yahong Han, Yueting Zhuang |
J. Comput. Sci. Technol. | 3 |
| 2010 | CMSOF: a structured data organization framework for scanned Chinese medicine books in digital librariesabstractOrganizing unstructured information from books into a well-defined structure is a significant challenge in digital libraries. Most digital libraries can provide only search services at the granularity of books and few libraries allow books to be accessed at the granularity of chapters, as manually constructing directory information for books is time-consuming. Extracting structured data from scanned books thus remains an urgent and important work. In this paper, we propose a novel structured data organization framework called CMSOF to organize scanned data automatically, and apply it to a Chinese medicine digital library. In the framework, image blocks and text blocks on the scanned page of books are separated based on the gray histogram projection method or a hybrid method of region growth and the Ada-Boosting classifier at first, and then the text structure is obtained from text blocks by text size and font type recognition. Finally, image blocks and structured OCRed text are correlated at the semantic level. By integrating the structured data into a Chinese medicine information system (CMIS), we can organize the Chinese medicine books well and users can access the books with flexibility, which indicates that CMSOF is an efficient framework to organize books mixed with images and text. Baogang Wei, Weiming Lu 0001, Yueting Zhuang |
J. Zhejiang Univ. Sci. C | 5 |
| 2010 | Javelin: an access and manipulation interface for large displaysabstractWe describe a user interface and interaction technique, named ‘Javelin’, designed for large display environments. It provides quick access to random screen regions and manipulation methods for screen widgets which are difficult or impossible to reach. It consists of a dynamic global thumbnail, a touchpad widget that drives the screen cursor, and a teleport widget in which interactions are transferred to its target screen region. Javelin can be easily integrated into many programs to optimize their interaction performance in large screens. The experiment and user study show that Javelin can extend user access field and enhance widget manipulation in large displays. Zhenkun Zhou, Jiangqin Wu, Yin Zhang 0006, Da-wei Xie, Yueting Zhuang |
J. Zhejiang Univ. Sci. C | 5 |
| 2010 | A group of novel approaches and a toolkit for motion capture data reusing
Jun Xiao 0001, Yueting Zhuang, Fei Wu 0001, Tongqiang Guo, Zhang Liang |
Multim. Tools Appl. | 2 |
| 2010 | Cross-media retrieval using query dependent search methods
Yi Yang 0001, Fei Wu 0001, Dong Xu 0001, Yueting Zhuang, Liang-Tien Chia |
Pattern Recognit. | 4 |
| 2010 | Multi-Label Transfer Learning With Sparse RepresentationabstractDue to the visually polysemous barrier, videos and images may be annotated by multiple tags. Discovering the correlations among different tags can significantly help predicting precise labels for videos and images. Many of recent studies toward multi-label learning construct a linear subspace embedding with encoded multi-label information, such that data points sharing many common labels tend to be close to each other in the embedded subspace. Motivated by the advances of compressive sensing research, a sparse representation that selects a compact subset to describe the input data can be more discriminative. In this paper, we propose a sparse multi-label learning method to circumvent the visually polysemous barrier of multiple tags. Our approach learns a multi-label encoded sparse linear embedding space from a related dataset, and maps the target data into the learned new representation space to achieve better annotation performance. Instead of using l1-norm penalty (lasso) to induce sparse representation, we propose to formulate the multi-label learning as a penalized least squares optimization problem with elastic-net penalty. By casting the video concept detection and image annotation tasks into a sparse multi-label transfer learning framework in this paper, ridge regression, lasso, elastic net, and the multi-label extended sparse discriminant analysis methods are, respectively, well explored and compared. Yahong Han, Fei Wu 0001, Yueting Zhuang, Xiaofei He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Recognizing Cartoon Image Gestures for Retrieval and Interactive Cartoon Clip SynthesisabstractIn this paper, we propose a new method to recognize gestures of cartoon images with two practical applications, i.e., content-based cartoon image retrieval and interactive cartoon clip synthesis. Upon analyzing the unique properties of four types of features including global color histogram, local color histogram (LCH), edge feature (EF), and motion direction feature (MDF), we propose to employ different features for different purposes and in various phases. We use EF to define a graph and then refine its local structure by LCH. Based on this graph, we adopt a transductive learning algorithm to construct local patches for each cartoon image. A spectral method is then proposed to optimize the local structure of each patch and then align these patches globally. MDF is fused with EF and LCH and a cartoon gesture space is constructed for cartoon image gesture recognition. We apply the proposed method to content-based cartoon image retrieval and interactive cartoon clip synthesis. The experiments demonstrate the effectiveness of our method. Yi Yang 0001, Yueting Zhuang, Dacheng Tao, Dong Xu 0001, Jun Yu 0002, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Image Clustering Using Local Discriminant Models and Global IntegrationabstractIn this paper, we propose a new image clustering algorithm, referred to as clustering using local discriminant models and global integration (LDMGI). To deal with the data points sampled from a nonlinear manifold, for each data point, we construct a local clique comprising this data point and its neighboring data points. Inspired by the Fisher criterion, we use a local discriminant model for each local clique to evaluate the clustering performance of samples within the local clique. To obtain the clustering result, we further propose a unified objective function to globally integrate the local models of all the local cliques. With the unified objective function, spectral relaxation and spectral rotation are used to obtain the binary cluster indicator matrix for all the samples. We show that LDMGI shares a similar objective function with the spectral clustering (SC) algorithms, e.g., normalized cut (NCut). In contrast to NCut in which the Laplacian matrix is directly calculated based upon a Gaussian function, a new Laplacian matrix is learnt in LDMGI by exploiting both manifold structure and local discriminant information. We also prove that K-means and discriminative K-means (DisKmeans) are both special cases of LDMGI. Extensive experiments on several benchmark image datasets demonstrate the effectiveness of LDMGI. We observe in the experiments that LDMGI is more robust to algorithmic parameter, when compared with NCut. Thus, LDMGI is more appealing for the real image clustering applications in which the ground truth is generally not available for tuning algorithmic parameters. Yi Yang 0001, Dong Xu 0001, Feiping Nie 0001, Shuicheng Yan, Yueting Zhuang |
IEEE Trans. Image Process. | 5 |
| 2009 | Web image interpretation: semi-supervised mining annotated wordsabstractAn image is worth of thousand words. Automatic Web image annotation is a practical and effective way for both Web image retrieval and image understanding. However, current annotation techniques are very difficult to get natural language interpretation for images such as ldquopandas eat bamboordquo. In this paper, we proposed an approach to interpret image semantics through semi-supervised mining annotated words. The idea in this approach mainly consists of three parts: at first, the visibility of annotated words of target image is calculated by semi-supervised learning approach from the landmark words in WordNet; then the annotated words are used as queries to retrieve matched Web pages; at last, the meaningful sentences in the matched Web pages are ranked as the interpretation of target image by semi-supervised learning approach. Experiments conducted on real-world Web images demonstrate the effectiveness of the proposed approach. Fei Wu 0001, Dingyin Xia, Yueting Zhuang, Hanwang Zhang |
ICME | 3 |
| 2009 | Face Inpainting by Feature GuidanceabstractFace image partially occluded or damaged can be repaired automatically. We proposed a new inpainting algorithm, based on patch guidance deduced from an existing face database, to recover the damaged portions. This newly proposed concept of guided inpainting method produces seamless faces which are hardly seen drawbacks. Examples of our results can be retrieved from http://member.mine.tku.edu.tw/www/ISCAS09/. Nick C. Tang, Yueting Zhuang, Yushun Wang, Timothy K. Shih, Joseph C. Tsai |
ISCAS | 2 |
| 2009 | Ranking with local regression and global alignment for cross media retrievalabstractRich multimedia content including images, audio and text are frequently used to describe the same semantics in E-Learning and Ebusiness web pages, instructive slides, multimedia cyclopedias, and so on. In this paper, we present a framework for cross-media retrieval, where the query example and the retrieved result(s) can be of different media types. We first construct Multimedia Correlation Space (MMCS) by exploring the semantic correlation of different multimedia modalities, during which multimedia content and co-occurrence information is utilized. We propose a novel ranking algorithm, namely ranking with Local Regression and Global Alignment (LRGA), which learns a robust Laplacian matrix for data ranking. In LRGA, for each data point, a local linear regression model is used to predict the ranking values of its neighboring points. We propose a unified objective function to globally align the local models from all the data points so that an optimal ranking value can be assigned to each data point. LRGA is insensitive to parameters, making it particularly suitable for data ranking. A relevance feedback algorithm is proposed to improve the retrieval performance. Comprehensive experiments have demonstrated the effectiveness of our methods. Yi Yang 0001, Dong Xu 0001, Feiping Nie 0001, Jiebo Luo 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2009 | Retrieval based interactive cartoon synthesis via unsupervised bi-distance metric learningabstractCartoons play important roles in many areas, but it requires a lot of labor to produce new cartoon clips. In this paper, we propose a gesture recognition method for cartoon character images with two applications, namely content-based cartoon image retrieval and cartoon clip synthesis. We first define Edge Features (EF) and Motion Direction Features (MDF) for cartoon character images. The features are classified into two different groups, namely intra-features and inter-features. An Unsupervised Bi-Distance Metric Learning (UBDML) algorithm is proposed to recognize the gestures of cartoon character images. Different from the previous research efforts on distance metric learning, UBDML learns the optimal distance metric from the heterogeneous distance metrics derived from intra-features and inter-features. Content-based cartoon character image retrieval and cartoon clip synthesis can be carried out based on the distance metric learned by UBDML. Experiments show that the cartoon character image retrieval has a high precision and that the cartoon clip synthesis can be carried out efficiently. Yi Yang 0001, Yueting Zhuang, Dong Xu 0001, Yunhe Pan, Dacheng Tao, Stephen J. Maybank |
ACM Multimedia | 2 |
| 2009 | Perceptual 3D pose distance estimation by boosting relational geometric featuresabstractAbstract Traditional pose similarity functions based on joint coordinates or rotations often do not conform to human perception. We propose a new perceptual pose distance:Relational Geometric Distancethat accumulates the differences over a set of features that reflects the geometric relations between different body parts. An extensive relational geometric feature pool that contains a large number of potential features is defined, and the features effective for pose similarity estimation are selected using a set of labeled data by Adaboost. The extensive feature pool guarantees that a wide diversity of features is considered, and the boosting ensures that the selected features are optimized when used jointly. Finally, the selected features form a pose distance function that can be used for novel poses. Experiments show that our method outperforms others in emulating human perception in pose similarity. Our method can also adapt to specific motion types and capture the features that are important for pose similarity of a certain motion type. Copyright © 2009 John Wiley & Sons, Ltd. Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001, Zhang Liang |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Competitive motion synthesis based on hybrid controlabstractAbstract We propose a simple and effective framework to deal with the problem of synthesizing interactive and competitive motions while reflecting the interactions. Two nontrivial issues are addressed in synthesizing two‐character motions in competitive environment: how to reveal the embedded routines while keeping visual reality and how to build interactive models based on singly captured motions. To solve these issues, we employ a hierarchical framework: the finite state machine (FSM) controls the state transition in the higher layer, and the hybrid approach controls the action selection in the lower layer. The proposed approach contains two folds: first, a rule‐based control scheme is proposed to simulate routine steps based on statistical analysis. Second, the interactive models are designed for simulating dense interactions between two players. The Relevance Vector Machine (RVM) algorithm is adopted to select attack styles and coupled with motion transition graph to determine combination blows. Here we apply the proposed framework of hybrid paradigm to boxing sport as an example. Copyright © 2009 John Wiley & Sons, Ltd. Liang Zhang 0045, Jun Xiao 0001, Yueting Zhuang, Cheng Chen 0023 |
Comput. Animat. Virtual Worlds | 3 |
| 2009 | Latent Style Model: Discovering writing styles for calligraphy works
Yueting Zhuang, Weiming Lu 0001, Jiangqin Wu |
J. Vis. Commun. Image Represent. | 1 |
| 2009 | Discovering calligraphy style relationships by Supervised Learning Weighted Random Walk Model
Weiming Lu 0001, Yueting Zhuang, Jiangqin Wu |
Multim. Syst. | 2 |
| 2009 | Tensor-Based Transductive Learning for Multimodality Video Semantic Concept DetectionabstractInteraction and integration of multimodality media types such as visual, audio, and textual data in video are the essence of video semantic analysis. Contextual information propagation is useful for both intra- and inter-shot correlations. However, the traditional concatenated vector representation of videos weakens the power of the propagation and compensation among the multiple modalities. In this paper, we introduce a higher-order tensor framework for video analysis. We represent image frame, audio, and text in video shots as data points by the 3rd-order tensor. Then we propose a novel dimension reduction algorithm which explicitly considers the manifold structure of the tensor space from contextual temporal associated cooccurring multimodal media data. Our algorithm inherently preserves the intrinsic structure of the submanifold where tensorshots are sampled and is also able to map out-of-sample data points directly. We propose a new transductive support tensor machines algorithm to train effective classifier using large amount of unlabeled data together with the labeled data. Experiment results on TREVID 2005 data set show that our method improves the performance of video semantic concept detection. Fei Wu 0001, Yueting Zhuang |
IEEE Trans. Multim. | 3 |
| 2008 | Adaptive and compact shape descriptor by progressive feature combination and selection with boostingabstractMany types of shape descriptors have been proposed for 2D shape analysis, but most of them consist of component features that are not adapted to specific problems. This has two drawbacks. First, computation is wasted on the irrelevant components; second, the accuracy is impaired. This paper proposes an effective method that generates compact descriptors adapted to specific problems in hand, where each component of the new descriptor is a linear combination of the components in some classic descriptors. A progressive strategy is used to construct and select the most suitable linear combinations in successive rounds, where a variant of Adaboost is employed to ensure the optimum of the selected combinations in each round. Experiments show that our method effectively generates adaptive and compact descriptors for typical applications such as shape classification and retrieval. Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001, Fei Wu 0001 |
CVPR | 2 |
| 2008 | Indexing high-dimensional data in dual distance spaces: a symmetrical encoding approachabstractDue to the well-known dimensionality curse problem, search in a high-dimensional space is considered as a "hard" problem. In this paper, a novel symmetrical encoding-based index structure, which is called EHD-Tree (for symmetrical Encoding-based Hybrid Distance Tree), is proposed to support fast k-Nearest-Neighbor (k-NN) search in high-dimensional spaces. In an EHD-Tree, all data points are first grouped into clusters by a k-Means clustering algorithm. Then the uniform ID number of each data point is obtained by a dual-distance-driven encoding scheme in which each cluster sphere is partitioned twice according to the dual distances of start- and centroid-distance. Finally, the uniform ID number and the centroid-distance of each data point are combined to get a uniform index key, the latter is then indexed through a partition-based B+-tree. Thus, given a query point, its k-NN search in high-dimensional spaces can be transformed into search in a single dimensional space with the aid of the EHD-Tree index. Extensive performance studies are conducted to evaluate the effectiveness and efficiency of our proposed scheme, and the results demonstrate that this method outperforms the state-of-the-art high dimensional search techniques such as the X-Tree, VA-file, iDistance and NB-Tree, especially when the query radius is not very large. Yi Zhuang 0001, Yueting Zhuang, Qing Li 0001, Lei Chen 0002, Yi Yu 0001 |
EDBT | 2 |
| 2008 | Clustering by evidence accumulation on affinity propagationabstractAffinity propagation (AP) is a clustering algorithm which has much better performance than traditional clustering approach such as k-means algorithm. In this paper, we present an algorithm called voting partition affinity propagation (voting-PAP) which is a method for clustering using evidence accumulation based on AP. Resulting clusters by voting-PAP are not constrained to be hyper-spherically shaped. Voting-PAP consists of three parts: Partition Affinity propagation (PAP), relaxed multi-root minimum spanning tree (MST) and majority voting. PAP is a method which can produce different exemplar set based on AP. Relaxed multi-root MST is a data point assign algorithm which has better performance than nearest assign rule. Majority voting is a scheme used to find a consistent clustering result of different partitions based on the idea of evidence accumulation. We also discuss how to find an appropriate threshold corresponding to an approximate ideal consistent partition in this paper. Xuqing Zhang, Fei Wu 0001, Yueting Zhuang |
ICPR | 3 |
| 2008 | Active post-refined multimodality video semantic concept detection with tensor representationabstractIn this paper, we resolve the problem of multi-modality video representation and semantic concept detection. Interaction and integration of multi-modality media types such as visual, audio and textual data in video are essential to video semantic analysis. Traditionally, videos are represented as vectors in the Euclidean space. Many learning algorithms are then taken to these vectors in a high dimensional space for dimension reduction, classification, clustering and so on. However, the multiple modalities in video not only have their own properties, but also have correlations among them; whereas the simple vector representation weakens the power of these relatively independent modalities and even ignores their relations to some extent. In this paper, we introduce a higher-order tensor framework for video analysis, in which we represent image, video and text three modalities in video shots as data points by the 3rd-order tensor called tensorshots. We propose a novel dimension reduction method that explicitly considers the manifold structure of the tensor space from multimodal media data which is temporal associated co-occurrence and then detect video semantic concepts through powerful classifiers which take tensor as input. Our algorithm preserves the intrinsic structure of the submanifold where tensorshots are sampled, and is also able to map out-of-sample data points directly. Moreover we apply an active learning based contextual and temporal post-refining strategy to enhance detection accuracy. Experiment results show that our method improves the performance of video semantic concept detection. Fei Wu 0001, Yueting Zhuang, Jun Xiao 0001 |
ACM Multimedia | 3 |
| 2008 | Heterogeneous multimedia data semantics mining using content and location contextabstractBecause it is very common that the heterogeneous multimedia data of the same semantics always exist jointly in many domain and application specific databases, it is very helpful to consider the location information when analyzing multimedia data. In this paper we propose a method of integrating the content and location context for multimedia data mining to enable the cross-media retrieval, by which the query examples and the returned results can be of different modalities, e.g. to query audios by an example of image. We construct a graph model by combing the multimedia content and location information. The graph model is then refined according to different strategies. The semantic correlations among multimedia data are calculated by learning the high-order neighborhood structure of the graph and the Multimedia Correlation Space is constructed in which the cross-media retrieval can be performed. We also propose different methods of Relevance Feedback to improve the search results. Experiments demonstrate the promise of the proposed method. Yi Yang 0001, Yueting Zhuang |
ACM Multimedia | 2 |
| 2008 | An encoding-based dual distance tree high-dimensional index
Yi Zhuang 0001, Yueting Zhuang, Fei Wu 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2008 | Perspective-aware cartoon clips synthesisabstractAbstract In this paper we propose an approach, which allows the users to synthesize cartoon clips according to the perspective of the background image. In order to construct the cartoons smoothly, the character's edge distance and motion direction distance are demonstrated to be the factors affecting the human perception in similarity evaluation, and utilized in cartoon clips synthesis. When applying the generated cartoons to the background image, in which the perspective exists, the size of the character is coordinated according to the scaling factor calculated from the vanishing line. The experiment results demonstrate that our approach can synthesize the cartoon clips more smoothly compared with other single frame reusing strategies. The generated cartoons, which are applied to the background image, can be accepted by the human perception well. Copyright © 2008 John Wiley & Sons, Ltd. Yueting Zhuang, Jun Yu 0002, Jun Xiao 0001, Cheng Chen 0023 |
Comput. Animat. Virtual Worlds | 1 |
| 2008 | Harmonizing Hierarchical Manifolds for Multimedia Document Semantics Understanding and Cross-Media RetrievalabstractIn this paper, we consider the problem of multimedia document (MMD) semantics understanding and content-based cross-media retrieval. An MMD is a set of media objects of different modalities but carrying the same semantics and the content-based cross-media retrieval is a new kind of retrieval method by which the query examples and search results can be of different modalities. Two levels of manifolds are learned to explore the relationships among all the data in the level of MMD and in the level of media object respectively. We first construct a Laplacian media object space for media object representation of each modality and an MMD semantic graph to learn the MMD semantic correlations. The characteristics of media objects propagate along the MMD semantic graph and an MMD semantic space is constructed to perform cross-media retrieval. Different methods are proposed to utilize relevance feedback and experiment shows that the proposed approaches are effective. Yi Yang 0001, Yueting Zhuang, Fei Wu 0001, Yunhe Pan |
IEEE Trans. Multim. | 2 |
| 2008 | Mining Semantic Correlation of Heterogeneous Multimedia Data for Cross-Media RetrievalabstractAlthough multimedia objects such as images, audios and texts are of different modalities, there are a great amount of semantic correlations among them. In this paper, we propose a method of transductive learning to mine the semantic correlations among media objects of different modalities so that to achieve the cross-media retrieval. Cross-media retrieval is a new kind of searching technology by which the query examples and the returned results can be of different modalities, e.g., to query images by an example of audio. First, according to the media objects features and their co-existence information, we construct a uniform cross-media correlation graph, in which media objects of different modalities are represented uniformly. To perform the cross-media retrieval, a positive score is assigned to the query example; the score spreads along the graph and media objects of target modality or MMDs with the highest scores are returned. To boost the retrieval performance, we also propose different approaches of long-term and short-term relevance feedback to mine the information contained in the positive and negative examples. Yueting Zhuang, Yi Yang 0001, Fei Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2007 | A Novel Scalable Texture Video Coding Scheme with GPCAabstractThis paper proposes a novel SNR scalable coding method with the support of generalized principle component analysis (GPCA). This method encodes the low-pass and high-pass pictures generated by the MCTF decomposition with a hybrid linear model instead of traditional block-based DCT transform. GPCA is a powerful tool to identify the hybrid linear model in the textures, which segment the texture into heterogeneous regions, and then encode each region with PCA method. By keeping various proportions of PCA coefficients, and altering the quantization step sizes for different layers, a better scalable coding result can be achieved. Yueting Zhuang, Fei Wu 0001 |
ICASSP (1) | 2 |
| 2007 | Efficient Silhouette Extraction with Dynamic ViewpointabstractA novel approach is proposed that extends the classical background subtraction method to extract silhouettes from videos in real time with dynamic viewpoint variation caused by camera movement. First, manifold learning is used to model the background under viewpoint variations. Then, for each new frame, the background image corresponding to the same viewpoint is synthesized on the fly by examining the local neighborhood on the manifold, and the silhouette is extracted via background subtraction. An extension is also presented to generate stabilized silhouettes at any fixed viewpoint within the training range. Experiments show that our approach can efficiently extract accurate silhouettes in complex situations while maintaining a low noise level. Yueting Zhuang, Cheng Chen 0023 |
ICCV | 1 |
| 2007 | Video Motion Capture by Silhouette Analysis and Pose OptimizationabstractVideo based 3D reconstruction of human motion plays an important role in many applications. We implement a system that robustly reconstructs 3D human motion from markerless videos taken by a single camera. Our system only requires a desktop PC and a mainstream camera, and doesn't involve complex camera calibration, making it easy to implement and widely accessible in daily uses such as human computer interaction or entertainment. Cheng Chen 0023, Yueting Zhuang, Shicong Zhao, Yin Cheng |
ICME | 2 |
| 2007 | Adaptive Weight Selection for Incremental Eigen-Background ModelingabstractBackground modeling is an important approach for motion detection. The background model should adapt to dynamic change of the environment in time and generate background image with no moving foreground. Accordingly, we propose to incorporate the adaptive weight selection mechanism for roughly detected motion regions into the incremental eigen-background method. Comparing with existing works, we originally provide a way to reasonably design and adaptively compute the weight for each frame. Experiments show that the proposed adaptive incremental eigen-background method not only models the dynamic background scene well but also generates better background image with no ghost effect when salient motion occurs. Yueting Zhuang |
ICME | 2 |
| 2007 | Cross-modal correlation learning for clustering on image-audio datasetabstractIt is interesting and challenging to explore correlations between different datasets and utilize such correlations for the clustering on these datasets. Cross-modal correlation between images and audios can help identify images (or audios) of certain semantics. However, the heterogeneous problem makes it difficult to learn cross-modal correlation between visual and auditory features. In this paper, we analyze canonical correlation between feature matrices of images and audios during subspace mapping; then we design correlation-based similarity reinforcement for images and audios; thirdly we implement image clustering and audio clustering with affinity propagation. Experiment results on image-audio dataset are encouraging and show that the performance of our approach is effective. We give an interesting application of querying images by audio examples. Yueting Zhuang, Fei Wu 0001 |
ACM Multimedia | 2 |
| 2007 | 3D Facial Modeling for Animation: A Nonlinear Approach
Yushun Wang, Yueting Zhuang |
MMM (1) | 2 |
| 2007 | Visual Verification of Historical Chinese Calligraphy Works
Xiafen Zhang, Yueting Zhuang |
MMM (1) | 2 |
| 2007 | Boosting Cross-Media Retrieval by Learning with Positive and Negative Examples
Yueting Zhuang, Yi Yang 0001 |
MMM (2) | 1 |
| 2007 | Hierarchical Approximate Matching for Retrieval of Chinese Historical Calligraphy Character
Xiafen Zhang, Yueting Zhuang, Jiangqin Wu, Fei Wu 0001 |
J. Comput. Sci. Technol. | 2 |
| 2007 | Composite Distance Transformation for Indexing and k -Nearest-Neighbor Searching in High-Dimensional Spaces
Yi Zhuang 0001, Yueting Zhuang, Fei Wu 0001 |
J. Comput. Sci. Technol. | 2 |
| 2007 | Adaptive control in cartoon data reusingabstractAbstract In this paper, we propose a novel approach, which reuses Traditional Chinese Cartoon to create new animations. In order to extract the cartoon character precisely, a segmentation method based on edge detection is implemented. Before reusing the data, a lower‐ dimensional space of the cartoon data is constructed by ISOmap. The character's gesture difference calculated by optical flow is combined with character's edge difference through a novel distance function, which is controlled by a weight parameter. The animation is created by reordering the existing data into a sequence, which is the shortest path between two designated data in the space. Our approach utilizes image processing, computer vision, and machine learning in cartoon creation and the experiment results demonstrate that the animation's quality can be effectively improved by the fusion of these techniques. Copyright © 2007 John Wiley & Sons, Ltd. Jun Yu 0002, Yueting Zhuang, Jun Xiao 0001, Cheng Chen 0023 |
Comput. Animat. Virtual Worlds | 2 |
| 2007 | Content-based retrieval of FlashTM movies: research issues, generic framework, and future directions
Jun Yang 0003, Qing Li 0001, Wenyin Liu, Yueting Zhuang |
Multim. Tools Appl. | 4 |
| 2007 | Hallucinating faces: LPH super-resolution and neighbor reconstruction for residue compensation
Yueting Zhuang, Fei Wu 0001 |
Pattern Recognit. | 1 |
| 2007 | Interactive high-dimensional index for large Chinese calligraphic character databasesabstractThe large numbers of Chinese calligraphic scripts in existence are valuable part of the Chinese cultural heritage. However, due to the shape complexity of these characters, it is hard to employ existing techniques to effectively retrieve and efficiently index them. In this article, using a novel shape-similarity- based retrieval method in which shapes of calligraphic characters are represented by their contour points extracted from the character images, we propose an interactive partial-distance-map (PDM)- based high-dimensional indexing scheme which is designed specifically to speed up the retrieval performance of the large Chinese calligraphic character databases effectively. Specifically, we use the approximate minimal bounding sphere of a query character and utilize users' relevance feedback to refine the query gradually. Comprehensive experiments are conducted to testify the efficiency and effectiveness of this method. In addition, a new k -NN search called Pseudo k -NN (P k -NN) search is presented to better facilitate the PDM-based character retrieval. Yi Zhuang 0001, Yueting Zhuang, Qing Li 0001, Lei Chen 0002 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2006 | Video-Based Facial Expression Hallucination: A Two- Level Hierarchical Fusion Approach
Yueting Zhuang, Fei Wu 0001 |
ACIVS | 2 |
| 2006 | An Efficient Keyframe Extraction from Motion Capture Data
Jun Xiao 0001, Yueting Zhuang, Fei Wu 0001 |
Computer Graphics International | 2 |
| 2006 | Towards interactive indexing for large Chinese calligraphic character databasesabstractIn this paper, based on a novel shape-similarity-based retrieval method, we propose an interactive partial-distance-map (PDM)- based high-dimensional indexing scheme to speed up the retrieval performance of the large Chinese calligraphic character databases. Specifically, we use the approximate minimal bounding hyper- sphere of query character to search the PDM and utilize the users' relevance feedback to refine the search process. We conduct comprehensive experiments to testify the efficiency and effectiveness of the proposed method. Yi Zhuang 0001, Yueting Zhuang, Qing Li 0001, Lei Chen 0002 |
CIKM | 2 |
| 2006 | A Web-Based Examination and Evaluation System for Computer EducationabstractA Web-based operational skills examination and evaluation system is designed and implemented for computer courses. It consists of four systems, including preparation, examination, monitor and auto-grading subsystem. Various techniques involving DCOM, mark-method and fuzzy-match are adopted in this system, and a universal approach is generalized to enable auto-grading system suitable for different operated results. This system has been successfully applied in operational skills' evaluation and training, such as programming, editing documents, using Microsoft Windows Liang Zhang 0045, Yueting Zhuang, Zhenming Yuan, Guo-hua Zhan |
ICALT | 2 |
| 2006 | Learning Semantic Correlations for Cross-Media RetrievalabstractThis paper proposes a novel cross-media retrieval approach. First, an isomorphic subspace is constructed based on canonical correlation analysis (CCA) to learn multi-modal correlations of media objects; second, polar coordinates are used to judge the general distance of media objects with different modalities in the subspace. Since the integrity of semantic correlations is not likely learned from limited training samples, users' relevance feedback is used to accurately refine cross-media similarities. We also propose methods to map new media objects into the learned subspace, and any new media object would be taken as query example. Experiment results show that our approaches are effective for cross-media retrieval, and meanwhile achieve a significant improvement over content-based image retrieval and content-based audio retrieval. Fei Wu 0001, Yueting Zhuang |
ICIP | 3 |
| 2006 | Web based Chinese Calligraphy Learning with 3-D Visualization MethodabstractChinese calligraphy is pictographic and each calligraphist has his own writing style. People often feel difficult in writing a demanded beautiful calligraphy style. In order to help people enjoy the art of calligraphy and learn how it is written step-by-step we present a new approach to animate its writing process by 3-D visualization method. In this paper some novel algorithms used in the approach are presented to solve the following problems: 1) estimate varied stroke's thickness 2) extract strokes order from an offline Chinese calligraphic writing. Through this approach we implement a system. Experimental result is given to demonstrate the application finally. Yingfei Wu, Yueting Zhuang, Yunhe Pan, Jiangqin Wu |
ICME | 2 |
| 2006 | An approach for cross-media retrieval with cross-reference graph and PageRankabstractIn this paper, we propose a novel cross-media retrieval method. The most important feature of it is to integrate the multi-modal data seamlessly via a cross-reference graph, and then based on the graph, it is able to use improved personalized PageRank to calculate how close the media object associates with the query on semantic and content level. It is also able to adjust the cross-reference graph according to user's relevance feedback, which refines the semantic relationship between the media objects, so as to improve the retrieval accuracy progressively. As demonstrated by the experiments, our method achieves satisfactory retrieval efficiency on multi-modal datasets. Yueting Zhuang, Hanhuai Shan, Fei Wu 0001 |
MMM | 1 |
| 2006 | Filling Holes in Meshes and Recovering Sharp EdgesabstractAn efficient approach to recover implicit sharp edges with the help of filling holes in meshes is presented in this paper. First, the edges near the hole are classified as smooth and sharp edges by the dihedral angel between their adjacent triangles. The hole-nodes are classified as smooth and sharp nodes according to the types of their adjacent edges. Second, different methods are adopted to shrinkage the holes for different types of hole-nodes. A triangulation method is used for smooth hole-nodes, while an extension method is used for sharp hole-nodes. Third, the filled patch is refined for a smooth and uniform mesh preserving sharp edges logically. Tongqiang Guo, Jijun Li, Jianguang Weng, Yueting Zhuang |
SMC | 4 |
| 2006 | Data-driven Generation of Decision Tree based on Ensemble Multiple-instance Learning for Motion RetrievalabstractIn this paper, a motion retrieval system is investigated from a multiple-instance learning view. In order to retrieve similar motion data, each human joint's motion clip is regarded as a bag, while each of its segments is regarded as an instance. First 3D temporal-spatial features and their keyspaces of each human joint are extracted. Then data driven decision trees based on ensemble multiple-instance are automatically constructed to reflect the influence of each point during the comparison of motion similarity. At last the method of multiple-instance retrieval is used to complete motion retrieval. Experimental results show that our approaches are effective for motion data retrieval. Yueting Zhuang, Fei Wu 0001 |
SMC | 2 |
| 2006 | A hierarchical clustering algorithm based on fuzzy graph connectedness
Yi-hong Dong, Yueting Zhuang, Xiaoying Tai |
Fuzzy Sets Syst. | 2 |
| 2005 | Multi-Modal Information Retrieval with a Semantic View MechanismabstractThe explosive growth of multimedia information on the Web in recent years calls for an elegant means to model and manage multimedia content to facilitate semantic-level access and sharing across diversified applications. From the perspective of retrieval, the semantics of multimedia data features context-dependency and media-independency; both are inadequately supported by the state-of-the-art data modeling technology. In this paper, we address this problem by advocating MediaView as an extended object-oriented view mechanism to bridge the "semantic gap" between conventional databases and semantics-intensive multimedia applications. This mechanism captures the dynamic semantics of multimedia using a modeling construct named media view (MV), which formulates a customized context where heterogeneous media objects with similar/related semantics are characterized by additional properties and user-defined semantic relationships. View operators are proposed for the manipulation and derivation of individual MVs, which can be fit into the desired real-life scenarios automatically. The usefulness and elegancy of MediaView are demonstrated by its applications in various (subjective) activities supporting multi-modal retrieval. Qing Li 0001, Jun Yang 0003, Yueting Zhuang |
AINA | 3 |
| 2005 | Segmenting Layers in Automated Visual SurveillanceabstractDetecting objects of interest from a video sequence is a fundamental and critical task in automated visual surveillance. Those objects can either be moving or stationary. However, most of current approaches only focus on discriminating moving objects by background subtraction. In this work, we propose layers segmentation to detect both of moving and stationary target objects from surveillance video. We first construct a codebook with set of codewords for each pixel and then extend the Matrix Entropy statistical model to segment layers with codewords features. Our experimental results are presented in terms of success layer segmentation rate. Lijuan Qin, Yueting Zhuang, Yunhe Pan, Fei Wu 0001 |
ICME | 2 |
| 2005 | Sketch-based retrieval on Flash movies via primary sceneabstractAs a multimedia format, Flash is becoming more and more popular over the Web. The typical structure of Flash can benefit from both image retrieval and video retrieval methods. In this paper, we present an approach of sketch-based retrieval on flash movies with analysis on directional and motional relations. Via the selection of primary scenes, query result can be displayed to users in an ideal way. Experiment of the proposed approach is evaluated on a test set with different genres of flash movies, and it shows the usefulness of the approach. Qing Li 0001, Minhao Yu, Yueting Zhuang |
ISM | 4 |
| 2005 | Automatic generation of human animation based on motion programmingabstractAbstract In motion simulations, video games and animation films, lots of interactions between characters and virtual environments are needed. Even though realistic motion data can be derived from MoCap system, motion editing and synthesis, animators must adapt these motion data to specific virtual environment manually, which is a boring and time‐consuming job. Here we propose a framework to program the movements of characters and generate navigation animations in virtual environment. Given a virtual environment, a visual user interface is provided for animators to interactively generate motion scripts, describing the characters' movements in this scene and finally used to retrieve motion clips from MoCap database and generate navigation animations automatically. This framework also provides flexible mechanism for animators to get varied resulting animations by configurable table of motion bias coefficients and interactive visual user interface. Copyright © 2005 John Wiley & Sons, Ltd. Yueting Zhuang, Jun Xiao 0001, Yizi Wu, Fei Wu 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2005 | Steerable pyramid-based face hallucination
Congyong Su, Yueting Zhuang, Fei Wu 0001 |
Pattern Recognit. | 2 |
| 2005 | Searching for Flash Movies on the Web: A Content and Context Based Framework
Jun Yang 0003, Qing Li 0001, Wenyin Liu, Yueting Zhuang |
World Wide Web | 4 |
| 2004 | Towards Data-Adaptive and User-Adaptive Image Retrieval by Peer Indexing
Jun Yang 0003, Qing Li 0001, Yueting Zhuang |
Int. J. Comput. Vis. | 3 |
| 2003 | Modeling Data and User Characteristics by Peer Indexing in Content-based Image Retrieval
Jun Yang 0003, Qing Li 0001, Yueting Zhuang |
MMM | 3 |
| 2003 | Popular music retrieval by detecting moodabstractNo abstract available. Yazhong Feng, Yueting Zhuang, Yunhe Pan |
SIGIR | 2 |
| 2003 | Music Information Retrieval by Detecting Mood via Computational Media AestheticsabstractIt is well known that music can convey emotion and modulate mood, to retrieve music by mood is sometimes the exclusive manner people select music to enjoy. We concentrate on music retrieval by detecting mood. Mood detection is implemented on the viewpoint of computational media aesthetics, that is, by analyzing two music dimensions, tempo and articulation, in the procedure of making music, we derive four categories of mood, happiness, anger, sadness and fear. Concretely, with regard to music in the format of raw audio, after tempo is detected using a multiple agent approach, a feature called relative tempo is calculated, and after the mean and standard deviation of the feature called average silence ratio in the presented computational articulation model are calculated, a simple BP neural network classifier is trained to detect mood. Users retrieval music by browsing the 3D visualization of feature space associated with specific mood. We report the experimental result on a test corpus of 353 pieces of popular music with various genres. Yazhong Feng, Yueting Zhuang, Yunhe Pan |
Web Intelligence | 2 |
| 2003 | 3D motion retrieval with motion index tree
Feng Liu 0015, Yueting Zhuang, Fei Wu 0001, Yunhe Pan |
Comput. Vis. Image Underst. | 2 |
| 2002 | Incomplete motion feature tracking algorithm in video sequencesabstractTo effectively track incomplete motion features, a novel feature tracking algorithm for motion capture is presented. According to feature attributes and relationship among features, extracted features are classified as four types of features. Then different strategies are applied to track different kinds of features. To verify the tracks, cross correlation test and predicted 3D model based test are used to test and remove outliers. Experimental results demonstrate the effectiveness of our algorithm. Zhongxiang Luo, Yueting Zhuang, Feng Liu 0015, Yunhe Pan |
ICIP (3) | 2 |
| 2002 | A graphic-theoretic model for incremental relevance feedback in image retrievalabstractMany traditional relevance feedback approaches for content-based image retrieval (CBIR) can only achieve limited short-term performance improvement without benefiting long-term performance. To remedy this limitation, we propose a graphic-theoretic model for incremental relevance feedback in image retrieval. Firstly, a two-layered graph model is introduced that describes the correlations between images. A teaming strategy is then suggested to enrich the graph model with semantic correlations between images derived from user feedback. Based on the graph model, we propose a link analysis approach for image retrieval and relevance feedback. Experiments conducted on real-world images have demonstrated the advantage of our approach over traditional approaches in both short-term and long-term performance. Yueting Zhuang, Jun Yang 0003, Qing Li 0001, Yunhe Pan |
ICIP (1) | 1 |
| 2002 | Image retrieval and relevance feedback using peer indexingabstractWe present the idea of peer indexing - indexing an image by semantically correlated images - and its application in image retrieval. A learning strategy is suggested for automatic acquisition of peer indices from user feedback, and the similarity metric for the peer index is formulated. A cooperative framework is proposed under which the peer index is integrated with low-level features for image retrieval and relevance feedback. Encouraging results on both short-term and long-term retrieval performance of our approach are shown by experiments. Jun Yang 0003, Qing Li 0001, Yueting Zhuang |
ICME (2) | 3 |
| 2002 | A hierarchical approach: query large music database by acoustic inputabstractNo abstract available. Yazhong Feng, Yueting Zhuang, Yunhe Pan |
SIGIR | 2 |
| 2002 | OCTOPUS: aggressive search of multi-modality data using multifaceted knowledge baseabstractAn important trend in Web information processing is the support of multimedia retrieval. However, the most prevailing paradigm for multimedia retrieval, content-based retrieval (CBR), is a rather conservative one whose performance depends on a set of specifically defined low-level features and a carefully chosen sample object. In this paper, an aggressive search mechanism called Octopus is proposed which addresses the retrieval of multi-modality data using multifaceted knowledge. In particular, Octopus promotes a novel scenario in which the user supplies seed objects of arbitrary modality as the hint of his information need, and receives a set of multi-modality objects satisfying his need. The foundation of Octopus is a multifaceted knowledge base constructed on a layered graph model (LGM), which describes the relevance between media objects from various perspectives. Link analysis based retrieval algorithm is proposed based on the LGM. A unique relevance feedback technique is developed to update the knowledge base by learning from user behaviors, and to enhance the retrieval performance in a progressive manner. A prototype implementing the proposed approach has been developed to demonstrate its feasibility and capability through illustrative examples. Jun Yang 0003, Qing Li 0001, Yueting Zhuang |
WWW | 3 |
| 2002 | Multiple animated characters motion fusionabstractAbstract One of the major problems of the motion capture‐based computer animation technique is the relatively high cost of equipment and low reuse rate of data. To overcome this problem, many motion‐editing methods have been developed. However, most of them can only handle one character whose motions are preset, and hence cannot interact with its environment automatically. In this paper, we construct a new architecture of multiple animated character motion fusion, which not only enables the characters to perceive and respond to the virtual environment, but also allows them to interact with each other. We will also discuss in detail the key issues, such as motion planning, coordination of multiple animated characters and generation of vivid continuous motions. Our experimental results will further testify to the effectiveness of the new methodology. Copyright © 2002 John Wiley & Sons, Ltd. Zhongxiang Luo, Yueting Zhuang, Feng Liu 0015, Yunhe Pan |
Comput. Animat. Virtual Worlds | 2 |
| 2002 | Accommodating hybrid retrieval in a comprehensive video database management systemabstractA comprehensive video retrieval system should be able to accommodate and utilize various (complementary) description data in facilitating effective retrieval. We advocate a hybrid retrieval approach by integrating a query-based (database) mechanism with content-based retrieval (CBR) functions. We describe the VideoMAP/sup +/ architecture, discuss issues related to developing such a comprehensive video database management system, and its specific language mechanism (CAROL/ST with CBR) which provides an improved expressive power than pure query-based or CBR methods currently offer. We also describe an experimental prototype being developed based on a commercial object-oriented toolkit using VC++ and Java. Shermann S.-M. Chan, Qing Li 0001, Yi Wu 0012, Yueting Zhuang |
IEEE Trans. Multim. | 4 |
| 2001 | Web-Based Image Retrieval: A Hybrid ApproachabstractIn recent years, image retrieval has received tremendous attention and some progress has been made. However, most existing work on image retrieval focuses on specific issues and techniques local to image computing or access. Little work has been done to combine various aspects into a single framework. We describe our approach of developing a general-purpose image retrieval system over the Web. A main feature of our system is its hybrid approach by integrating both semantic and visual feature retrieval methods. By combining keyword-based query selection with content-based retrieval techniques, the system is able to provide effective image searching and retrieval. The results are further improved through a relevance feedback process. A research prototype system has been constructed on the Web environment, and experimental results demonstrate the effectiveness of the new approach. Yueting Zhuang, Qing Li 0001, Rynson W. H. Lau |
Computer Graphics International | 1 |
| 2001 | A Hybrid Approach to Video Retrieval in a Generic Video Management and Application Processing FrameworkabstractVideos are multi-faceted objects which can have different kinds of information descriptions. A comprehensive video retrieval system should be able to accommodate and utilize such various (complementary) description data in facilitating effective retrieval. In this paper, we present a hybrid approach by integrating a query-based (database) mechanism with content-based retrieval (CBR) functions. We describe the VideoMAP + architecture and its specific language mechanism (CAROL/ST with CBR) which provides an improved expressive power than what pure query-based or CBR methods currently offer. An experimental prototype is being developed using a commercial object-oriented toolkit on a PC platform. Shermann S.-M. Chan, Qing Li 0001, Yi Wu 0012, Yueting Zhuang |
ICME | 4 |
| 2001 | Thesaurus-Aided Approach For Image Browsing And RetrievalabstractThe current trend of image retrieval is to incorporate image semantics with visual features to enhance retrieval performance. Although many approaches annotate images with keywords and process query at the semantic level, they fail to explore the full potentials of semantics. This paper proposes thesaurus-aided approaches to facilitate semantics-based access to images. The contribution of our work are two-fold: constructing a dynamic semantic hierarchy (DSH) which supports flexible image browsing by semantic subjects, as well as formulating a semantic similarity metric to improve the accuracy of semantic matching. Both approaches are seamlessly integrated into a unified framework for semantics- and feature-based image retrieval. Experiments conducted on the real-world images demonstrate the effectiveness of our approaches. 1. Jun Yang 0003, Wenyin Liu, HongJiang Zhang, Yueting Zhuang |
ICME | 4 |
| 2001 | Web-Based Multimedia Retrieval: Balancing Out between Common Knowledge and Personalized ViewsabstractThe major challenges of multimedia retrieval are the difficulty of generating semantic indexes, as well as the incapability of identifying personalized user interests. This paper attempts to address both problems by suggesting a collaborative yet personalized approach for Web-based multimedia retrieval, which employs a synergy between the relevance feedback technique from the information retrieval community, and the user profiling technique from the information filtering community. Specifically, a "common profile" is established to represent the common knowledge on the semantics of multimedia data, which allow a user to "learn from others" in the retrieval process. On the other hand, for each user a "user profile" is constructed to characterize his/her personal views, which allow a user to "learn from own history". Both types of profiles can be learned and updated incrementally from user feedback. By using an integrated retrieval algorithm based on profiles, this approach strikes the balance between exploiting the common knowledge of most users and catering for the personalized interest of a particular user. The results of some preliminary experiments have demonstrated the effectiveness of the proposed approach. Qing Li 0001, Jun Yang 0003, Yueting Zhuang |
WISE (1) | 3 |
| 2000 | Hierarchical Model Based Human Motion TrackingabstractImage sequence based tracking is the pivotal technique of human motion. In the paper, we first propose an appropriate human model and color model. Second, two approaches are proposed aiming at two different levels of color-block information including the boundary and the inner-area: the extraction of the boundary algorithm is based on the Robert operator and the clustering algorithm is based on self-adaptation. Third, we unite two regions, which are processed by different approaches. This step counteracts the ambiguity of obtaining information from each level. Finally, after the rough area of the color-block is obtained, the boundary of the color-block is determined by calculating the histogram of the X,Y coordinate of every point on the boundary of the block. Experimental results are presented. Yueting Zhuang, Yunhe Pan |
ICIP | 1 |
| 2000 | Content-based video similarity model
Yi Wu 0012, Yueting Zhuang, Yunhe Pan |
ACM Multimedia | 2 |
| 1999 | Video Motion Capture Using Feature Tracking and Skeleton ReconstructionabstractIn the domain of computer vision, there exists a very wide application for the research of human motion capture. This paper proposes a new approach to do motion capture in video. It is composed of image sequence based tracking of human feature points and the reconstruction of three dimension (3D) motion skeleton. First, we track every part of human body from top to bottom on the basis of a human model. The Kalman filter and a morph-block similarity algorithm based on subpixel are used. Then we do camera calibration using the line correspondences between the 3D model and the image. Finally the 3D motion skeleton is established by using the model knowledge. This approach does not aim at a given mode of human motion. Rather, it analyzes large motion from frame to frame in complex, variational background, and sets up a 3D motion skeleton under the perspective projection. We also present the experimental result at the end of the paper. Yueting Zhuang, Xiaoming Liu 0002, Yunhe Pan |
ICIP (4) | 1 |
| 1999 | Video based human animation techniqueabstractHuman animation is a challenging domain in computer animation. To aim at many shortcomings in conventional techniques, this paper proposes a new video based human animation technique. Given a clip of video, firstly human joints are tracked with the support of Kalman filter and morph-block based match in the image sequence. Then corresponding sequence of three-dimension (3D) human motion skeleton is constructed under the perspective projection using camera calibration and human anatomy knowledge. Finally a motion library is established automatically by annotating multiform motion attributes, which can be browsed and queried by the animator. This approach has the characteristic of rich source material, low computing cost, efficient production, and realistic animation result. We demonstrate it on several video clips of people doing full body movements, and visualize the results by re-animating a 3D human skeleton model. Xiaoming Liu 0002, Yueting Zhuang, Yunhe Pan |
ACM Multimedia (1) | 2 |
| 1999 | A new approach to retrieve video by example video clipabstractThe similarity measure between video clips is a key issue in video retrieval. In the developing of our video retrieval system, we propose a new video similarity model. In contrast to existing algorithms, it proposes many influencing factors, such as order factor, speed factor, disturbance factor, etc, based on the subjective visual judgement of human. So this algorithm embodies the degree of similarity completely and systematically. On the other hand, it has resolution adaptation because it can be applied to every level of video structure. In the retrieval system, it can be used to process video query by example clip. This paper introduces it in detail and presents experiment results at the end of the paper. Xiaoming Liu 0002, Yueting Zhuang, Yunhe Pan |
ACM Multimedia (2) | 2 |
| 1999 | Video based human motion captureabstractProposes a new approach to capture human motion in video. This approach does not aim at a given human motion mode, but instead analyzes large-scale motion from frame to frame in a complex variational background and sets up a 3D motion skeleton under perspective projection. This approach is composed of two steps. First, we track every part of a human body from top to bottom on the basis of a human model. Then we perform a camera calibration using the line correspondences between the 3D model and the image, and establish the 3D motion skeleton by using the human model knowledge. The experimental results are presented at the end of the paper. Xiaoming Liu 0002, Yueting Zhuang, Yi Wu 0012, Yunhe Pan |
MMSP | 2 |
| 1999 | Video key frame extraction by unsupervised clustering and feedback adjustment
Yueting Zhuang, Yong Rui, Thomas S. Huang |
J. Comput. Sci. Technol. | 1 |
| 1998 | Adaptive Key Frame Extraction using Unsupervised ClusteringabstractKey frame extraction has been recognized as one of the important research issues in video information retrieval. Although progress has been made in key frame extraction, the existing approaches are either computationally expensive or ineffective in capturing salient visual content. We first discuss the importance of key frame selection; and then review and evaluate the existing approaches. To overcome the shortcomings of the existing approaches, we introduce a new algorithm for key frame extraction based on unsupervised clustering. The proposed algorithm is both computationally simple and able to adapt to the visual content. The efficiency and effectiveness are validated by large amount of real-world videos. Yueting Zhuang, Yong Rui, Thomas S. Huang, Sharad Mehrotra |
ICIP (1) | 1 |
| 1996 | OOADS: An object-oriented design model for advertising CAD system
Yueting Zhuang, Yunhe Pan, Zhijun He 0001 |
J. Comput. Sci. Technol. | 1 |