VLDB 2026 Research / reviewers in the wild / expert
Shuqiang Jiang
dblp:90/3651
· DBLP profile ↗
234ranked-venue papers
22as first author
70since 2021 · last 2026
0000-0002-1596-4326ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 203 · 16 first-author · 54 since 2021Artificial intelligence and machine learning · 66 · 5 first-author · 30 since 2021Computer networks · 11 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VPN: Visual Prompt NavigationabstractWhile natural language is commonly used to guide embodied agents, the inherent ambiguity and verbosity of language often hinder the effectiveness of language-guided navigation in complex environments. To this end, we propose Visual Prompt Navigation (VPN), a novel paradigm that guides agents to navigate using only user-provided visual prompts within 2D top-view maps. This visual prompt primarily focuses on marking the visual navigation trajectory on a top-down view of a scene, offering intuitive and spatially grounded guidance without relying on language instructions. It is more friendly for non-expert users and reduces interpretive ambiguity. We build VPN tasks in both discrete and continuous navigation settings, constructing two new datasets, R2R-VP and R2R-CE-VP, by extending existing R2R and R2R-CE episodes with corresponding visual prompts. Furthermore, we introduce VPNet, a dedicated baseline network to handle the VPN tasks, with two data augmentation strategies: view-level augmentation (altering initial headings and prompt orientations) and trajectory-level augmentation (incorporating diverse trajectories from large-scale 3D scenes), to enhance navigation performance. Extensive experiments evaluate how visual prompt forms, top-view map formats, and data augmentation strategies affect the performance of visual prompt navigation. Yuchen Li 0006, Hengyi Cai, Shuaiqiang Wang, Gim Hee Lee, Piji Li, Shuqiang Jiang |
AAAI | 9 |
| 2026 | RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image RetrievalabstractFine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44% to 17.20% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication. Weiqing Min, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ICMR | 6 |
| 2026 | Self Model for Embodied Artificial Intelligence
Shuqiang Jiang, Si-Xian Zhang, Shi-Da Tao, Xi-Hong Zhu, Tian-Liang Qi, Xin-Hang Song |
J. Comput. Sci. Technol. | 1 |
| 2026 | Continual novel class discovery under domain shift with entropy-based selection and representation evolution
Feifei Shi, Xiangyang Li 0002, Shuqiang Jiang, Yong Rui |
Multim. Syst. | 3 |
| 2026 | DPFA-net: a lightweight hybrid neural network with dual path feature aggregation for food image recognition
Xiangyi Zhu, Yingnan Sheng, Congrui Lv, Guorui Sheng, Weiqing Min, Shuqiang Jiang |
Multim. Syst. | 7 |
| 2026 | Large-Scale Logo DetectionabstractLogo detection is crucial for trademark compliance and media monitoring, enabling companies to monitor online trademark usage and evaluate brand visibility on social media and advertisements. The use of large datasets significantly improves accuracy and generalization, emphasizing the need for high-quality datasets to optimize performance and enhance reasoning abilities in visual detection models. This drove us to create Logo4500, an unparalleled dataset featuring 4,500 logo categories and over 293,000 meticulously labeled images. To ensure the dataset's quality, we meticulously designed the construction and annotation process, with detailed information provided in our paper. Compared to existing logo datasets, Logo4500 offers greater diversity and class imbalance, making it more reflective of real-world distribution. Leveraging this high-quality dataset, we introduce a benchmark called Frequency-Aware Learnable Dual Reweighting Network (FALDR-Net), which enhances the representation of ambiguous features and addresses class imbalance for large-scale logo detection. We conducted extensive experiments, evaluating various recent methods on this new dataset and several existing publicly available logo datasets, demonstrating its effectiveness. Additionally, we verified Logo4500's generalization ability in several tasks. We anticipate that Logo4500 and the benchmark will inspire further exploration in the logo-related research community, facilitating the advancement of visual foundation models. Sujuan Hou, Weiqing Min, Jianxin Zhan, Mengmeng Zhang 0008, Peng Li 0081, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Goal-Oriented Dynamic Weight Optimization for Multi-Object NavigationabstractMulti-object navigation (MON) tasks involve sequentially locating multiple targets in an unknown environment, requiring global long-term planning under incomplete information. This necessitates that the agent dynamically balance immediate actions and long-term rewards while considering both local adaptability and global foresight. However, current methods overly focus on local path optimization, which leads to slower convergence in sparse reward settings and increases the risk of deadlocks or trap states. The core challenge of MON lies in the deformation of the shared decision space, where independent optimization leads to redundant and overlapping paths. Thus, path planning requires dynamic, cross-task optimization rather than simple subtask aggregation. To minimize overall effort, the optimization process should adaptively balance task contributions through weight adjustment. Thus, we propose the Goal-oriented Dynamic Weight Optimization (GDWO) algorithm. GDWO integrates target-specific value loss functions into a unified optimization framework and dynamically adjusts weights through gradient-based updates. To prevent over-optimization, weights are normalized during training according to navigation success rates, prioritizing more challenging targets. This adaptive mechanism effectively addresses the challenge of sparse rewards and improves convergence efficiency. By leveraging this mechanism, GDWO unifies multiple objectives within a unified decision space, achieving efficient optimization and balancing short-term gains with long-term goals. Additionally, we introduce two auxiliary modules: prior knowledge-based navigation and frontier-aware exploration to further enhance GDWO's performance. Experimental results on the Gibson and Matterport3D datasets demonstrate that GDWO achieves improvements in key metrics for MON tasks. It optimizes path planning, reduces exploration costs, and enhances navigation efficiency, enabling the agent to perform tasks more effectively in complex environments. Haitao Zeng, Xinhang Song, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Structured-condensed prompt tuning in vision-language models for fine-grained image recognition
Xinda Liu, Weiqing Min, Guohua Geng, Shuqiang Jiang |
Pattern Recognit. | 5 |
| 2026 | LLM-informed global-local contextualization for zero-shot food detection
Weiqing Min, Guorui Sheng, Jingru Song, Yancun Yang, Shuqiang Jiang |
Pattern Recognit. | 7 |
| 2026 | SyMFood: Synergistic Multi-Modal Prompting for Fine-Grained Zero-Shot Food DetectionabstractFine-grained object detection in food computing is severely constrained by the vast diversity of food items and the high cost of data annotation. Existing Zero-Shot Food Detection (ZSFD) methods attempt to solve this by leveraging semantic information, but they suffer from two critical bottlenecks: (1) a "Semantic Dilemma" (SD) where textual descriptions are too ambiguous to distinguish visually similar food categories, and (2) an "Architectural Bottleneck" (AB) due to the granularity mismatch between high-level semantics and low-level visual features. In this article, we propose SyMFood (Synergistic Multi-modal Framework for Food Generalization), a novel ZSFD framework designed to systematically overcome these challenges. To resolve the SD, SyMFood employs a multi-modal prompt system, which combines rich descriptions from Large Language Models (LLMs) with unambiguous visual exemplars to provide precise semantic grounding. To break the AB, SyMFood introduces a “Refine-then-Fuse” architecture. This design first utilizes a Context-Aware Spatial-Channel Refinement (CaSC) block to enhance visual features independently. Subsequently, a Progressive Food Knowledge Fusion (ProFus) module performs bi-directional, iterative co-refinement between the enhanced visual features and multi-modal prompts across all scales. Extensive experiments across four challenging datasets, including food-specific (UEC FOOD 256, FOWA) and general-purpose (PASCAL VOC, MS COCO) benchmarks, validate our approach. The proposed method outperforms baselines, yielding a notable 8.5% improvement in Harmonic Mean on the FOWA dataset in the genearl ZSD (GZSD) setting. The source code will be available at https://github.com/Niko000202/SymFood0202. Weiqing Min, Shoulong Liu, Guorui Sheng, Shuqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | CondFoodGen: A Conditional Two-Stream Network for Controllable Food Image GenerationabstractFood image generation is an important research direction in food computing, aiming to produce highly realistic images that accurately capture the visual characteristics of various dishes while adhering to specified input conditions. Existing methods that rely solely on textual descriptions struggle to handle the large intra-class variability of food, often resulting in limited diversity and accuracy. Although some approaches incorporate additional conditions, they generally lack optimizations for food-specific challenges, leading to inconsistencies in texture, shape, and color fidelity. To address these limitations, we propose CondFoodGen, a diffusion-based two-stream network for controllable food image generation. The architecture consists of a control stream and a generation stream, where the control stream provides conditional guidance to regulate the generation process. To optimize bidirectional interactions between the two streams, we introduce the Bidirectional Adaptive Gating (BAG) mechanism, which not only guides synthesis but also adaptively refines control representations through feedback from the generation stream. In addition, we propose the Wavelet-Guided Hierarchical Attention (WGHA) module, which combines wavelet-based multi-frequency analysis with hierarchical attention to enhance fine-grained texture fidelity and structural realism. A progressive multi-stage training strategy further stabilizes optimization and enables seamless integration of conditional guidance with bidirectional interaction. Extensive experiments on three food image datasets demonstrate that CondFoodGen consistently generates high-quality and diverse images. Compared with the best existing food image generation methods, our approach achieves an average improvement of about 11.0% across three evaluation metrics and compared to the leading conditional generation approaches, the average improvement reaches 16.2%. The source code, trained models, and supplementary materials are publicly available at https://github.com/housujuan123/CondFoodGen. Mengyao Zhao, Hao Xiong 0001, Weiqing Min, Sujuan Hou, Mengmeng Zhang 0008, Shuqiang Jiang |
IEEE Trans. Image Process. | 6 |
| 2026 | FoodHash: Context-Aware Proxy Interaction and Fusion for Food Image RetrievalabstractVision-based food image retrieval has garnered significant attention due to its potential for critical applications in dietary and health management. However, food images exhibit more complex feature distributions and lack the geometric regularity and structured patterns typically observed in general image retrieval tasks. This complexity poses a challenge for existing models to extract fine-grained features and semantic information, thereby compromising retrieval performance. To address this challenge, we propose FoodHash, a context-aware proxy interaction and fusion hashing method for food image retrieval. The method incorporates an Aggregation–Interaction–Propagation (AIP) module that facilitates contextual information exchange among patch tokens within the same feature map, guided by proxy tokens, thereby effectively capturing the intricate details of food images. Furthermore, to leverage the rich semantic information in food images, a Cross-Fusion Module is introduced to efficiently integrate multi-scale information and enhance feature representation. Additionally, we employ a novel loss function to optimize hash learning by ensuring consistency between hash codes and the semantic space, thereby enhancing the learning capability of hash coding. Extensive experiments on three publicly available food datasets demonstrate that FoodHash significantly surpasses existing models in retrieval performance. Specifically, on the ETH Food-101 dataset, FoodHash achieves improvements of 18.1%, 6.7%, 5.2%, and 4.5% over the suboptimal method PTLCH for 16-bits, 32-bits, 48-bits, and 64-bits hash codes, respectively. The source code will be made publicly available upon publication of the article. Pindan Cao, Weiqing Min, Guorui Sheng, Yongqiang Song, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Trial-Oriented Visual Rearrangement
Yuyi Liu, Xinhang Song, Tianliang Qi, Shuqiang Jiang |
ICCV | 4 |
| 2025 | Learning on the Go: A Meta-Learning Object Navigation Model
Xiaorong Qin, Xinhang Song, Sixian Zhang, Xinyao Yu 0002, Xinmiao Zhang 0004, Shuqiang Jiang |
ICCV | 6 |
| 2025 | Function-Centric Bayesian Network for Zero-Shot Object Goal Navigation
Sixian Zhang, Xinyao Yu 0002, Xinhang Song, Yiyao Wang, Shuqiang Jiang |
ICCV | 5 |
| 2025 | DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language ModelabstractGenerating robot demonstrations through simulation is widely recognized as an effective way to scale up robot data. Previous work often trained reinforcement learning agents to generate expert policies, but this approach lacks sample efficiency. Recently, a line of work has attempted to generate robot demonstrations via differentiable simulation, which is promising but heavily relies on reward design, a labor-intensive process. In this paper, we propose DiffGen, a novel framework that integrates differentiable physics simulation, differentiable rendering, and a vision-language model to enable automatic and efficient generation of robot demonstrations. Given a simulated robot manipulation scenario and a natural language instruction, DiffGen can generate realistic robot demonstrations by minimizing the distance between the embedding of the language instruction and the embedding of the simulated observation after manipulation in representation space. The embeddings are obtained from the vision-language model, and the optimization is achieved by calculating and descending gradients through the differentiable simulation, differentiable rendering, and vision-language model components. Experiments demonstrate that with DiffGen, we could efficiently and effectively generate robot data with minimal human effort or training time. The videos of the results can be accessed at https://sites.google.com/view/diffgen. Shuqiang Jiang, Cewu Lu |
IROS | 3 |
| 2025 | DSDGF-Nutri: A Decoupled Self-Distillation Network with Gating Fusion For Food Nutritional AssessmentabstractAccurate assessment of food nutrition is essential for promoting healthy eating habits. While recent deep learning approaches have enhanced vision-based nutritional estimation through RGB-D multi-modal fusion, they often overlook fine-grained surface components (e.g., oil and sugar) that significantly influence nutritional values. Some recent approaches have improved accuracy by incorporating ingredient data, but their reliance on such input during inference limits practical applicability, as ingredient details are often unavailable in real-world settings. To address this limitation, we propose DSDGF-Nutri, a novel Decoupled Self-Distillation network with Gating Fusion for food Nutri tional assessment. Our method leverages ingredient knowledge during training but relies solely on RGB-D inputs at inference. Specifically, DSDGF-Nutri introduces: (1) a self-distillation mechanism with gating fusion that transfers ingredient-aware features to the RGB-D network, enabling robust prediction without test-time ingredient input, and (2) a multi-task decoupling architecture with task-specific decoders to minimize cross-task interference. Extensive evaluations on two benchmark datasets demonstrate DSDGF-Nutri outperforms existing methods, achieving state-of-the-art results. This work establishes a new paradigm of multimodal fusion in nutritional assessment by unifying scientific measurements with scalable computer vision applications. Sujuan Hou, Zhihui Feng, Hao Xiong 0001, Weiqing Min, Peng Li 0081, Shuqiang Jiang |
ACM Multimedia | 6 |
| 2025 | Spatial-Aware Multi-Modal Information Fusion for Food Nutrition EstimationabstractFood nutrition assessment plays a crucial role in maintaining health, preventing diseases, and promoting scientific dietary habits. However, existing nutrition assessment methods often fail to fully consider the relationships between tasks, leading to limited overall performance. Specifically, these methods suffer from three major challenges: (1) task conflicts, where different tasks compete during joint optimization, leading to suboptimal overall performance; (2) varying training difficulties among tasks, leading to imbalanced learning and subpar model generalization; and (3) the small-scale and complex distribution of datasets, which limits the robustness of learned representations. To address these issues, we propose a novel method that reduces interference between tasks, dynamically focuses on more challenging tasks, and incorporates 3D spatial awareness to enhance multi-modal feature representation. First, we decouple the prediction network from the backbone and introduce a CAMTH (Cross-Attention-Based Multi-Task Head Module), effectively mitigating task interference and fully leveraging each task's learning potential. Second, we improve the loss function to adaptively focus on more challenging tasks, improving overall model performance. Third, we design a 3D-FEM (3D Feature Extraction Module) and MMFF (Multi-Modal Feature Fusion Module), enabling the model to fully exploit the spatial information of food and enhance the food's multi-modal feature representation. We validate our method through extensive experiments on the Nutrition5K dataset, comparing it with state-of-the-art (SOTA) models. The results show that our method achieves superior performance in nutrition estimation, demonstrating the effectiveness of our method. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2025 | Cross-Layer and Selective Distillation for Asymmetric Image Retrieval
Weiqing Min, Fangyuan Yao, Guorui Sheng, Shuqiang Jiang |
PRCV (12) | 5 |
| 2025 | Channel grouping vision transformer for lightweight fruit and vegetable recognition
Chengxu Liu 0002, Weiqing Min, Jingru Song, Yancun Yang, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
Expert Syst. Appl. | 8 |
| 2025 | HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object NavigationabstractThe goal of object navigation task is to reach the expected objects using visual information in unseen environments. Previous works typically implement deep models as agents that are trained to predict actions based on visual observations. Despite extensive training, agents often fail to make wise decisions when navigating in unseen environments toward invisible targets. In contrast, humans demonstrate a remarkable talent to navigate toward targets even in unseen environments. This superior capability is attributed to the cognitive map in the hippocampus, which enables humans to recall past experiences in similar situations and anticipate future occurrences during navigation. It is also dynamically updated with new observations from unseen environments. The cognitive map equips humans with a wealth of prior knowledge, significantly enhancing their navigation capabilities. Inspired by human navigation mechanisms, we propose the Hierarchical Object-to-Zone (HOZ++) graph, which encapsulates the regularities among objects, zones, and scenes. The HOZ++ graph helps the agent to identify the current zone and the target zone, and computes an optimal path between them, then selects the next zone along the path as the guidance for the agent. Moreover, the HOZ++ graph continuously updates based on real-time observations in new environments, thereby enhancing its adaptability to new environments. Our HOZ++ graph is versatile and can be integrated into existing methods, including end-to-end RL and modular methods. Our method is evaluated across four simulators, including AI2-THOR, RoboTHOR, Gibson, and Matterport 3D. Additionally, we build a realistic environment to evaluate our method in the real world. Experimental results demonstrate the effectiveness and efficiency of our proposed method. Sixian Zhang, Xinhang Song, Xinyao Yu 0002, Yubing Bai, Xinlong Guo, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | SSC-PPI: A Subspace Structure Consistency-Based Method for Protein-Protein Interactions PredictionabstractProtein-protein interactions (PPIs) play an indispensable role in understanding disease-causing mechanisms, and the basic laws of food and drugs on life. Contemporary research on this issue, however, is incapable of guaranteeing structure consistency between extracted features and raw data, and fails to fully investigate the interconnection information of features. Thus, this paper proposes a subspace structure consistency-based method for protein-protein interactions prediction. SSC-PPI is not only capable of investigating the coherent relations between the encoded features generated from amino acid composition and conjoint triad numeric composition of F-vector, composition and transition descriptors, but also fully maintains the latent geometrical structure consistency between feature subspace and data space. Numerous comparative experiments demonstrate its excellent predictable performance with significant accuracies of 100$\%$, 99.95$\%$, 99.98$\%$, 100$\%$ and 100$\%$ respectively on Helicobacter pylori, Human, Saccharomyces cerevisiae (core subset), Human-Bacillus Anthracis and Human-Yersinia pestis datasets, significantly outperforming the comparative models by average increases of 14.39$\%$, 5.45$\%$, 8.10$\%$, 6.05$\%$ and 8.79$\%$ respectively. Additionally, SSC-PPI offers an efficient and reliable framework for large-scale prediction tasks such as drug-drug and drug-food interactions. Ziping Ma 0001, Weiqing Min, Huanpu Zhang, Yulei Huang, Shuqiang Jiang |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Food3D: Text-Driven Customizable 3D Food Generation With Gaussian SplattingabstractRealistic 3D food creation generation plays a critical role in applications such as nutritional assessment, advertising, and virtual content creation. The existing text-to-3D models typically begin by initializing a 3D representation, which is subsequently refined using supervision from a text-to-image model to obtain the final 3D output. In this work, we present Food3D, a novel framework for 3D food generation designed to address two main limitations of current models. First, the limitation of initialization in 3D generation: poor initialization can result in the generated 3D food lacking crucial details and realism, thereby reducing its quality. To address this issue, we propose a generalized method named Food3D-G, which uses Mamba-based initialization to improve the starting point of the initialization process, thereby enhancing the visual fidelity and quality of the generated 3D food. Second, the limitation of text-to-image models: current text-to-3D models often rely on text-to-image models for supervision. However, a considerable gap persists between the generated images and real-world visuals, particularly when modeling complex food structures. These models fail to accurately capture the fine details and textures, which negatively impacts the quality and realism of the generated 3D food models. To address this limitation, we propose a customizable method for personalized 3D food generation, termed Food3D-C. This method employs a dual-branch diffusion model that effectively captures intricate details, particularly in complex food structures. Within the Food3D framework, both proposed methods incorporate 3D Gaussian splatting (3D GS) and a schedulable interval score matching (S-ISM) algorithm to enhance shape and texture generation. Extensive experiments demonstrate that Food3D achieves state-of-the-art performance, with substantial improvements in detail, shape accuracy, and overall visual realism. Project page and source codes: https://yudongjian.github.io/Food3D/. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shaowen Yao 0001, Shuqiang Jiang |
IEEE Trans. Image Process. | 6 |
| 2025 | Guest Editorial: When Multimedia Meets Food: Multimedia Computing for Food Data Analysis and Applications
Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic 0001, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Multimodal Food LearningabstractFood-centered study has received more attention in the multimedia community for its profound impact on our survival, nutrition and health, pleasure, and enjoyment. Our experience of food is typically multi-sensory: We see food objects, smell its odors, taste its flavors, feel its texture, and hear sounds when chewing. Therefore, multimodal food learning is vital in food-centered study, which aims to relate information from multiple food modalities to support various multimedia tasks, ranging from recognition, retrieval, generation, recommendation, and interaction, enabling applications in different fields like healthcare and agriculture. However, there is no surveys on this topic to our knowledge. To fill this gap, this article formalizes multimodal food learning and comprehensively surveys its typical tasks, technical achievements, existing datasets, and applications to provide the blueprint with researchers and practitioners. Based on the current state of the art, we identify both open research issues and promising research directions, such as multimodal food learning benchmark construction, multimodal food foundation model construction, and multimodality diet estimation. We also point out that closer cooperation from researchers between multimedia and food science can handle some existing challenges and meanwhile open up more new opportunities to advance the fast development of multimodal food learning. This is the first comprehensive survey in this topic and we anticipate about 170 reviewed research articles can benefit academia and industry in this community and beyond. Weiqing Min, Xingjian Hong, Yuxin Liu 0009, Leyi Xu, Yilin Wang 0012, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 9 |
| 2025 | Diverse and High-Quality Food Image Generation from Only Food NamesabstractFood image generation holds promising application prospects in food design, advertising, and food education. However, the existing methods rely on information such as recipes, ingredients, or food names, which leads to generated food images with less intra-class diversity. When recipes, ingredients, and food names are identical for the same food, the real-world images may vary significantly in appearance. The question of how to simultaneously ensure the quality and diversity of the generated images is a key issue. To this end, we employ pre-trained diffusion model and Transformer to propose a method for generating diverse and high-quality images of both Chinese and Western food, named CW-Food. Different from previous works that utilize an overall food feature to generate new images, CW-Food first decouples the food images to obtain common intra-class features and private instance features. Additionally, we design a Transformer-based feature fusion module to integrate the common and private features, in order to avoid the shortcomings of conventional methods. Moreover, we also utilize a pre-trained diffusion model as our backbone, which is fine-tuned using LoRA with the fused multi-variate features. Extensive experiments on four datasets demonstrate the advantages of our proposed method, producing diverse and high-quality food images encompassing both Chinese and Western cuisines. To the best of our knowledge, our work is the first attempt to generate Chinese food images using only food names. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | A Category Agnostic Model for Visual RearrangmentabstractThis paper presents a novel category agnostic model for visual rearrangement task, which can help an embodied agent to physically recover the shuffled scene configuration without any category concepts to the goal configuration. Previous methods usually follow a similar architecture, completing the rearrangement task by aligning the scene changes of the goal and shuffled configuration, according to the semantic scene graphs. However, constructing scene graphs requires the inference of category labels, which not only causes the accuracy drop of the entire task but also limits the application in real world scenario. In this paper, we delve deep into the essence of visual re-arrangement task and focus on the two most essential issues, scene change detection and scene change matching. We utilize the movement and the protrusion of point cloud to accurately identify the scene changes and match these changes depending on the similarity of category agnostic appearance feature. Moreover, to assist the agent to explore the environment more efficiently and comprehensively, we propose a closer-aligned-retrace exploration policy, aiming to observe more details of the scene at a closer distance. We conduct extensive experiments onAI2THOR Rearrangement Challenge based on RoomR dataset and a new multi-room multi-instance dataset MrMiR collected by us. The experimental results demonstrate the effectiveness of our proposed method. Yuyi Liu, Xinhang Song, Shuqiang Jiang |
CVPR | 5 |
| 2024 | An Interactive Navigation Method with Effect-oriented AffordanceabstractVisual navigation is to let the agent reach the target according to the continuous visual input. In most previous works, visual navigation is usually assumed to be done in a static and ideal environment: the target is always reachable with no need to alter the environment. However, the “messy” environments are more general and practical in our daily lives, where the agent may get blocked by obstacles. Thus Interactive Navigation (InterNav) is introduced to navigate to the objects in more realistic “messy” environments according to the object interaction. Prior work on InterNav learns shortterm interaction through extensive trials with reinforcement learning. However, interaction does not guarantee efficient navigation, that is, plan-ning obstacle interactions that make shorter paths and con-sume less effort is also crucial. In this paper, we introduce an effect-oriented affordance map to enable longterm interactive navigation, extending the existing map-based nav-igation framework to the domain of dynamic environment. We train a set of affordance functions predicting available interactions and the time cost of removing obstacles, which informatively support an interactive modular system to ad-dress interaction and longterm planning. Experiments on the ProcTHOR simulator demonstrate the capability of our affordance-driven system in longterm navigation in complex dynamic environments. Yuehu Liu, Xinhang Song, Yuyi Liu, Sixian Zhang, Shuqiang Jiang |
CVPR | 6 |
| 2024 | Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language NavigationabstractVision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step, the agent selects from possible candidate locations and then makes the move. For better navigation planning, the lookahead exploration strategy aims to effectively evaluate the agent's next action by accurately anticipating the future environment of candidate locations. To this end, some existing works predict RGB images for future environments, while this strategy suffers from image distortion and high computational cost. To address these issues, we propose the pre-trained hierarchical neural radiance representation model (HNR) to produce multi-level semantic features for future environments, which are more robust and efficient than pixel-wise RGB reconstruction. Furthermore, with the predicted future environmental representations, our lookahead VLN model is able to construct the navigable future path tree and select the optimal path via efficient parallel evaluation. Extensive experiments on the VLN-CE datasets confirm the effectiveness of our method. The code is available at https://github.com/MrZihan/HNR-VLN Xiangyang Li 0002, Yeqi Liu, Junjie Hu 0001, Ming Jiang 0018, Shuqiang Jiang |
CVPR | 7 |
| 2024 | Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationabstractThe Object Goal navigation (ObjectNav) task requires the agent to navigate to a specified target in an unseen environment. Since the environment layout is unknown, the agent needs to infer the unknown contextual objects from partially observations, thereby deducing the likely location of the target. Previous end-to-end RL methods capture contextual relationships through implicit representations while they lack notion of geometry. Alternatively, modular methods construct local maps for recording the observed geometric structure of unseen environment, however, lacking the reasoning of contextual relation limits the exploration efficiency. In this work, we propose the self-supervised generative map (SGM), a modular method that learns the explicit context relation via self-supervised learning. The SGM is trained to leverage both episodic observations and general knowledge to reconstruct the masked pixels of a cropped global map. During navigation, the agent maintains an incomplete local semantic map, meanwhile, the unknown regions of the local map are generated by the pretrained SGM. Based on the generated map, the agent sets the predicted location of the target as the goal and moves towards it. Experiments on Gibson, MP3D and HM3D show the effectiveness of our method. The code is available at https://github.com/sx-zhang/SGM. Sixian Zhang, Xinyao Yu 0002, Xinhang Song, Shuqiang Jiang |
CVPR | 5 |
| 2024 | Trajectory Diffusion for ObjectGoal NavigationabstractObject goal navigation requires an agent to navigate to a specified object in an unseen environment based on visual observations and user-specified goals.
Human decision-making in navigation is sequential, planning a most likely sequence of actions toward the goal.
However, existing ObjectNav methods, both end-to-end learning methods and modular methods, rely on single-step planning. They output the next action based on the current model input, which easily overlooks temporal consistency and leads to myopic planning.
To this end, we aim to learn sequence planning for ObjectNav. Specifically, we propose trajectory diffusion to learn the distribution of trajectory sequences conditioned on the current observation and the goal.
We utilize DDPM and automatically collected optimal trajectory segments to train the trajectory diffusion.
Once the trajectory diffusion model is trained, it can generate a temporally coherent sequence of future trajectory for agent based on its current observations.
Experimental results on the Gibson and MP3D datasets demonstrate that the generated trajectories effectively guide the agent, resulting in more accurate and efficient navigation. Xinyao Yu 0002, Sixian Zhang, Xinhang Song, Xiaorong Qin, Shuqiang Jiang |
NeurIPS | 5 |
| 2024 | Multi-state Ingredient Recognition via Adaptive Multi-centric NetworkabstractIngredient recognition has received significant attention due to its numerous industrial applications, such as intelligent retail terminals and intelligent cooking devices. However, ingredient recognition has the following challenges: 1) dynamic changes in the number of categories; 2) greater diversity and regionality of ingredients; and 3) large visual differences among different states of ingredients. In this article, we propose an adaptive multi-centric network (AdMNet) to solve the problem of ingredient recognition. AdMNet is based on the idea of retrieval, which consists of two main parts, the adaptive multi-centric nearest-neighbor central mean (AdM-NCM) classifier, and the context-aware attentional pooling (CAP) module. The AdM-NCM classifier adaptively establishes category-centric vector groups to recognize ingredients via optimizing the minimum clustering variance, where each state of the ingredient has its corresponding centric vector. The CAP module combines contextual information and multiple attention mechanisms. It captures more focused and discriminative features with higher weights assigned to fine-grained features, which results in better feature representation. In addition, we collect a large-scale ingredient dataset, ISIA Ingredient-201 with 201 classes and 100 442 images. To prove the greater robustness and generalization of our method, we compare the metrics in basic scenarios and realistic scenarios with those of other methods. Specifically, the base scenario is the regular setup, and the real scenario is similar to the class incremental learning setup. The experimental results show that our method reaches the state of the art on both basic scenarios and realistic scenarios with small samples. Jiajun Song, Weiqing Min, Weimin Xiao, Shuqiang Jiang |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Convolution-Enhanced Bi-Branch Adaptive Transformer With Cross-Task Interaction for Food Category and Ingredient RecognitionabstractRecently, visual food analysis has received more and more attention in the computer vision community due to its wide application scenarios, e.g., diet nutrition management, smart restaurant, and personalized diet recommendation. Considering that food images are unstructured images with complex and unfixed visual patterns, mining food-related semantic-aware regions is crucial. Furthermore, the ingredients contained in food images are semantically related to each other due to the cooking habits and have significant semantic relationships with food categories under the hierarchical food classification ontology. Therefore, modeling the long-range semantic relationships between ingredients and the categories-ingredients semantic interactions is beneficial for ingredient recognition and food analysis. Taking these factors into consideration, we propose a multi-task learning framework for food category and ingredient recognition. This framework mainly consists of a food-orient Transformer named Convolution-Enhanced Bi-Branch Adaptive Transformer (CBiAFormer) and a multi-task category-ingredient recognition network called Structural Learning and Cross-Task Interaction (SLCI). In order to capture the complex and unfixed fine-grained patterns of food images, we propose a query-aware data-adaptive attention mechanism called Bi-Branch Adaptive Attention (BiA-Attention) in CBiAFormer, which consists of a local fine-grained branch and a global coarse-grained branch to mine local and global semantic-aware regions for different input images through an adaptive candidate key/value sets assignment for each query. Additionally, a convolutional patch embedding module is proposed to extract the fine-grained features which are neglected by Transformers. To fully utilize the ingredient information, we propose SLCI, which consists of cross-layer attention to model the semantic relationships between ingredients and two cross-task interaction modules to mine the semantic interactions between categories and ingredients. Extensive experiments show that our method achieves competitive performance on three mainstream food datasets (ETH Food-101, Vireo Food-172, and ISIA Food-200). Visualization analyses of CBiAFormer and SLCI on two tasks prove the effectiveness of our method. Codes will be released upon publication. Code and models are available at https://github.com/Liuyuxinict/CBiAFormer. Yuxin Liu 0009, Weiqing Min, Shuqiang Jiang, Yong Rui |
IEEE Trans. Image Process. | 3 |
| 2024 | Synthesizing Knowledge-Enhanced Features for Real-World Zero-Shot Food DetectionabstractFood computing brings various perspectives to computer vision like vision-based food analysis for nutrition and health. As a fundamental task in food computing, food detection needs Zero-Shot Detection (ZSD) on novel unseen food objects to support real-world scenarios, such as intelligent kitchens and smart restaurants. Therefore, we first benchmark the task of Zero-Shot Food Detection (ZSFD) by introducing FOWA dataset with rich attribute annotations. Unlike ZSD, fine-grained problems in ZSFD like inter-class similarity make synthesized features inseparable. The complexity of food semantic attributes further makes it more difficult for current ZSD methods to distinguish various food categories. To address these problems, we propose a novel framework ZSFDet to tackle fine-grained problems by exploiting the interaction between complex attributes. Specifically, we model the correlation between food categories and attributes in ZSFDet by multi-source graphs to provide prior knowledge for distinguishing fine-grained features. Within ZSFDet, Knowledge-Enhanced Feature Synthesizer (KEFS) learns knowledge representation from multiple sources (e.g., ingredients correlation from knowledge graph) via the multi-source graph fusion. Conditioned on the fusion of semantic knowledge representation, the region feature diffusion model in KEFS can generate fine-grained features for training the effective zero-shot detector. Extensive evaluations demonstrate the superior performance of our method ZSFDet on FOWA and the widely-used food dataset UECFOOD-256, with significant improvements by 1.8% and 3.7% ZSD mAP compared with the strong baseline RRFS. Further experiments on PASCAL VOC and MS COCO prove that enhancement of the semantic knowledge can also improve the performance on general ZSD. Code and dataset are available at https://github.com/LanceZPF/KEFS. Weiqing Min, Jiajun Song, Yang Zhang 0117, Shuqiang Jiang |
IEEE Trans. Image Process. | 5 |
| 2024 | Deep Learning for Logo Detection: A SurveyabstractLogo detection has gradually become a research hotspot in the field of computer vision and multimedia for its various applications, such as social media monitoring, intelligent transportation, and video advertising recommendation. Recent advances in this area are dominated by deep learning-based solutions, where many datasets, learning strategies, network architectures, and loss functions have been employed. This article reviews the advance in applying deep learning techniques to logo detection. First, we discuss a comprehensive account of public datasets designed to facilitate performance evaluation of logo detection algorithms, which tend to be more diverse, more challenging, and more reflective of real life. Next, we perform an in-depth analysis of the existing logo detection strategies and their strengths and weaknesses of each learning strategy. Subsequently, we summarize the applications of logo detection in various fields, from intelligent transportation and brand monitoring to copyright and trademark compliance. Finally, we analyze the potential challenges and present the future directions for the development of logo detection. This study aims better to inform readers about the current state of logo detection and encourage more researchers to get involved in logo detection. Sujuan Hou, Weiqing Min, Yanna Zhao, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Towards Food Image Retrieval via Generalization-Oriented Sampling and Loss Function DesignabstractFood computing has increasingly received widespread attention in the multimedia field. As a basic task of food computing, food image retrieval has wide applications, that is, food image retrieval can help users to find the desired food from a large number of food images. Besides, the retrieved information can be applied to establish a richer database for the subsequent food content-related recommendation. Food image retrieval aims to achieve better performance on novel categories. Thus, it is worth studying to transfer the embedding ability from the training set to the unseen test set, that is, the generalization of the model. Food is influenced by various factors, such as culture and geography, leading to great differences between domains, such as Asian food and western food. Therefore, it is challenging to study the generalization of the model in food image retrieval. In this article, we improve the classical metric learning framework and propose a generalization-oriented sampling strategy, which boosts the generalization of the model by maximizing the intra-class distance from a proportion of positive pairs to avoid the excessive distance compression in the embedding space. Considering that the existing optimization process is in an opposite direction to our proposed sampling strategy, we further propose an adaptive gradient assignment policy named gradient-adaptive optimization , which can alleviate the intra-class distance compression during optimization by assigning different gradients to different samples. Extensive evaluation on three popular food image datasets demonstrates the effectiveness of the proposed method. We also experiment on three popular general datasets to prove that solving the problem from the generalization can also improve the performance of general image retrieval. Code is available at https://github.com/Jiajun-ISIA/Generalization-oriented-Sampling-and-Loss . Jiajun Song, Weiqing Min, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Lightweight Food Recognition via Aggregation Block and Feature EncodingabstractFood image recognition has recently been given considerable attention in the multimedia field in light of its possible implications on health. The characteristics of the dispersed distribution of ingredients in food images put forward higher requirements on the long-range information extraction ability of neural networks, leading to more complex and deeper models. Nevertheless, the lightweight version of food image recognition is essential for improved implementation on end devices and sustained server-side expansion. To address this issue, we present Aggregation Feature Net (AFNet), a lightweight network that is capable of effectively capturing both global and local features from food images. In AFNet, we develop a novel convolution based on a residual model by encoding global features through row-wise and column-wise information integration. Merging aggregation block with classic local convolution yields a framework that works as the backbone of the network. Based on the efficient use of parameters by the aggregation block, we constructed a lightweight food image recognition network with fewer layers and a smaller scale, assisted by a new type of activation function. Experimental results on four popular food recognition datasets demonstrate that our approach achieves state-of-the-art performance with higher accuracy and fewer FLOPs and parameters. For example, in comparison to the current state-of-the-art model of MobileViTv2, AFNet achieved 88.4% accuracy of the top-1 level on the ETHZ Food-101 dataset, with similar parameters and FLOPs but 1.4% more accuracy. The source code will be provided in supplementary materials. Yancun Yang, Weiqing Min, Jingru Song, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Toward Egocentric Compositional Action Anticipation with Adaptive Semantic DebiasingabstractPredicting the unknown from the first-person perspective is expected as a necessary step toward machine intelligence, which is essential for practical applications including autonomous driving and robotics. As a human-level task, egocentric action anticipation aims at predicting an unknown action seconds before it is performed from the first-person viewpoint. Egocentric actions are usually provided as verb-noun pairs; however, predicting the unknown action may be trapped in insufficient training data for all possible combinations. Therefore, it is crucial for intelligent systems to use limited known verb-noun pairs to predict new combinations of actions that have never appeared, which is known as compositional generalization. In this article, we are the first to explore the egocentric compositional action anticipation problem, which is more in line with real-world settings but neglected by existing studies. Whereas prediction results are prone to suffer from semantic bias considering the distinct difference between training and test distributions, we further introduce a general and flexible adaptive semantic debiasing framework that is compatible with different deep neural networks. To capture and mitigate semantic bias, we can imagine one counterfactual situation where no visual representations have been observed and only semantic patterns of observation are used to predict the next action. Instead of the traditional counterfactual analysis scheme that reduces semantic bias in a mindless way, we devise a novel counterfactual analysis scheme to adaptively amplify or penalize the effect of semantic experience by considering the discrepancy both among categories and among examples. We also demonstrate that the traditional counterfactual analysis scheme is a special case of the devised adaptive counterfactual analysis scheme. We conduct experiments on three large-scale egocentric video datasets. Experimental results verify the superiority and effectiveness of our proposed solution. Weiqing Min, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | KERM: Knowledge Enhanced Reasoning for Vision-and-Language NavigationabstractVision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the previous approaches utilize the entire features or object-centric features to represent navigable candidates. However, these representations are not efficient enough for an agent to perform actions to arrive the target location. As knowledge provides crucial information which is complementary to visible content, in this paper, we propose a Knowledge Enhanced Reasoning Model (KERM) to leverage knowledge to improve agent navigation ability. Specifically, we first retrieve facts (i.e., knowledge described by language descriptions) for the navigation views based on local regions from the constructed knowledge base. The re-trieved facts range from properties of a single object (e.g., color, shape) to relationships between objects (e.g., action, spatial position), providing crucial information for VLN. We further present the KERM which contains the purification, fact-aware interaction, and instruction-guided aggregation modules to integrate visual, history, instruction, and fact features. The proposed KERM can automatically select and gather crucial and relevant cues, obtaining more accurate action prediction. Experimental results on the REVERIE, R2R, and SOON datasets demonstrate the effectiveness of the proposed method. The source code is available at https://github.com/XiangyangLi20/KERM. Xiangyang Li 0002, Yaowei Wang 0001, Shuqiang Jiang |
CVPR | 5 |
| 2023 | Bi-Level Meta-Learning for Few-Shot Domain GeneralizationabstractThe goal of few-shot learning is to learn the generalization from seen to unseen data with only a few samples. Most previous few-shot learning methods focus on learning the generalization within particular domains. However, the more practical scenarios may also require the generalization ability across domains. In this paper, we study the problem of few-shot domain generalization (FSDG), which is a more challenging variant of few-shot classification. FSDG requires additional generalization with larger gap from seen domains to unseen domains. We address FSDG problem by meta-learning two levels of meta-knowledge, where the lower-level meta-knowledge is domain-specific embedding spaces as subspaces of a base space for intra-domain generalization, and the upper-level meta-knowledge is the base space and a prior subspace over domain-specific spaces for inter-domain generalization. We formulate the two levels of meta-knowledge learning problem with bi-level optimization, and further develop an optimization algorithm without higher-order derivative information to solve it. We demonstrate our method is significantly superior to the previous works by evaluating it on the widely used benchmark Meta-Dataset. Xiaorong Qin, Xinhang Song, Shuqiang Jiang |
CVPR | 3 |
| 2023 | Data-Free Knowledge Distillation via Feature Exchange and Activation Region ConstraintabstractDespite the tremendous progress on data-free knowledge distillation (DFKD) based on synthetic data generation, there are still limitations in diverse and efficient data synthesis. It is naive to expect that a simple combination of generative network-based data synthesis and data augmentation will solve these issues. Therefore, this paper proposes a novel data-free knowledge distillation method (Spaceship-Net) based on channel-wise feature exchange (CFE) and multi-scale spatial activation region consistency (mSARC) constraint. Specifically, CFE allows our generative network to better sample from the feature space and efficiently synthesize diverse images for learning the student network. However, using CFE alone can severely amplify the unwanted noises in the synthesized images, which may result in failure to improve distillation learning and even have negative effects. Therefore, we propose mSARC to assure the student network can imitate not only the logit output but also the spatial activation region of the teacher network in order to alleviate the influence of unwanted noises in diverse synthetic images on distillation learning. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, Imagenette, and ImageNet100 show that our method can work well with different backbone networks, and outperform the state-of-the-art DFKD methods. Code will be available at: https://github.com/skgyu/Spaceship-Net. Shikang Yu, Hu Han 0001, Shuqiang Jiang |
CVPR | 4 |
| 2023 | Layout-based Causal Inference for Object NavigationabstractPrevious works for ObjectNav task attempt to learn the association (e.g. relation graph) between the visual inputs and the goal during training. Such association contains the prior knowledge of navigating in training environments, which is denoted as the experience. The experience performs a positive effect on helping the agent infer the likely location of the goal when the layout gap between the unseen environments of the test and the prior knowledge obtained in training is minor. However, when the layout gap is significant, the experience exerts a negative effect on navigation. Motivated by keeping the positive effect and removing the negative effect of the experience, we propose the layout-based soft Total Direct Effect (L-sTDE) framework based on the causal inference to adjust the prediction of the navigation policy. In particular, we propose to calculate the layout gap which is defined as the KL divergence between the posterior and the prior distribution of the object layout. Then the sTDE is proposed to appropriately control the effect of the experience based on the layout gap. Experimental results on AI2THOR, RoboTHOR, and Habitat demonstrate the effectiveness of our method. The code is available at https://github.com/sx-zhang/Layout-based-sTDE.git. Sixian Zhang, Xinhang Song, Yubing Bai, Xinyao Yu 0002, Shuqiang Jiang |
CVPR | 6 |
| 2023 | GridMM: Grid Memory Map for Vision-and-Language NavigationabstractVision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. To represent the previously visited environment, most approaches for VLN implement memory using recurrent states, topological maps, or top-down semantic maps. In contrast to these approaches, we build the top-down egocentric and dynamically growing Grid Memory Map (i.e., GridMM) to structure the visited environment. From a global perspective, historical observations are projected into a unified grid map in a top-down view, which can better represent the spatial relations of the environment. From a local perspective, we further propose an instruction relevance aggregation method to capture fine-grained visual clues in each grid region. Extensive experiments are conducted on both the REVERIE, R2R, SOON datasets in the discrete environments, and the R2R-CE dataset in the continuous environments, showing the superiority of our proposed method. The source code is available at https://github.com/MrZihan/GridMM. Xiangyang Li 0002, Yeqi Liu, Shuqiang Jiang |
ICCV | 5 |
| 2023 | A Cross-direction Task Decoupling Network for Small Logo DetectionabstractLogo detection plays an integral role in many applications. However, handling small logos is still difficult since they occupy too few pixels in the image, which burdens the extraction of discriminative features. The aggregation of small logos also brings a great challenge to the classification and localization of logos. To solve these problems, we creatively propose Cross-direction Task Decoupling Network (CTDNet) for small logo detection. We first introduce Cross-direction Feature Pyramid (CFP) to realize cross-direction feature fusion by adopting horizontal transmission and vertical transmission. In addition, Multi-frequency Task Decoupling Head (MTDH) decouples the classification and localization tasks into two branches. A multi-frequency attention convolution branch is designed to achieve more accurate regression by combining discrete cosine transform and convolution creatively. Comprehensive experiments on four logo datasets demonstrate the effectiveness and efficiency of the proposed method. Sujuan Hou, Xingzhuo Li, Weiqing Min, Jing Wang 0138, Yuanjie Zheng, Shuqiang Jiang |
ICME | 7 |
| 2023 | Long-Short Term Policy for Visual Object NavigationabstractThe goal of visual object navigation for an agent is to find the target objects accurately. Recent works mainly focus on the feature of embedding, attempting to learn better features with different variants, such as object distribution and graph representations. However, some typical navigation problems in complex environments, such as partially known and obstacle problems, may not be effectively addressed by previous feature embedding methods. In this paper, we propose a framework with a long-short objective policy, where the hidden states are classified according to the navigation objectives at that moment and separately rewarded. Specifically, we consider two objectives: the long-term objective is to go closer to the target, and the short-term objective is for obstacle avoidance and exploration. To alleviate the effect of long-term and short-term alternation, we build a state memory and propose an adjustment gate to update the state memory. Finally, all past hidden states are reweighted and combined for action prediction with an action-boosting gate. Experimental results on RoboTHOR show that the proposed method can significantly outperform the state-of-the-art. Yubing Bai, Xinhang Song, Sixian Zhang, Shuqiang Jiang |
IROS | 5 |
| 2023 | Generating Explanations for Embodied Action Decision from Visual ObservationabstractGetting trust is crucial for embodied agents (such as robots and autonomous vehicles) to collaborate with human beings, especially non-experts. The most direct way for mutual understanding is through natural language explanation. Existing researches consider generating visual explanations for object recognition, while the exploration of explaining embodied decisions remains vacant. In this paper, we study generating action decisions and explanations based on visual observation. Distinct to explanations for recognition, justifying an action needs to show why it's better than other actions. Besides, the understanding of scene structure is required since the agent needs to interact with the environment (e.g. navigation, moving objects). We introduce a new dataset THOR-EAE (Embodied Action Explanation) collected based on AI2-THOR simulator. The dataset consists of over 840,000 egocentric images of indoor embodied observation which are annotated with the optimal action labels and explanation sentences. An explainable decision-making criterion is developed considering scene layout and action attributes for efficient annotation. We propose a graph action justification model, exploiting graph neural networks for obstacle-surroundings relations representation and justifying the actions under the guidance of decision results. Experimental results on THOR-EAE dataset showcase its challenge and the effectiveness of the proposed method. Yuehu Liu, Xinhang Song, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2023 | SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food DetectionabstractFood detection is becoming a fundamental task in food computing that supports various multimedia applications, including food recommendation and dietary monitoring. To deal with real-world scenarios, food detection needs to localize and recognize novel food objects that are not seen during training, demanding Zero-Shot Detection (ZSD). However, the complexity of semantic attributes and intra-class feature diversity poses challenges for ZSD methods in distinguishing fine-grained food classes. To tackle this, we propose the Semantic Separable Diffusion Synthesizer (SeeDS) framework for Zero-Shot Food Detection (ZSFD). SeeDS consists of two modules: a Semantic Separable Synthesizing Module (S3M) and a Region Feature Denoising Diffusion Model (RFDDM). The S3M learns the disentangled semantic representation for complex food attributes from ingredients and cuisines, and synthesizes discriminative food features via enhanced semantic information. The RFDDM utilizes a novel diffusion model to generate diversified region features and enhances ZSFD via fine-grained synthesized features. Extensive experiments show the state-of-the-art ZSFD performance of our proposed method on two food datasets, ZSFooD and UECFOOD-256. Moreover, SeeDS also maintains effectiveness on general ZSD datasets, PASCAL VOC and MS COCO. The code and dataset can be found at https://github.com/LanceZPF/SeeDS https://github.com/LanceZPF/SeeDS. Weiqing Min, Yang Zhang 0117, Jiajun Song, Shuqiang Jiang |
ACM Multimedia | 6 |
| 2023 | CaMP: Causal Multi-policy Planning for Interactive Navigation in Multi-room ScenesabstractVisual navigation has been widely studied under the assumption that there may be several clear routes to reach the goal. However, in more practical scenarios such as a house with several messy rooms, there may not. Interactive Navigation (InterNav) considers agents navigating to their goals more effectively with object interactions, posing new challenges of learning interaction dynamics and extra action space. Previous works learn single vision-to-action policy with the guidance of designed representations. However, the causality between actions and outcomes is prone to be confounded when the attributes of obstacles are diverse and hard to measure. Learning policy for long-term action planning in complex scenes also leads to extensive inefficient exploration. In this paper, we introduce a causal diagram of InterNav clarifying the confounding bias caused by obstacles. To address the problem, we propose a multi-policy model that enables the exploration of counterfactual interactions as well as reduces unnecessary exploration. We develop a large-scale dataset containing 600k task episodes in 12k multi-room scenes based on the ProcTHOR simulator and showcase the effectiveness of our method with the evaluations on our dataset. Yuehu Liu, Xinhang Song, Shuqiang Jiang |
NeurIPS | 5 |
| 2023 | Dataset Bias in Few-Shot Image RecognitionabstractThe goal of few-shot image recognition (FSIR) is to identify novel categories with a small number of annotated samples by exploiting transferable knowledge from training data (base categories). Most current studies assume that the transferable knowledge can be well used to identify novel categories. However, such transferable capability may be impacted by the dataset bias, and this problem has rarely been investigated before. Besides, most of few-shot learning methods are biased to different datasets, which is also an important issue that needs to be investigated deeply. In this paper, we first investigate the impact of transferable capabilities learned from base categories. Specifically, we use the relevance to measure relationships between base categories and novel categories. Distributions of base categories are depicted via the instance density and category diversity. The FSIR model learns better transferable knowledge from relevant training data. In the relevant data, dense instances or diverse categories can further enrich the learned knowledge. Experimental results on different sub-datasets of Imagenet demonstrate category relevance, instance density and category diversity can depict transferable bias from distributions of base categories. Second, we investigate performance differences on different datasets from the aspects of dataset structures and different few-shot learning methods. Specifically, we introduce image complexity, intra-concept visual consistency, and inter-concept visual similarity to quantify characteristics of dataset structures. We use these quantitative characteristics and eight few-shot learning methods to analyze performance differences on multiple datasets. Based on the experimental analysis, some insightful observations are obtained from the perspective of both dataset structures and few-shot learning methods. We hope these observations are useful to guide future few-shot learning research on new datasets or tasks. Our data is available at http://123.57.42.89/dataset-bias/dataset-bias.html. Shuqiang Jiang, Chenlong Liu, Xinhang Song, Xiangyang Li 0002, Weiqing Min |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Large Scale Visual Food RecognitionabstractFood recognition plays an important role in food choice and intake, which is essential to the health and well-being of humans. It is thus of importance to the computer vision community, and can further support many food-oriented vision and multimodal tasks, e.g., food detection and segmentation, cross-modal recipe retrieval and generation. Unfortunately, we have witnessed remarkable advancements in generic visual recognition for released large-scale datasets, yet largely lags in the food domain. In this paper, we introduce Food2K, which is the largest food recognition dataset with 2,000 categories and over 1 million images. Compared with existing food recognition datasets, Food2K bypasses them in both categories and images by one order of magnitude, and thus establishes a new challenging benchmark to develop advanced models for food visual representation learning. Furthermore, we propose a deep progressive region enhancement network for food recognition, which mainly consists of two components, namely progressive local feature learning and region feature enhancement. The former adopts improved progressive training to learn diverse and complementary local features, while the latter utilizes self-attention to incorporate richer context with multiple scales into local features for further local feature enhancement. Extensive experiments on Food2K demonstrate the effectiveness of our proposed method. More importantly, we have verified better generalization ability of Food2K in various tasks, including food image recognition, food image retrieval, cross-modal recipe retrieval, food detection and segmentation. Food2K can be further explored to benefit more food-relevant tasks including emerging and more complex ones (e.g., nutritional understanding of food), and the trained models on Food2K can be expected as backbones to improve the performance of more food-relevant tasks. We also hope Food2K can serve as a large scale fine-grained visual recognition benchmark, and contributes to the development of large scale fine-grained visual analysis. Weiqing Min, Yuxin Liu 0009, Mengjiang Luo, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | TransWeaver: Weave Image Pairs for Class Agnostic Common Object DetectionabstractMeasuring the similarity of two images is of crucial importance in computer vision. Class agnostic common object detection is a nascent research topic about mining image similarity, which aims to detect common object pairs from two images without category information. This task is general and less restrictive which explores the similarity between objects and can further describe the commonality of image pairs at the object level. However, previous works suffer from features with low discrimination caused by the lack of category information. Moreover, most existing methods compare objects extracted from two images in a simple and direct way, ignoring the internal relationships between objects in the two images. To overcome these limitations, in this paper, we propose a new framework called TransWeaver, which learns intrinsic relationships between objects. Our TransWeaver takes image pairs as input and flexibly captures the inherent correlation between candidate objects from two images. It consists of two modules (i.e., the representation-encoder and the weave-decoder) and captures efficient context information by weaving image pairs to make them interact with each other. The representation-encoder is used for representation learning, which can obtain more discriminative representations for candidate proposals. Furthermore, the weave-decoder weaves the objects from two images and is able to explore the inter-image and intra-image context information at the same time, bringing a better object matching ability. We reorganize the PASCAL VOC, COCO, and Visual Genome datasets to obtain training and testing image pairs. Extensive experiments demonstrate the effectiveness of the proposed TransWeaver which achieves state-of-the-art performance on all datasets. Xiaoqian Guo, Xiangyang Li 0002, Yaowei Wang 0001, Shuqiang Jiang |
IEEE Trans. Image Process. | 4 |
| 2023 | Ingredient Prediction via Context Learning Network With Class-Adaptive Asymmetric LossabstractIngredient prediction has received more and more attention with the help of image processing for its diverse real-world applications, such as nutrition intake management and cafeteria self-checkout system. Existing approaches mainly focus on multi-task food category-ingredient joint learning to improve final recognition by introducing task relevance, while seldom pay attention to making good use of inherent characteristics of ingredients independently. Actually, there are two issues for ingredient prediction. First, compared with fine-grained food recognition, ingredient prediction needs to extract more comprehensive features of the same ingredient and more detailed features of various ingredients from different regions of the food image. Because it can help understand various food compositions and distinguish the differences within ingredient features. Second, the ingredient distributions are extremely unbalanced. Existing loss functions can not simultaneously solve the imbalance between positive-negative samples belonging to each ingredient and significant differences among all classes. To solve these problems, we propose a novel framework named Class-Adaptive Context Learning Network (CACLNet) for ingredient prediction. In order to extract more comprehensive and detailed features, we introduce Ingredient Context Learning (ICL) to reduce the negative impact of complex background in food images and construct internal spatial connections among ingredient regions of food objects in a self-supervised manner, which can strengthen the contacts of the same ingredients through region interactions. In order to solve the imbalance of different classes among ingredients, we propose one novel Class-Adaptive Asymmetric Loss (CAAL) to focus on various ingredient classes adaptively. Besides, considering that the over-suppression of negative samples will over-fit positive samples of those rare ingredients, CAAL alleviates this continuous suppression according to the imbalanced ratios based on gradients while maintaining the contribution of positive samples by lesser suppression. Extensive evaluation on two popular benchmark datasets (Vireo Food-172, UEC Food-100) demonstrates our proposed method achieves the state-of-the-art performance. Further qualitative analysis and visualization show the effectiveness of our method. Code and models are available at https://123.57.42.89/codes/CACLNet/index.html. Mengjiang Luo, Weiqing Min, Jiajun Song, Shuqiang Jiang |
IEEE Trans. Image Process. | 5 |
| 2023 | Composite Object Relation Modeling for Few-Shot Scene RecognitionabstractThe goal of few-shot image recognition is to classify different categories with only one or a few training samples. Previous works of few-shot learning mainly focus on simple images, such as object or character images. Those works usually use a convolutional neural network (CNN) to learn the global image representations from training tasks, which are then adapted to novel tasks. However, there are many more abstract and complex images in real world, such as scene images, consisting of many object entities with flexible spatial relations among them. In such cases, global features can hardly obtain satisfactory generalization ability due to the large diversity of object relations in the scenes, which may hinder the adaptability to novel scenes. This paper proposes a composite object relation modeling method for few-shot scene recognition, capturing the spatial structural characteristic of scene images to enhance adaptability on novel scenes, considering that objects commonly co- occurred in different scenes. In different few-shot scene recognition tasks, the objects in the same images usually play different roles. Thus we propose a task-aware region selection module (TRSM) to further select the detected regions in different few-shot tasks. In addition to detecting object regions, we mainly focus on exploiting the relations between objects, which are more consistent to the scenes and can be used to cleave apart different scenes. Objects and relations are used to construct a graph in each image, which is then modeled with graph convolutional neural network. The graph modeling is jointly optimized with few-shot recognition, where the loss of few-shot learning is also capable of adjusting graph based representations. Typically, the proposed graph based representations can be plugged in different types of few-shot architectures, such as metric-based and meta-learning methods. Experimental results of few-shot scene recognition show the effectiveness of the proposed method. Xinhang Song, Chenlong Liu, Haitao Zeng, Gongwei Chen, Xiaorong Qin, Shuqiang Jiang |
IEEE Trans. Image Process. | 7 |
| 2023 | MemBridge: Video-Language Pre-Training With Memory-Augmented Inter-Modality BridgeabstractVideo-language pre-training has attracted considerable attention recently for its promising performance on various downstream tasks. Most existing methods utilize the modality-specific or modality-joint representation architectures for the cross-modality pre-training. Different from previous methods, this paper presents a novel architecture named Memory-augmented Inter-Modality Bridge (MemBridge), which uses the learnable intermediate modality representations as the bridge for the interaction between videos and language. Specifically, in the transformer-based cross-modality encoder, we introduce the learnable bridge tokens as the interaction approach, which means the video and language tokens can only perceive information from bridge tokens and themselves. Moreover, a memory bank is proposed to store abundant modality interaction information for adaptively generating bridge tokens according to different cases, enhancing the capacity and robustness of the inter-modality bridge. Through pre-training, MemBridge explicitly models the representations for more sufficient inter-modality interaction. Comprehensive experiments show that our approach achieves competitive performance with previous methods on various downstream tasks including video-text retrieval, video captioning, and video question answering on multiple datasets, demonstrating the effectiveness of the proposed method. The code has been available at https://github.com/jahhaoyang/MemBridge. Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Yuchen Yuan, Zifeng Chai, Shuqiang Jiang |
IEEE Trans. Image Process. | 9 |
| 2023 | Multi-Object Navigation Using Potential Target Position Policy FunctionabstractVisual object navigation is an essential task of embodied AI, which is letting the agent navigate to the goal object under the user's demand. Previous methods often focus on single-object navigation. However, in real life, human demands are generally continuous and multiple, requiring the agent to implement multiple tasks in sequence. These demands can be addressed by repeatedly performing previous single task methods. However, by dividing multiple tasks into several independent tasks to perform, without the global optimization between different tasks, the agents' trajectories may overlap, reducing the efficiency of navigation. In this paper, we propose an efficient reinforcement learning framework with a hybrid policy for multi-object navigation, aiming to maximally eliminate noneffective actions. First, the visual observations are embedded to detect the semantic entities (such as objects). And the detected objects are memorized and projected into semantic maps, which can also be regarded as a long-term memory of the observed environment. Then a hybrid policy consisting of exploration and long-term planning strategies is proposed to predict the potential target position. In particular, when the target is directly oriented, the policy function makes long-term planning for the target based on the semantic map, which is implemented by a sequence of motion actions. In the alternative, when the target is not oriented, the policy function estimates an object's potential position toward exploring the most possible objects (positions) that have close relations to the target. The relation between different objects is obtained with prior knowledge, which is used to predict the potential target position by integrating with the memorized semantic map. And then a path to the potential target is planned by the policy function. We evaluate our proposed method on two large-scale 3D realistic environment datasets, Gibson and Matterport3D, and the experimental results demonstrate the effectiveness and generalization of the proposed method. Haitao Zeng, Xinhang Song, Shuqiang Jiang |
IEEE Trans. Image Process. | 3 |
| 2023 | Focus and Align: Learning Tube Tokens for Video-Language Pre-TrainingabstractVideo-language pre-training (VLP) has attracted increasing attention for cross-modality understanding tasks. To enhance visual representations, recent works attempt to adopt transformer-based architectures as video encoders. These works usually focus on the visual representations of the sampled frames. Compared with frame representations, frame patches incorporate more fine-grained spatio-temporal information, which could lead to a better understanding of video contents. However, how to exploit the spatio-temporal information within frame patches for VLP has been less investigated. In this work, we propose a method to learn tube tokens to model the key spatio-temporal information from frame patches. To this end, multiple semantic centers are introduced to focus on the underlying patterns of frame patches. Based on each semantic center, the spatio-temporal information within frame patches is integrated into a unique tube token. Complementary to frame representations, tube tokens provide detailed clues of video contents. Furthermore, to better align the generated tube tokens and the contents of descriptions, a local alignment mechanism is introduced. The experiments based on a variety of downstream tasks demonstrate the effectiveness of the proposed method. Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Zifeng Chai, Yuchen Yuan, Shuqiang Jiang |
IEEE Trans. Multim. | 9 |
| 2022 | Rethinking the Optimization of Average Precision: Only Penalizing Negative Instances before Positive Ones Is EnoughabstractOptimising the approximation of Average Precision (AP) has been widely studied for image retrieval. Limited by the definition of AP, such methods consider both negative and positive instances ranking before each positive instance. However, we claim that only penalizing negative instances before positive ones is enough, because the loss only comes from these negative instances. To this end, we propose a novel loss, namely Penalizing Negative instances before Positive ones (PNP), which can directly minimize the number of negative instances before each positive one. In addition, AP-based methods adopt a fixed and sub-optimal gradient assignment strategy. Therefore, we systematically investigate different gradient assignment solutions via constructing derivative functions of the loss, resulting in PNP-I with increasing derivative functions and PNP-D with decreasing ones. PNP-I focuses more on the hard positive instances by assigning larger gradients to them and tries to make all relevant instances closer. In contrast, PNP-D pays less attention to such instances and slowly corrects them. For most real-world data, one class usually contains several local clusters. PNP-I blindly gathers these clusters while PNP-D keeps them as they were. Therefore, PNP-D is more superior. Experiments on three standard retrieval datasets show consistent results with the above analysis. Extensive evaluations demonstrate that PNP-D achieves the state-of-the-art performance. Code is available at https://github.com/interestingzhuo/PNPloss Weiqing Min, Jiajun Song, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
AAAI | 8 |
| 2022 | Generative Meta-Adversarial Network for Unseen Object Navigation
Sixian Zhang, Xinhang Song, Yubing Bai, Shuqiang Jiang |
ECCV (39) | 5 |
| 2022 | Ingredient-Guided Region Discovery and Relationship Modeling for Food Category-Ingredient PredictionabstractRecognizing the category and its ingredient composition from food images facilitates automatic nutrition estimation, which is crucial to various health relevant applications, such as nutrition intake management and healthy diet recommendation. Since food is composed of ingredients, discovering ingredient-relevant visual regions can help identify its corresponding category and ingredients. Furthermore, various ingredient relationships like co-occurrence and exclusion are also critical for this task. For that, we propose an ingredient-oriented multi-task food category-ingredient joint learning framework for simultaneous food recognition and ingredient prediction. This framework mainly involves learning an ingredient dictionary for ingredient-relevant visual region discovery and building an ingredient-based semantic-visual graph for ingredient relationship modeling. To obtain ingredient-relevant visual regions, we build an ingredient dictionary to capture multiple ingredient regions and obtain the corresponding assignment map, and then pool the region features belonging to the same ingredient to identify the ingredients more accurately and meanwhile improve the classification performance. For ingredient-relationship modeling, we utilize the visual ingredient representations as nodes and the semantic similarity between ingredient embeddings as edges to construct an ingredient graph, and then learn their relationships via the graph convolutional network to make label embeddings and visual features interact with each other to improve the performance. Finally, fused features from both ingredient-oriented region features and ingredient-relationship features are used in the following multi-task category-ingredient joint learning. Extensive evaluation on three popular benchmark datasets (ETH Food-101, Vireo Food-172 and ISIA Food-200) demonstrates the effectiveness of our method. Further visualization of ingredient assignment maps and attention maps also shows the superiority of our method. Weiqing Min, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
IEEE Trans. Image Process. | 7 |
| 2022 | Amorphous Region Context Modeling for Scene RecognitionabstractScene images are usually composed of foreground and background regional contents. Some existing methods propose to extract regional contents with dense grids or objectness region proposals. However, dense grids may split the object into several discrete parts, learning semantic ambiguity for the patches. The objectness methods may focus on particular objects but only pay attention to the foreground contents and do not exploit the background that is key to scene recognition. In contrast, we propose a novel scene recognition framework with amorphous region detection and context modeling. In the proposed framework, discriminative regions are first detected with amorphous contours that can tightly surround the targets through semantic segmentation techniques. In addition, both foreground and background regions are jointly embedded to obtain the scene representations with the graph model. Based on the graph modeling module, we explore the contextual relations between the regions in geometric and morphology aspects, and generate the discriminative representations for scene recognition. Experimental results on MIT67 and SUN397 demonstrate the effectiveness and generality of the proposed method. Haitao Zeng, Xinhang Song, Gongwei Chen, Shuqiang Jiang |
IEEE Trans. Multim. | 4 |
| 2022 | LogoDet-3K: A Large-scale Image Dataset for Logo DetectionabstractLogo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media. In this article, we introduce LogoDet-3K, the largest logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects, and 158,652 images. LogoDet-3K creates a more challenging benchmark for logo detection, for its higher comprehensive coverage and wider variety in both logo categories and annotated objects compared with existing datasets. We describe the collection and annotation process of our dataset and analyze its scale and diversity in comparison to other datasets for logo detection. We further propose a strong baseline method Logo-Yolo, which incorporates Focal loss and CIoU loss into the basic YOLOv3 framework for large-scale logo detection. It obtains about 4% improvement on the average performance compared with YOLOv3, and greater improvements compared with reported several deep detection models on LogoDet-3K. We perform extensive evaluation on three other existing datasets to further verify on both logo detection and retrieval tasks, and we demonstrate better generalization ability of LogoDet-3K on logo detection and retrieval tasks. The LogoDet-3K dataset is used to promote large-scale logo-related research. The code and LogoDet-3K can be found at https://github.com/Wangjing1551/LogoDet-3K-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Hierarchical Object-to-Zone Graph for Object NavigationabstractThe goal of object navigation is to reach the expected objects according to visual information in the unseen environments. Previous works usually implement deep models to train an agent to predict actions in real-time. However, in the unseen environment, when the target object is not in egocentric view, the agent may not be able to make wise decisions due to the lack of guidance. In this paper, we propose a hierarchical object-to-zone (HOZ) graph to guide the agent in a coarse-to-fine manner, and an online-learning mechanism is also proposed to update HOZ according to the real-time observation in new environments. In particular, the HOZ graph is composed of scene nodes, zone nodes and object nodes. With the pre-learned HOZ graph, the real-time observation and the target goal, the agent can constantly plan an optimal path from zone to zone. In the estimated path, the next potential zone is regarded as sub-goal, which is also fed into the deep reinforcement learning model for action prediction. Our methods are evaluated on the AI2-Thor simulator. In addition to widely used evaluation metrics SR and SPL, we also propose a new evaluation metric of SAE that focuses on the effective action rate. Experimental results demonstrate the effectiveness and efficiency of our proposed method. The code is available at https://github.com/sx-zhang/HOZ.git. Sixian Zhang, Xinhang Song, Yubing Bai, Yakui Chu, Shuqiang Jiang |
ICCV | 6 |
| 2021 | What If We Could Not See? Counterfactual Analysis for Egocentric Action AnticipationabstractEgocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacity. In the factual case, we can predict the upcoming action based on visual features and semantic labels from past observation. Imagining one counterfactual situation where no visual representation had been observed, we would obtain a counterfactual predicted action only using past semantic labels. In this way, we can reduce the side-effect caused by semantic labels via a comparison between factual and counterfactual outcomes, which moves a step towards unbiased prediction for egocentric action anticipation. We conduct experiments on two large-scale egocentric video datasets. Qualitative and quantitative results validate the effectiveness of our proposed CA-EAA. Weiqing Min, Shuqiang Jiang, Yong Rui |
IJCAI | 5 |
| 2021 | AIxFood'21: 3rd Workshop on AIxFoodabstractFood and cooking analysis present exciting research and application challenges for modern AI systems, particularly in the context of multimodal data such as images or video. A meal that appears in a food image is a product of a complex progression of cooking stages, often described in the accompanying textual recipe form. In the cooking process, individual ingredients change their physical properties, become combined with other food components, all to produce a final, yet highly variable, appearance of the meal. Recognizing food items or meals on a plate from images or videos, their physical properties such as the amount, nutritional content such as the caloric value, food attributes such as the flavor, elucidating the cooking process behind it, or creating robotic assistants that help users complete that cooking process, is of essential scientific and technological value yet technically extremely challenging. The 3rd AIxFood workshop was held as a half-day workshop in conjunction with the 29th ACM International Conference on Multimedia (ACM MM 2021), in Chengdu, China and virtually. Ricardo Guerrero, Michael Spranger, Shuqiang Jiang, Chong-Wah Ngo |
ACM Multimedia | 3 |
| 2021 | FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkabstractFood logo detection plays an important role in the multimedia for its wide real-world applications, such as food recommendation of the self-service shop and infringement detection on e-commerce platforms. A large-scale food logo dataset is urgently needed for developing advanced food logo detection algorithms. However, there are no available food logo datasets with food brand information. To support efforts towards food logo detection, we introduce the dataset FoodLogoDet-1500, a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects. We describe the collection and annotation process of FoodLogoDet-1500, analyze its scale and diversity, and compare it with other logo datasets. To the best of our knowledge, FoodLogoDet-1500 is the first largest publicly available high-quality dataset for food logo detection. The challenge of food logo detection lies in the large-scale categories and similarities between food logo categories. For that, we propose a novel food logo detection method Multi-scale Feature Decoupling Network (MFDNet), which decouples classification and regression into two branches and focuses on the classification branch to solve the problem of distinguishing multiple food logo categories. Specifically, we introduce the feature offset module, which utilizes the deformation-learning for optimal classification offset and can effectively obtain the most representative features of classification in detection. In addition, we adopt a balanced feature pyramid in MFDNet, which pays attention to global information, balances the multi-scale feature maps, and enhances feature extraction capability. Comprehensive experiments on FoodLogoDet-1500 and other two popular benchmark logo datasets demonstrate the effectiveness of the proposed method. The code and FoodLogoDet-1500 can be found at https://github.com/hq03/FoodLogoDet-1500-Dataset. Weiqing Min, Jing Wang 0138, Sujuan Hou, Yuanjie Zheng, Shuqiang Jiang |
ACM Multimedia | 6 |
| 2021 | ION: Instance-level Object NavigationabstractVisual object navigation is a fundamental task in Embodied AI. Previous works focus on the category-wise navigation, in which navigating to any possible instance of target object category is considered a success. Those methods may be effective to find the general objects. However, it may be more practical to navigate to the specific instance in our real life, since our particular requirements are usually satisfied with specific instances rather than all instances of one category. How to navigate to the specific instance has been rarely researched before and is typically challenging to current works. In this paper, we introduce a new task of Instance Object Navigation (ION), where instance-level descriptions of targets are provided and instance-level navigation is required. In particular, multiple types of attributes such as colors, materials and object references are involved in the instance-level descriptions of the targets. In order to allow the agent to maintain the ability of instance navigation, we propose a cascade framework with Instance-Relation Graph (IRG) based navigator and instance grounding module. To specify the different instances of the same object categories, we construct instance-level graph instead of category-level one, where instances are regarded as nodes, encoded with the representation of colors, materials and locations (bounding boxes). During navigation, the detected instances can activate corresponding nodes in IRG, which are updated with graph convolutional neural network (GCNN). The final instance prediction is obtained with the grounding module by selecting the candidates (instances) with maximum probability (a joint probability of category, color and material, obtained by corresponding regressors with softmax). For the task evaluation, we build a benchmark for instance-level object navigation on AI2-Thor simulator, where over 27,735 object instance descriptions and navigation groundtruth are automatically obtained through the interaction with the simulator. The proposed model outperforms the baseline in instance-level metrics, showing that our proposed graph model can guide instance object navigation, as well as leaving promising room for further improvement. The project is available at https://github.com/LWJ312/ION. Xinhang Song, Yubing Bai, Sixian Zhang, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2021 | See More for Scene: Pairwise Consistency Learning for Scene ClassificationabstractScene classification is a valuable classification subtask and has its own characteristics which still needs more in-depth studies. Basically, scene characteristics are distributed over the whole image, which cause the need of “seeing” comprehensive and informative regions. Previous works mainly focus on region discovery and aggregation, while rarely involves the inherent properties of CNN along with its potential ability to satisfy the requirements of scene classification. In this paper, we propose to understand scene images and the scene classification CNN models in terms of the focus area. From this new perspective, we find that large focus area is preferred in scene classification CNN models as a consequence of learning scene characteristics. Meanwhile, the analysis about existing training schemes helps us to understand the effects of focus area, and also raises the question about optimal training method for scene classification. Pursuing the better usage of scene characteristics, we propose a new learning scheme with a tailored loss in the goal of activating larger focus area on scene images. Since the supervision of the target regions to be enlarged is usually lacked, our alternative learning scheme is to erase already activated area, and allow the CNN models to activate more area during training. The proposed scheme is implemented by keeping the pairwise consistency between the output of the erased image and its original one. In particular, a tailored loss is proposed to keep such pairwise consistency by leveraging category-relevance information. Experiments on Places365 show the significant improvements of our method with various CNNs. Our method shows an inferior result on the object-centric dataset, ImageNet, which experimentally indicates that it captures the unique characteristics of scenes. Gongwei Chen, Xinhang Song, Shuqiang Jiang |
NeurIPS | 4 |
| 2021 | Plant Disease Recognition: A Large-Scale Benchmark Dataset and a Visual Region and Loss Reweighting ApproachabstractPlant disease diagnosis is very critical for agriculture due to its importance for increasing crop production. Recent advances in image processing offer us a new way to solve this issue via visual plant disease analysis. However, there are few works in this area, not to mention systematic researches. In this paper, we systematically investigate the problem of visual plant disease recognition for plant disease diagnosis. Compared with other types of images, plant disease images generally exhibit randomly distributed lesions, diverse symptoms and complex backgrounds, and thus are hard to capture discriminative information. To facilitate the plant disease recognition research, we construct a new large-scale plant disease dataset with 271 plant disease categories and 220,592 images. Based on this dataset, we tackle plant disease recognition via reweighting both visual regions and loss to emphasize diseased parts. We first compute the weights of all the divided patches from each image based on the cluster distribution of these patches to indicate the discriminative level of each patch. Then we allocate the weight to each loss for each patch-label pair during weakly-supervised training to enable discriminative disease part learning. We finally extract patch features from the network trained with loss reweighting, and utilize the LSTM network to encode the weighed patch feature sequence into a comprehensive feature representation. Extensive evaluations on this dataset and another public dataset demonstrate the advantage of the proposed method. We expect this research will further the agenda of plant disease recognition in the community of image processing. Xinda Liu, Weiqing Min, Shuhuan Mei, Lili Wang 0006, Shuqiang Jiang |
IEEE Trans. Image Process. | 5 |
| 2021 | Hybrid-Attention Enhanced Two-Stream Fusion Network for Video Venue PredictionabstractVideo venue category prediction has been drawing more attention in the multimedia community for various applications such as personalized location recommendation and video verification. Most of existing works resort to the information from either multiple modalities or other platforms for strengthening video representations. However, noisy acoustic information, sparse textual descriptions and incompatible cross-platform data could limit the performance gain and reduce the universality of the model. Therefore, we focus on discriminative visual feature extraction from videos by introducing a hybrid-attention structure. Particularly, we propose a novel Global-Local Attention Module (GLAM), which can be inserted to neural networks to generate enhanced visual features from video content. In GLAM, the Global Attention (GA) is used to catch contextual scene-oriented information via assigning channels with various weights while the Local Attention (LA) is employed to learn salient object-oriented features via allocating different weights for spatial regions. Moreover, GLAM can be extended to ones with multiple GAs and LAs for further visual enhancement. These two types of features respectively captured by GAs and LAs are integrated via convolution layers, and then delivered into convolutional Long Short-Term Memory (convLSTM) to generate spatial-temporal representations, constituting the content stream. In addition, video motions are explored to learn long-term movement variations, which also contributes to video venue prediction. The content and motion stream constitute our proposed Hybrid-Attention Enhanced Two-Stream Fusion Network (HA-TSFN). HA-TSFN finally merges the features from two streams for comprehensive representations. Extensive experiments demonstrate that our method achieves the state-of-the-art performance in the large-scale dataset Vine. The visualization also shows that the proposed GLAM can capture complementary scene-oriented and object-oriented visual features from videos. Our code is available at:https://github.com/zhangyanchao1014/HA-TSFN. Weiqing Min, Liqiang Nie, Shuqiang Jiang |
IEEE Trans. Multim. | 4 |
| 2021 | Attribute-Guided Feature Learning for Few-Shot Image RecognitionabstractFew-shot image recognition has become an essential problem in the field of machine learning and image recognition, and has attracted more and more research attention. Typically, most few-shot image recognition methods are trained across tasks. However, these methods are apt to learn an embedding network for discriminative representations of training categories, and thus could not distinguish well for novel categories. To establish connections between training and novel categories, we use attribute-related representations for few-shot image recognition and propose an attribute-guided two-layer learning framework, which is capable of learning general feature representations. Specifically, few-shot image recognition trained over tasks and attribute learning trained over images share the same network in a multi-task learning framework. In this way, few-shot image recognition learns feature representations guided by attributes, and is thus less sensitive to novel categories compared with feature representations only using category supervision. Meanwhile, the multi-layer features associated with attributes are aligned with category learning on multiple levels respectively. Therefore we establish a two-layer learning mechanism guided by attributes to capture more discriminative representations, which are complementary compared with a single-layer learning mechanism. Experimental results on CUB-200, AWA and MiniImageNet datasets demonstrate our method effectively improves the performance. Weiqing Min, Shuqiang Jiang |
IEEE Trans. Multim. | 3 |
| 2020 | Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo ClassificationabstractLogo classification has gained increasing attention for its various applications, such as copyright infringement detection, product recommendation and contextual advertising. Compared with other types of object images, the real-world logo images have larger variety in logo appearance and more complexity in their background. Therefore, recognizing the logo from images is challenging. To support efforts towards scalable logo classification task, we have curated a dataset, Logo-2K+, a new large-scale publicly available real-world logo dataset with 2,341 categories and 167,140 images. Compared with existing popular logo datasets, such as FlickrLogos-32 and LOGO-Net, Logo-2K+ has more comprehensive coverage of logo categories and larger quantity of logo images. Moreover, we propose a Discriminative Region Navigation and Augmentation Network (DRNA-Net), which is capable of discovering more informative logo regions and augmenting these image regions for logo classification. DRNA-Net consists of four sub-networks: the navigator sub-network first selected informative logo-relevant regions guided by the teacher sub-network, which can evaluate its confidence belonging to the ground-truth logo class. The data augmentation sub-network then augments the selected regions via both region cropping and region dropping. Finally, the scrutinizer sub-network fuses features from augmented regions and the whole image for logo classification. Comprehensive experiments on Logo-2K+ and other three existing benchmark datasets demonstrate the effectiveness of proposed method. Logo-2K+ and the proposed strong baseline DRNA-Net are expected to further the development of scalable logo image recognition, and the Logo-2K+ dataset can be found at https://github.com/msn199959/Logo-2k-plus-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Haishuai Wang, Shuqiang Jiang |
AAAI | 7 |
| 2020 | Multi-attention Meta Learning for Few-shot Fine-grained Image RecognitionabstractThe goal of few-shot image recognition is to distinguish different categories with only one or a few training samples. Previous works of few-shot learning mainly work on general object images. And current solutions usually learn a global image representation from training tasks to adapt novel tasks. However, fine-gained categories are distinguished by subtle and local parts, which could not be captured by global representations effectively. This may hinder existing few-shot learning approaches from dealing with fine-gained categories well. In this work, we propose a multi-attention meta-learning (MattML) method for few-shot fine-grained image recognition (FSFGIR). Instead of using only base learner for general feature learning, the proposed meta-learning method uses attention mechanisms of the base learner and task learner to capture discriminative parts of images. The base learner is equipped with two convolutional block attention modules (CBAM) and a classifier. The two CBAM can learn diverse and informative parts. And the initial weights of classifier are attended by the task learner, which gives the classifier a task-related sensitive initialization. For adaptation, the gradient-based meta-learning approach is employed by updating the parameters of two CBAM and the attended classifier, which facilitates the updated base learner to adaptively focus on discriminative parts. We experimentally analyze the different components of our method, and experimental results on four benchmark datasets demonstrate the effectiveness and superiority of our method. Chenlong Liu, Shuqiang Jiang |
IJCAI | 3 |
| 2020 | Expressional Region RetrievalabstractImage retrieval is a long-standing topic in the multimedia community due to its various applications, e.g., product search and artworks retrieval in museum. The regions in images contain a wealth of information. Users may be interested in the objects presented in the image regions or the relationships between them. But previous retrieval methods are either limited to the single object of images, or tend to the entire visual scene. In this paper, we introduce a new task called expressional region retrieval, in which the query is formulated as a region of image with the associated description. The goal is to find images containing the similar content with the query and localize the regions within them. As far as we know, this task has not been explored yet. We propose a framework to address this issue. The region proposals are first generated based on region detectors and language features are extracted. Then the Gated Residual Network (GRN) takes language information as a gate to control the transformation of visual features. In this way, the combined visual and language representation is more specific and discriminative for expressional region retrieval. We evaluate our method on a new established benchmark which is constructed based on the Visual Genome dataset. Experimental results demonstrate that our model effectively utilizes both visual and language information, outperforming the baseline methods. Xiaoqian Guo, Xiangyang Li 0002, Shuqiang Jiang |
ACM Multimedia | 3 |
| 2020 | Food Computing for MultimediaabstractFood computing applies computational approaches for acquiring and analyzing heterogeneous food data from disparate sources for perception, recognition, retrieval, recommendation, prediction and monitoring of food to address food-related issues in multimedia and beyond. It has received more attention from both academia and industry as one emerging interdiscipline for its various applications, such as improving human health and understanding the culinary culture. Recently, there are more studies on food computing in the multimedia, such as food recognition and multimodal recipe analysis. This tutorial will provide a basic understanding of food computing, and discuss its use in various multimedia tasks, ranging from food recognition, retrieval, recommendation, recipe analysis to cooking behavior understanding. Specifically, we will first introduce food computing, including its method, task and applications. Then we will discuss several typical tasks of food computing in the multimedia including food image recognition, food retrieval and recommendation, multimodal recipe analysis and cooking action anticipation. Finally, we will point out future research directions on food computing in the multimedia. Shuqiang Jiang, Weiqing Min |
ACM Multimedia | 1 |
| 2020 | ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkabstractFood recognition has received more and more attention in the multimedia community for its various real-world applications, such as diet management and self-service restaurants. A large-scale ontology of food images is urgently needed for developing advanced large-scale food recognition algorithms, as well as for providing the benchmark dataset for such algorithms. To encourage further progress in food recognition, we introduce the dataset ISIA Food-500 with 500 categories from the list in the Wikipedia and 399,726 images, a more comprehensive food dataset that surpasses existing popular benchmark datasets by category coverage and data volume. Furthermore, we propose a stacked global-local attention network, which consists of two sub-networks for food recognition. One sub-network first utilizes hybrid spatial-channel attention to extract more discriminative features, and then aggregates these multi-scale discriminative features from multiple layers into global-level representation (e.g., texture and shape information about food). The other one generates attentional regions (e.g., ingredient relevant regions) from different regions via cascaded spatial transformers, and further aggregates these multi-scale regional features from different layers into local-level representation. These two types of features are finally fused as comprehensive representation for food recognition. Extensive experiments on ISIA Food-500 and other two popular benchmark datasets demonstrate the effectiveness of our proposed method, and thus can be considered as one strong baseline. The dataset, code and models can be found at http://123.57.42.89/FoodComputing-Dataset/ISIA-Food500.html. Weiqing Min, Linhu Liu, Zhengdong Luo, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
ACM Multimedia | 7 |
| 2020 | Generalized Zero-shot Learning with Multi-source Semantic Embeddings for Scene RecognitionabstractRecognizing visual categories from semantic descriptions is a promising way to extend the capability of a visual classifier beyond the concepts represented in the training data (i.e. seen categories). This problem is addressed by (generalized) zero-shot learning methods (GZSL), which leverage semantic descriptions that connect them to seen categories (e.g. label embedding, attributes). Conventional GZSL are designed mostly for object recognition. In this paper we focus on zero-shot scene recognition, a more challenging setting with hundreds of categories where their differences can be subtle and often localized in certain objects or regions. Conventional GZSL representations are not rich enough to capture these local discriminative differences. Addressing these limitations, we propose a feature generation framework with two novel components: 1) multiple sources of semantic information (i.e. attributes, word embeddings and descriptions), 2) region descriptions that can enhance scene discrimination. To generate synthetic visual features we propose a two-step generative approach, where local descriptions are sampled and used as conditions to generate visual features. The generated features are then aggregated and used together with real features to train a joint classifier. In order to evaluate the proposed method, we introduce a new dataset for zero-shot scene recognition with multi-semantic annotations. Experimental results on the proposed dataset and SUN Attribute dataset illustrate the effectiveness of the proposed method. Xinhang Song, Haitao Zeng, Sixian Zhang, Luis Herranz, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2020 | An Egocentric Action Anticipation Framework via Fusing Intuition and AnalysisabstractIn this paper, we focus on egocentric action anticipation from videos, which enables various applications, such as helping intelligent wearable assistants understand users' needs and enhance their capabilities in the interaction process. It requires intelligent systems to observe from the perspective of the first person and predict an action before it occurs. Owing to the uncertainty of future, it is insufficient to perform action anticipation relying on visual information especially when there exists salient visual difference between past and future. In order to alleviate this problem, which we call visual gap in this paper, we propose one novel Intuition-Analysis Integrated (IAI) framework inspired by psychological research, which mainly consists of three parts: Intuition-based Prediction Network (IPN), Analysis-based Prediction Network (APN) and Adaptive Fusion Network (AFN). To imitate the implicit intuitive thinking process, we model IPN as an encoder-decoder structure and introduce one procedural instruction learning strategy implemented by textual pre-training. On the other hand, we allow APN to process information under designed rules to imitate the explicit analytical thinking, which is divided into three steps: recognition, transitions and combination. Both the procedural instruction learning strategy in IPN and the transition step of APN are crucial to improving the anticipation performance via mitigating the visual gap problem. Considering the complementarity of intuition and analysis, AFN adopts attention fusion to adaptively integrate predictions from IPN and APN to produce the final anticipation results. We conduct experiments on the largest egocentric video dataset. Qualitative and quantitative evaluation results validate the effectiveness of our IAI framework, and demonstrate the advantage of bridging visual gap by utilizing multi-modal information, including both visual features of observed segments and sequential instructions of actions. Weiqing Min, Yong Rui, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2020 | Deep neural networks for emerging multimedia computing and applications
Shuqiang Jiang, Weiqing Min, Yonggang Wen 0001, Qingming Huang, Shuicheng Yan |
Neurocomputing | 1 |
| 2020 | Scene Recognition With Prototype-Agnostic Scene LayoutabstractExploiting the spatial structure in scene images is a key research direction for scene recognition. Due to the large intra-class structural diversity, building and modeling flexible structural layout to adapt various image characteristics is a challenge. Existing structural modeling methods in scene recognition either focus on predefined grids or rely on learned prototypes, which all have limited representative ability. In this paper, we propose Prototype-agnostic Scene Layout (PaSL) construction method to build the spatial structure for each image without conforming to any prototype. Our PaSL can flexibly capture the diverse spatial characteristic of scene images and have considerable generalization capability. Given a PaSL, we build Layout Graph Network (LGN) where regions in PaSL are defined as nodes and two kinds of independent relations between regions are encoded as edges. The LGN aims to incorporate two topological structures (formed in spatial and semantic similarity dimensions) into image representations through graph convolution. Extensive experiments show that our approach achieves state-of-the-art results on widely recognized MIT67 and SUN397 datasets without multi-model or multi-scale fusion. Moreover, we also conduct the experiments on one of the largest scale datasets, Places365. The results demonstrate the proposed method can be well generalized and obtains competitive performance. Gongwei Chen, Xinhang Song, Haitao Zeng, Shuqiang Jiang |
IEEE Trans. Image Process. | 4 |
| 2020 | Multi-Scale Multi-View Deep Feature Aggregation for Food RecognitionabstractRecently, food recognition has received more and more attention in image processing and computer vision for its great potential applications in human health. Most of the existing methods directly extracted deep visual features via convolutional neural networks (CNNs) for food recognition. Such methods ignore the characteristics of food images and are, thus, hard to achieve optimal recognition performance. In contrast to general object recognition, food images typically do not exhibit distinctive spatial arrangement and common semantic patterns. In this paper, we propose a multi-scale multi-view feature aggregation (MSMVFA) scheme for food recognition. MSMVFA can aggregate high-level semantic features, mid-level attribute features, and deep visual features into a unified representation. These three types of features describe the food image from different granularity. Therefore, the aggregated features can capture the semantics of food images with the greatest probability. For that solution, we utilize additional ingredient knowledge to obtain mid-level attribute representation via ingredient-supervised CNNs. High-level semantic features and deep visual features are extracted from class-supervised CNNs. Considering food images do not exhibit distinctive spatial layout in many cases, MSMVFA fuses multi-scale CNN activations for each type of features to make aggregated features more discriminative and invariable to geometrical deformation. Finally, the aggregated features are more robust, comprehensive, and discriminative via two-level fusion, namely multi-scale fusion for each type of features and multi-view aggregation for different types of features. In addition, MSMVFA is general and different deep networks can be easily applied into this scheme. Extensive experiments and evaluations demonstrate that our method achieves state-of-the-art recognition performance on three popular large-scale food benchmark datasets in Top-1 recognition accuracy. Furthermore, we expect this paper will further the agenda of food recognition in the community of image processing and computer vision. Shuqiang Jiang, Weiqing Min, Linhu Liu, Zhengdong Luo |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-Task Deep Relative Attribute Learning for Visual Urban PerceptionabstractVisual urban perception aims to quantify perceptual attributes (e.g., safe and depressing attributes) of physical urban environment from crowd-sourced street-view images and their pairwise comparisons. It has been receiving more and more attention in computer vision for various applications, such as perceptive attribute learning and urban scene understanding. Most existing methods adopt either (i) a regression model trained using image features and ranked scores converted from pairwise comparisons for perceptual attribute prediction or (ii) a pairwise ranking algorithm to independently learn each perceptual attribute. However, the former fails to directly exploit pairwise comparisons while the latter ignores the relationship among different attributes. To address them, we propose a Multi-Task Deep Relative Attribute Learning Network (MTDRALN) to learn all the relative attributes simultaneously via multi-task Siamese networks, where each Siamese network will predict one relative attribute. Combined with deep relative attribute learning, we utilize the structured sparsity to exploit the prior from natural attribute grouping, where all the attributes are divided into different groups based on semantic relatedness in advance. As a result, MTDRALN is capable of learning all the perceptual attributes simultaneously via multi-task learning. Besides the ranking sub-network, MTDRALN further introduces the classification sub-network, and these two types of losses from two sub-networks jointly constrain parameters of the deep network to make the network learn more discriminative visual features for relative attribute learning. In addition, our network can be trained in an end-to-end way to make deep feature learning and multi-task relative attribute learning reinforce each other. Extensive experiments on the large-scale Place Pulse 2.0 dataset validate the advantage of our proposed network. Our qualitative results along with visualization of saliency maps also show that the proposed network is able to learn effective features for perceptual attributes. Weiqing Min, Shuhuan Mei, Linhu Liu, Shuqiang Jiang |
IEEE Trans. Image Process. | 5 |
| 2020 | Image Representations With Spatial Object-to-Object Relations for RGB-D Scene RecognitionabstractScene recognition is challenging due to the intra-class diversity and inter-class similarity. Previous works recognize scenes either with global representations or with the intermediate representations of objects. In contrast, we investigate more discriminative image representations of object-to-object relations for scene recognition, which are based on the triplets of obtained with detection techniques. Particularly, two types of representations, including co-occurring frequency of object-to-object relation (denoted as COOR) and sequential representation of object-to-object relation (denoted as SOOR), are proposed to describe objects and their relative relations in different forms. COOR is represented as the intermediate representation of co-occurring frequency of objects and their relations, with a three order tensor that can be fed to scene classifier without further embedding. SOOR is represented in a more explicit and freer form that sequentially describe image contents with local captions. And a sequence encoding model (e.g., recurrent neural network (RNN)) is implemented to encode SOOR to the features for feeding the classifiers. In order to better capture the spatial information, the proposed COOR and SOOR are adapted to RGB-D data, where a RGB-D proposal fusion method is proposed for RGB-D object detection. With the proposed approaches COOR and SOOR, we obtain the state-of-the-art results of RGB-D scene recognition on SUN RGB-D and NYUD2 datasets. Xinhang Song, Shuqiang Jiang, Chengpeng Chen, Gongwei Chen |
IEEE Trans. Image Process. | 2 |
| 2020 | Food Recommendation: Framework, Existing Solutions, and ChallengesabstractA growing proportion of the global population is becoming overweight or obese, leading to various diseases (e.g., diabetes, ischemic heart disease and even cancer) due to unhealthy eating patterns, such as increased intake of food with high energy and high fat. Food recommendation is of paramount importance to alleviate this problem. Unfortunately, modern multimedia research has enhanced the performance and experience of multimedia recommendation in many fields such as movies and POI, yet largely lags in the food domain. This article proposes a unified framework for food recommendation, and identifies main issues affecting food recommendation including incorporating various context and domain knowledge, building the personal model, and analyzing unique food characteristics. We then review existing solutions for these issues, and finally elaborate research challenges and future directions in this field. To our knowledge, this is the first survey that targets the study of food recommendation in the multimedia field and offers a collection of research studies and technologies to benefit researchers in this field. Weiqing Min, Shuqiang Jiang, Ramesh Jain 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | A Two-Stage Triplet Network Training Framework for Image RetrievalabstractIn this paper, we propose a novel framework for instance-level image retrieval. Recent methods focus on fine-tuning the Convolutional Neural Network (CNN) via a Siamese architecture to improve off-the-shelf CNN features. They generally use the ranking loss to train such networks, and do not take full use of supervised information for better network training, especially with more complex neural architectures. To solve this, we propose a two-stage triplet network training framework, which mainly consists of two stages. First, we propose a Double-Loss Regularized Triplet Network (DLRTN), which extends basic triplet network by attaching the classification sub-network, and is trained via simultaneously optimizing two different types of loss functions. Double-loss functions of DLRTN aim at specific retrieval task and can jointly boost the discriminative capability of DLRTN from different aspects via supervised learning. Second, considering feature maps of the last convolution layer extracted from DLRTN and regions detected from the region proposal network as the input, we then introduce the Regional Generalized-Mean Pooling (RGMP) layer for the triplet network, and re-train this network to learn pooling parameters. Through RGMP, we pool feature maps for each region and aggregate features of different regions from each image to Regional Generalized Activations of Convolutions (R-GAC) as final image representation. R-GAC is capable of generalizing existing Regional Maximum Activations of Convolutions (R-MAC) and is thus more robust to scale and translation. We conduct the experiment on six image retrieval datasets including standard benchmarks and recently introduced INSTRE dataset. Extensive experimental results demonstrate the effectiveness of the proposed framework. Weiqing Min, Shuhuan Mei, Shuqiang Jiang |
IEEE Trans. Multim. | 4 |
| 2020 | Learning Scene Attribute for Scene RecognitionabstractScene recognition has been a challenging task in the field of computer vision and multimedia for a long time. The current scene recognition works often extract object features and scene features through CNN, and combine these two types of features to obtain complementary and discriminative scene representations. However, when the scene categories are visually similar, the object features might lack of discriminations. Therefore, it may be debatable to consider only object features. In contrast to the existing works, in this paper, we discuss the discrimination of scene attributes in local regions and utilize scene attributes as the complementary features of object and scene features. We extract these visual features from two individual CNN branches, one extracting the global features of the image while the other extracting the features of local regions. Through contextual modeling framework, we aggregate these features and generate more discriminative scene representations, which achieve better performance than the feature aggregation of object and scene. Moreover, we achieve the new state-of-the-art performance on both standard scene recognition benchmarks by aggregating more complementary visual features: MIT67 (88.06%) and SUN397 (74.12%). Haitao Zeng, Xinhang Song, Gongwei Chen, Shuqiang Jiang |
IEEE Trans. Multim. | 4 |
| 2020 | Few-shot Food Recognition via Multi-view Representation LearningabstractThis article considers the problem of few-shot learning for food recognition. Automatic food recognition can support various applications, e.g., dietary assessment and food journaling. Most existing works focus on food recognition with large numbers of labelled samples, and fail to recognize food categories with few samples. To address this problem, we propose a Multi-View Few-Shot Learning (MVFSL) framework to explore additional ingredient information for few-shot food recognition. Besides category-oriented deep visual features, we introduce ingredient-supervised deep network to extract ingredient-oriented features. As general and intermediate attributes of food, ingredient-oriented features are informative and complementary to category-oriented features, and thus they play an important role in improving food recognition. Particularly in few-shot food recognition, ingredient information can bridge the gap between disjoint training categories and test categories. To take advantage of ingredient information, we fuse these two kinds of features by first combining their feature maps from their respective deep networks and then convolving combined feature maps. Such convolution is further incorporated into a multi-view relation network, which is capable of comparing pairwise images to enable fine-grained feature learning. MVFSL is trained in an end-to-end fashion for joint optimization on two types of feature learning subnetworks and relation subnetworks. Extensive experiments on different food datasets have consistently demonstrated the advantage of MVFSL in multi-view feature fusion. Furthermore, we extend another two types of networks, namely, Siamese Network and Matching Network, by introducing ingredient information for few-shot food recognition. Experimental results have also shown that introducing ingredient information into these two networks can improve the performance of few-shot food recognition. Shuqiang Jiang, Weiqing Min, Yongqiang Lyu 0002, Linhu Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Learning Object Context for Dense CaptioningabstractDense captioning is a challenging task which not only detects visual elements in images but also generates natural language sentences to describe them. Previous approaches do not leverage object information in images for this task. However, objects provide valuable cues to help predict the locations of caption regions as caption regions often highly overlap with objects (i.e. caption regions are usually parts of objects or combinations of them). Meanwhile, objects also provide important information for describing a target caption region as the corresponding description not only depicts its properties, but also involves its interactions with objects in the image. In this work, we propose a novel scheme with an object context encoding Long Short-Term Memory (LSTM) network to automatically learn complementary object context for each caption region, transferring knowledge from objects to caption regions. All contextual objects are arranged as a sequence and progressively fed into the context encoding module to obtain context features. Then both the learned object context features and region features are used to predict the bounding box offsets and generate the descriptions. The context learning procedure is in conjunction with the optimization of both location prediction and caption generation, thus enabling the object context encoding LSTM to capture and aggregate useful object context. Experiments on benchmark datasets demonstrate the superiority of our proposed approach over the state-of-the-art methods. Xiangyang Li 0002, Shuqiang Jiang, Jungong Han |
AAAI | 2 |
| 2019 | A Real-Time Scene Recognition System Based on RGB-D Video StreamsabstractDepth data captured by the cameras such as Microsoft Kinect can bring depth information than traditional RGB data, which is also more robust to different environments, such as dim or dark lighting conditions. In this technical demonstration, we build a scene recognition system based on real-time processing of RGB-D video streams. Our system recognizes the scenes with video clips, where three types of threads are implemented to ensure the realtime. This system first buffers the frames of both RGB and depth videos with the capturing threads. When the buffered videos reach the certain length, the frames will be packed into clips and forwarded in a pre-trained C3D model to predict scene labels with the scene recognition thread. Finally, the predicted scene labels and captured videos are illustrated in our user interface with illustration thread. Yuyun Hua, Sixian Zhang, Xinhang Song, Jia'ning Li, Shuqiang Jiang |
ICMI | 5 |
| 2019 | Ingredient-Guided Cascaded Multi-Attention Network for Food RecognitionabstractRecently, food recognition is gaining more attention in the multimedia community due to its various applications, e.g., multimodal foodlog and personalized healthcare. Most of existing methods directly extract visual features of the whole image using popular deep networks for food recognition without considering its own characteristics. Compared with other types of object images, food images generally do not exhibit distinctive spatial arrangement and common semantic patterns, and thus are very hard to capture discriminative information. In this work, we achieve food recognition by developing an Ingredient-Guided Cascaded Multi-Attention Network (IG-CMAN), which is capable of sequentially localizing multiple informative image regions with multi-scale from category-level to ingredient-level guidance in a coarse-to-fine manner. At the first level, IG-CMAN generates the initial attentional region from the category-supervised network with Spatial Transformer (ST). Taking this localized attentional region as the reference, IG-CMAN combined ST with LSTM to sequentially discover diverse attentional regions with fine-grained scales from ingredient-guided sub-network in the following levels. Furthermore, we introduce a new dataset ISIA Food-200 with 200 food categories from the list in the Wikipedia, about 200,000 food images and 319 ingredients. We conducted extensive experiment on two popular food datasets and newly proposed ISIA Food-200, and verified the effectiveness of our method. Qualitative results along with visualization further show that IG-CMAN can introduce the explainability for localized regions, and is able to learn relevant regions for ingredients. Weiqing Min, Linhu Liu, Zhengdong Luo, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2019 | MUCH: Mutual Coupling Enhancement of Scene Recognition and Dense CaptioningabstractDue to the abstraction of scenes, comprehensive scene understanding requires semantic modeling in both global and local aspects. Scene recognition is usually researched from a global point of view, while dense captioning is typically studied for local regions. Previous works separately research on the modeling of scene recognition and dense captioning. In contrast, we propose a joint learning framework that benefits from the mutual coupling of scene recognition and dense captioning models. Generally, these two tasks are coupled through two steps, 1) fusing the supervision by considering the contexts between scene labels and local captions, and 2) jointly optimizing semantically symmetric LSTM models. Particularly, in order to balance bias between dense captioning and scene recognition, a scene adaptive non-maximum suppression (NMS) method is proposed to emphasize the scene related regions in region proposal procedure, and a region-wise and category-wise weighted pooling method is proposed to avoid over attention on particular regions in local to global pooling procedure. For the model training and evaluation, scene labels are manually annotated for Visual Genome database. The experimental results on Visual Genome show the effectiveness of the proposed method. Moreover, the proposed method also can improve previous CNN based works on public scene databases, such as MIT67 and SUN397. Xinhang Song, Gongwei Chen, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2019 | Aberrance-aware Gradient-sensitive Attentions for Scene Recognition with RGB-D VideosabstractWith the developments of deep learning, previous approaches have made successes in scene recognition with massive RGB data obtained from the ideal environments. However, scene recognition in real world may face various types of aberrant conditions caused by different unavoidable factors, such as the lighting variance of the environments and the limitations of cameras, which may damage the performance of previous models. In addition to ideal conditions, our motivation is to investigate researches on robust scene recognition models for unconstrained environments. In this paper, we propose an aberrance-aware framework for RGB-D scene recognition, where several types of attentions, such as temporal, spatial and modal attentions are integrated to spatio-temporal RGB-D CNN models to avoid the interference of RGB frame blurring, depth missing, and light variance. All the attentions are homogeneously obtained by projecting the gradient-sensitive maps of visual data into corresponding spaces. Particularly, the gradient maps are captured with the convolutional operations with the typically designed kernels, which can be seamlessly integrated into end-to-end CNN training. The experiments under different challenging conditions demonstrate the effectiveness of the proposed method. Xinhang Song, Sixian Zhang, Yuyun Hua, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2019 | Attention-based Densely Connected LSTM for Video CaptioningabstractRecurrent Neural Networks (RNNs), especially the Long Short-Term Memory (LSTM), have been widely used for video captioning, since they can cope with the temporal dependencies within both video frames and the corresponding descriptions. However, as the sequence gets longer, it becomes much harder to handle the temporal dependencies within the sequence. And in traditional LSTM, previously generated hidden states except the last one do not work directly to predict the current word. This may lead to the predicted word highly related to the last generated hidden state other than the overall context. To better capture long range dependencies and directly leverage early generated hidden states, in this work, we propose a novel model named Attention-based Densely Connected Long Short-Term Memory (DenseLSTM). In DenseLSTM, to ensure maximum information flow, all previous cells are connected to the current cell, which makes the updating of the current state directly related to all its previous states. Furthermore, an attention mechanism is designed to model the impacts of different hidden states. Because each cell is directly connected with all its successive cells, each cell has direct access to the gradients from later ones. In this way, the long-range dependencies are more effectively captured. We perform experiments on two publicly used video captioning datasets: the Microsoft Video Description Corpus (MSVD) and the MSR-VTT, and experimental results illustrate the effectiveness of DenseLSTM. Shuqiang Jiang |
ACM Multimedia | 2 |
| 2019 | Hybrid incremental learning of new data and new classes for hand-held object recognition
Chengpeng Chen, Weiqing Min, Shuqiang Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Instance-level object retrieval via deep region CNN
Shuhuan Mei, Weiqing Min, Hua Duan, Shuqiang Jiang |
Multim. Tools Appl. | 4 |
| 2019 | Class Agnostic Image Common Object DetectionabstractLearning similarity of two images is an important problem in computer vision and has many potential applications. Most of previous works focus on generating image similarities in three aspects: global feature distance computing, local feature matching and image concepts comparison. However, the task of directly detecting class agnostic common objects from two images has not been studied before, which goes one step further to capture image similarities at region level. In this paper, we propose an end-to-end Image Common Object Detection Network (CODN) to detect class agnostic common objects from two images. The proposed method consists of two main modules: locating module and matching module. The locating module generates candidate proposals of each two images. The matching module learns the similarities of the candidate proposal pairs from two images, and refines the bounding boxes of the candidate proposals. The learning procedure of CODN is implemented in an integrated way and a multi-task loss is designed to guarantee both region localization and common object matching. Experiments are conducted on PASCAL VOC 2007 and COCO 2014 datasets. Experimental results validate the effectiveness of the proposed method. Shuqiang Jiang, Sisi Liang, Chengpeng Chen, Xiangyang Li 0002 |
IEEE Trans. Image Process. | 1 |
| 2019 | Learning Effective RGB-D Representations for Scene RecognitionabstractDeep convolutional networks (CNN) can achieve impressive results on RGB scene recognition thanks to large datasets such as Places. In contrast, RGB-D scene recognition is still underdeveloped in comparison, due to two limitations of RGB-D data we address in this paper. The first limitation is the lack of depth data for training deep learning models. Rather than fine tuning or transferring RGB-specific features, we address this limitation by proposing an architecture and a twostep training approach that directly learns effective depth-specific features using weak supervision via patches. The resulting RGBD model also benefits from more complementary multimodal features. Another limitation is the short range of depth sensors (typically 0.5m to 5.5m), resulting in depth images not capturing distant objects in the scenes that RGB images can. We show that this limitation can be addressed by using RGB-D videos, where more comprehensive depth information is accumulated as the camera travels across the scenes. Focusing on this scenario, we introduce the ISIA RGB-D video dataset to evaluate RGB-D scene recognition with videos. Our video recognition architecture combines convolutional and recurrent neural networks (RNNs) that are trained in three steps with increasingly complex data to learn effective features (i.e. patches, frames and sequences). Our approach obtains state-of-the-art performances on RGB-D image (NYUD2 and SUN RGB-D) and video (ISIA RGB-D) scene recognition. Xinhang Song, Shuqiang Jiang, Luis Herranz, Chengpeng Chen |
IEEE Trans. Image Process. | 2 |
| 2019 | Hierarchy-Dependent Cross-Platform Multi-View Feature Learning for Venue Category PredictionabstractIn this paper, we focus on visual venue category prediction, which can facilitate various applications for location-based service and personalization. Considering the complementarity of different media platforms, it is reasonable to leverage venue-relevant media data from different platforms to boost the prediction performance. Intuitively, recognizing one venue category involves multiple semantic cues, especially objects and scenes and, thus, they should contribute together to venue category prediction. In addition, these venues can be organized in a natural hierarchical structure, which provides prior knowledge to guide venue category estimation. Taking these aspects into account, we propose a Hierarchy-dependent Cross-platform Multi-view Feature Learning (HCM-FL) framework for venue category prediction from videos by leveraging images from other platforms. HCM-FL includes two major components, namely Cross-Platform Transfer Deep Learning (CPTDL) and Multi-View Feature Learning with the Hierarchical Venue Structure (MVFL-HVS). CPTDL is capable of reinforcing the learned deep network from videos using images from other platforms. Specifically, CPTDL first trained a deep network using videos. These images from other platforms are filtered by the learnt network and these selected images are then fed into this learnt network to enhance it. Two kinds of pre-trained networks on the ImageNet and Places dataset are employed. Therefore, we can harness both object-oriented and scene-oriented deep features through these enhanced deep networks. MVFL-HVS is then developed to enable multi-view feature fusion. It is capable of embedding the hierarchical structure ontology to support more discriminative joint feature learning. We conduct the experiment on videos from Vine and images from Foursquare. These experimental results demonstrate the advantage of our proposed framework in jointly utilizing multi-platform data, multi-view deep features, and hierarchical venue structure knowledge. Shuqiang Jiang, Weiqing Min, Shuhuan Mei |
IEEE Trans. Multim. | 1 |
| 2019 | Know More Say Less: Image Captioning Based on Scene GraphsabstractAutomatically describing the content of an image has been attracting considerable research attention in the multimedia field. To represent the content of an image, many approaches directly utilize convolutional neural networks (CNNs) to extract visual representations, which are fed into recurrent neural networks to generate natural language. Recently, some approaches have detected semantic concepts from images and then encoded them into high-level representations. Although substantial progress has been achieved, most of the previous methods treat entities in images individually, thus lacking structured information that provides important cues for image captioning. In this paper, we propose a framework based on scene graphs for image captioning. Scene graphs contain abundant structured information because they not only depict object entities in images but also present pairwise relationships. To leverage both visual features and semantic knowledge in structured scene graphs, we extract CNN features from the bounding box offsets of object entities for visual representations, and extract semantic relationship features from triples (e.g.,man riding bike) for semantic representations. After obtaining these features, we introduce a hierarchical-attention-based module to learn discriminative features for word generation at each time step. The experimental results on benchmark datasets demonstrate the superiority of our method compared with several state-of-the-art methods. Xiangyang Li 0002, Shuqiang Jiang |
IEEE Trans. Multim. | 2 |
| 2019 | Deep Patch Representations with Shared Codebook for Scene ClassificationabstractScene classification is a challenging problem. Compared with object images, scene images are more abstract, as they are composed of objects. Object and scene images have different characteristics with different scales and composition structures. How to effectively integrate the local mid-level semantic representations including both object and scene concepts needs to be investigated, which is an important aspect for scene classification. In this article, the idea of a sharing codebook is introduced by organically integrating deep learning, concept feature, and local feature encoding techniques. More specifically, the shared local feature codebook is generated from the combined ImageNet1K and Places365 concepts (Mixed1365) using convolutional neural networks. As the Mixed1365 features cover all the semantic information including both object and scene concepts, we can extract a shared codebook from the Mixed1365 features, which only contain a subset of the whole 1,365 concepts with the same codebook size. The shared codebook can not only provide complementary representations without additional codebook training but also be adaptively extracted toward different scene classification tasks. A method of fusing the encoded features with both the original codebook and the shared codebook is proposed for scene classification. In this way, more comprehensive and representative image features can be generated for classification. Extensive experimentations conducted on two public datasets validate the effectiveness of the proposed method. Besides, some useful observations are also revealed to show the advantage of shared codebook. Shuqiang Jiang, Gongwei Chen, Xinhang Song, Linhu Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Deep Structured Learning for Visual Relationship DetectionabstractIn the research area of computer vision and artificial intelligence, learning the relationships of objects is an important way to deeply understand images. Most of recent works detect visual relationship by learning objects and predicates respectively in feature level, but the dependencies between objects and predicates have not been fully considered. In this paper, we introduce deep structured learning for visual relationship detection. Specifically, we propose a deep structured model, which learns relationship by using feature-level prediction and label-level prediction to improve learning ability of only using feature-level predication. The feature-level prediction learns relationship by discriminative features, and the label-level prediction learns relationships by capturing dependencies between objects and predicates based on the learnt relationship of feature level. Additionally, we use structured SVM (SSVM) loss function as our optimization goal, and decompose this goal into the subject, predicate, and object optimizations which become more simple and more independent. Our experiments on the Visual Relationship Detection (VRD) dataset and the large-scale Visual Genome (VG) dataset validate the effectiveness of our method, which outperforms state-of-the-art methods. Shuqiang Jiang |
AAAI | 2 |
| 2018 | Attentive Recurrent Neural Network for Weak-supervised Multi-label Image ClassificationabstractMulti-label image classification is a fundamental and challenging task in computer vision, and recently achieved significant progress by exploiting semantic relations among labels. However, the spatial positions of labels for multi-labels images are usually not provided in real scenarios, which brings insuperable barrier to conventional models. In this paper, we propose an end-to-end attentive recurrent neural network for multi-label image classification under only image-level supervision, which learns the discriminative feature representations and models the label relations simultaneously. First, inspired by attention mechanism, we propose a recurrent highlight network (RHN) which focuses on the most related regions in the image to learn the discriminative feature representations for different objects in an iterative manner. Second, we develop a gated recurrent relation extractor (GRRE) to model the label relations using multiplicative gates in a recurrent fashion, which learns to decide how multiple labels of the image influence the relation extraction. Extensive experiments on three benchmark datasets show that our model outperforms the state-of-the-arts, and performs better on small-object categories and under the scenario with large number of labels. Liang Li 0003, Shuhui Wang, Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 3 |
| 2018 | Session details: Deep-3 (Image Processing-Inpainting, Super-Resolution, Deblurring)
Shuqiang Jiang |
ACM Multimedia | 1 |
| 2018 | Session details: Grand Challenge-1
Shuqiang Jiang |
ACM Multimedia | 1 |
| 2018 | Session details: Grand Challenge-2
Shuqiang Jiang |
ACM Multimedia | 1 |
| 2018 | Focal Loss for Region Proposal Network
Chengpeng Chen, Xinhang Song, Shuqiang Jiang |
PRCV (2) | 3 |
| 2018 | Bundled Object Context for Referring ExpressionsabstractReferring expressions are natural language descriptions of objects within a given scene. Context is of crucial importance for a referring expression, as the description not only depicts the properties of the object but also involves the relationships of the referred object with other ones. Most of previous work uses either the whole image or one particular contextual object as the context. However, the context of these approaches is holistic and insufficient, as a referring expression often describes relationships of multiple objects in an image. To leverage rich context information from all objects in an image, in this paper, we propose a novel scheme that is composed of a visual context long short-term memory (LSTM) module and a sentence LSTM module to model bundled object context for referring expressions. All contextual objects are arranged with their spatial locations and progressively fed into the visual context LSTM module to acquire and aggregate the context features. Then the concatenation of the learned context features and the features of the referred object are put into the sentence LSTM module to learn the probability of a referring expression. The feedback connections and internal gating mechanism of the LSTM cells enable our model to selectively propagate relevant contextual information through the whole network. Experiments on three benchmark datasets show that our methods can achieve promising results compared to state-of-the-art methods. Moreover, visualization of the internal states of the visual context LSTM cells also shows that our method can automatically select the pertinent context objects. Xiangyang Li 0002, Shuqiang Jiang |
IEEE Trans. Multim. | 2 |
| 2018 | You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food AnalysisabstractCuisine is a style of cooking and usually associated with a specific geographic region. Recipes from different cuisines shared on the web are an indicator of culinary cultures in different countries. Therefore, analysis of these recipes can lead to deep understanding of food from the cultural perspective. In this paper, we perform the first cross-region recipe analysis by jointly using the recipe ingredients, food images, and attributes such as the cuisine and course (e.g., main dish and dessert). For that solution, we propose a culinary culture analysis framework to discover the topics of ingredient bases and visualize them to enable various applications. We first propose a probabilistic topic model to discover cuisine-course specific topics. The manifold ranking method is then utilized to incorporate deep visual features to retrieve food images for topic visualization. At last, we applied the topic modeling and visualization method for three applications: 1) multimodal cuisine summarization with both recipe ingredients and images, 2) cuisine-course pattern analysis including topic-specific cuisine distribution and cuisine-specific course distribution of topics, and 3) cuisine recommendation for both cuisine-oriented and ingredient-oriented queries. Through these three applications, we can analyze the culinary cultures at both macro and micro levels. We conduct the experiment on a recipe database Yummly-66K with 66,615 recipes from 10 cuisines in Yummly. Qualitative and quantitative evaluation results have validated the effectiveness of topic modeling and visualization, and demonstrated the advantage of the framework in utilizing rich recipe information to analyze and interpret the culinary cultures from different regions. Weiqing Min, Bing-Kun Bao, Shuhuan Mei, Yong Rui, Shuqiang Jiang |
IEEE Trans. Multim. | 6 |
| 2017 | Depth CNNs for RGB-D Scene Recognition: Learning from Scratch Better than Transferring from RGB-CNNsabstractScene recognition with RGB images has been extensively studied and has reached very remarkable recognition levels, thanks to convolutional neural networks (CNN) and large scene datasets. In contrast, current RGB-D scene data is much more limited, so often leverages RGB large datasets, by transferring pretrained RGB CNN models and fine-tuning with the target RGB-D dataset. However, we show that this approach has the limitation of hardly reaching bottom layers, which is key to learn modality-specific features. In contrast, we focus on the bottom layers, and propose an alternative strategy to learn depth features combining local weakly supervised training from patches followed by global fine tuning with images. This strategy is capable of learning very discriminative depth-specific features with limited depth images, without resorting to Places-CNN. In addition we propose a modified CNN architecture to further match the complexity of the model and the amount of data available. For RGB-D scene recognition, depth and RGB features are combined by projecting them in a common space and further leaning a multilayer classifier, which is jointly optimized in an end-to-end network. Our framework achieves state-of-the-art accuracy on NYU2 and SUN RGB-D in both depth only and combined RGB-D data. Xinhang Song, Luis Herranz, Shuqiang Jiang |
AAAI | 3 |
| 2017 | Keyword-driven image captioning via Context-dependent Bilateral LSTMabstractImage captioning has recently received much attention. Existing approaches, however, are limited to describing images with simple contextual information, which typically generate one sentence to describe each image with only a single contextual emphasis. In this paper, we address this limitation from a user perspective with a novel approach. Given some keywords as additional inputs, the proposed method would generate various descriptions according to the provided guidance. Hence, descriptions with different focuses can be generated for the same image. Our method is based on a new Context-dependent Bilateral Long Short-Term Memory (CDB-LSTM) model to predict a keyword-driven sentence by considering the word dependence. The word dependence is explored externally with a bilateral pipeline, and internally with a unified and joint training process. Experiments on the MS COCO dataset demonstrate that the proposed approach not only significantly outperforms the baseline method but also shows good adaptation and consistency with various keywords. Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Pengxu Wei, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao, Rynson W. H. Lau |
ICME | 5 |
| 2017 | Visual relationship detection with object spatial distributionabstractRecently, object recognition techniques have been rapidly developed. Most of existing object recognition focused on recognizing several independent concepts. The relationship of objects is also an important problem, which shows in-depth semantic information of images. In this work, toward general visual relationship detection, we propose a method to integrate spatial distribution of object to facilitate visual relation detection. Spatial distribution can not only reflect positional relation of object but also describe structural information between objects. Spatial distributions are described with different features such as positional relation, size relation, shape relation, and so on. By combing spatial distribution features with visual and concept features, we establish a modeling method to make these three aspects working together to facilitate visual relationship detection. To evaluate the proposed method, we conduct experiments on two datasets, which are the Stanford VRD dataset, and a newly proposed larger new dataset which contains 15k images. Experimental results demonstrate that our approach is effective. Shuqiang Jiang, Xiangyang Li 0002 |
ICME | 2 |
| 2017 | Dual Track Multimodal Automatic Learning through Human-Robot InteractionabstractHuman beings are constantly improving their cognitive ability via automatic learning from the interaction with the environment. Two important aspects of automatic learning are the visual perception and knowledge acquisition. The fusion of these two aspects is vital for improving the intelligence and interaction performance of robots. Many automatic knowledge extraction and recognition methods have been widely studied. However, little work focuses on integrating automatic knowledge extraction and recognition into a unified framework to enable jointly visual perception and knowledge acquisition. To solve this problem, we propose a Dual Track Multimodal Automatic Learning (DTMAL) system, which consists of two components: Hybrid Incremental Learning (HIL) from the vision track and Multimodal Knowledge Extraction (MKE) from the knowledge track. HIL can incrementally improve recognition ability of the system by learning new object samples and new object concepts. MKE is capable of constructing and updating the multimodal knowledge items based on the recognized new objects from HIL and other knowledge by exploring the multimodal signals. The fusion of the two tracks is a mutual promotion process and jointly devote to the dual track learning. We have conducted the experiments through human-machine interaction and the experimental results validated the effectiveness of our proposed system. Shuqiang Jiang, Weiqing Min, Huayang Wang, Jiaqi Zhou 0001 |
IJCAI | 1 |
| 2017 | Combining Models from Multiple Sources for RGB-D Scene RecognitionabstractDepth can complement RGB with useful cues about object volumes and scene layout. However, RGB-D image datasets are still too small for directly training deep convolutional neural networks (CNNs), in contrast to the massive monomodal RGB datasets. Previous works in RGB-D recognition typically combine two separate networks for RGB and depth data, pretrained with a large RGB dataset and then fine tuned to the respective target RGB and depth datasets. These approaches have several limitations: 1) only use low-level filters learned from RGB data, thus not being able to exploit properly depth-specific patterns, and 2) RGB and depth features are only combined at high-levels but rarely at lower-levels. In this paper, we propose a framework that leverages both knowledge acquired from large RGB datasets together with depth-specific cues learned from the limited depth data, obtaining more effective multi-source and multi-modal representations. We propose a multi-modal combination method that selects discriminative combinations of layers from the different source models and target modalities, capturing both high-level properties of the task and intrinsic low-level properties of both modalities. Xinhang Song, Shuqiang Jiang, Luis Herranz |
IJCAI | 2 |
| 2017 | A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various AttributesabstractHuman beings have developed a diverse food culture. Many factors like ingredients, visual appearance, courses (e.g., breakfast and lunch), flavor and geographical regions affect our food perception and choice. In this work, we focus on multi-dimensional food analysis based on these food factors to benefit various applications like summary and recommendation. For that solution, we propose a delicious recipe analysis framework to incorporate various types of continuous and discrete attribute features and multi-modal information from recipes. First, we develop a Multi-Attribute Theme Modeling (MATM) method, which can incorporate arbitrary types of attribute features to jointly model them and the textual content. We then utilize a multi-modal embedding method to build the correlation between the learned textual theme features from MATM and visual features from the deep learning network. By learning attribute-theme relations and multi-modal correlation, we are able to fulfill different applications, including (1) flavor analysis and comparison for better understanding the flavor patterns from different dimensions, such as the region and course, (2) region-oriented multi-dimensional food summary with both multi-modal and multi-attribute information and (3) multi-attribute oriented recipe recommendation. Furthermore, our proposed framework is flexible and enables easy incorporation of arbitrary types of attributes and modalities. Qualitative and quantitative evaluation results have validated the effectiveness of the proposed method and framework on the collected Yummly dataset. Weiqing Min, Shuqiang Jiang, Shuhui Wang, Shuhuan Mei |
ACM Multimedia | 2 |
| 2017 | RGB-D Scene Recognition with Object-to-Object RelationabstractA scene is usually abstract that consists of several less abstract entities such as objects or themes. It is very difficult to reason scenes from visual features due to the semantic gap between the abstract scenes and low-level visual features. Some alternative works recognize scenes with a two-step framework by representing images with intermediate representations of objects or themes. However, the object co-occurrences between scenes may lead to ambiguity for scene recognition. In this paper, we propose a framework to represent images with intermediate (object) representations with spatial layout, i.e., object-to-object relation (OOR) representation. In order to better capture the spatial information, the proposed OOR is adapted to RGB-D data. In the proposed framework, we first apply object detection technique on RGB and depth images separately. Then the detected results of both modalities are combined with a RGB-D proposal fusion process. Based on the detected results, we extract semantic feature OOR and regional convolutional neural network (CNN) features located by bounding boxes. Finally, different features are concatenated to feed to the classifier for scene recognition. The experimental results on SUN RGB-D and NYUD2 datasets illustrate the efficiency of the proposed method. Xinhang Song, Chengpeng Chen, Shuqiang Jiang |
ACM Multimedia | 3 |
| 2017 | Guest editorial: mobile visual tagging with mobile context
Shuqiang Jiang, Liangliang Cao, Jiebo Luo 0001, Ramesh Jain 0001 |
Multim. Syst. | 1 |
| 2017 | A survey on context-aware mobile visual recognition
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Ruihan Xu 0001, Yushan Cao, Luis Herranz, Zhiqiang He 0002 |
Multim. Syst. | 2 |
| 2017 | Guest Editorial: Knowledge-Based Multimedia Computing
Liang Li 0003, Zi Huang, Zhengjun Zha, Shuqiang Jiang |
Multim. Tools Appl. | 4 |
| 2017 | Modality-specific and hierarchical feature learning for RGB-D hand-held object recognition
Xiong Lv, Xinda Liu, Xiangyang Li 0002, Shuqiang Jiang, Zhiqiang He 0002 |
Multim. Tools Appl. | 5 |
| 2017 | Multi-Scale Multi-Feature Context Modeling for Scene Recognition in the Semantic ManifoldabstractBefore the big data era, scene recognition was often approached with two-step inference using localized intermediate representations (objects, topics, and so on). One of such approaches is the semantic manifold (SM), in which patches and images are modeled as points in a semantic probability simplex. Patch models are learned resorting to weak supervision via image labels, which leads to the problem of scene categories co-occurring in this semantic space. Fortunately, each category has its own co-occurrence patterns that are consistent across the images in that category. Thus, discovering and modeling these patterns are critical to improve the recognition performance in this representation. Since the emergence of large data sets, such as ImageNet and Places, these approaches have been relegated in favor of the much more powerful convolutional neural networks (CNNs), which can automatically learn multi-layered representations from the data. In this paper, we address many limitations of the original SM approach and related works. We propose discriminative patch representations using neural networks and further propose a hybrid architecture in which the semantic manifold is built on top of multiscale CNNs. Both representations can be computed significantly faster than the Gaussian mixture models of the original SM. To combine multiple scales, spatial relations, and multiple features, we formulate rich context models using Markov random fields. To solve the optimization problem, we analyze global and local approaches, where a top-down hierarchical algorithm has the best performance. Experimental results show that exploiting different types of contextual relations jointly consistently improves the recognition accuracy. Xinhang Song, Shuqiang Jiang, Luis Herranz |
IEEE Trans. Image Process. | 2 |
| 2017 | Modeling Restaurant Context for Food RecognitionabstractFood photos are widely used in food logs for diet monitoring and in social networks to share social and gastronomic experiences. A large number of these images are taken in restaurants. Dish recognition in general is very challenging, due to different cuisines, cooking styles, and the intrinsic difficulty of modeling food from its visual appearance. However, contextual knowledge can be crucial to improve recognition in such scenario. In particular, geocontext has been widely exploited for outdoor landmark recognition. Similarly, we exploit knowledge about menus and location of restaurants and test images. We first adapt a framework based on discarding unlikely categories located far from the test image. Then, we reformulate the problem using a probabilistic model connecting dishes, restaurants, and locations. We apply that model in three different tasks: dish recognition, restaurant recognition, and location refinement. Experiments on six datasets show that by integrating multiple evidences (visual, location, and external knowledge) our system can boost the performance in all tasks. Luis Herranz, Shuqiang Jiang, Ruihan Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Being a Supercook: Joint Food Attributes and Multimodal Content Modeling for Recipe Retrieval and ExplorationabstractThis paper considers the problem of recipe-oriented image-ingredient correlation learning with multi-attributes for recipe retrieval and exploration. Existing methods mainly focus on food visual information for recognition while we model visual information, textual content (e.g., ingredients), and attributes (e.g., cuisine and course) together to solve extended recipe-oriented problems, such as multimodal cuisine classification and attribute-enhanced food image retrieval. As a solution, we propose a multimodal multitask deep belief network ($\mathrm{M}^{3}$TDBN) to learn joint image-ingredient representation regularized by different attributes. By grouping ingredients into visible ingredients (which are visible in the food image, e.g., “chicken” and “mushroom”) and nonvisible ingredients (e.g., “salt” and “oil”),$\mathrm{M}^{3}$TDBN is capable of learning both midlevel visual representation between images and visible ingredients and nonvisual representation. Furthermore, in order to utilize different attributes to improve the intermodality correlation,$\mathrm{M}^{3}$TDBN incorporates multitask learning to make different attributes collaborate each other. Based on the proposed$\mathrm{M}^{3}$TDBN, we exploit the derived deep features and the discovered correlations for three extended novel applications: 1) multimodal cuisine classification; 2) attribute-augmented cross-modal recipe image retrieval; and 3) ingredient and attribute inference from food images. The proposed approach is evaluated on the constructed Yummly dataset and the evaluation results have validated the effectiveness of the proposed approach. Weiqing Min, Shuqiang Jiang, Huayang Wang, Xinda Liu, Luis Herranz |
IEEE Trans. Multim. | 2 |
| 2016 | Scene Recognition with CNNs: Objects, Scales and Dataset BiasabstractSince scenes are composed in part of objects, accurate recognition of scenes requires knowledge about both scenes and objects. In this paper we address two related problems: 1) scale induced dataset bias in multi-scale convolutional neural network (CNN) architectures, and 2) how to combine effectively scene-centric and object-centric knowledge (i.e. Places and ImageNet) in CNNs. An earlier attempt, Hybrid-CNN[23], showed that incorporating ImageNet did not help much. Here we propose an alternative method taking the scale into account, resulting in significant recognition gains. By analyzing the response of ImageNet-CNNs and Places-CNNs at different scales we find that both operate in different scale ranges, so using the same network for all the scales induces dataset bias resulting in limited performance. Thus, adapting the feature extractor to each particular scale (i.e. scale-specific CNNs) is crucial to improve recognition, since the objects in the scenes have their specific range of scales. Experimental results show that the recognition accuracy highly depends on the scale, and that simple yet carefully chosen multi-scale combinations of ImageNet-CNNs and Places-CNNs, can push the stateof-the-art recognition accuracy in SUN397 up to 66.26% (and even 70.17% with deeper architectures, comparable to human performance). Luis Herranz, Shuqiang Jiang, Xiangyang Li 0002 |
CVPR | 2 |
| 2016 | RGB-D scene classification via heterogeneous model fusionabstractWe study the problem of scene classification for RGB-D images in this paper. Firstly we analyze the difference between the RGB and depth images. And then based on the difference, an efficient method is implemented to make use of the RGB and depth images and make a well fusion for the RGB and depth features. Focusing on the difference of modality between the RGB and depth images, we propose a method to learn features from color and depth separately using the heterogeneous model. Especially we use the deep ConvNet model with shallow finetuning for RGB images and the relatively shallow ConvNet model with deep finetuning, which can adequately extract different characteristics of the two modalities. After obtaining the discriminative features for each modality, a multiple fully-connected layers connected with a soft-max classifier is trained to harness the complementary relationship between the two modalities. Experimental evaluations on two publicly RGB-D datasets validate the effectiveness of the proposed method. Xinda Liu, Xueming Wang, Shuqiang Jiang |
ICIP | 3 |
| 2016 | Image Captioning with both Object and Scene InformationabstractRecently, automatic generation of image captions has attracted great interest not only because of its extensive applications but also because it connects computer vision and natural language processing. By combining convolutional neural networks (CNNs), which learn visual representations from images, and recurrent neural networks (RNNs), which translate the learned features into text sequences, the content of a image can be transformed into linguistic sequences. Existing approaches typically focus on visual features extracted form an object-oriented CNN (train on ImageNet) and then decode them into natural language. In this paper, we propose a novel model using not only object-related, but also scene-related information extracted from the images. To make full use of both object and scene information, we first combine object information and scene information (extracted from a scene-oriented CNN), and then using as inputs to RNNs. Both types of information provide complementary aspects that help in generating a more complete description of the image. Qualitative and quantitative evaluation results validate the effectiveness of our method. Xiangyang Li 0002, Xinhang Song, Luis Herranz, Shuqiang Jiang |
ACM Multimedia | 5 |
| 2016 | Online web video topic detection and tracking with semi-supervised learning
Guorong Li, Shuqiang Jiang, Weigang Zhang, Junbiao Pang, Qingming Huang |
Multim. Syst. | 2 |
| 2016 | Guest Editorial: Image Analysis and Processing Leveraging Additional Information
Luis Herranz, Jian Cheng 0001, Yue Gao 0002, Shuqiang Jiang |
Multim. Tools Appl. | 4 |
| 2016 | Scalable storyboards in handheld devices: applications and evaluation metrics
Luis Herranz, Shuqiang Jiang |
Multim. Tools Appl. | 2 |
| 2016 | Category co-occurrence modeling for large scale scene recognition
Xinhang Song, Shuqiang Jiang, Luis Herranz, Yan Kong, Kai Zheng 0001 |
Pattern Recognit. | 2 |
| 2015 | Joint multi-feature spatial context for scene recognition in the semantic manifoldabstractIn the semantic multinomial framework patches and images are modeled as points in a semantic probability simplex. Patch theme models are learned resorting to weak supervision via image labels, which leads the problem of scene categories co-occurring in this semantic space. Fortunately, each category has its own co-occurrence patterns that are consistent across the images in that category. Thus, discovering and modeling these patterns is critical to improve the recognition performance in this representation. In this paper, we observe that not only global co-occurrences at the image-level are important, but also different regions have different category co-occurrence patterns. We exploit local contextual relations to address the problem of discovering consistent co-occurrence patterns and removing noisy ones. Our hypothesis is that a less noisy semantic representation, would greatly help the classifier to model consistent co-occurrences and discriminate better between scene categories. An important advantage of modeling features in a semantic space is that this space is feature independent. Thus, we can combine multiple features and spatial neighbors in the same common space, and formulate the problem as minimizing a context-dependent energy. Experimental results show that exploiting different types of contextual relations consistently improves the recognition accuracy. In particular, larger datasets benefit more from the proposed method, leading to very competitive performance. Xinhang Song, Shuqiang Jiang, Luis Herranz |
CVPR | 2 |
| 2015 | A probabilistic model for food image recognition in restaurantsabstractA large amount of food photos are taken in restaurants for diverse reasons. This dish recognition problem is very challenging, due to different cuisines, cooking styles and the intrinsic difficulty of modeling food from its visual appearance. Contextual knowledge is crucial to improve recognition in such scenario. In particular, geocontext has been widely exploited for outdoor landmark recognition. Similarly, we exploit knowledge about menus and geolocation of restaurants and test images. We first adapt a framework based on discarding unlikely categories located far from the test image. Then we reformulate the problem using a probabilistic model connecting dishes, restaurants and geolocations. We apply that model in three different tasks: dish recognition, restaurant recognition and geolocation refinement. Experiments on a dataset including 187 restaurants and 701 dishes show that combining multiple evidences (visual, geolocation, and external knowledge) can boost the performance in all tasks. Luis Herranz, Ruihan Xu 0001, Shuqiang Jiang |
ICME | 3 |
| 2015 | Hand-Object Sense: A Hand-held Object Recognition System Based on RGB-D InformationabstractHand-held objects play an important role in human-human and human-machine interaction. It can be used as a reference for understanding user intentions or user requirements. In this technical demonstration, we introduce an object recognition system called Hand-Object Sense that can automatically recognize the object held by user. This system first detects and segments the hand-held object by exploiting skeleton information combined with depth information. Second, in the object recognition stage, this system exploits features computed in different ways and fuses them to improve the recognition accuracy. Our system can recognize objects in real-time and have a good tolerance to angle and scale transformation. Furthermore, it has a good generalization capability for unknown objects. Xiong Lv, Shuqiang Jiang, Luis Herranz |
ACM Multimedia | 2 |
| 2015 | Rich Image Description Based on RegionsabstractAbstract Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In contrast to the previous image description methods that focus on describing the whole image, this paper presents a method of generating rich image descriptions from image regions. First, we detect regions with R-CNN (regions with convolutional neural network features) framework. We then utilize the RNN (recurrent neural networks) to generate sentences for image regions. Finally, we propose an optimization method to select one suitable region. The proposed model generates several sentence description of regions in an image, which has sufficient representative power of the whole image and contains more detailed information. Comparing to general image level description, generating more specific and accurate sentences on the different regions can satisfy more personal requirements for different people. Experimental evaluations validate the effectiveness of the proposed method. Xiaodan Zhang 0003, Xinhang Song, Xiong Lv, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao |
ACM Multimedia | 4 |
| 2015 | Online learning affinity measure with CovBoost for multi-target tracking
Guorong Li, Qingming Huang, Shuqiang Jiang, Yingkun Xu, Weigang Zhang |
Neurocomputing | 3 |
| 2015 | Cluster-sensitive Structured Correlation Analysis for Web cross-modal retrieval
Shuhui Wang, Fuzhen Zhuang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 3 |
| 2015 | RGB-D Hand-Held Object Recognition Based on Heterogeneous Feature Fusion
Xiong Lv, Shuqiang Jiang, Luis Herranz |
J. Comput. Sci. Technol. | 2 |
| 2015 | LSH-based semantic dictionary learning for large scale image understanding
Liang Li 0003, Chenggang Yan 0001, Bo-Wei Chen, Shuqiang Jiang, Qingming Huang |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Polysemious visual representation based on feature aggregation for large scale image applications
Xinhang Song, Shuqiang Jiang, Shuhui Wang, Liang Li 0003, Qingming Huang |
Multim. Tools Appl. | 2 |
| 2015 | Geolocalized Modeling for Dish RecognitionabstractFood-related photos have become increasingly popular , due to social networks, food recommendations, and dietary assessment systems. Reliable annotation is essential in those systems, but unconstrained automatic food recognition is still not accurate enough. Most works focus on exploiting only the visual content while ignoring the context. To address this limitation, in this paper we explore leveraging geolocation and external information about restaurants to simplify the classification problem. We propose a framework incorporating discriminative classification in geolocalized settings and introduce the concept of geolocalized models, which, in our scenario, are trained locally at each restaurant location. In particular, we propose two strategies to implement this framework: geolocalized voting and combinations of bundled classifiers. Both models show promising performance, and the later is particularly efficient and scalable. We collected a restaurant-oriented food dataset with food images, dish tags, and restaurant-level information, such as the menu and geolocation. Experiments on this dataset show that exploiting geolocation improves around 30% the recognition performance, and geolocalized models contribute with an additional 3-8% absolute gain, while they can be trained up to five times faster. Ruihan Xu 0001, Luis Herranz, Shuqiang Jiang, Xinhang Song, Ramesh Jain 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | INSTRE: A New Benchmark for Instance-Level Object Retrieval and RecognitionabstractOver the last several decades, researches on visual object retrieval and recognition have achieved fast and remarkable success. However, while the category-level tasks prevail in the community, the instance-level tasks (especially recognition) have not yet received adequate focuses. Applications such as content-based search engine and robot vision systems have alerted the awareness to bring instance-level tasks into a more realistic and challenging scenario. Motivated by the limited scope of existing instance-level datasets, in this article we propose a new benchmark for INSTance-level visual object REtrieval and REcognition (INSTRE). Compared with existing datasets, INSTRE has the following major properties: (1) balanced data scale, (2) more diverse intraclass instance variations, (3) cluttered and less contextual backgrounds, (4) object localization annotation for each image, (5) well-manipulated double-labelled images for measuring multiple object (within one image) case. We will quantify and visualize the merits of INSTRE data, and extensively compare them against existing datasets. Then on INSTRE, we comprehensively evaluate several popular algorithms to large-scale object retrieval problem with multiple evaluation metrics. Experimental results show that all the methods suffer a performance drop on INSTRE, proving that this field still remains a challenging problem. Finally we integrate these algorithms into a simple yet efficient scheme for recognition and compare it with classification-based methods. Importantly, we introduce the realistic multiobjects recognition problem. All experiments are conducted in both single object case and multiple objects case. Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Accuracy and Specificity Trade-off in k -nearest Neighbors Classification
Luis Herranz, Shuqiang Jiang |
ACCV (2) | 2 |
| 2014 | Graph-Density-based visual word vocabulary for image retrievalabstractDescriptive visual word vocabulary serves as the foundation of large scale image retrieval systems. However, the visual word descriptive power is limited by the construction mechanisms based on either cluster center or partitioned feature space, since such mechanisms may merge the sparsely distributed features and split the densely distributed features. Besides, there are a large number of outlier features that are not similar with any visual word. Quantizing such features into visual words inevitably decreases the visual word descriptive power. In this paper, we propose a novel Graph-Density-based visual word Vocabulary (GDV), which constructs the visual word by dense feature subgraph and directly measures the intra-word similarity by the corresponding graph density. Our method remarkably enhances the visual word descriptive power from the following three aspects: 1) GDV guarantees the high intra-word similarity by constructing visual words under the criterion of large graph density; 2) GDV improves the inter-word dissimilarity by alleviating the unexpected effect of subgraph splitting; 3) GDV suppresses the influence of outlier features by selectively quantizing only the features that are similar enough with the visual words. Extensive experiments demonstrate GDV's advanced descriptive power over traditional visual word vocabularies in enhancing both the retrieval accuracy and efficiency, which provides a higher level starting point for most image retrieval systems. Lingyang Chu, Shuhui Wang, Shuqiang Jiang, Qingming Huang |
ICME | 4 |
| 2014 | Cross media topic analytics based on synergetic content and user behavior modelingabstractHot topic in cross media, defined as a set of Web documents containing similar semantic information, describes the same real world event with significant social impact. However, cross media topic detection still remains a challenging issue since it is unclear how users interact with the cross media topics, leading to difficulties in identifying the influence of different topics on different user communities. In this paper, we propose a solution framework for cross media topic analysis based on synergetic modeling of multi-modal content and user behavior. First, we detect atom topics by multi-modal topic detection. Second, we propose a multi-resolution user behavior modeling method to discover communities on the active Web users by considering the distribution of related atom topics along the temporal axis. We analyze the topic-topic, topic-community and community-community relations by using the proposed method. Consequently, the macro-topic and macro-community structures can be obtained for better understanding of the interaction between topics and communities. Experiments show the capability of our method in knowledge discovery on cross media topics. Shuhui Wang, Zhenjun Wang, Shuqiang Jiang, Qingming Huang |
ICME | 3 |
| 2014 | Fusing multi-cues description for partial-duplicate image retrieval
Chenggang Yan 0001, Liang Li 0003, Jian Yin 0003, Hailong Shi, Shuqiang Jiang, Qingming Huang |
J. Vis. Commun. Image Represent. | 6 |
| 2014 | Relative image similarity learning with contextual information for Internet cross-media retrieval
Shuqiang Jiang, Xinhang Song, Qingming Huang |
Multim. Syst. | 1 |
| 2014 | Preface: Internet multimedia computing and service
Shuqiang Jiang, Changsheng Xu, Yong Rui, Alberto Del Bimbo, Hongxun Yao |
Multim. Tools Appl. | 1 |
| 2013 | Multi-level Discriminative Dictionary Learning towards Hierarchical Visual CategorizationabstractFor the task of visual categorization, the learning model is expected to be endowed with discriminative visual feature representation and flexibilities in processing many categories. Many existing approaches are designed based on a flat category structure, or rely on a set of pre-computed visual features, hence may not be appreciated for dealing with large numbers of categories. In this paper, we propose a novel dictionary learning method by taking advantage of hierarchical category correlation. For each internode of the hierarchical category structure, a discriminative dictionary and a set of classification models are learnt for visual categorization, and the dictionaries in different layers are learnt to exploit the discriminative visual properties of different granularity. Moreover, the dictionaries in lower levels also inherit the dictionary of ancestor nodes, so that categories in lower levels are described with multi-scale visual information using our dictionary learning approach. Experiments on Image Net object data subset and SUN397 scene dataset demonstrate that our approach achieves promising performance on data with large numbers of classes compared with some state-of-the-art methods, and is more efficient in processing large numbers of categories. Li Shen 0005, Shuhui Wang, Gang Sun 0005, Shuqiang Jiang, Qingming Huang |
CVPR | 4 |
| 2013 | ObjectSense: a scalable multi-objects recognition system based on partial-duplicate image retrievalabstractIn this demo, we present ObjectSense, a scalable object recognition system that recognizes multiple objects present in a static image or in the camera frames. Instead of applying learning based recognition framework, this system identifies objects through Partial-Duplicate Image Retrieval (PDIR) based method. First, objects are identified by measuring the similarity between an incoming image and reference image corpus that are labeled with the objects. To compute image similarities, we explore the Consistency Graph Model (CGM), which robustly rejects spatially inconsistent feature matches with the advantage of orientations and positions of local features. Then a kNN voting method is used to decide the object category based on the quantized image similarities. ObjectSense is scalable with promisingly high recall and accuracy, which fits well into recognition-guided shopping and human computer interaction. We built ObjectSense on two platforms, PC and Android. Yunfeng Xue, Lingyang Chu, Shuqiang Jiang |
ICMR | 5 |
| 2013 | Flexible navigation in smartphones and tablets using scalable storyboardsabstractIn this demo paper we present a multiscale browsing interface for handheld devices, in which the user can interactively change the scale of the storyboard to easily adjust the amount of information desired. Conventional and hierarchical storyboards provide one or very few possible lengths. In contrast, scalable storyboards allow the number of images and the storyboard itself to be adapted to the device constraints (e.g. aspect ratio, resolution) and navigation state with much finer granularity. Several levels and modes, including segment of interest, are provided for more intuitive and convenient navigation. Shuai Zheng 0004, Luis Herranz, Shuqiang Jiang |
ICMR | 3 |
| 2013 | Cross Concept Local Fisher Discriminant Analysis for Image Classification
Xinhang Song, Shuqiang Jiang, Shuhui Wang, Jinhui Tang 0001, Qingming Huang |
MMM (2) | 2 |
| 2013 | Weighted visual vocabulary to balance the descriptive ability on general dataset
Shuqiang Jiang, Qingming Huang |
Neurocomputing | 2 |
| 2013 | SSOCBT: A Robust Semisupervised Online CovBoost Tracker That Uses Samples DifferentlyabstractMost existing feature selection methods for object tracking assume that the samples in the previous frames are governed by the same distribution of the labeled samples obtained in the current frame and unlabeled samples collected in the next frame. However, according to our statistical analysis on very common videos, this assumption is not true in many scenarios. As a result, the selected features are not suitable for discriminating between the target from the background in the next frame. A tracking error accumulates and finally the drift problem happens. In this paper, we consider data distribution in tracking from a new perspective to adapt to target's and background's changes. We classify the samples into three categories: auxiliary samples (samples in the previous frames), target samples (samples collected in the current frame), and unlabeled samples (samples obtained in the next frame). To make the best use of them for tracking, we propose a novel semisupervised transfer learning approach that treats samples differently. Specifically, we assume that only target samples follow the same distribution as the unlabeled samples that we want to classify. Then, a novel and interesting semisupervised CovBoost method is developed utilizing the information provided by the three kinds of samples effectively when training the best strong classifier for tracking. Furthermore, we develop a new online updating algorithm for semisupervised CovBoost, making our tracker handle with significant variations of the tracked target and background successfully. Our experimental results demonstrate the advantages of treating samples differently during tracking. Our tracker outperforms state-of-the-art trackers on the benchmark datasets. Guorong Li, Qingming Huang, Shuqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Robust Spatial Consistency Graph Model for Partial Duplicate Image RetrievalabstractPartial duplicate images often have large non-duplicate regions and small duplicate regions with random rotation, which lead to the following problems: 1) large number of noisy features from the non-duplicate regions; 2) small number of representative features from the duplicate regions; 3) randomly rotated or deformed duplicate regions. These problems challenge many content based image retrieval (CBIR) approaches, since most of them cannot distinguish the representative features from a large proportion of noisy features in a rotation invariant way. In this paper, we propose a rotation invariant partial duplicate image retrieval (PDIR) approach, which effectively and efficiently retrieves the partial duplicate images by accurately matching the representative SIFT features. Our method is based on the Combined-Orientation-Position (COP) consistency graph model, which consists of the following two parts: 1) The COP consistency, which is a rotation invariant measurement of the relative spatial consistency among the candidate matches of SIFT features; it uses a coarse-to-fine family of evenly sectored polar coordinate systems to softly quantize and combine the orientations and positions of the SIFT features. 2) The consistency graph model, which robustly rejects the spatially inconsistent noisy features by effectively detecting the group of candidate feature matches with the largest average COP consistency. Extensive experiments on five large scale image data sets show promising retrieval performances. Lingyang Chu, Shuqiang Jiang, Shuhui Wang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2012 | Multi-feature metric learning with knowledge transfer among semantics and social taggingabstractPrevious metric learning approaches learn a unified metric for all the classes on single feature representation, thus cannot be directly transplanted to applications involving multiple features, hundreds to thousands of hierarchical structured semantics and abundant social tagging. In this paper, we propose a novel multi-task multi-feature metric learning method which models the information sharing mechanism among different learning tasks. We decompose the real world multi-class problems such as semantic categorization or automatic tagging into a set of tasks where each task corresponds to several classes with strong visual correlation. We conduct metric learning to learn a set of (hyper)category-specific metrics for all the tasks. By encouraging model sharing among tasks, more generalization power is acquired. Another advantage is the capability of simultaneous learning with semantic information and social tagging based on the multi-task learning framework, and thus they both benefit from the information provided by each other. Experiments demonstrate the advantages on applications including semantic categorization and automatic tagging compared with other popular metric learning approaches. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
CVPR | 2 |
| 2012 | Color Maximal-Dissimilarity Pattern for pedestrian detection
Junbiao Pang, Guoyi Liu, Qingming Huang, Shuqiang Jiang |
ICPR | 6 |
| 2012 | Nearest-neighbor method using multiple neighborhood similarities for social media data mining
Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
Neurocomputing | 3 |
| 2012 | Online selection of the best k-feature subset for object tracking
Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2012 | @ICT: attention-based virtual content insertion
Qingming Huang, Changsheng Xu, Shuqiang Jiang |
Multim. Syst. | 4 |
| 2012 | Learning Hierarchical Semantic Description Via Mixed-Norm Regularization for Image UnderstandingabstractThis paper proposes a new perspective-Vicept representation to solve the problem of visual polysemia and concept polymorphism in the large-scale semantic image understanding. Vicept characterizes the membership probability distribution between visual appearances and semantic concepts, and forms a hierarchical representation of image semantic from local to global. In the implementation, incorporating group sparse coding, visual appearance is encoded as a weighted sum of dictionary elements, which could obtain more accurate image representation with sparsity at the image level. To obtain discriminative Vicept descriptions with structural sparsity, mixed-norm regularization is adopted in the optimization problem for learning the concept membership distribution of visual appearance. Furthermore, we introduce a novel image distance measurement based on the hierarchical Vicept description, where different levels of Vicept distance are fused together by multi-level separability analysis. Finally, the wide applications of Vicept description are validated in our experiments, including large-scale semantic image search, image annotation, and semantic image re-ranking. Liang Li 0003, Shuqiang Jiang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2012 | S3MKL: Scalable Semi-Supervised Multiple Kernel Learning for Real-World Image ApplicationsabstractWe study the visual learning models that could work efficiently with little ground-truth annotation and a mass of noisy unlabeled data for large scale Web image applications, following the subroutine of semi-supervised learning (SSL) that has been deeply investigated in various visual classification tasks. However, most previous SSL approaches are not able to incorporate multiple descriptions for enhancing the model capacity. Furthermore, sample selection on unlabeled data was not advocated in previous studies, which may lead to unpredictable risk brought by real-world noisy data corpse. We propose a learning strategy for solving these two problems. As a core contribution, we propose a scalable semi-supervised multiple kernel learning method$({\rm S}^{3}{\rm MKL})$to deal with the first problem. The aim is to minimize an overall objective function composed of log-likelihood empirical loss, conditional expectation consensus (CEC) on the unlabeled data and group LASSO regularization on model coefficients. We further adapt CEC into a group-wise formulation so as to better deal with the intrinsic visual property of real-world images. We propose a fast block coordinate gradient descent method with several acceleration techniques for model solution. Compared with previous approaches, our model better makes use of large scale unlabeled images with multiple feature representation with lower time complexity. Moreover, to address the issue of reducing the risk of using unlabeled data, we design a multiple kernel hashing scheme to identify the “informative” and “compact” unlabeled training data subset. Comprehensive experiments are conducted and the results show that the proposed learning framework provides promising power for real-world image applications, such as image categorization and personalized Web image re-ranking with very little user interaction. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2011 | Efficient lp-norm multiple feature metric learning for image categorizationabstractPrevious metric learning approaches are only able to learn the metric based on single concatenated multivariate feature representation. However, for many real world problems with multiple feature representation such as image categorization, the model trained by previous approaches will degrade because of sparsity brought by significant dimension growth and uncontrolled influence from each feature channel. In this paper, we propose an efficient distance metric learning model which adapts Distance Metric Learning on multiple feature representations. The aim is to learn the Mahalanobis matrices for each independent feature and their non-sparse lp-norm weight coefficients simultaneously by maximizing the margin of the overall learned distance metric among the pairs from the same class and the distance of pairs from different classes. We further extend this method to nonlinear kernel learning and category specific metric learning, which demonstrate the applicability of using many existing kernels for image data and exploring the hierarchical semantic structures for large scale image datasets. Experiments on various datasets demonstrate the promising power of our method. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
CIKM | 3 |
| 2011 | Learning image Vicept description via mixed-norm regularization for large scale semantic image searchabstractThe paradox of visual polysemia and concept polymorphism has been a great challenge in the large scale semantic image search. To address this problem, our paper proposes a new method to generate image Vicept representation. Vicept characterizes the membership distribution between elementary visual appearances and semantic concepts, and forms a hierarchical representation of image semantic from local to global. To obtain discriminative Vicept descriptions with structural sparsity, we adopt mixed-norm regularization in the optimization problem for learning the concept membership distribution of visual word. Furthermore, considering the structure of BOV in images, visual descriptor is encoded as a weighted sum of dictionary elements using group sparse coding, which could obtain sparse representation at the image level. The wide applications of Vicept are validated in our experiments, including large scale semantic image search, image annotation, and semantic image re-ranking. Liang Li 0003, Shuqiang Jiang, Qingming Huang |
CVPR | 2 |
| 2011 | Treat samples differently: Object tracking with semi-supervised online CovBoostabstractMost feature selection methods for object tracking assume that the labeled samples obtained in the next frames follow the similar distribution with the samples in the previous frame. However, this assumption is not true in some scenarios. As a result, the selected features are not suitable for tracking and the “drift” problem happens. In this paper, we consider data's distribution in tracking from a new perspective. We classify the samples into three categories: auxiliary samples (samples in the previous frames), target samples (collected in the current frame) and unlabeled samples (obtained in the next frame). To make the best use of them for tracking, we propose a novel semi-supervised transfer learning approach. Specifically, we assume only target samples follow the same distribution as the unlabeled samples and develop a novel semi-supervised CovBoost method. It could utilize auxiliary samples and unlabeled samples effectively when training the best strong classifier for tracking. Furthermore, we develop a new online updating algorithm for semi-supervised CovBoost, making our tracker handle with significant variations of the tracked target and background successfully. We demonstrate the excellent performance of the proposed tracker on several challenging test videos. Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang |
ICCV | 5 |
| 2011 | Fast common visual pattern detection via radiate geometric modelabstractIn this paper, we propose a novel method to implement fast detection of Common Visual Pattern (CVP). The purpose of CVP detection is to find the correspondences between the common visual regions of two given partial duplicate images. There are two major components of the proposed method which guarantee the good performance. First, we establish the Radiate-Geometric-Model (RGM). The RGM is represented by a set of radiate structures, and each structure is geometrically made up of a group of matched feature pairs. By utilizing the statistical information gained from the radiate structures, the RGM can not only quickly estimate the potential pairs of common regions but also organize the scale relationship between matched pairs into a compact form, hence increase the detection speed substantially. Second, we formulize the Radiate-Geometric-Model (RGM) into a graph optimization problem which could be solved by the method of graph-shift, thus make our algorithm capable of detecting the CVPs of all kinds of correspondences. Experimental results prove that the speed of our algorithm is at least 40 times faster than the state-of-the-art, while achieving a better detection performance at the same time. Lingyang Chu, Shuqiang Jiang, Qingming Huang |
ICIP | 2 |
| 2011 | Online Vicept learning for web-scale image understandingabstractWeb-scale image understanding is a challenging but significant task to comprehend image contents on the internet. The de-facto standard methods based on machine learning or computer vision still suffer from a phenomenon of visual pol-ysemia and concept polymorphism (VPCP). To resolve the VPCP, Vicept has been proposed to characterize the membership distribution between visual appearances and semantic concepts. In this paper, we propose an online Vicept learning algorithm on the base of stochastic approximations, which can scale up to large scale datasets with millions of training samples. With the help of the Vicept, we develop an extension of the spatial pyramid matching (SPM) kernel method by generalizing the Vicept as a basic semantic description. The efficiency of our approach is validated in the experiments of web-scale semantic image search and image classification on the ImageNet dataset and Caltech-256 dataset. Liang Li 0003, Shuqiang Jiang, Qingming Huang |
ICIP | 2 |
| 2011 | Query sensitive dynamic web video thumbnail generationabstractWith the fast rising of the video sharing websites, the online video becomes an important media for people to share messages, interests, ideas, beliefs, etc. In this paper, we propose a novel approach to dynamically generate the web video thumbnails according to user's query. Two issues are addressed: the video content representativeness of the selected video thumbnail, and the relationship between the selected video thumbnail and the user's query. For the first issue the reinforcement based algorithm is adopted to rank the frames in each video. For the second issue the relevance model based method is employed to calculate the similarity between the video frames and the query keywords. The final video thumbnail is generated by linear fusion of the above two scores. Compared with the existing web video thumbnails, which only reflect the preference of the video owner, the thumbnails generated in our approach not only consider the video content representativeness of the frame, but also reflect the intention of the video searcher. In order to show the effectiveness of the proposed method, experiments are conducted on the videos selected from the video sharing website. Experimental results and subjective evaluations demonstrate that the proposed method is effective and can meet the user's intention requirement. Chunxi Liu, Qingming Huang, Shuqiang Jiang |
ICIP | 3 |
| 2011 | Human tracking by structured body partsabstractTracking non-rigid objects with significant shape variation in complex scenario is a difficult problem. Human tracking is a special case of this problem since human body has good local rigid properties. In this paper, we propose a novel human tracking method which explores the local rigid properties while keeping the global structure very well. This method consists of three stages. First, the human body is represented by structured rigid parts extracted using patches clustering method. Then, the rigid parts are tracked through a structured constraint method. Finally, the optimal estimated state of the object is obtained through global similarity. Experimental results show that the proposed method has good performance for human tracking with big posture and shape change. Yingkun Xu, Shuqiang Jiang, Qingming Huang |
ICIP | 3 |
| 2011 | Matching Content-based Saliency Regions for partial-duplicate image retrievalabstractIn traditional partial-duplicate image retrieval, images are commonly represented using the Bag-of-Visual-Words (BOV) model built from image local features, such as SIFT. Actually, there is only a small similar portion between partial-duplicate images so that such representation on the whole image is not adequate for the partial-duplicate image retrieval task. In this paper, we propose a novel perspective to retrieval partial-duplicate images with Contented-based Saliency Region (CSR). CSRs are such sub-regions with abundant visual content and high visual attention in the image. The content of CSR is represented with the BOV model while saliency analysis is employed to ensure the high visual attention of CSR. Each CSR is regarded as an independent unit to be retrieved in the dataset. To effectively retrieve the CSRs, we design a relative saliency ordering constraint, which captures a weak saliency relative layout among interest points in the CSR. Comparison experiments with four state-of-the-art methods on the standard partial-duplicate image dataset clearly verify the effectiveness of our scheme. Further, our approach can provide a more diverse retrieval result, which facilitates the interaction of portable-device users. Liang Li 0003, Zhengjun Zha, Shuqiang Jiang, Qingming Huang |
ICME | 4 |
| 2011 | News video story sentiment classification and rankingabstractIn this paper, we present a novel approach for news video story sentiment analysis. Two research challenges are addressed: news video story sentiment classification and ranking. For classification, a graph based semi-supervised learning approach is utilized to classify the news stories into sentiment classes. Graph based semi-supervised learning is able to tackle the problem of lacking labeled data. After classification, two sentiment classes are obtained: positive and negative. In order to project the news videos into sentiment space, a multimodal approach by fusing the text sentiment and visual representation scores is adopted to rank the videos in each class. For sentiment representation, inter and intra sentiment class analysis is conducted based on affinity propagation clustering and PageRank algorithm. A user study is conducted to evaluate the video ranking performance. The experimental results on the selected topics are promising and demonstrate the proposed approach is effective. Chunxi Liu, Li Su 0003, Qingming Huang, Shuqiang Jiang |
ICME | 4 |
| 2011 | Detection and location of near-duplicate video sub-clips by finding dense subgraphsabstractRobust and fast near-duplicate video detection is an important task with many potential applications. Most existing systems focus on the comparison between full copy videos or partial near-duplicate videos. While it is more challenging to find similar content for videos containing multiple near-duplicate segments at random locations with various connections. In this paper, we propose a new graph based method to detect complex near-duplicate video sub-clips. First, we develop a new succinct video descriptor for keyframe match. Then a graph is established to exploit temporal consistency of matched keyframes. The nodes of the graph are the matched frame pairs; the edge weights are computed from the temporal alignment and frame pair similarities. In this way, the validly matched keyframes would form a dense subgraph whose nodes are strongly connected. This graph model also preserves the complex connections of sub-clips. Thus detecting complex near-duplicate sub-clips is transformed to the problem of finding all the dense subgraphs. We employ the optimization method of graph shift to solve this problem due to its robust performance. The experiments are conducted on the dataset with various transformations and complex temporal relations. The results demonstrate the effectiveness and efficiency of the proposed method. Tianlong Chen 0003, Shuqiang Jiang, Lingyang Chu, Qingming Huang |
ACM Multimedia | 2 |
| 2011 | Human group activity analysis with fusion of motion and appearance informationabstractHuman activity analysis is an important and challenging task in video content analysis and understanding. In this paper, we focus on the activity of small human group, which involves countable persons and complex interactions. To cope with the variant number of participants and inherent interactions within the activity, we propose a hierarchical model with three layers to depict the characteristics at different granularities. In traditional methods, group activity is represented mainly based on motion information, such as human trajectories, but ignoring discriminative appearance information, e.g. the rough sketch of a pose style. In our approach, we take advantage of both the motion and the appearance information in the spatiotemporal activity context under the hierarchical model. These features are inhomogeneous. Therefore, we employ multiple kernel learning methods to fuse the features for group activity recognition. Experiments on a surveillance-like human group activity database demonstrate the validity of our approach and the recognition performance is promising. Zhongwei Cheng, Qingming Huang, Shuqiang Jiang, Shuicheng Yan, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2011 | Special edition on semi-supervised learning for visual content analysis and understanding
Jian Cheng 0001, Jinjun Wang, Shuqiang Jiang, Zhi-Hua Zhou, Edwin R. Hancock |
Pattern Recognit. | 3 |
| 2011 | Transferring Boosted Detectors Towards Viewpoint and Scene AdaptivenessabstractIn object detection, disparities in distributions between the training samples and the test ones are often inevitable, resulting in degraded performance for application scenarios. In this paper, we focus on the disparities caused by viewpoint and scene changes and propose an efficient solution to these particular cases by adapting generic detectors, assuming boosting style. A pretrained boosting-style detector encodes a priori knowledge in the form of selected features and weak classifier weighting. Towards adaptiveness, the selected features are shifted to the most discriminative locations and scales to compensate for the possible appearance variations. Moreover, the weighting coefficients are further adapted with covariate boost, which maximally utilizes the related training data to enrich the limited new examples. Extensive experiments validate the proposed adaptation mechanism towards viewpoint and scene adaptiveness and show encouraging improvement on detection accuracy over state-of-the-art methods. Junbiao Pang, Qingming Huang, Shuicheng Yan, Shuqiang Jiang |
IEEE Trans. Image Process. | 4 |
| 2010 | Novel observation model for probabilistic object trackingabstractTreating visual object tracking as foreground and background classification problem has attracted much attention in the past decade. Most methods adopt mean shift or brute force search to perform object tracking on the generated probability map, which is obtained from the classification results; however, performing probabilistic object tracking on the probability map is almost unexplored. This paper proposes a novel observation model which is suitable to perform this task. The observation model considers both region and boundary cues on the probability map, and can be computed very efficiently by using the integral image data structure. Extensive experiments are carried out on several challenging image sequences, which include abrupt motion change, background clutter, partial occlusion, and significant appearance change. Quantitative experiments are further performed with several related trackers on a public benchmark dataset. The experimental results demonstrate the effectiveness of the proposed approach. Dawei Liang, Qingming Huang, Hongxun Yao, Shuqiang Jiang, Rongrong Ji, Wen Gao 0001 |
CVPR | 4 |
| 2010 | Multi-description of local interest point for partial-duplicate image retrievalabstractIn partial-duplicate image retrieval, images are commonly represented using Bag-of-visual-Words (BoW) built from image local features, such as SIFT. Therefore, the discriminative power of the local features is closely related with the BoW image representation and its performance in different applications. In this paper, we first propose a rotation-invariant Local Self-Similarity Descriptor (LSSD), which captures the internal geometric layouts in the local textural self-similar regions around interest points. Then we combine LSSD with SIFT to develop a multi-description of images for retrieving partial-duplicate. Finally, we formulate the Semi-Relative Entropy as the distance metric. Retrieval performance of this multi-description evaluated in the Oxford building dataset and an image corpus crawled from Google shows that the average precision achieves 11.1% and 2.8% improvement, respectively, comparing with state-of-the-art bundling feature. Liang Li 0003, Shuqiang Jiang, Qingming Huang |
ICIP | 2 |
| 2010 | A close-up detection method for moviesabstractClose-up (CU) is a photographic technique which tightly frames a person or an object. In movies, it is applied to guide audience attention and to evoke audience emotion. In this paper, we detect face CU, object CU, and lean of movies, which are widely used to romance emotions. A lean consists of shots in a sequence, with a close-up shot as focus. A set of features are extracted by considering movie making techniques and human attention for CU detection. The features are average saliency, color entropy, color variance, face height, skin area, and texture scales. These features are tested through statistical hypothesis test to be significantly discriminating for CUs. Then, Support Vector Machine (SVM) is applied on these features to detect face CU and object CU. Based on the face CU and object CU detection result, lean is further detected by investigating the changing of the face/object size. Lean detection is of challenge due to the technique of montage. We solve this problem through color similarity estimation and SIFT point matching. Experimental results on four full length movies verify the effectiveness of the proposed method. Min Xu 0001, Qingming Huang, Jesse S. Jin, Shuqiang Jiang, Changsheng Xu |
ICIP | 5 |
| 2010 | Fast copy detection based on Slice Entropy ScattergraphabstractWith the exponential growth of digital video resources, huge amount of videos are uploaded onto the Internet. Therefore, the Content Based Copy Detection (CBCD) issue becomes a hot research topic and has been extensively studied recently. However, most of the approaches lack the power to efficiently handle large data corpus while maintaining a good detection quality. In this paper, we propose a fast CBCD approach based on the Slice Entropy Scattergraph (SES). SES employs video spatio-temporal slices which can greatly decrease the storage and computational complexity. It is based on entropy and its deviation so as to preserve as much as the video information. Besides, SES takes advantage of a scattergraph which is succinct and efficient to plot the distribution of video content. To effectively describe SES, we introduce three descriptors: Projection Histograms, Shape Contexts and Polynomial Coefficients. The experiments on CIVR'07 Copy Detection Corpus and Video Transformation Corpus show the performance improvement of our approach both on efficiency and effectiveness. Shuqiang Jiang, Qingming Huang |
ICME | 3 |
| 2010 | Event based news video people classification and ranking using multimodality featuresabstractExisting research on news video analysis mainly concentrates on structure analysis, semantic concept detection, annotation and search. However, little work has been contributed to news video people community analysis, which is helpful for users to understand the event. In this paper, we propose a novel approach to classify the people appearing in the news video into different communities. In our approach, the people appearing in the news video are first identified by associating their faces with names. The faces are detected from the video frames, and the names are obtained from the text. Then, the people belonging to the same organization are clustered. After that, the relationships between these organizations are determined using sentiment analysis. The sentiment words are diverse in each news story and contain both positive and negative ones. However, we have news title, which is the summary of the story and the sentiment of which is clear, to help us to mine the relationships between the organizations. At last, social networks are built to classify those people/organizations into different classes, and the people/organizations are ranked in each community according to their influence. The main contributions of the paper are two folds: 1) we propose a novel approach to present the news video event according to communities; 2) we propose to use the sentiment analysis and social network to classify the news people/organizations. The experimental results on the selected news topics demonstrate that the proposed approach is effective. Chunxi Liu, Qingming Huang, Shuqiang Jiang, Changsheng Xu |
ICME | 3 |
| 2010 | Bridging the gap between objective score and subjective preference in video quality assessmentabstractNowadays, the issue of objective video quality assessment has been extensively studied. However, the human visual system (HVS) is the ultimate receiver for videos thus leading to a gap between objective scores calculated by computers and subjective preferences given by observers. In this paper, we focus on bridging this gap by introducing a psychological criterion called contrast effect. That is, because of the impression about the quality of previous frame still remaining in observers' minds, they tend to underestimate or overestimate the quality of the current one. Noticing this fact, we propose a video quality assessment system with an additional revision module to bridge the gap mentioned above. Firstly, the video is described by several representative clips with large entropy values. Then, we present Quality Words (including luminance, contrast, structure and spatio-temporal texture) to evaluate the quality of distorted video. To characterize the spatio-temporal texture, a new descriptor called Rotation Sensitive 3D Texture Pattern (RS-3D) is proposed. Finally, we revise the result in the revision module motivated by contrast effect. Experiments on VQEG Phase I FR-TV test dataset verify the effectiveness of our method. Qianqian Xu 0001, Li Su 0003, Shuqiang Jiang, Qingming Huang |
ICME | 5 |
| 2010 | Group Activity Recognition by Gaussian Processes EstimationabstractHuman action recognition has been well studied recently, but recognizing the activities of more than three persons remains a challenging task. In this paper, we propose a motion trajectory based method to classify human group activities. Gaussian Processes are introduced to represent human motion trajectories from a probabilistic perspective to handle the variability of people's activities in group. With respect to the relationships of persons in group activities, three discriminative descriptors are designed, which are Individual, Dual and Unitized Group Activity Pattern. We adopt the Bag of Words approach to solve the problem of unbalanced number of persons in different activities. Experiments are conducted on the human group-activity video database, and the results show that our approach outperforms the state-of-the-art. Zhongwei Cheng, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
ICPR | 4 |
| 2010 | Action Recognition Using Spatial-Temporal ContextabstractThe spatial-temporal local features and the bag of words representation have been widely used in the action recognition field. However, this framework usually neglects the internal spatial-temporal relations between video-words, resulting in ambiguity in action recognition task, especially for videos “in the wild”. In this paper, we solve this problem by utilizing the volumetric context around a video-word. Here, a local histogram of video-words distribution is calculated, which is referred as the “context” and further clustered into contextual words. To effectively use the contextual information, the descriptive video-phrases (ST-DVPs) and the descriptive video-cliques (ST-DVCs) are proposed. A general framework for ST-DVP and ST-DVC generation is described, and then action recognition can be done based on all these representations and their combinations. The proposed method is evaluated on two challenging human action datasets: the KTH dataset and the YouTube dataset. Experiment results confirm the validity of our approach. Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
ICPR | 4 |
| 2010 | Multiple Kernel Learning with High Order KernelsabstractPrevious Multiple Kernel Learning approaches (MKL) employ different kernels by their linear combination. Though some improvements have been achieved over methods using single kernel, the advantages of employing multiple kernels for machine learning are far from being fully developed. In this paper, we propose to use “high order kernels” to enhance the learning of MKL when a set of original kernels are given. High order kernels are generated by the products of real power of the original kernels. We incorporate the original kernels and high order kernels into a unified localized kernel logistic regression model. To avoid over-fitting, we apply group LASSO regularization to the kernel coefficients of each training sample. Experiments on image classification prove that our approach outperforms many of the existing MKL approaches. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
ICPR | 2 |
| 2010 | Adding Affine Invariant Geometric Constraint for Partial-Duplicate Image RetrievalabstractThe spring up of large numbers of partial-duplicate images on the internet brings a new challenge to the image retrieval systems. Rather than taking the image as a whole, researchers bundle the local visual words by MSER detector into groups and add simple relative ordering geometric constraint to the bundles. Experiments show that bundled features become much more discriminative than single feature. However, the weak geometric constraint is only applicable when there is no significant rotation between duplicate images and it couldn't handle the circumstances of image flip or large rotation transformation. In this paper, we improve the bundled features with an affine invariant geometric constraint. It employs area ratio invariance property of affine transformation to build the affine invariant matrix for bundled visual words. Such affine invariant geometric constraint can cope well with flip, rotation or other transformations. Experimental results on the internet partial-duplicate image database verify the promotion it brings to the original bundled features approach. Since currently there is no available public corpus for partial-duplicate image retrieval, we also publish our dataset for future studies. Qianqian Xu 0001, Shuqiang Jiang, Qingming Huang, Liang Li 0003 |
ICPR | 3 |
| 2010 | The third eye: mining the visual cognition across multi-language communitiesabstractExisting research work in the multimedia domain mainly focuses on image/video indexing, retrieval, annotation, tagging, re-ranking, etc. However, little work has been contributed to people's visual cognition. In this paper, we propose a novel framework to mine people's visual cognition across multi-language communities. Two challenges are addressed: the visual cognition representation for a specific language community, and the visual cognition comparison between different language communities. We call it "the third eye", which means that through this way people with different backgrounds can better understand the cognition of each other, and can view the concept more objectively to avoid culture conflict. In this study, we utilize the image search engine to mine the visual cognition of the different communities. The assumption is that the image semantic distribution over the search results can reflect the visual cognition of the community. When a user submits a text query, it is first translated into different languages, and fed into the corresponding image search engine ports to retrieve images from these communities. After retrieval, the obtained images are categorized into different semantic clusters automatically. Finally, inter semantic cluster ranking is employed to rank the semantic clusters according to their relationship to the query, and intra cluster ranking is used to rank the images according to their representativeness. The visual cognition difference among these language communities is achieved by comparing the different community image distributions over these semantic clusters. The experimental results are promising and show that the proposed visual cognition mining approach is effective. Chunxi Liu, Qingming Huang, Shuqiang Jiang, Changsheng Xu |
ACM Multimedia | 3 |
| 2010 | Nearest-neighbor classification using unlabeled data for real world image applicationabstractCurrently, Nearest-Neighbor approaches (NN) have been widely applied to real world image data mining. These approaches have the following three disadvantages: (i) the performance is inferior on small datasets; (ii) the performance of approximated nearest neighbor search will degrade for data with high dimensions; (iii) they are heavily dependent on the chosen feature and distance measure. To overcome these intrinsic weaknesses, we propose a novel Nearest-Neighbor method, which improves the original NN approaches from three aspects. Firstly, we propose a novel neighborhood similarity measure, where the similarity between test images and labeled images in the database is calculated jointly by the original image-to-image similarity and the average similarity of their neighboring unlabeled data. Secondly, we adopt the kernelized locality sensitive hashing to effectively conduct the nearest neighbor search for high dimensional data. Finally, to enhance the robustness of the method on different genres of images, we propose to fuse the discrimination power of different features by considering all the retrieved nearest neighbors via hashing systems using different features/kernels. Experimental result shows the advantage over traditional Nearest-Neighbor methods using the labeled data only. Even when the ratio of labeled data is very small, our method could also achieve remarkable results, thanks to the help of unlabeled data and multiple features. Shuhui Wang, Qingming Huang, Shuqiang Jiang, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2010 | S3MKL: scalable semi-supervised multiple kernel learning for image data miningabstractFor large scale image data mining, a challenging problem is to design a method that could work efficiently under the situation of little ground-truth annotation and a mass of unlabeled or noisy data. As one of the major solutions, semi-supervised learning (SSL) has been deeply investigated and widely used in image classification, ranking and retrieval. However, most SSL approaches are not able to incorporate multiple information sources. Furthermore, no sample selection is done on unlabeled data, leading to the unpredictable risk brought by uncontrolled unlabeled data and heavy computational burden that is not suitable for learning on real world dataset. In this paper, we propose a scalable semi-supervised multiple kernel learning method (S3MKL) to deal with the first problem. Our method imposes group LASSO regularization on the kernel coefficients to avoid over-fitting and conditional expectation consensus for regularizing the behaviors of different kernel on the unlabeled data. To reduce the risk of using unlabeled data, we also design a hashing system where multiple kernel locality sensitive hashing (MKLSH) are constructed with respect to different kernels to identify a set of "informative" and "compact" unlabeled training subset from a large unlabeled data corpus. Combining S3MKL with MKLSH, the method is suitable for real world image classification and personalized web image re-ranking with very little user interaction. Comprehensive experiments are conducted to test the performance of our method, and the results show that our method provides promising powers for large scale real world image classification and retrieval. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2010 | Vicept: link visual features to concepts for large-scale image understandingabstractOn noticing the paradox of visual polysemia and concept poly-morphism, this paper proposes a new perspective called "Vicept" to associate elementary visual features and cognitive concepts. Firstly, a carefully prepared large image dataset and associate concepts are established. Secondly, we extract local interest points as the ele-mentary visual features, cluster them into visual words, and use Fuzzy Concept Membership Updating (FCMU) to build the link between codebook and concept membership distributions. This bottommost feature is called "Vicept word". Then, the global level Vicept features are established to correlate concepts with (partial) images. Finally, we validate our Vicept approach and show its effectiveness in concept detection task. Our approach is independent of case-specific training data and thus can be extended to web-scale scenarios. Shuqiang Jiang, Liang Li 0003, Qingming Huang, Wen Gao 0001 |
ACM Multimedia | 2 |
| 2010 | Memory matrix: a novel user experience for home videoabstractNowadays, various efforts have sprung up aiming to automatically analyze home videos and provide users satisfactory experiences. In this paper, we present a novel user experience for home video called Memory Matrix, which could facilitate users to re-experience the joy of their memories, travelling along not only the time axis but also the space axis. In other words, the video clips (sub-shots) are organized both by taken times and taken locations, which further allows the user to browse home videos taken at similar locations. Moreover, given a specific query in Memory Matrix (row, column), it can also provide the user optional summaries along the time axis or space axis. The summarization scheme in this paper is based on a top-down interest score generation algorithm which automatically propagates the pre-labeled video level interest scores to sub-shot level interest scores. Firstly, the user is asked to provide interest scores to all the video sequences in the home video collection. Then, the video sequences are decomposed into sub-shots which are represented by keyframes. Consequently, we employ multi-scale spatial saliency analysis to remove the foregrounds and model the background scenes based on histogram of visual words. Finally, the interest scores are propagated from video level to sub-shot level by using gradient descent algorithm. Experimental results demonstrate the effectiveness, efficiency, and robustness of our framework. Qianqian Xu 0001, Guorong Li, Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 5 |
| 2010 | Building contextual visual vocabulary for large-scale image applicationsabstractNot withstanding its great success and wide adoption in Bag-of-visual Words representation, visual vocabulary created from single image local features is often shown to be ineffective largely due to three reasons. First, many detected local features are not stable enough, resulting in many noisy and non-descriptive visual words in images. Second, single visual word discards the rich spatial contextual information among the local features, which has been proven to be valuable for visual matching. Third, the distance metric commonly used for generating visual vocabulary does not take the semantic context into consideration, which renders them to be prone to noise. To address these three confrontations, we propose an effective visual vocabulary generation framework containing three novel contributions: 1) we propose an effective unsupervised local feature refinement strategy; 2) we consider local features in groups to model their spatial contexts; 3) we further learn a discriminant distance metric between local feature groups, which we call discriminant group distance. This group distance is further leveraged to induce visual vocabulary from groups of local features. We name it contextual visual vocabulary, which captures both the spatial and semantic contexts. We evaluate the proposed local feature refinement strategy and the contextual visual vocabulary in two large-scale image applications: large-scale near-duplicate image retrieval on a dataset containing 1.5 million images and image search re-ranking tasks. Our experimental results show that the contextual visual vocabulary shows significant improvement over the classic visual vocabulary. Moreover, it outperforms the state-of-the-art Bundled Feature in the terms of retrieval precision, memory consumption and efficiency. Shiliang Zhang, Qingming Huang, Gang Hua 0001, Shuqiang Jiang, Wen Gao 0001, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2010 | Affective Visualization and Retrieval for Music VideoabstractIn modern times, music video (MV) has become an important favorite pastime to people because of its conciseness, convenience, and the ability to bring both audio and visual experiences to audiences. As the amount of MVs is explosively increasing, it has become an important task to develop new techniques for effective MV analysis, retrieval, and management. By stimulating the human affective response mechanism, affective video content analysis extracts the affective information contained in videos, and, with the affective information, natural, user-friendly, and effective MV access strategies could be developed. In this paper, a novel integrated system (i.MV) is proposed for personalized MV affective analysis, visualization, and retrieval. In i.MV, we not only perform the personalized MV affective analysis, which is a challenging and insufficiently covered problem in current affective content analysis field, but also propose novel affective visualization to convert the abstract affective states intuitive and friendly to users. Based on the affective analysis and visualization, affective information based MV retrieval is achieved. Both comprehensive experiments and subjective user studies on a large MV dataset demonstrate that our personalized affective analysis is more effective than the previous algorithms. In addition, affective visualization is proved to be more suitable for affective information-based MV retrieval than the commonly used affective state representation strategies. Shiliang Zhang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2009 | Advertise gently - in-image advertising with low intrusivenessabstractThe new trend of online advertisement is in-image advertising, which is facing the risk of being intrusive. Several works have been done to reduce the intrusiveness. However, intrusiveness is a subjective concept and is difficult to be measured objectively. In this paper, by considering the fact that gentle advertising will not disturb audiences' attention too much but the intrusive ones will, we investigate the relationship between intrusiveness and audience attention. By experiment, we find that two aspects of attention will affect intrusiveness. Firstly, if the inserted advertisement covers the Region of Interest (ROI), it is truly very intrusive. Secondly, if the advertisement distracts audience attention from the original attending point, it is also very intrusive. We measure intrusiveness from the above two aspects. Using this measurement, we insert advertisements into online image collections gently. Given a pair of an image and an advertisement, we detect the suitable place, using attention analysis and visual consistency, to reduce intrusiveness. Given an image set and an advertisement set, we minimize the intrusiveness by searching for an optimal match. Experimental results verify the effectiveness of the proposed measurement of intrusiveness and of the advertising approach. Xuekan Qiu, Qingming Huang, Shuqiang Jiang, Changsheng Xu |
ICIP | 4 |
| 2009 | Spatial-temporal video browsing for mobile environment based on visual attention analysisabstractIn this paper, we propose an attention based method to provide convenient browsing experience on mobile devices. To achieve this, we extract video summary in both temporal and spatial domains. Firstly, an improved video attention analysis method is performed to obtain attention distribution of video. Then, temporal summarization is carried out by extracting attended frames from attended shots to make the summary representative and fluent. And spatial summary is further generated by detecting the Regions of Interest (ROIs) of temporal summary frames. At this stage, to improve the viewing experience, a curve fitting based smoothing method is proposed to avoid drastic change of ROIs' positions and sizes. Finally, we will obtain the spatial-temporal summary to be displayed on screens of mobile devices. Experimental results verify the effectiveness of this method. Xuekan Qiu, Shuqiang Jiang, Qingming Huang |
ICME | 2 |
| 2009 | Robust copy detection by mining temporal self-similaritiesabstractThis paper introduces a self-similarity matrix (SSM) based video copy detection scheme and a visual character-string (VCS) descriptor for SSM matching. SSM, which exploits the spatial and temporal information in a video clip, is extracted from exhaustive calculation of distances between the frames. The SSM based method treats the video clip as a whole and transforms the temporal self-similarity into a matrix. Moreover, by implementing the proposed VCS descriptor, the problem of SSM alignment failure and size variation can also be solved properly. Experimental evaluations based on CIVR07 copy detection corpus validate the effectiveness of the proposed solution. Qingming Huang, Shuqiang Jiang |
ICME | 3 |
| 2009 | Near-duplicate video matching with transformation recognitionabstractNowadays, the issue of near-duplicate video matching has been extensively studied. However, transformation, which is one of the major causes of near-duplicates, has been little discussed. In this paper, we focus on the fact that a certain kind of feature may per-form excellently to deal with one type of transformation while not be that good on another. We present a self-similarity matrix based near-duplicate video matching scheme with an additional transformation recognition module. By detecting the type of transformations, the near-duplicates can be treated with the 'best' feature which is decided experimentally. Thus, we obtain an enhanced matching result by employing the selected feature. Our work includes seven features and ten transformations respectively, and experimental results show the effectiveness of transformation recognition and the promotion it brings to boost the near-duplicate matching scheme. Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 2 |
| 2009 | Friend recommendation according to appearances on photosabstractUnlike the questionnaire based friend recommendation scheme used in Social Network Service (SNS) websites nowadays (e.g. online dating sites, online matchmaking sites), we focus on the fact that most of the online users may be interested in the strangers whose appearances are somehow attractive according to their own preferences. In this paper, we present a friend recommendation system based on the appearances on photos. The system is built upon 5000 portraits photos as source dataset with another 50 photos as training set. Once the user provides rating to several photos in the training set, we first build his/her appearance prefe-rence model based on face detection and multi-features cooperation. Then, the images in the source are ranked according to different features respectively. Finally, the results of multi-features are fused via the method of Borda count. The system is a useful complement to the conventional psychological tests based friend rec-ommendation scheme. It is easy to play with and of a lot of fun. Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 2 |
| 2009 | A framework for flexible summarization of racquet sports video using multiple modalities
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Liyuan Xing, Qixiang Ye, Wen Gao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2009 | Event Tactic Analysis Based on Broadcast Sports VideoabstractMost existing approaches on sports video analysis have concentrated on semantic event detection. Sports professionals, however, are more interested in tactic analysis to help improve their performance. In this paper, we propose a novel approach to extract tactic information from the attack events in broadcast soccer video and present the events in a tactic mode to the coaches and sports professionals. We extract the attack events with far-view shots using the analysis and alignment of web-casting text and broadcast video. For a detected event, two tactic representations, aggregate trajectory and play region sequence, are constructed based on multi-object trajectories and field locations in the event shots. Based on the multi-object trajectories tracked in the shot, a weighted graph is constructed via the analysis of temporal-spatial interaction among the players and the ball. Using the Viterbi algorithm, the aggregate trajectory is computed based on the weighted graph. The play region sequence is obtained using the identification of the active field locations in the event based on line detection and competition network. The interactive relationship of aggregate trajectory with the information of play region and the hypothesis testing for trajectory temporal-spatial distribution are employed to discover the tactic patterns in a hierarchical coarse-to-fine framework. Extensive experiments on FIFA World Cup 2006 show that the proposed approach is highly effective. Guangyu Zhu 0002, Changsheng Xu, Qingming Huang, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
IEEE Trans. Multim. | 5 |
| 2008 | Multiple Instance Boost Using Graph Embedding Based Decision Stump for Pedestrian Detection
Junbiao Pang, Qingming Huang, Shuqiang Jiang |
ECCV (4) | 3 |
| 2008 | Visual-aural attention modeling for talk show video highlight detectionabstractIn this paper, we propose a visual-aural attention modeling based video content analysis approach, which can be used to automatically detect the highlights of the popular TV program—talk show video. First, the visual and aural affective features are extracted to represent and model the human attention of highlight. For efficiency consideration, the adopted affective features are kept as few as possible. Then, a specific fusion strategy called ordinal-decision is used to combine the visual, aural attention models and form the attention curve for a video. This curve can reflect the change of human attention while watching TV. Finally, highlight segments are located at the peaks of the attention curve. Moreover, sentence boundary detection is used to refine the highlight boundaries in order to keep the segments’ integrality and fluency. This framework is extensible and flexible in integrating more affective features with a variety of fusion schemes. Experimental results demonstrate our proposed visual-aural attention analysis approach is effective for talk show video highlight detection. Guangyu Zhu 0002, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICASSP | 3 |
| 2008 | People re-detection using Adaboost with sift and color correlogramabstractPeople re-detection aims at performing re-identification of people who leave the scene and reappear after some time. This is an important problem especially in video surveillance scenarios. In this paper, we present a method of people re-detection within the context of visual sequence in single-camera setup. We consider re-detection as a binary classification problem, where both global and local descriptors are employed for training strong classifier on-line with Adaboost to distinguish a newly detected people as tracked or new occurrence. The strong classifier will be updated while match is ascertained. A predetermined classifier with well-chosen threshold is employed as assistant of training examples collection. We test the performance of our approach on 4 different scenes including 51 video sequences taken from the CAVIAR database and 4 video sequences shot by ourselves. The results show that our re-detection algorithm can robustly handle variations in illumination, pose, scale, and camera-view. Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICIP | 2 |
| 2008 | Object tracking using incremental 2D-LDA learning and Bayes inferenceabstractThe appearances of the tracked object and its surrounding background usually change during tracking. As for tracking methods using subspace analysis, fixed subspace basis tends to cause tracking failure. In this paper, a novel tracking method is proposed by using incremental 2D-LDA learning and Bayes inference. Incremental 2D-LDA formulates object tracking as online classification between foreground and background. It updates the row- or/and column-projected matrix efficiently. Based on the current object location and the prior knowledge, the possible locations of the object (candidates) in the next frame are predicted using simple sampling method. Applying 2D-LDA projection matrix and Bayes inference, candidate that maximizes the posterior probability is selected as the target object. Moreover, informative background samples are selected to update the subspace basis. Experiments are performed on image sequences with the object’s appearance variations due to pose, lighting, etc. We also make comparison to incremental 2D-PCA and incremental FDA. The experimental results demonstrate that the proposed method is efficient and outperforms both the compared methods. Guorong Li, Dawei Liang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICIP | 4 |
| 2008 | Fast and effective text detectionabstractText in images and videos is a significant cue for visual content understanding and retrieval. In this paper, we present a fast and effective approach to locate text lines even under complex background. First, our algorithm uses the stroke filter to calculate the stroke maps in horizontal, vertical, left-diagonal, right-diagonal directions. Then a 24- dimensional feature is extracted for each sliding window and a SVM is used to obtain rough text regions. The rough text regions are further refined through a group of rules. And candidate text lines were localized more accurately through projection profile of the refined text regions. Finally another SVM classifier based on a 6-dimensional feature is used to verify the candidate text lines. The experimental results on challenging databases show that this approach can fast and effectively detect and localize text lines. Weiqiang Wang 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICIP | 3 |
| 2008 | Pedestrian detection via logistic multiple instance boostingabstractPedestrian detection in still image should handle the large appearance and pose variations arising from the articulated structure and various clothing of human bodies as well as view points. So it is difficult to design effective classifier for this problem. In this paper, we address these variations in detection via multiple instance learning, specifically logistic multiple instance boosting (LMIB). In LMIB, a example is represented as a set of instances, which implicitly encode the variations. Giving different confidence to the instances in a bag, the LMIB will automatically reduce the influence of the variations at training stage. To obtain rapid detection speed, the LMIBs are grouped into the cascaded structure. The proposed detection algorithm is tested on MIT and INRIA human datasets where promising detection results are comparable with the baseline algorithms. Junbiao Pang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICIP | 3 |
| 2008 | Shot classification for action movies based on motion characteristicsabstractIn this paper, we propose a shot classification method for action movies. Considering that motion characteristic is very important for semantic movie analysis, and it contains abundant information in action movies, the structure tensor analysis is used for feature extraction due to its capability of representing both spatial and temporal characteristics of a shot. Firstly, the movie shots with known labels are decomposed into a set of overlapped fixed-length segments and their structure tensor histogram are computed. The labels of segments are identical to the shots they belong to. Then Adaboost is used to train the semantic classifier with these structure tensor histogram sets. In testing procedure, the unknown shot are decomposed in the same way, and feature vector of each segment is extracted and classified by the classifier. Finally, the label of the shot is generalized by the segment label voting scheme. Experimental results show that this scheme could effectively deal with multiple motion patterns within shots and promising results are achieved. Shuhui Wang, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICIP | 2 |
| 2008 | Lower attentive region detection for virtual content insertion in broadcast videoabstractVirtual Content Insertion (VCI) is an emerging application of video analysis. For VCI the spatial position is very important as improper placement will make the insertion intrusive. To choose the spatial position, we propose the notation of Lower Attentive Region (LAR) and provide a generic framework of LAR detection for broadcast video. An LAR is defined, from the cognition point of view, as a region of the video frame which attracts less audiencepsilas attention. It can be changed with little interruption to the main content of the original video. The proposed LAR detection framework includes both bottom-up and top-down modules and can be adapted to all types of videos. Finally we apply the proposed LAR detection approach to broadcast sports video by integrating domain knowledge. The Experiments on LAR detection and VCI in broadcast video demonstrate the effectiveness of the proposed method. Shuqiang Jiang, Qingming Huang, Changsheng Xu |
ICME | 2 |
| 2008 | Coarse-to-fine video text detectionabstractIn this paper, we propose an effective coarse-to-fine algorithm to detect text in video. Firstly, in coarse-detection section, stroke filter is employed to detect all candidate stroke pixels, and then a fast region growing method is developed to connect these pixels into regions which are further separated into candidate text lines by projection operation. Secondly, in fine-detection section, correct text regions are selected from candidate ones by support vector machine (SVM) model and stroke features, and text regions in multi-resolution are integrated. Finally, the result is optimized significantly according to temporal correlation information. Experimental results show that our algorithm achieves real-time performance and is robust for the variation of language, font, size, color and noise of text caused by low frame resolution in video. Guangyi Miao, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICME | 3 |
| 2008 | Spatial-temporal attention analysis for home videoabstractIn this paper, by considering the multiple spatial-temporal characteristic of visual perception system, we propose a novel home video attention analysis method. Firstly, each frame of the video is segmented into regions which are more informative than pixels and image blocks. Then the saliency of each region is analyzed by combining static, motion and location attentions. Finally a region based saliency map is generated for each frame, and an attention score curve is obtained for the video clip by combining attention scores of all regions in each frame. Both of them can be utilized in wide applications. This method takes advantage of the properties of human visual perception and can well present the attention information of home videos. Experimental results show the effectiveness of this approach. Xuekan Qiu, Shuqiang Jiang, Qingming Huang, Longbing Cao |
ICME | 2 |
| 2008 | A pixel-wise local information-based background subtraction approachabstractBackground subtraction is a widely used method for moving object detection in computer vision. It is usually applied in video surveillance systems. There are two major kinds of background subtraction approaches: pixel-based and blockbased. Yet there are three problems that can not be simultaneously solved by either method: the robustness to illumination changes, the effectiveness in suppressing shadows, and the smoothness of foreground’s boundary. In order to solve these problems, a pixel-wise local information-based background subtraction method is proposed in this paper. In the proposed method, Gabor filters are performed to extract the spatial feature vectors for each pixel from the source image sequence. Then, the spatial feature vectors are modeled by Gaussian Mixture Model, and then moving objects are detected. Experiments show the validity of the proposed method. Shuqiang Jiang, Qingming Huang |
ICME | 2 |
| 2008 | Affective MTV analysis based on arousal and valence featuresabstractNowadays, MTV has become an important favorite pastime to modern people because of its conciseness, convenience to play and the characteristic that can bring both audio and visual experiences to audiences. In this paper, we propose an affective MTV analysis framework, which realizes MTV affective state extraction, representation and clustering. Firstly, affective features are extracted from both audio and visual signals. Then, the affective state of each MTV is modeled with 2D dimensional affective model and visualized in the Arousal-Valence space. Finally the MTVs having similar affective states are clustered into same categories. The validity of proposed framework is proved by subjective user study. The comparisons between our selected features and those in related work prove that our features improve the performance by a significant margin. Shiliang Zhang, Qi Tian 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICME | 3 |
| 2008 | Human reappearance detection based on on-line learningabstractMany video surveillance applications require detecting human reappearances in a scene monitored by a camera or over a network of cameras. This is the human reappearance detection (HRD) problem. Studying this problem is important for analyzing a surveillance scenario at semantic level. In this paper, we propose a novel online learning framework for solving HRD problem. Both generative model and discriminative model are employed in this framework and a voting scheme is presented to fuse the decisions of both models for determining whether a just entered person is one of those who have shown up, i.e. whether a reappearance happens. Both models will be updated based on mistake-driven online learning strategy. Our experimental results show that the adopted online learning framework not only improves the reappearance detection accuracy but also achieves high robustness in various surveillance scenes. Yizhou Wang 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICPR | 3 |
| 2008 | Effective scene matching with local feature representativesabstractScene matching measures the similarity of scenes in photos and is of central importance in applications where we have to properly organize large amount of digital photos by scene categories. In this paper, we present a novel scene matching method using local features representatives. For a given image, its scene is compactly represented as a set of cluster centers, called local feature representatives, where the clusters are obtained using the affinity propagation (AP) algorithm to aggregate local features according to their spatial closeness and appearance similarity. The similarity of scenes in two images is then measured by a modified Earth Mover Distance (EMD) between their corresponding sets of local feature representatives. Empirical experiments on real world photos shows that our method is comparable to the state-of-the-arts. Shugao Ma, Weiqiang Wang 0001, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICPR | 4 |
| 2008 | Matching images more efficiently with local descriptorsabstractImage matching is a fundamental task for many applications of computer vision. Today it is very popular to represent two matched images as two bags of local descriptors, and the classic RANSAC based matching procedure is always exploited in the task. In this paper, we present a much efficient image matching approach based on sets of any local descriptors. A block-to-block strategy is devised to speed up the establishment of local correspondences. Additionally, the weighted RANSAC (w-RANSAC) technique is proposed to make the search of optimal global models converge faster. Comparative experiments with the RANSAC based paradigm show our approach can not only generate more accurate correspondences, but also double the matching speed. Dong Zhang 0001, Weiqiang Wang 0001, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICPR | 4 |
| 2008 | Naming faces in broadcast news video by image googleabstractNaming faces is important for news videos browsing and indexing. Although some research efforts have been contributed to it, they only use the concurrent information between the face and name or employ some clues as features and use simple heuristic method or machine learning approach to finish the task. They use little extra knowledge about the names and faces. Different from previous work, in this paper we present a novel approach to name the faces by exploring extra knowledge obtained from image google. The behind assumption is that the faces of those important persons will turn out many times in the web images and could be retrieved from image google easily. Firstly, faces are detected in the video frames; and the name entities of candidate persons are extracted from the textual information by automatic speech recognition and close caption detection. Then, these candidate person names are used as queries to find the name related person images through image google. After that, the retrieved result is analyzed and some typical faces are selected through feature density estimation. Finally, the detected faces in the news video are matched with the faces selected from the result returned by image google to label each face. Experimental results on MSNBC news and CNN news demonstrate that the proposed approach is effective. Chunxi Liu, Shuqiang Jiang, Qingming Huang |
ACM Multimedia | 2 |
| 2008 | A generic virtual content insertion system based on visual attention analysisabstractThis paper presents a generic Virtual Content Insertion (VCI) system based on visual attention analysis. VCI is an emerging application of video analysis and has been used in video augmentation and advertisement insertion. There are three critical issues for a VCI system: when (time), where (place) and how (method) to insert the Virtual Content (VC) into the video. Our system selects the insertion time and place by performing temporal and spatial attention analysis, which predicts the attention change along time and the attended region over space. In order to enable the inserted VC to be noticed by audience while not to interrupt the audience's viewing experience to the original content, the VC should be inserted at the time when the video content attracts much audience attention and at the place where attracts less. Dynamic insertion is performed by using Global Motion Estimation (GME) and affine transformation. Our VCI system is able to obtain an optimal balance between the notice of the VC by audience and disruption of viewing experience to the original content. Extensive subjective evaluations based on user study on the VCI result have verified the effectiveness of the system. Shuqiang Jiang, Qingming Huang, Changsheng Xu |
ACM Multimedia | 2 |
| 2008 | i.MTV: an integrated system for mtv affective analysisabstractIn modern time, MTV has become an important favorite pastime to people because of its conciseness, convenience and the ability to bring both audio and visual experiences to audiences. It has become an significant task to develop new techniques for natural, user-friendly, and effective MTV access. In this demo, an integrated system (i.MTV) is constructed for MTV Affective Analysis, Visualization, Retrieval, and User Profile Analysis. We not only perform the effective MTV affective analysis, but also propose novel Affective Visualization techniques to make the abstract affective states intuitive and friendly to users. Based on the affective analysis and visualization, MTV affective retrieval and management are achieved. Furthermore, novel methods are proposed for user affective preferences analysis and MTV recommendation. Shiliang Zhang, Qingming Huang, Qi Tian 0001, Shuqiang Jiang, Wen Gao 0001 |
ACM Multimedia | 4 |
| 2008 | Unsupervised texture classification: Automatically discover and classify texture patterns
Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
Image Vis. Comput. | 3 |
| 2007 | Mean-Shift Blob Tracking with Adaptive Feature Selection and Scale AdaptationabstractWhen the appearances of the tracked object and surrounding background change during tracking, fixed feature space tends to cause tracking failure. To address this problem, we propose a method to embed adaptive feature selection into mean shift tracking framework. From a feature set, the most discriminative features are selected after ranking these features based on their Bayes error rates, which are estimated from object and background samples. For the selected features, a criterion is proposed to evaluate their stability for tracking and to guide feature reselection. The selected features are used to generate a weight image, in which mean shift is employed to locate the object. Moreover, a simple yet effective scale adaptation method is proposed to deal with object changing in size. Experiments on several video sequences show the effectiveness of the proposed method. Dawei Liang, Qingming Huang, Shuqiang Jiang, Hongxun Yao, Wen Gao 0001 |
ICIP (3) | 3 |
| 2007 | Monocular Tracking 3D People By Gaussian Process Spatio-Temporal Variable ModelabstractTracking 3D people from monocular video is often poorly constrained. To mitigate this problem, prior knowledge should be exploited. In this paper, the Gaussian process spatio-temporal variable model (GPSTVM), a novel dynamical system modeling method is proposed for learning human pose and motion priors. The GPSTVM provides a low dimensional embedding of human motion data, with a smooth density function that provides higher probability to the poses and motions close to the training data. The low dimensional latent space is optimized directly to retain the spatio-temporal structure of the high dimensional pose space. After the prior on human pose is learned, the particle filtering can be used tracking articulated human pose; particle filtering propagates over time in the embedding space, avoiding the curse of dimensionality. Experiments demonstrate that our approach tracks 3D people accurately. Junbiao Pang, Laiyun Qing, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICIP (5) | 4 |
| 2007 | Mining Information of Attack-Defense Status from Soccer Video Based on Scene AnalysisabstractVideo content is always huge by itself with abundant information. Extracting explicit semantic information has been extensively investigated such as object detection, structure analysis and event detection. However, little work has been devoted on the problem of discovering global or inexplicit information from the huge video stream. As an implementation in this topic, this paper proposes a solution to mining the statistical global attack-defense status information from soccer video by scene analysis. Semantic scene information of play field detection, view classification, midline detection and global motion are extracted as the mid level information, and then they are fed into the finite state machine based status mining model to generate the statistical results, which will be of much usefulness for users. Experimental results reveal the feasibility of the method and more research work on the topic of discovering high-level inexplicit information from video are expected. Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICME | 1 |
| 2007 | Generating Video Sequence from Photo Image for Mobile Screens by Content AnalysisabstractTo bridge the gap between the high resolution digital images and limited display capability of mobile devices, this paper proposes a method to automatically transform static images to dynamic video clips to fit the requirement of mobile devices. Face detection and attention region detection techniques are applied to extract important regions that users may be interested in. Then a procedure of video resolution adaptation is performed for different screen size and aspect ratio. An algorithm to imitate camera motions is applied to visit the detected important regions to generate the motion effects. Finally the video clip is encoded with AVS-M standard as an implementation to be displayed on different mobile platforms. The solution provided in this paper can enable users viewing photographs in a dynamic mode using mobile devices that could not only glance over the whole picture of the image automatically, but also enjoy the important details of the photo. Furthermore, vivid dynamic effect enhances the viewing experience of users. Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICME | 1 |
| 2007 | A Real-Time Score Detection and Recognition Approach for Broadcast Basketball VideoabstractFor broadcast sports video, score information is an effective mid-level representation to facilitate high-level video content analysis. In this paper, we propose a real-time approach to detect score region and recognize scores in broadcast basketball video. First, score region is automatically detected using frame difference and texture information without any prior knowledge. Then, score digit is recognized using a coarse-to-fine scheme based on score spatial structure, temporal correlation and changing rules. Compared with traditional video text recognition method, our approach is computing-insensitive and independent of digit model training. Experimental results show that our approach achieves real-time performance and is robust for the variation of digit font, size, color and noise caused by low frame resolution in broadcast basketball video. Guangyi Miao, Guangyu Zhu 0002, Shuqiang Jiang, Qingming Huang, Changsheng Xu, Wen Gao 0001 |
ICME | 3 |
| 2007 | The Demo: A Real-Time Score Detection and Recognition Approach in Broadcast Basketball Sports VideoabstractFor broadcast sports video, score information is an effective mid-level representation to facilitate high-level video content analysis. In this paper, we propose a real-time approach to detect score region and recognize scores in broadcast basketball sports video. The flow chart of our proposed approach has two major modules: score detection and score recognition. Using our approach, we can locate score regions automatically and reliably without detecting all the texts in the frame, and recognize the scores without training process. The running speed of our algorithm is fast. Guangyi Miao, Guangyu Zhu 0002, Shuqiang Jiang, Qingming Huang, Changsheng Xu, Wen Gao 0001 |
ICME | 3 |
| 2007 | An Effective Local Invariant Descriptor Combining Luminance and Color InformationabstractExtraction of stable local invariant features is very important in many computer vision applications, such as image matching, object recognition and image retrieval. Most existing local invariant features mainly characterize luminance information, and neglect color information. In this paper, we present a new local invariant descriptor characterizing both of them, which combines three photometric invariant color descriptors with the famous SIFT descriptor. To reduce the dimension of the combined high-dimensional invariant feature the principal component analysis (PCA) is used. Our experiments show the proposed local descriptor through combining luminance and color information outperforms the descriptors that only utilize a single category of information, and combining the three color feature representations is more effective than only one. Dong Zhang 0001, Weiqiang Wang 0001, Wen Gao 0001, Shuqiang Jiang |
ICME | 4 |
| 2007 | Highlight Ranking for Racquet Sports Video in User Attention Subspaces Based on Relevance FeedbackabstractIn this paper, we propose a method to rank the highlights of broadcast racquet sports videos. Compared with previous work, we integrate relevance feedback into highlight ranking framework to effectively capture the user's interest in attention subspaces and generate personalized ranking result. First, we establish three user attention subspaces and extract audio, visual, temporal affective features to represent the human perception of highlight in each subspace. Then, the highlight ranking models are constructed using support vector regression (SVR) for the three subspaces respectively. Finally, the three submodels are linearly combined to generate the final ranking model. Relevance feedback technique is employed to adjust the weights of each submodel to obtain the result which is suitable to the user's preference. Experimental results demonstrate our approach is effective. Guangyu Zhu 0002, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICME | 3 |
| 2007 | Region-based visual attention analysis with its application in image browsing on small displaysabstractVisual attention has been a hot research point for many years and many new applications are emerging especially for wireless multimedia services. In this paper a novel region-based visual attention is proposed to detect the Regions of Interest (ROI) of images. In the proposed method, density based image segmentation is first performed by regarding region as the perceptive unit, which makes the model robust to the scale of ROIs and contains more perceptive information. To generate region saliency map to detect ROI, global effect and contextual difference are covered in the form of distance factor and adjacency factor respectively. Since different ROIs may have different importance for different purposes, a ROI ranking algorithm is designed for browsing large images on small displays. Experimental results and evaluation reveal that our method works effectively to detect ROIs from images and the users are satisfied with the browsing sequence on small displays. Shuqiang Jiang, Qingming Huang, Changsheng Xu, Wen Gao 0001 |
ACM Multimedia | 2 |
| 2007 | Trajectory based event tactics analysis in broadcast sports videoabstractMost of existing approaches on event detection in sports video are general audience oriented. The extracted events are then presented to the audience without further analysis. However, professionals, such as soccer coaches, are more interested in the tactics used in the events. In this paper, we present a novel approach to extract tactic information from the goal event in broadcast soccer video and present the goal event in a tactic mode to the coaches and sports professionals. We first extract goal events with far-view shots based on analysis and alignment of web-casting text and broadcast video. For a detected goal event, we employ a multi-object detection and tracking algorithm to obtain the players and ball trajectories in the shot. Compared with existing work, we proposed an effective tactic representation called aggregate trajectory which is constructed based on multiple trajectories using a novel analysis of temporal-spatial interaction among the players and the ball. The interactive relationship with play region information and hypothesis testing for trajectory temporal-spatial distribution are exploited to analyze the tactic patterns in a hierarchical coarse-to-fine framework. The experimental results on the data of FIFA World Cup 2006 are promising and demonstrate our approach is effective. Guangyu Zhu 0002, Qingming Huang, Changsheng Xu, Yong Rui, Shuqiang Jiang, Wen Gao 0001, Hongxun Yao |
ACM Multimedia | 5 |
| 2006 | Extracting Story Units in Sports Video Based on Unsupervised Video Scene ClusteringabstractMany sports videos such as archery, diving and tennis have repetitive structure patterns. They are reliable clues to generate highlights, summarization and automatic annotation. In this paper, we present a novel approach to analyze these structure patterns in sports video to extract story units. First, an unsupervised scene clustering method for sports video is adopted to automatically categorize the video shots into several disparate scenes. Then, the clustering results are modeled by a transition matrix. Finally, the key scene shots are detected to analyze the structure patterns and extract the story units. Experimental results on several types of broadcast sports video demonstrate that our approach is effective Chunxi Liu, Qingming Huang, Shuqiang Jiang, Weigang Zhang |
ICME | 3 |
| 2006 | Highlight Summarization in Sports Video Based on Replay DetectionabstractHighlight summarization technology has been studied widely in sports video analysis. In this paper, we propose a highlight summarization system based on replays. First the replay clips in the sports video are extracted as the highlight candidates. Then the features including audio energy and motion activity are employed to rank the arousal level of the replay clips. Finally we model the highlight with the arousal rank to generate summarization. The contribution of this paper concentrates on two aspects. Firstly, Event-Replay (ER) structure is proposed and some new features are employed to represent the arousal levels of ER for general sports video. Secondly a novel highlight model is proposed considering the inter-relation of ERs. The experiments evaluate the rationality of the system. 1. Shuqiang Jiang, Qingming Huang, Guangyu Zhu 0002 |
ICME | 2 |
| 2006 | An effective method to detect and categorize digitized traditional Chinese paintings
Shuqiang Jiang, Qingming Huang, Qixiang Ye, Wen Gao 0001 |
Pattern Recognit. Lett. | 1 |
| 2005 | Playfield Detection Using Adaptive GMM and Its ApplicationabstractPlayfield detection is a key step in sports video content analysis, since many semantic clues could be inferred from it. In this paper we propose an adaptive GMM based algorithm for playfield detection. Its advantages are twofold. First, it can update model parameters by the incremental expectation maximization (IEM) algorithm, which enables the model to adapt to the playfield variation with time; Second, online training is performed, which saves buffer for training samples. Then, the playfield detection results are applied in recognizing the key zone of the current playfield in soccer video, in which a fast algorithm based on playfield contour and least square is proposed. Experimental results show that the proposed algorithms are encouraging. Yang Liu 0006, Shuqiang Jiang, Qixiang Ye, Wen Gao 0001, Qingming Huang |
ICASSP (2) | 2 |
| 2005 | Video2Cartoon: generating 3D cartoon from broadcast soccer videoabstractIn this demonstration, a prototype system for generating 3D cartoon from broadcast soccer video is proposed. This system takes advantage of computer vision (CV) and computer graphics (CG) techniques to provide users new experience that can not be obtained from original video. Firstly, it uses CV techniques to obtain 3D positions of the players and ball. Then, CG techniques are applied to model the playfield, players, and ball. Finally, 3D cartoon is generated. Our system allows users to watch the game at any point of view using a 3D viewer based on OpenGL. Dawei Liang, Yang Liu 0006, Qingming Huang, Guangyu Zhu 0002, Shuqiang Jiang, Zhebin Zhang, Wen Gao 0001 |
ACM Multimedia | 5 |
| 2005 | Exciting event detection in broadcast soccer video with mid-level description and incremental learningabstractIn this paper, we propose a method for exciting event detection in broadcast soccer video with mid-level description and SVM-based incremental learning. In the method, video frames are firstly classified and grouped into views in terms of low-level playfield features. Mid-level description including view label, motion descriptor and shot descriptor are then extracted to present the characteristics of a view. By using the fixed temporal structure of views, SVM classification models are constructed to detected exciting events in a soccer match. In the view classification and event detection procedures, SVM-based incremental learning method is explored to improve the extensibility of view classification and event detection. Experiments on real soccer video programs demonstrate encouraging results. Qixiang Ye, Qingming Huang, Wen Gao 0001, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2005 | Visual Ontology Construction for Digitized Art Image Retrieval
Shuqiang Jiang, Qingming Huang, Tiejun Huang 0001, Wen Gao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2004 | A new method to segment playfield and its applications in match analysis in sports videoabstractWith the growing popularity of digitized sports video, automatic analysis of them need be processed to facilitate semantic summarization and retrieval. Playfield plays the fundamental role in automatically analyzing many sports programs. Many semantic clues could be inferred from the results of playfield segmentation. In this paper, a novel playfield segmentation method based on Gaussian mixture models (GMMs) is proposed. Firstly, training pixels are automatically sampled from frames. Then, by supposing that field pixels are the dominant components in most of the video frames, we build the GMMs of the field pixels and use these models to detect playfield pixels. Finally region-growing operation is employed to segment the playfield regions from the background. Experimental results show that the proposed method is robust to various sports videos even for very poor grass field conditions. Based on the results of playfield segmentation, match situation analysis is investigated, which is also desired for sports professionals and longtime fanners. The results are encouraging. Shuqiang Jiang, Qixiang Ye, Wen Gao 0001, Tiejun Huang 0001 |
ACM Multimedia | 1 |
| 2004 | An Ontology-based Approach to Retrieve Digitized Art ImagesabstractAlthough much progress has been made, current low-level based visual information retrieval technology does not allow users to formulate queries through high-level semantics. More and more digitized art images appear on the Internet, and techniques need to be established on how to organize and retrieve them. In this work, a framework for retrieving art images using an ontology-based method is introduced. The proposed ontology describes images in various aspects. Non-objectionable semantics are first introduced, and how to express these semantics is given. Concepts in the ontology could be automatically derived. The retrieval scheme makes users more naturally find visual information and experimental implementation demonstrates good potential on retrieving art images in a human-centered manner. Shuqiang Jiang, Tiejun Huang 0001, Wen Gao 0001 |
Web Intelligence | 1 |