Hao Zhang 0047

dblp:55/2270-47 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-1664-8739ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module Learning
abstract
We introduce PKR-QA (Procedural Knowledge Reasoning Question Answering), a new benchmark for question answering over procedural tasks that require structured reasoning. PKR-QA is constructed semi-automatically using a procedural knowledge graph (PKG), which encodes task-specific knowledge across diverse domains. The PKG is built by curating and linking information from the COIN instructional video dataset and the ontology, enriched with commonsense knowledge from ConceptNet and structured outputs from Large Language Models (LLMs), followed by manual verification. To generate question-answer pairs, we design graph traversal templates where each template is applied systematically over PKG. To enable interpretable reasoning, we propose a neurosymbolic approach called Knowledge Module Learning (KML), which learns procedural relations via neural modules and composes them for structured reasoning with LLMs. Experiments demonstrate that this paradigm improves reasoning performance on PKR-QA and enables step-by-step reasoning traces that facilitate interpretability.
Thanh-Son Nguyen 0001, Tzeh Yuan Neoh, Hao Zhang 0047, Ee Yeo Keat, Basura Fernando
AAAI4
2026 Improving Temporal Action Segmentation via Constraint-Aware Decoding
Ee Yeo Keat, Debaditya Roy, Hao Zhang 0047, Basura Fernando
ICPR (11)4
2026 PointTFA$^{m}$: Multi-Modal, Training-Free Adaptation for Point Cloud Understanding
Jinmeng Wu, Youxiang Hu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong
IEEE Trans. Multim.4
2025 Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment
abstract
Open-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understanding. However, the performance of current methods on novel predicates remains significantly inferior to that on base categories. We attribute this discrepancy to two primary factors: (1) A significant task misalignment between the Visual Relation Detection (VRD) task and the pre-trained models’ visual feature extractors, which are often designed for tasks like video-text retrieval and image-text retrieval, resulting in poor generalization to the novel set. (2) The relatively small size and limited vocabulary of open-vocabulary datasets, which create a substantial gap between base and novel predicates. Consequently, text prompts trained on the base set fail to generalize effectively to the novel set. To address these issues, we propose two improvement measures: (1) We decompose base and novel relations into actional and spatial patterns and introduce an innovative text prompt learning method that leverages the shared patterns between base and novel relations. (2) We develop a relation probability adjustment mechanism that utilizes reliable base relation predictions to adjust the probabilities of relations in novel classes by considering their overlaps in either actional or spatial contents. Experimental results on the benchmark dataset demonstrate significant performance improvements.
Ming Pei, Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Jinmeng Wu, Basura Fernando, Xun Yang 0001
ICASSP4
2025 SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
abstract
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video grounding, which segments object regions based on natural language descriptions. However, most existing approaches tackle these tasks in isolation, limiting progress toward unified, referentially grounded video interaction. We identify a key bottleneck in the lack of high-quality, unified video instruction data and a comprehensive benchmark for evaluating referentially grounded video chat. To address these challenges, we contribute in three core aspects: dataset, model, and benchmark. First, we introduce SAMA-239K, a large-scale dataset comprising 15K videos specifically curated to enable joint learning of video referring understanding, grounding, and multi-turn video chat. Second, we propose the SAMA model, which incorporates a versatile spatio-temporal context aggregator and a Segment Anything Model to jointly enhance fine-grained video comprehension and precise grounding capabilities. Finally, we establish SAMA-Bench, a meticulously designed benchmark consisting of 5,067 questions from 522 videos, to comprehensively evaluate the integrated capabilities of Video LMMs in multi-turn, spatio-temporal referring understanding and grounded dialogue. Extensive experiments and benchmarking results show that SAMA not only achieves strong performance on SAMA-Bench but also sets a new state-of-the-art on general grounding benchmarks, while maintaining highly competitive performance on standard visual understanding benchmarks.
Ye Sun 0004, Hao Zhang 0047, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang 0001
NeurIPS2
2025 Deduce and Select Evidences with Language Models for Training-Free Video Goal Inference
abstract
We introduce ViDSE, a Video framework that Deduce and Selects visual Evidence for training-free video goal inference using language models. Unlike approaches that directly apply vision-language models (VLM) or combine VLM+LLM to process dense video visuals, ViDSE explicitly selects relevant visual evidence (e.g., frames) based on the hypothesis deduced by the LLM. This approach not only im-proves accuracy but also reveals the logical process behind the model's decisions, enhancing explainability. Our exper-iments demonstrate that this selection process significantly reduces ambiguity in the subsequent inference reasoning stage and outperforms VLM-only and VLM+LLM models on goal inference tasks such as CrossTask and COIN. We further validate ViDSE’ s generalizability and robustness on action recognition tasks, such as ActivityNet and UCF-101, under training-free and open-vocabulary conditions. We observe that ViDSE easily generalizes to other video tasks (e.g., action recognition) requiring filtering of redundant and irrelevant information.
Ee Yeo Keat, Hao Zhang 0047, Alexander Matyasko, Basura Fernando
WACV2
2024 PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models
Jinmeng Wu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong
IJCAI3
2024 RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
Hao Zhang 0047, Ee Yeo Keat, Basura Fernando
ACM Multimedia1
2024 UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
abstract
Image segmentation is a crucial vision task that groups pixels within an image into semantically meaningful segments, which is pivotal in obtaining a fine-grained understanding of real-world scenes. However, an increasing privacy concern exists regarding training large-scale image segmentation models on unauthorized private data. In this work, we exploit the concept of unlearnable examples to make images unusable to model training by generating and adding unlearnable noise into the original images. Particularly, we propose a novel Unlearnable Segmentation (UnSeg) framework to train a universal unlearnable noise generator that is capable of transforming any downstream images into their unlearnable version. The unlearnable noise generator is finetuned from the Segment Anything Model (SAM) via bilevel optimization on an interactive segmentation dataset towards minimizing the training error of a surrogate model that shares the same architecture with SAM (but trains from scratch). We empirically verify the effectiveness of UnSeg across 6 mainstream image segmentation tasks, 10 widely used datasets, and 7 different network architectures, and show that the unlearnable images can reduce the segmentation performance by a large margin. Our work provides useful insights into how to leverage foundation models in a data-efficient and computationally affordable manner to protect images against image segmentation models.
Ye Sun 0004, Hao Zhang 0047, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang 0001
NeurIPS2
2023 Knowledge driven weights estimation for large-scale few-shot image recognition
Jingjing Chen 0001, Linhai Zhuo, Zhipeng Wei 0001, Hao Zhang 0047, Huazhu Fu, Yu-Gang Jiang 0001
Pattern Recognit.4
2022 Group Contextualization for Video Recognition
abstract
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that decomposes the feature channels into several groups and separately refines them with different axial contexts in parallel. We refer this lightweight feature calibration as group contextualization (GC). Specifically, we design a family of efficient element-wise calibrators, i.e., ECal-G/S/T/L, where their axial contexts are information dynamics aggregated from other axes either globally or locally, to contextualize feature channel groups. The GC module can be densely plugged into each residual layer of the off-the-shelf video networks. With little computational overhead, consistent improvement is observed when plugging in GC on different networks. By utilizing calibrators to embed feature with four different kinds of contexts in parallel, the learnt representation is expected to be more resilient to diverse types of activities. On videos with rich temporal variations, empirically GC can boost the performance of 2D-CNN (e.g., TSN and TSM) to a level comparable to the state-of-the-art video networks. Code is available at https://github.com/haoyanbin918/Group-Contextualization.
Yanbin Hao, Hao Zhang 0047, Chong-Wah Ngo, Xiangnan He 0001
CVPR2
2022 Adaptive Temporal Grouping for Black-box Adversarial Attacks on Videos
abstract
Deep-learning based video models, which have remarkable performance on action recognition tasks, are recently proved to be vulnerable to adversarial samples, even those generated in the black-box setting. However, these black-box attack methods are insufficient to attack videos models in real-world applications due to the requirement of lots of queries. To this end, we propose to boost the efficiency of black-box attacks on video recognition models. Although videos carry rich temporal information, they include redundant spatial information from adjacent frames. This motivates us to introduce the adaptive temporal grouping (ATG) method, which groups video frames by the similarity of their features extracted from the ImageNet-pretrained image model. By selecting one key-frame from each group, ATG helps any black-box attack methods to optimize the adversarial perturbations over key-frames instead of all frames, where the estimated gradient of key-frame is shared with other frames in each group. To balance the efficiency and precision of estimated gradients, ATG adaptively adjusts the group number by the magnitude of the current perturbation and the current query number. Through extensive experiments on the HMDB-51 dataset and the UCF-101 dataset, we demonstrate that ATG can significantly reduce the number of queries by more than 10% for the targeted attack.
Zhipeng Wei 0001, Jingjing Chen 0001, Hao Zhang 0047, Linxi Jiang, Yu-Gang Jiang 0001
ICMR3
2022 Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure Preservation
abstract
Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely integrate them into MLP-Mixer for feature contextualization, leading to a novel model MC-MLP. To facilitate unsupervised learning, we investigate three kinds of data structures, including clusters, local neighborhood similarity structure, and inter/intra-class variations, and design a multi-objective task to train MC-MLP. These data structures show high complementarities in hash code learning. We conduct extensive experiments using three video retrieval benchmark datasets, demonstrating that our MCMSH not only boosts the performance of the backbone MLP-Mixer significantly but also outperforms the competing methods notably. Code is available at: https://github.com/haoyanbin918/MCMSH.
Yanbin Hao, Jingru Duan, Hao Zhang 0047, Bin Zhu 0006, Peng Yuan Zhou, Xiangnan He 0001
ACM Multimedia3
2022 Hierarchical Hourglass Convolutional Network for Efficient Video Classification
abstract
Videos naturally contain dynamic variation over the temporal axis, which will result in the same visual clues (e.g., semantics, objects) changing their scale, position, and perspective patterns between adjacent frames. A primary trend in video CNN is adopting spatial-2D convolution for spatial semantics and temporal-1D convolution for temporal dynamics. Though the direction achieves a favorable balance between efficiency and efficacy, it suffers from misalignment of visual clues with large displacements. Particularly, rigid temporal convolution would fail to capture correct motions when a specific target moves out of the reception field of temporal convolution between adjacent frames.
Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Shuo Wang 0008, Xiangnan He 0001
ACM Multimedia3
2022 Parameterization of Cross-token Relations with Relative Positional Encoding for Vision MLP
abstract
Vision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixing layers to capture cross-token interactions, as opposed to the multi-head self-attention mechanism used by Transformers. However, the heavily parameterized token-mixing layers naturally lack mechanisms to capture local information and multi-granular non-local relations, thus their discriminative power is restrained. To tackle this issue, we propose a new positional spacial gating unit (PoSGU). It exploits the attention formulations used in the classical relative positional encoding (RPE), to efficiently encode the cross-token relations for token mixing. It can successfully reduce the current quadratic parameter complexity O(N2) of vision MLPs to $O(N)$ and O(1). We experiment with two RPE mechanisms, and further propose a group-wise extension to improve their expressive power with the accomplishment of multi-granular contexts. These then serve as the key building blocks of a new type of vision MLP, referred to as PosMLP. We evaluate the effectiveness of the proposed approach by conducting thorough experiments, demonstrating an improved or comparable performance with reduced parameter complexity. For instance, for a model trained on ImageNet1K, we achieve a performance improvement from 72.14% to 74.02% and a learnable parameter reduction from 19.4M to 18.2M. Code could be found at https://github.com/Zhicaiwww/PosMLP https://github.com/Zhicaiwww/PosMLP.
Zhicai Wang, Yanbin Hao, Xingyu Gao 0001, Hao Zhang 0047, Shuo Wang 0008, Tingting Mu, Xiangnan He 0001
ACM Multimedia4
2022 Long-term Leap Attention, Short-term Periodic Shift for Video Classification
abstract
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes T times longer sequence than the latter under the current attention of quadratic complexity (T2N2). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy.
Hao Zhang 0047, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo
ACM Multimedia1
2022 Adversarial Multi-Grained Embedding Network for Cross-Modal Text-Video Retrieval
abstract
Cross-modal retrieval between texts and videos has received consistent research interest in the multimedia community. Existing studies follow a trend of learning a joint embedding space to measure the distance between text and video representations. In common practice, video representation is constructed by feeding clips into 3D convolutional neural networks for a coarse-grained global visual feature extraction. In addition, several studies have attempted to align the local objects of video with the text. However, these representations share a drawback of neglecting rich fine-grained relation features capturing spatial-temporal object interactions that benefits mapping textual entities in the real-world retrieval system. To tackle this problem, we propose an adversarial multi-grained embedding network (AME-Net), a novel cross-modal retrieval framework that adopts both fine-grained local relation and coarse-grained global features in bridging text-video modalities. Additionally, with the newly proposed visual representation, we also integrate an adversarial learning strategy into AME-Net, to further narrow the domain gap between text and video representations. In summary, we contribute AME-Net with an adversarial learning strategy for learning a better joint embedding space, and experimental results on MSR-VTT and YouCook2 datasets demonstrate that our proposed framework consistently outperforms the state-of-the-art method.
Ning Han 0005, Jingjing Chen 0001, Hao Zhang 0047, Huanwen Wang, Hao Chen 0051
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Fine-grained Cross-modal Alignment Network for Text-Video Retrieval
abstract
Despite the recent progress of cross-modal text-to-video retrieval techniques, their performance is still unsatisfactory. Most existing works follow a trend of learning a joint embedding space to measure the distance between global-level or local-level textual and video representation. The fine-grained interactions between video segments and phrases are usually neglected in cross-modal learning, which results in suboptimal retrieval performances. To tackle the problem, we propose a novel Fine-grained Cross-modal Alignment Network (FCA-Net), which considers the interactions between visual semantic units (i.e., sub-actions/sub-events) in videos and phrases in sentences for cross-modal alignment. Specifically, the interactions between visual semantic units and phrases are formulated as a link prediction problem optimized by a graph auto-encoder to obtain the explicit relations between them and enhance the aligned feature representation for fine-grained cross-modal alignment. Experimental results on MSR-VTT, YouCook2, and VATEX datasets demonstrate the superiority of our model as compared to the state-of-the-art method.
Ning Han 0005, Jingjing Chen 0001, Guangyi Xiao 0001, Hao Zhang 0047, Yawen Zeng, Hao Chen 0051
ACM Multimedia4
2021 Token Shift Transformer for Video Classification
abstract
Transformer achieves remarkable successes in understanding 1 and 2-dimensional signals (e.g., NLP and Image Content Understanding). As a potential alternative to convolutional neural networks, it shares merits of strong interpretability, high discriminative power on hyper-scale data, and flexibility in processing varying length inputs. However, its encoders naturally contain computational intensive operations such as pair-wise self-attention, incurring heavy computational burden when being applied on the complex 3-dimensional video signals. This paper presents Token Shift Module (i.e., TokShift), a novel, zero-parameter, zero-FLOPs operator, for modeling temporal relations within each transformer encoder. Specifically, the TokShift barely temporally shifts partial [Class] token features back-and-forth across adjacent frames. Then, we densely plug the module into each encoder of a plain 2D vision transformer for learning 3D video representation. It is worth noticing that our TokShift transformer is a pure convolutional-free video transformer pilot with computational efficiency for video understanding. Experiments on standard benchmarks verify its robustness, effectiveness, and efficiency. Particularly, with input clips of 8/12 frames, the TokShift transformer achieves SOTA precision: 79.83%/80.40% on the Kinetics-400, 66.56% on EGTEA-Gaze+, and 96.80% on UCF-101 datasets, comparable or better than existing SOTA convolutional counterparts. Our code is open-sourced in: https://github.com/VideoNetworks/TokShift-Transformer.
Hao Zhang 0047, Yanbin Hao, Chong-Wah Ngo
ACM Multimedia1
2021 Chinese White Dolphin Detection in the Wild
abstract
For ecological protection of the ocean, biologists usually conduct line-transect vessel surveys to measure sea species’ population density within their habitat (such as dolphins). However, sea species observation via vessel surveys consumes a lot of manpower resources and is more challenging compared to observing common objects, due to the scarcity of the object in the wild, tiny-size of the objects, and similar-sized distracter objects (e.g., floating trash). To reduce the human experts’ workload and improve the observation accuracy, in this paper, we develop a practical system to detect Chinese White Dolphins in the wild automatically. First, we construct a dataset named Dolphin-14k with more than 2.6k dolphin instances. To improve the dataset annotation efficiency caused by the rarity of dolphins, we design an interactive dolphin box annotation strategy to annotate sparse dolphin instances in long videos efficiently. Second, we compare the performance and efficiency of three off-the-shelf object detection algorithms, including Faster-RCNN, FCOS, and YoloV5, on the Dolphin-14k dataset and pick YoloV5 as the detector, where a new category (Distracter) is added to the model training to reject the false positives. Finally, we incorporate the dolphin detector into a system prototype, which detects dolphins in video frames at 100.99 FPS per GPU with high accuracy (i.e., 90.95 [email protected]).
Hao Zhang 0047, Qi Zhang 0041, Phuong Anh Nguyen 0002, Victor C. S. Lee, Antoni B. Chan
MMAsia1
2020 Visual Relations Augmented Cross-modal Retrieval
abstract
Retrieving relevant samples across multiple-modalities is a primary topic that receives consistently research interests in multimedia communities, and has benefited various real-world multimedia applications (e.g., text-based image searching). Current models mainly focus on learning a unified visual semantic embedding space to bridge visual contents & text query, targeting at aligning relevant samples from different modalities as neighbors in the embedding space. However, these models did not consider relations between visual components in learning visual representations, resulting in their incapability of distinguishing images with the same visual components but different relations (i.e., Figure 1). To precisely modeling visual contents, we introduce a novel framework that enhanced visual representation with relations between components. Specifically, visual relations are represented by the scene graph extracted from an image, then encoded by the graph convolutional neural networks for learning visual relational features. We combine the relational and compositional representation together for image-text retrieval. Empirical results conducted on the challenging MS-COCO and Flicker 30K datasets demonstrate the effectiveness of our proposed model for cross-modal retrieval task.
Yutian Guo, Jingjing Chen 0001, Hao Zhang 0047, Yu-Gang Jiang 0001
ICMR3
2020 Person-level Action Recognition in Complex Events via TSD-TSM Networks
abstract
The task of person-level action recognition in complex events aims to densely detect pedestrians and individually predict their actions from surveillance videos. In this paper, we present a simple yet efficient pipeline for this task, referred to as TSD-TSM networks. Firstly, we adopt the TSD detector for the pedestrian localization on each single keyframe. Secondly, we generate the sequential ROIs for a person proposal by replicating the adjusted bounding box coordinates around the keyframe. Particularly, we propose to conduct straddling expansion and region squaring on the original bounding box of a person proposal to widen the potential space of motion and interaction and lead to a square box for ROI detection. Finally, we adapt the TSM classifier on the generated ROI sequences to perform action classification and further adopt late fusion to promote the prediction. Our proposed pipeline achieved the 3rd place in the ACM-MM 2020 grand challenge, i.e., Large-scale Human-centric Video Analysis in Complex Events (Track-4), obtaining final 15.31% [email protected] and 20.63% [email protected] on the testing set.
Yanbin Hao, Zi-Niu Liu, Hao Zhang 0047, Bin Zhu 0006, Jingjing Chen 0001, Yu-Gang Jiang 0001, Chong-Wah Ngo
ACM Multimedia3
2020 Compact Bilinear Augmented Query Structured Attention for Sport Highlights Classification
abstract
Understanding fine-grained activities, such as sport highlights, is a problem being overlooked and receives considerably less research attention. Potential reasons include absences of specific fine-grained action benchmark datasets, research preferences to general super-categorical activities classification, and challenges of large visual similarities between fine-grained actions. To tackle these, we collect and manually annotate two sport highlights datasets, i.e., Basketball-8 & Soccer-10, for fine-grained action classification. Sample clips in the datasets are annotated with professional sub-categorical actions like "dunk", "goalkeeping" and etc. We also propose a Compact Bilinear Augmented Query Structured Attention (CBA-QSA) module and stack it on top of general three-dimensional neural networks in a plug-and-play manner to emphasize important spatio-temporal clues in highlight clips. Specifically, we adapt the hierarchical attention neural networks, which contain learnable query-scheme, on the video to identify discriminative spatial/temporal visual clues within highlight clips. We name this altered attention which separately learns a query for spatial/temporal feature as query structured attention (QSA). Furthermore, we inflate bilinear mapping, which is a mature technique to represent local pairwise interactions for image-level fine-grained classification, on video understanding. In detail, we extend its compact version (i.e., compact bilinear mapping (CBM) based on TensorSketch) to deal with the three-dimensional video signal for modeling local pairwise motion information. We eventually incorporate CBM and QSA together to form CBA-QSA neural networks for fine-grained sport highlights classifications. Experimental results demonstrate that CBA-QSA improves the general state-of-the-arts on Basketball-8 and Soccer-10 datasets.
Yanbin Hao, Hao Zhang 0047, Chong-Wah Ngo, Xiaojun Hu
ACM Multimedia2
2019 A Fine Granularity Object-Level Representation for Event Detection and Recounting
abstract
Multimedia events such as “birthday party” usually involve the complex interaction between humans and objects. Unlike actions and sports, these events rarely contain unique motion patterns to be vividly explored for recognition. To encode rich objects in the events, a common practice is to tag an individual video frame with object labels, represented as a vector signifying probabilities of object appearances. These vectors are then pooled across frames to obtain a video-level representation. The current practices suffer from two deficiencies due to the direct employment of deep convolutional neural network (DCNN) and standard feature pooling techniques. First, the use of max-pooling and softmax layers in DCNN overemphasize the primary object or scene in a frame, producing a sparse vector that overlooks the existence of secondary or small-size objects. Second, feature pooling by max or average operator over sparse vectors makes the video-level feature unpredictable in modeling the object composition of an event. To address these problems, this paper proposes a new video representation, named Object-VLAD, which treats each object equally and encodes them into a vector for multimedia event detection. Furthermore, the vector can be flexibly decoded to identify evidences such as key objects to recount the reason why a video is retrieved for an event of interest. Experiments conducted on MED13 and MED14 datasets verify the merit of Object-VLAD by consistently outperforming several state-of-the-arts in both event detection and recounting.
Hao Zhang 0047, Chong-Wah Ngo
IEEE Trans. Multim.1
2018 Enhanced VIREO KIS at VBS 2018
Phuong Anh Nguyen 0002, Yi-Jie Lu, Hao Zhang 0047, Chong-Wah Ngo
MMM (2)3
2017 On the Selection of Anchors and Targets for Video Hyperlinking
abstract
A problem not well understood in video hyperlinking is what qualifies a fragment as an anchor or target. Ideally, anchors provide good starting points for navigation, and targets supplement anchors with additional details while not distracting users with irrelevant, false and redundant information. The problem is not trivial for intertwining relationship between data characteristics and user expectation. Imagine that in a large dataset, there are clusters of fragments spreading over the feature space. The nature of each cluster can be described by its size (implying popularity) and structure (implying complexity). A principle way of hyperlinking can be carried out by picking centers of clusters as anchors and from there reach out to targets within or outside of clusters with consideration of neighborhood complexity. The question is which fragments should be selected either as anchors or targets, in one way to reflect the rich content of a dataset, and meanwhile to minimize the risk of frustrating user experience. This paper provides some insights to this question from the perspective of hubness and local intrinsic dimensionality, which are two statistical properties in assessing the popularity and complexity of data space. Based these properties, two novel algorithms are proposed for low-risk automatic selection of anchors and targets.
Zhi-Qi Cheng, Hao Zhang 0047, Xiao Wu 0001, Chong-Wah Ngo
ICMR2
2017 Concept-Based Interactive Search System
Yi-Jie Lu, Phuong Anh Nguyen 0002, Hao Zhang 0047, Chong-Wah Ngo
MMM (2)3
2017 Semantic Reasoning in Zero Example Video Event Retrieval
abstract
Searching in digital video data for high-level events, such as a parade or a car accident, is challenging when the query is textual and lacks visual example images or videos. Current research in deep neural networks is highly beneficial for the retrieval of high-level events using visual examples, but without examples it is still hard to (1) determine which concepts are useful to pre-train ( Vocabulary challenge ) and (2) which pre-trained concept detectors are relevant for a certain unseen high-level event ( Concept Selection challenge ). In our article, we present our Semantic Event Retrieval System which (1) shows the importance of high-level concepts in a vocabulary for the retrieval of complex and generic high-level events and (2) uses a novel concept selection method ( i-w2v ) based on semantic embeddings. Our experiments on the international TRECVID Multimedia Event Detection benchmark show that a diverse vocabulary including high-level concepts improves performance on the retrieval of high-level events in videos and that our novel method outperforms a knowledge-based concept selection method.
Maaike de Boer, Yi-Jie Lu, Hao Zhang 0047, Klamer Schutte, Chong-Wah Ngo, Wessel Kraaij
ACM Trans. Multim. Comput. Commun. Appl.3
2016 Event Detection with Zero Example: Select the Right and Suppress the Wrong Concepts
abstract
Complex video event detection without visual examples is a very challenging issue in multimedia retrieval. We present a state-of-the-art framework for event search without any need of exemplar videos and textual metadata in search corpus. To perform event search given only query words, the core of our framework is a large, pre-built bank of concept detectors which can understand the content of a video in the perspective of object, scene, action and activity concepts. Leveraging such knowledge can effectively narrow the semantic gap between textual query and the visual content of videos. Besides the large concept bank, this paper focuses on two challenges that largely affect the retrieval performance when the size of the concept bank increases: (1) How to choose the right concepts in the concept bank to accurately represent the query; (2) if noisy concepts are inevitably chosen, how to minimize their influence. We share our novel insights on these particular problems, which paves the way for a practical system that achieves the best performance in NIST TRECVID 2015.
Yi-Jie Lu, Hao Zhang 0047, Maaike de Boer, Chong-Wah Ngo
ICMR2