EDBT 2026 Demo / reviewers in the wild / expert
Yifan Zhang 0001
dblp:57/4707-1
· DBLP profile ↗
61ranked-venue papers
10as first author
14since 2021 · last 2025
0000-0002-9190-3509ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 35 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Generalize Without Bias for Open-Vocabulary Action RecognitionabstractLeveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary action recognition in-context. However, due to the static bias of CLIP, such video learners tend to overfit on shortcut static features, thereby compromising their generalizability, especially to novel out-of-context actions. To address this issue, we introduce Open-MeDe, a novel Meta-optimization framework with static Debiasing for Open-vocabulary action recognition. From a fresh perspective of generalization, Open-MeDe adopts a meta-learning approach to improve known-to-open generalizing and image-to-video debiasing in a cost-effective manner. Specifically, Open-MeDe introduces a cross-batch meta-optimization scheme that explicitly encourages video learners to quickly generalize to arbitrary subsequent data via virtual evaluation, steering a smoother optimization landscape. In effect, the free of CLIP regularization during optimization implicitly mitigates the inherent static bias of the video meta-learner. We further apply self-ensemble over the optimization trajectory to obtain generic optimal parameters that can achieve robust generalization to both in-context and out-of-context novel data. Extensive evaluations show that Open-MeDe not only surpasses state-of-the-art regularization methods tailored for in-context open-vocabulary action recognition but also substantially excels in out-of-context scenarios.Code is released at https://github.com/Mia-YatingYu/Open-MeDe. Yating Yu, Congqi Cao, Yifan Zhang 0001, Yanning Zhang 0001 |
ICCV | 3 |
| 2025 | Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) has achieved remarkable success in various real-world scenarios, but its high cost of online training makes it impractical to learn each task from scratch.
To enable effective policy reuse, we consider the problem of zero-shot generalization from offline data across multiple tasks.
While prior work focuses on transferring individual skills of agents, we argue that the effective policy transfer across tasks should also capture the team-level coordination knowledge.
In this paper, we propose Bi-Level Knowledge Transfer (BiKT) for Multi-Task MARL, which performs knowledge transfer at both the individual and team levels.
At the individual level, we extract transferable individual skill embeddings from offline MARL trajectories.
At the team level, we define tactics as coordinated patterns of skill combinations and capture them by leveraging the learned skill embeddings.
We map skill combinations into compact tactic embeddings and then construct a tactic codebook.
To incorporate both skills and tactics into decision-making, we design a bi-level decision transformer that infers them in sequence.
Our BiKT leverages both the generalizability of individual skills and the diversity of tactics, enabling the learned policy to perform effectively across multiple tasks.
Extensive experiments on SMAC and MPE benchmarks demonstrate that BiKT achieves strong generalization to previously unseen tasks. Jinmin He, Yifan Zhang 0001, Yifan Zang 0001, Jian Cheng 0001 |
NeurIPS | 3 |
| 2024 | Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement LearningabstractEfficient collaboration in the centralized training with decentralized execution (CTDE) paradigm remains a challenge in cooperative multi-agent systems. We identify divergent action tendencies among agents as a significant obstacle to CTDE's training efficiency, requiring a large number of training samples to achieve a unified consensus on agents' policies. This divergence stems from the lack of adequate team consensus-related guidance signals during credit assignment in CTDE. To address this, we propose Intrinsic Action Tendency Consistency, a novel approach for cooperative multi-agent reinforcement learning. It integrates intrinsic rewards, obtained through an action model, into a reward-additive CTDE (RA-CTDE) framework. We formulate an action model that enables surrounding agents to predict the central agent's action tendency. Leveraging these predictions, we compute a cooperative intrinsic reward that encourages agents to align their actions with their neighbors' predictions. We establish the equivalence between RA-CTDE and CTDE through theoretical analyses, demonstrating that CTDE's training process can be achieved using N individual targets. Building on this insight, we introduce a novel method to combine intrinsic rewards and RA-CTDE. Extensive experiments on challenging tasks in SMAC, MPE, and GRF benchmarks showcase the improved performance of our method. Yifan Zhang 0001, Xi Sheryl Zhang, Yifan Zang 0001, Jian Cheng 0001 |
AAAI | 2 |
| 2024 | HGCN2SP: Hierarchical Graph Convolutional Network for Two-Stage Stochastic ProgrammingabstractTwo-stage Stochastic Programming (2SP) is a standard framework for modeling decision-making problems under uncertainty. While numerous methods exist, solving such problems with many scenarios remains challenging. Selecting representative scenarios is a practical method for accelerating solutions. However, current approaches typically rely on clustering or Monte Carlo sampling, failing to integrate scenario information deeply and overlooking the significant impact of the scenario order on solving time. To address these issues, we develop HGCN2SP, a novel model with a hierarchical graph designed for 2SP problems, encoding each scenario and modeling their relationships hierarchically. The model is trained in a reinforcement learning paradigm to utilize the feedback of the solver. The policy network is equipped with a hierarchical graph convolutional network for feature encoding and an attention-based decoder for scenario selection in proper order. Evaluation of two classic 2SP problems demonstrates that HGCN2SP provides high-quality decisions in a short computational time. Furthermore, HGCN2SP exhibits remarkable generalization capabilities in handling large-scale instances, even with a substantial number of variables or scenarios that were unseen during the training phase. Yifan Zhang 0001, Zhenxing Liang, Jian Cheng 0001 |
ICML | 2 |
| 2023 | Asynchronous Event Processing with Local-Shift Graph Convolutional NetworkabstractEvent cameras are bio-inspired sensors that produce sparse and asynchronous event streams instead of frame-based images at a high-rate. Recent works utilizing graph convolutional networks (GCNs) have achieved remarkable performance in recognition tasks, which model event stream as spatio-temporal graph. However, the computational mechanism of graph convolution introduces redundant computation when aggregating neighbor features, which limits the low-latency nature of the events. And they perform a synchronous inference process, which can not achieve a fast response to the asynchronous event signals. This paper proposes a local-shift graph convolutional network (LSNet), which utilizes a novel local-shift operation equipped with a local spatio-temporal attention component to achieve efficient and adaptive aggregation of neighbor features. To improve the efficiency of pooling operation in feature extraction, we design a node-importance based parallel pooling method (NIPooling) for sparse and low-latency event data. Based on the calculated importance of each node, NIPooling can efficiently obtain uniform sampling results in parallel, which retains the diversity of event streams. Furthermore, for achieving a fast response to asynchronous event signals, an asynchronous event processing procedure is proposed to restrict the network nodes which need to recompute activations only to those affected by the new arrival event. Experimental results show that the computational cost can be reduced by nearly 9 times through using local-shift operation and the proposed asynchronous procedure can further improve the inference efficiency, while achieving state-of-the-art performance on gesture recognition and object recognition. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
AAAI | 2 |
| 2023 | On the Data-Efficiency with Contrastive Image Transformation in Reinforcement Learning
Xi Sheryl Zhang, Yushuo Li, Yifan Zhang 0001, Jian Cheng 0001 |
ICLR | 4 |
| 2023 | Efficient spatiotemporal context modeling for action recognition
Congqi Cao, Yue Lu 0008, Yifan Zhang 0001, Dongmei Jiang, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2023 | Learnable Locality-Sensitive Hashing for Video Anomaly DetectionabstractVideo anomaly detection (VAD) mainly refers to identifying anomalous events that have not occurred in the training set where only normal samples are available. Existing works usually formulate VAD as a reconstruction or prediction problem. However, the adaptability and scalability of these methods are limited. In this paper, we propose a novel distance-based VAD method to take advantage of all the available normal data efficiently and flexibly. In our method, the smaller the distance between a testing sample and normal samples, the higher the probability that the testing sample is normal. Specifically, we propose to use locality-sensitive hashing (LSH) to map the samples whose similarity exceeds a certain threshold into the same bucket in advance. To utilize multiple hashes and further alleviate the computation and memory usage, we propose to use the hash codes rather than the features as the representations of the samples. In this manner, the complexity of near neighbor search is cut down significantly. To make the samples that are semantically similar get closer and those not similar get further apart, we propose a novel learnable version of LSH that embeds LSH into a neural network and optimizes the hash functions with contrastive learning strategy. The proposed method is robust to data imbalance and can handle the large intra-class variations in normal data flexibly. Besides, it has a good ability of scalability. Extensive experiments demonstrate the superiority of our method, which achieves new state-of-the-art results on VAD benchmarks. Yue Lu 0008, Congqi Cao, Yifan Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | MENet: A Memory-Based Network with Dual-Branch for Efficient Event Stream Processing
Linhui Sun, Yifan Zhang 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu |
ECCV (24) | 2 |
| 2022 | PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientabstractKnowledge distillation(KD) is a widely-used technique to train compact models in object detection. However, there is still a lack of study on how to distill between heterogeneous detectors. In this paper, we empirically find that better FPN features from a heterogeneous teacher detector can help the student although their detection heads and label assignments are different. However, directly aligning the feature maps to distill detectors suffers from two problems. First, the difference in feature magnitude between the teacher and the student could enforce overly strict constraints on the student. Second, the FPN stages and channels with large feature magnitude from the teacher model could dominate the gradient of distillation loss, which will overwhelm the effects of other features in KD and introduce much noise. To address the above issues, we propose to imitate features with Pearson Correlation Coefficient to focus on the relational information from the teacher and relax constraints on the magnitude of the features. Our method consistently outperforms the existing detection KD methods and works for both homogeneous and heterogeneous student-teacher pairs. Furthermore, it converges faster. With a powerful MaskRCNN-Swin detector as the teacher, ResNet-50 based RetinaNet and FCOS achieve 41.5% and 43.9% $mAP$ on COCO2017, which are 4.1% and 4.8% higher than the baseline, respectively. Weihan Cao, Yifan Zhang 0001, Jianfei Gao 0003, Anda Cheng, Ke Cheng 0002, Jian Cheng 0001 |
NeurIPS | 2 |
| 2022 | FedFV: federated face verification via equivalent class embeddings
Lingyun Liu, Yifan Zhang 0001, Haoyuan Gao, Xingtao Yu, Jian Cheng 0001 |
Multim. Syst. | 2 |
| 2022 | Action recognition via pose-based graph convolutional networks with intermediate dense supervision
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
Pattern Recognit. | 2 |
| 2021 | AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action RecognitionabstractExisting methods for skeleton-based action recognition mainly focus on improving the recognition accuracy, whereas the efficiency of the model is rarely considered. Recently, there are some works trying to speed up the skeleton modeling by designing light-weight modules. However, in addition to the model size, the amount of the data involved in the calculation is also an important factor for the running speed, especially for the skeleton data where most of the joints are redundant or non-informative to identify a specific skeleton. Besides, previous works usually employ one fix-sized model for all the samples regardless of the difficulty of recognition, which wastes computations for easy samples. To address these limitations, a novel approach, called AdaSGN, is proposed in this paper, which can reduce the computational cost of the inference process by adaptively controlling the input number of the joints of the skeleton on-the-fly. Moreover, it can also adaptively select the optimal model size for each sample to achieve a better trade-off between the accuracy and the efficiency. We conduct extensive experiments on three challenging datasets, namely, NTU-60, NTU-120 and SHREC, to verify the superiority of the proposed approach, where AdaSGN achieves comparable or even higher performance with much lower GFLOPs compared with the baseline method. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICCV | 2 |
| 2021 | Extremely Lightweight Skeleton-Based Action Recognition With ShiftGCN++abstractIn skeleton-based action recognition, graph convolutional networks (GCNs) have achieved remarkable success. However, there are two shortcomings of current GCN-based methods. Firstly, the computation cost is pretty heavy, typically over 15 GFLOPs for one action sample. Some recent works even reach ~100 GFLOPs. Secondly, the receptive fields of both spatial graph and temporal graph are inflexible. Although recent works introduce incremental adaptive modules to enhance the expressiveness of spatial graph, their efficiency is still limited by regular GCN structures. In this paper, we propose a shift graph convolutional network (ShiftGCN) to overcome both shortcomings. ShiftGCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. To further boost the efficiency, we introduce four techniques and build a more lightweight skeleton-based action recognition model named ShiftGCN++. ShiftGCN++ is an extremely computation-efficient model, which is designed for low-power and low-cost devices with very limited computing power. On three datasets for skeleton-based action recognition, ShiftGCN notably exceeds the state-of-the-art methods with over 10× less FLOPs and 4× practical speedup. ShiftGCN++ further boosts the efficiency of ShiftGCN, which achieves comparable performance with 6× less FLOPs and 2× practical speedup. Ke Cheng 0002, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Image Process. | 2 |
| 2020 | Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ACCV (5) | 2 |
| 2020 | Skeleton-Based Action Recognition With Shift Graph Convolutional NetworkabstractAction recognition with skeleton data is attracting more attention in computer vision. Recently, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have obtained remarkable performance. However, the computational complexity of GCN-based methods are pretty heavy, typically over 15 GFLOPs for one action sample. Recent works even reach about 100 GFLOPs. Another shortcoming is that the receptive fields of both spatial graph and temporal graph are inflexible. Although some works enhance the expressiveness of spatial graph by introducing incremental adaptive modules, their performance is still limited by regular GCN structures. In this paper, we propose a novel shift graph convolutional network (Shift-GCN) to overcome both shortcomings. Instead of using heavy regular graph convolutions, our Shift-GCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. On three datasets for skeleton-based action recognition, the proposed Shift-GCN notably exceeds the state-of-the-art methods with more than 10 times less computational complexity. Ke Cheng 0002, Yifan Zhang 0001, Weihan Chen, Jian Cheng 0001, Hanqing Lu |
CVPR | 2 |
| 2020 | Decoupling GCN with DropGraph Module for Skeleton-Based Action Recognition
Ke Cheng 0002, Yifan Zhang 0001, Congqi Cao, Lei Shi 0018, Jian Cheng 0001, Hanqing Lu |
ECCV (24) | 2 |
| 2020 | Rethinking The Pid Optimizer For Stochastic Optimization Of Deep NetworksabstractStochastic gradient descent with momentum (SGD-Momentum) always causes the overshoot problem due to the integral action of the momentum term. Recently, an ID optimizer is proposed to solve the overshoot problem with the help of derivative information. However, the derivative term suffers from the interference of the high-frequency noise, especially for the stochastic gradient descent method that uses minibatch data in each update step. In this work, we propose a complete PID optimizer, which weakens the effect of the D term and adds a P term to more stably alleviate the overshoot problem. To further reduce the interference of the high-frequency noise, two effective and efficient methods are proposed to stabilize the training process. Extensive experiments on three widely used benchmark datasets with different scales, i.e., MNIST, Cifar10 and TinyImageNet, demonstrate the superiority of our proposed PID optimizer on various popular deep neural networks. Lei Shi 0018, Yifan Zhang 0001, Wanguo Wang, Jian Cheng 0001, Hanqing Lu |
ICME | 2 |
| 2020 | Motion Complementary Network for Efficient Action RecognitionabstractBoth two-stream ConvNet and 3D ConvNet are widely used in action recognition. However, both methods are not efficient for deployment: calculating optical flow is very slow, while 3D convolution is computationally expensive. Our key insight is that the motion information from optical flow maps is complementary to the motion information from 3D ConvNet. Instead of simply combining these two methods, we propose two novel techniques to enhance the performance with less computational cost: fixed-motion-accumulation and balanced-motion-policy. With these two techniques, we propose a novel framework called Efficient Motion Complementary Network(EMC-Net) that enjoys both high efficiency and high performance. We conduct extensive experiments on Kinetics, UCF101, and Jester datasets. We achieve notably higher performance while consuming 4.7× less computation than I3D, 11.6× less computation than ECO, 17.8× less computation than R(2+1)D. On Kinetics dataset, we achieve 2.6% better performance than the recent proposed TSM with 1.4× fewer FLOPs and 10ms faster on K80 GPU. Ke Cheng 0002, Yifan Zhang 0001, Chenghua Li, Jian Cheng 0001, Hanqing Lu |
ICPR | 2 |
| 2020 | PEAN: 3D Hand Pose Estimation Adversarial NetworkabstractDespite recent emerging research attention, 3D hand pose estimation still suffers from the problems of predicting inaccurate or invalid poses which conflict with physical and kinematic constraints. To address these problems, we propose a novel 3D hand pose estimation adversarial network (PEAN) which can implicitly utilize such constraints to regularize the prediction in an adversarial learning framework. PEAN contains two parts: a 3D hierarchical estimation network (3DHNet) to predict hand pose, which decouples the task into multiple subtasks with a hierarchical structure; a pose discrimination network (PDNet) to judge the reasonableness of the estimated 3D hand pose, which back-propagates the constraints to the estimation network. During the adversarial learning process, PDNet is expected to distinguish the estimated 3D hand pose and the ground truth, while 3DHNet is expected to estimate more valid pose to confuse PDNet. In this way, 3DHNet is capable of generating 3D poses with accurate positions and adaptively adjusting the invalid poses without additional prior knowledge. Experiments show that the proposed 3DHNet does a good job in predicting hand poses, and introducing PDNet to 3DHNet does further improve the accuracy and reasonableness of the predicted results. As a result, the proposed PEAN achieves the state-of-the-art performance on three public hand pose estimation datasets. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICPR | 2 |
| 2020 | Robust one-stage object detection with location-aware classifiers
Qiang Chen 0007, Peisong Wang 0001, Anda Cheng, Wanguo Wang, Yifan Zhang 0001, Jian Cheng 0001 |
Pattern Recognit. | 5 |
| 2020 | Gesture recognition based on deep deformable 3D convolutional neural networks
Yifan Zhang 0001, Lei Shi 0018, Yi Wu 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu |
Pattern Recognit. | 1 |
| 2020 | Skeleton-Based Action Recognition With Multi-Stream Adaptive Graph Convolutional NetworksabstractGraph convolutional networks (GCNs), which generalize CNNs to more generic non-Euclidean structures, have achieved remarkable performance for skeleton-based action recognition. However, there still exist several issues in the previous GCN-based models. First, the topology of the graph is set heuristically and fixed over all the model layers and input data. This may not be suitable for the hierarchy of the GCN model and the diversity of the data in action recognition tasks. Second, the second-order information of the skeleton data, i.e., the length and orientation of the bones, is rarely investigated, which is naturally more informative and discriminative for the human action recognition. In this work, we propose a novel multi-stream attention-enhanced adaptive graph convolutional neural network (MS-AAGCN) for skeleton-based action recognition. The graph topology in our model can be either uniformly or individually learned based on the input data in an end-to-end manner. This data-driven approach increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Besides, the proposed adaptive graph convolutional layer is further enhanced by a spatial-temporal-channel attention module, which helps the model pay more attention to important joints, frames and features. Moreover, the information of both the joints and bones, together with their motion information, are simultaneously modeled in a multi-stream framework, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Image Process. | 2 |
| 2019 | Skeleton-Based Action Recognition With Directed Graph Neural NetworksabstractThe skeleton data have been widely used for the action recognition tasks since they can robustly accommodate dynamic circumstances and complex backgrounds. In existing methods, both the joint and bone information in skeleton data have been proved to be of great help for action recognition tasks. However, how to incorporate these two types of data to best take advantage of the relationship between joints and bones remains a problem to be solved. In this work, we represent the skeleton data as a directed acyclic graph based on the kinematic dependency between the joints and bones in the natural human body. A novel directed graph neural network is designed specially to extract the information of joints, bones and their relations and make prediction based on the extracted features. In addition, to better fit the action recognition task, the topological structure of the graph is made adaptive based on the training process, which brings notable improvement. Moreover, the motion information of the skeleton sequence is exploited and combined with the spatial information to further enhance the performance in a two-stream framework. Our final model is tested on two large-scale datasets, NTU-RGBD and Skeleton-Kinetics, and exceeds state-of-the-art performance on both of them. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
CVPR | 2 |
| 2019 | Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have achieved remarkable performance. However, in existing GCN-based methods, the topology of the graph is set manually, and it is fixed over all layers and input samples. This may not be optimal for the hierarchical GCN and diverse samples in action recognition tasks. In addition, the second-order information (the lengths and directions of bones) of the skeleton data, which is naturally more informative and discriminative for action recognition, is rarely investigated in existing methods. In this work, we propose a novel two-stream adaptive graph convolutional network (2s-AGCN) for skeleton-based action recognition. The topology of the graph in our model can be either uniformly or individually learned by the BP algorithm in an end-to-end manner. This data-driven method increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Moreover, a two-stream framework is proposed to model both the first-order and the second-order information simultaneously, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
CVPR | 2 |
| 2019 | Gesture Recognition Using Spatiotemporal Deformable Convolutional RepresentationabstractDynamic gesture recognition, which plays an essential role in human-computer interaction, has been widely investigated but not yet addressed. The interference of the varied and complex background makes the classifier easily be misguided due to the relatively smaller size of the hands and arms compared with the full scenes. In this paper, we address the problem by proposing a novel spatiotemporal deformable convolutional neural network for end-to-end learning. To eliminate the background interference, a light-weight spatiotemporal deformable convolution module is specially designed to augment the spatiotemporal sampling locations of 3D convolution by learning additional offsets according to the preceding feature map. The proposed method is evaluated on two challenging datasets, EgoGesture and Jester, and achieves the state-of-the-art performance on both of the two datasets. The code and trained models will be released for better communication and future work. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICIP | 2 |
| 2019 | Skeleton-Based Action Recognition With Gated Convolutional Neural NetworksabstractFor skeleton-based action recognition, most of the existing works used recurrent neural networks. Using convolutional neural networks (CNNs) is another attractive solution considering their advantages in parallelization, effectiveness in feature learning, and model base sufficiency. Besides these, skeleton data are low-dimensional features. It is natural to arrange a sequence of skeleton features chronologically into an image, which retains the original information. Therefore, we solve the sequence learning problem as an image classification task using CNNs. For better learning ability, we build a classification network with stacked residual blocks and having a special design called linear skip gated connection which can benefit information propagation across multiple residual blocks. When arranging the coordinates of body joints in one frame into a skeleton feature, we systematically investigate the performance of part-based, chain-based, and traversal-based orders. Furthermore, a fully convolutional permutation network is designed to learn an optimized order for data rearrangement. Without any bells and whistles, our proposed model achieves state-of-the-art performance on two challenging benchmark datasets, outperforming existing methods significantly. Congqi Cao, Cuiling Lan, Yifan Zhang 0001, Wenjun Zeng 0001, Hanqing Lu, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Two-Step Quantization for Low-Bit Neural NetworksabstractEvery bit matters in the hardware design of quantized neural networks. However, extremely-low-bit representation usually causes large accuracy drop. Thus, how to train extremely-low-bit neural networks with high accuracy is of central importance. Most existing network quantization approaches learn transformations (low-bit weights) as well as encodings (low-bit activations) simultaneously. This tight coupling makes the optimization problem difficult, and thus prevents the network from learning optimal representations. In this paper, we propose a simple yet effective Two-Step Quantization (TSQ) framework, by decomposing the network quantization problem into two steps: code learning and transformation function learning based on the learned codes. For the first step, we propose the sparse quantization method for code learning. The second step can be formulated as a non-linear least square regression problem with low-bit constraints, which can be solved efficiently in an iterative manner. Extensive experiments on CIFAR-10 and ILSVRC-12 datasets demonstrate that the proposed TSQ is effective and outperforms the state-of-the-art by a large margin. Especially, for 2-bit activation and ternary weight quantization of AlexNet, the accuracy of our TSQ drops only about 0.5 points compared with the full-precision counterpart, outperforming current state-of-the-art by more than 5 points. Peisong Wang 0001, Qinghao Hu 0001, Yifan Zhang 0001, Chunjie Zhang 0001, Yang Liu 0021, Jian Cheng 0001 |
CVPR | 3 |
| 2018 | Training Binary Weight Networks via Semi-Binary Decomposition
Qinghao Hu 0001, Gang Li 0015, Peisong Wang 0001, Yifan Zhang 0001, Jian Cheng 0001 |
ECCV (13) | 4 |
| 2018 | Image Class Prediction by Joint Object, Context, and Background ModelingabstractState-of-the-art image classification methods often use spatial pyramid matching or its variants to make use of the spatial layout of visual features. However, objects may appear at various places with different scales and orientations. Besides, traditionally object-centric-based methods only consider objects and the background without fully exploring the context information. To solve these problems, in this paper we propose a novel image classification method by jointly modeling the object, context, and background information (OCB). OCB consists of three components: 1) locate the positions of objects; 2) determine the context areas of objects; and 3) treat the other areas as the background. We use objectness proposal techniques to select candidate bounding boxes. Boxes with high confidence scores are combined to determine objects' positions. To select the context areas, we use candidate boxes that have relatively lower confidence scores compared with boxes for object location selection. The other areas are viewed as the background. We jointly combine the object, context, and background for image representation and classification. Experiments on six data sets well demonstrate the superiority of the proposed OCB method over other spatial partition methods. Chunjie Zhang 0001, Guibo Zhu, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Body Joint Guided 3-D Deep Convolutional Descriptors for Action Recognitionabstract3-D convolutional neural networks (3-D CNNs) have been established as a powerful tool to simultaneously learn features from both spatial and temporal dimensions, which is suitable to be applied to video-based action recognition. In this paper, we propose not to directly use the activations of fully connected layers of a 3-D CNN as the video feature, but to use selective convolutional layer activations to form a discriminative descriptor for video. It pools the feature on the convolutional layers under the guidance of body joint positions. Two schemes of mapping body joints into convolutional feature maps for pooling are discussed. The body joint positions can be obtained from any off-the-shelf skeleton estimation algorithm. The helpfulness of the body joint guided feature pooling with inaccurate skeleton estimation is systematically evaluated. To make it end-to-end and do not rely on any sophisticated body joint detection algorithm, we further propose a two-stream bilinear model which can learn the guidance from the body joints and capture the spatio-temporal features simultaneously. In this model, the body joint guided feature pooling is conveniently formulated as a bilinear product operation. Experimental results on three real-world datasets demonstrate the effectiveness of body joint guided pooling which achieves promising performance. Congqi Cao, Yifan Zhang 0001, Chunjie Zhang 0001, Hanqing Lu |
IEEE Trans. Cybern. | 2 |
| 2018 | EgoGesture: A New Dataset and Benchmark for Egocentric Hand Gesture RecognitionabstractGesture is a natural interface in human-computer interaction, especially interacting with wearable devices, such as VR/AR helmet and glasses. However, in the gesture recognition community, it lacks of suitable datasets for developing egocentric (first-person view) gesture recognition methods, in particular in the deep learning era. In this paper, we introduce a new benchmark dataset named EgoGesture with sufficient size, variation, and reality to be able to train deep neural networks. This dataset contains more than 24 000 gesture samples and 3 000 000 frames for both color and depth modalities from 50 distinct subjects. We design 83 different static and dynamic gestures focused on interaction with wearable devices and collect them from six diverse indoor and outdoor scenes, respectively, with variation in background and illumination. We also consider the scenario when people perform gestures while they are walking. The performances of several representative approaches are systematically evaluated on two tasks: gesture classification in segmented data and gesture spotting and recognition in continuous data. Our empirical study also provides an in-depth analysis on input modality selection and domain adaptation between different scenes. Yifan Zhang 0001, Congqi Cao, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Multim. | 1 |
| 2017 | Fast K-means for Large Scale ClusteringabstractK-means algorithm has been widely used in machine learning and data mining due to its simplicity and good performance. However, the standard k-means algorithm would be quite slow for clustering millions of data into thousands of or even tens of thousands of clusters. In this paper, we propose a fast k-means algorithm named multi-stage k-means (MKM) which uses a multi-stage filtering approach. The multi-stage filtering approach greatly accelerates the k-means algorithm via a coarse-to-fine search strategy. To further speed up the algorithm, hashing is introduced to accelerate the assignment step which is the most time-consuming part in k-means. Extensive experiments on several massive datasets show that the proposed algorithm can obtain up to 600X speed-up over the k-means algorithm with comparable accuracy. Qinghao Hu 0001, Jiaxiang Wu 0001, Lu Bai 0001, Yifan Zhang 0001, Jian Cheng 0001 |
CIKM | 4 |
| 2017 | Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer ModulesabstractGesture is a natural interface in interacting with wearable devices such as VR/AR helmet and glasses. The main challenge of gesture recognition in egocentric vision arises from the global camera motion caused by the spontaneous head movement of the device wearer. In this paper, we address the problem by a novel recurrent 3D convolutional neural network for end-to-end learning. We specially design a spatiotemporal transformer module with recurrent connections between neighboring time slices which can actively transform a 3D feature map into a canonical view in both spatial and temporal dimensions. To validate our method, we introduce a new dataset with sufficient size, variation and reality, which contains 83 gestures designed for interaction with wearable devices, and more than 24,000 RGB-D gesture samples from 50 subjects captured in 6 scenes. On this dataset, we show that the proposed network outperforms competing state-of-the-art algorithms. Moreover, our method can achieve state-of-the-art performance on the challenging GTEA egocentric action dataset. Congqi Cao, Yifan Zhang 0001, Yi Wu 0001, Hanqing Lu, Jian Cheng 0001 |
ICCV | 2 |
| 2016 | Action Recognition with Joints-Pooled 3D Deep Convolutional Descriptors
Congqi Cao, Yifan Zhang 0001, Chunjie Zhang 0001, Hanqing Lu |
IJCAI | 2 |
| 2016 | Image Classification Using Spatial Difference Descriptor Under Spatial Pyramid Matching Framework
Jiucheng Xu, Yifan Zhang 0001, Chunjie Zhang 0001, Hongsheng Yin 0001, Hanqing Lu |
MMM (1) | 3 |
| 2016 | Socio-mobile landmark recognition using local features with adaptive region selection
Chunjie Zhang 0001, Yifan Zhang 0001, Xiaobin Zhu 0001, Zhe Xue, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 2 |
| 2016 | A Coupled Hidden Conditional Random Field Model for Simultaneous Face Clustering and Naming in VideosabstractFor face naming in TV series or movies, a typical way is using subtitles/script alignment to get the time stamps of the names, and tagging them to the faces. We study the problem of face naming in videos when subtitles are not available. To this end, we divide the problem into two tasks: face clustering which groups the faces depicting a certain person into a cluster, and name assignment which associates a name to each face. Each task is formulated as a structured prediction problem and modeled by a hidden conditional random field (HCRF) model. We argue that the two tasks are correlated problems whose outputs can provide prior knowledge of the target prediction for each other. The two HCRFs are coupled in a unified graphical model called coupled HCRF where the joint dependence of the cluster labels and face name association is naturally embedded in the correlation between the two HCRFs. We provide an effective algorithm to optimize the two HCRFs iteratively and the performance of the two tasks on real-world data set can be both improved. Yifan Zhang 0001, Baoyuan Wu, Hanqing Lu |
IEEE Trans. Image Process. | 1 |
| 2015 | Multi-modal learning for gesture recognitionabstractWith the development of sensing equipments, data from different modalities is available for gesture recognition. In this paper, we propose a novel multi-modal learning framework. A coupled hidden Markov model (CHMM) is employed to discover the correlation and complementary information across different modalities. In this framework, we use two configurations: one is multi-modal learning and multi-modal testing, where all the modalities used during learning are still available during testing; the other is multi-modal learning and single-modal testing, where only one modality is available during testing. Experiments on two real-world gesture recognition data sets have demonstrated the effectiveness of our multi-modal learning framework. Improvements on both of the multi-modal and single-modal testing have been observed. Congqi Cao, Yifan Zhang 0001, Hanqing Lu |
ICME | 2 |
| 2015 | Face Clustering in Videos with Proportion Prior
Yifan Zhang 0001, Zechao Li, Hanqing Lu |
IJCAI | 2 |
| 2015 | Spatio-Temporal Triangular-Chain CRF for Activity RecognitionabstractUnderstanding human activities in video is a fundamental problem in computer vision. In real life, human activities are composed of temporal and spatial arrangement of actions. Understanding such complex activities requires recognizing not only each individual action, but more importantly, capturing their spatio-temporal relationships. This paper addresses the problem of complex activity recognition with a unified hierarchical model. We expand triangular-chain CRFs (TriCRFs) to the spatial dimension. The proposed architecture can be perceived as a spatio-temporal version of the TriCRFs, in which the labels of actions and activity are modeled jointly and their complex dependencies are exploited. Experiments show that our model generates promising results, outperforming competing methods significantly. The framework also can be applied to model other structured sequential data. Congqi Cao, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 2 |
| 2015 | Automatic face annotation in TV series by video/script alignment
Yifan Zhang 0001, Chunjie Zhang 0001, Jing Liu 0001, Hanqing Lu |
Neurocomputing | 1 |
| 2015 | Joint image representation and classification in random semantic spaces
Chunjie Zhang 0001, Xiaobin Zhu 0001, Liang Li 0003, Yifan Zhang 0001, Jing Liu 0001, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 4 |
| 2015 | Image classification using boosted local features with random orientation and location selection
Chunjie Zhang 0001, Jian Cheng 0001, Yifan Zhang 0001, Jing Liu 0001, Chao Liang 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001 |
Inf. Sci. | 3 |
| 2014 | Video face naming using global sequence alignmentabstractThis paper explores the problem of automatically naming faces in TV series or films. A novel method is proposed to build association between the faces in the video and the names in the script by a global sequence alignment algorithm. We firstly build two heterogenous sequences: a face sequence and a name sequence. The elements of the two sequences are cluster labels, computed from the clustering process, and speaking names, respectively. Then the alignment of the two sequences is considered as a problem of surjection between the cluster set and the name set. The optimal solution is obtained by minimizing the Levenshtein Distance between the two sequences which is constrained by the temporal order information. Experiments on public videos demonstrate the effectiveness of our method. Yifan Zhang 0001, Shuang Qiu 0002, Hanqing Lu |
ICIP | 2 |
| 2014 | Beyond visual word ambiguity: Weighted local feature encoding with governing region
Chunjie Zhang 0001, Xian Xiao, Junbiao Pang, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | Undoing the codebook bias by linear transformation with sparsity and F-norm constraints for image classification
Chunjie Zhang 0001, Chao Liang 0001, Junbiao Pang, Yifan Zhang 0001, Jing Liu 0001, Qingming Huang |
Pattern Recognit. Lett. | 4 |
| 2013 | Constrained Clustering and Its Application to Face Clustering in VideosabstractIn this paper, we focus on face clustering in videos. Given the detected faces from real-world videos, we partition all faces into K disjoint clusters. Different from clustering on a collection of facial images, the faces from videos are organized as face tracks and the frame index of each face is also provided. As a result, many pair wise constraints between faces can be easily obtained from the temporal and spatial knowledge of the face tracks. These constraints can be effectively incorporated into a generative clustering model based on the Hidden Markov Random Fields (HMRFs). Within the HMRF model, the pair wise constraints are augmented by label-level and constraint-level local smoothness to guide the clustering process. The parameters for both the unary and the pair wise potential functions are learned by the simulated field algorithm, and the weights of constraints can be easily adjusted. We further introduce an efficient clustering framework specially for face clustering in videos, considering that faces in adjacent frames of the same face track are very similar. The framework is applicable to other clustering algorithms to significantly reduce the computational cost. Experiments on two face data sets from real-world videos demonstrate the significantly improved performance of our algorithm over state-of-the art algorithms. Baoyuan Wu, Yifan Zhang 0001, Bao-Gang Hu |
CVPR | 2 |
| 2013 | Event Detection in Complex Scenes Using Interval Temporal ConstraintsabstractIn complex scenes with multiple atomic events happening sequentially or in parallel, detecting each individual event separately may not always obtain robust and reliable result. It is essential to detect them in a holistic way which incorporates the causality and temporal dependency among them to compensate the limitation of current computer vision techniques. In this paper, we propose an interval temporal constrained dynamic Bayesian network to extend Allen's interval algebra network (IAN) [2] from a deterministic static model to a probabilistic dynamic system, which can not only capture the complex interval temporal relationships, but also model the evolution dynamics and handle the uncertainty from the noisy visual observation. In the model, the topology of the IAN on each time slice and the interlinks between the time slices are discovered by an advanced structure learning method. The duration of the event and the unsynchronized time lags between two correlated event intervals are captured by a duration model, so that we can better determine the temporal boundary of the event. Empirical results on two real world datasets show the power of the proposed interval temporal constrained model. Yifan Zhang 0001, Hanqing Lu |
ICCV | 1 |
| 2013 | Undo the codebook bias by linear transformation for visual applicationsabstractThe bag of visual words model (BoW) and its variants have demonstrate their effectiveness for visual applications and have been widely used by researchers. The BoW model first extracts local features and generates the corresponding codebook, the elements of a codebook are viewed as visual words. The local features within each image are then encoded to get the final histogram representation. However, the codebook is dataset dependent and has to be generated for each image dataset. This costs a lot of computational time and weakens the generalization power of the BoW model. To solve these problems, in this paper, we propose to undo the dataset bias by codebook linear transformation. To represent every points within the local feature space using Euclidean distance, the number of bases should be no less than the space dimensions. Hence, each codebook can be viewed as a linear transformation of these bases. In this way, we can transform the pre-learned codebooks for a new dataset. However, not all of the visual words are equally important for the new dataset, it would be more effective if we can make some selection using sparsity constraints and choose the most discriminative visual words for transformation. We propose an alternative optimization algorithm to jointly search for the optimal linear transformation matrixes and the encoding parameters. Image classification experimental results on several image datasets show the effectiveness of the proposed method. Chunjie Zhang 0001, Yifan Zhang 0001, Shuhui Wang, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2013 | Correlation consistency constrained probabilistic matrix factorization for social tag refinement
Jing Liu 0001, Yifan Zhang 0001, Zechao Li, Hanqing Lu |
Neurocomputing | 2 |
| 2013 | Modeling Temporal Interactions with Interval Temporal Bayesian Networks for Complex Activity RecognitionabstractComplex activities typically consist of multiple primitive events happening in parallel or sequentially over a period of time. Understanding such activities requires recognizing not only each individual event but, more importantly, capturing their spatiotemporal dependencies over different time intervals. Most of the current graphical model-based approaches have several limitations. First, time--sliced graphical models such as hidden Markov models (HMMs) and dynamic Bayesian networks are typically based on points of time and they hence can only capture three temporal relations: precedes, follows, and equals. Second, HMMs are probabilistic finite-state machines that grow exponentially as the number of parallel events increases. Third, other approaches such as syntactic and description-based methods, while rich in modeling temporal relationships, do not have the expressive power to capture uncertainties. To address these issues, we introduce the interval temporal Bayesian network (ITBN), a novel graphical model that combines the Bayesian Network with the interval algebra to explicitly model the temporal dependencies over time intervals. Advanced machine learning methods are introduced to learn the ITBN model structure and parameters. Experimental results show that by reasoning with spatiotemporal dependencies, the proposed model leads to a significantly improved performance when modeling and recognizing complex activities involving both parallel and sequential events. Yongmian Zhang, Yifan Zhang 0001, Eran Swears, Natalia Larios, Ziheng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Real-time multiple object instances detectionabstractIn this paper, we present a novel, real-time multiple object instance detection system via template matching and pairwise classification. Instance detection aims to find and locate exactly the same object instances as specified. Our system is composed of two heterogeneous stages. The first stage adopts instance-specific detection to generate candidates. And the second stage makes use of a pairwise-based classifier across instance categories to test and verify these candidates with respect to templates. Experiments show the superiority of our approach. Chengli Xie, Jinqiao Wang, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 3 |
| 2009 | Naming faces in films using hypergraph matchingabstractIn this paper, we aim to address the problem of naming faces in feature-length films using video and film script. Different from the state-of-the-art methods on naming faces in the videos, most of which used a local matching between a visible face and one of the names extracted from the local video transcript, we use a global matching between names and faces as it is not easy to obtain enough local name cues in the films. In the video, we cluster the faces into groups corresponding to characters and build a face network according to face co-occurrence relationship. Similarly in the film script, a name network is also built according to name cooccurrence relationship. The vertices of the two networks are finally matched by a hypergraph matching method. Experiments are conducted on five feature-length films and give encouraging results. Yifan Zhang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu |
ICME | 1 |
| 2009 | Personalized retrieval of sports video based on multi-modal analysis and user preference acquisition
Yifan Zhang 0001, Changsheng Xu, Xiaoyu Zhang 0002, Hanqing Lu |
Multim. Tools Appl. | 1 |
| 2009 | Character Identification in Feature-Length Films Using Global Face-Name MatchingabstractIdentification of characters in films, although very intuitive to humans, still poses a significant challenge to computer methods. In this paper, we investigate the problem of identifying characters in feature-length films using video and film script. Different from the state-of-the-art methods on naming faces in the videos, most of which used the local matching between a visible face and one of the names extracted from the temporally local video transcript, we attempt to do a global matching between names and clustered face tracks under the circumstances that there are not enough local name cues that can be found. The contributions of our work include: 1) A graph matching method is utilized to build face-name association between a face affinity network and a name affinity network which are, respectively, derived from their own domains (video and script). 2) An effective measure of face track distance is presented for face track clustering. 3) As an application, the relationship between characters is mined using social network analysis. The proposed framework is able to create a new experience on character-centered film browsing. Experiments are conducted on ten feature-length films and give encouraging results. Yifan Zhang 0001, Changsheng Xu, Hanqing Lu, Yueh-Min Huang |
IEEE Trans. Multim. | 1 |
| 2008 | Automatic character identification in feature-length filmsabstractThis paper presents a novel approach to automatically identify characters in films using audio visual cues and text analysis. The approach consists of three stages: (i) frontal face track detection and clustering, (ii) face track classification, (iii) name assignment. A Finite State Machine (FSM) method is utilized to filter faces detected on each frame and build face tracks. The face tracks are clustered using constrained K-Centers. The tracks located in the center area of each cluster are set as exemplars. The marginal points of each cluster and the newly detected non-frontal face tracks are classified to these exemplars using complementary cues of audio and visual. The names of characters are ranked based on their occurrences in the film script and the face track clusters are ranked based on track counts. The names are assigned to the clusters according to the ranking order. Experiments were conducted on two feature-length films and gave promising results. Yifan Zhang 0001, Changsheng Xu, Hanqing Lu |
ICME | 1 |
| 2008 | A Novel Framework for Semantic Annotation and Personalized Retrieval of Sports VideoabstractSports video annotation is important for sports video semantic analysis such as event detection and personalization. In this paper, we propose a novel approach for sports video semantic annotation and personalized retrieval. Different from the state of the art sports video analysis methods which heavily rely on audio/visual features, the proposed approach incorporates web-casting text into sports video analysis. Compared with previous approaches, the contributions of our approach include the following. 1) The event detection accuracy is significantly improved due to the incorporation of web-casting text analysis. 2) The proposed approach is able to detect exact event boundary and extract event semantics that are very difficult or impossible to be handled by previous approaches. 3) The proposed method is able to create personalized summary from both general and specific point of view related to particular game, event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis, text/video alignment, and personalized retrieval. The experimental results on event boundary detection in sports video are encouraging and comparable to the manually selected events. The evaluation on personalized retrieval is effective in helping meet users' expectations. Changsheng Xu, Jinjun Wang, Hanqing Lu, Yifan Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2008 | Using Webcast Text for Semantic Event Detection in Broadcast Sports VideoabstractSports video semantic event detection is essential for sports video summarization and retrieval. Extensive research efforts have been devoted to this area in recent years. However, the existing sports video event detection approaches heavily rely on either video content itself, which face the difficulty of high-level semantic information extraction from video content using computer vision and image processing techniques, or manually generated video ontology, which is domain specific and difficult to be automatically aligned with the video content. In this paper, we present a novel approach for sports video semantic event detection based on analysis and alignment of Webcast text and broadcast video. Webcast text is a text broadcast channel for sports game which is co-produced with the broadcast video and is easily obtained from the Web. We first analyze Webcast text to cluster and detect text events in an unsupervised way using probabilistic latent semantic analysis (pLSA). Based on the detected text event and video structure analysis, we employ a conditional random field model (CRFM) to align text event and video event by detecting event moment and event boundary in the video. Incorporation of Webcast text into sports video analysis significantly facilitates sports video semantic event detection. We conducted experiments on 33 hours of soccer and basketball games for Webcast analysis, broadcast video analysis and text/video semantic alignment. The results are encouraging and compared with the manually labeled ground truth. Changsheng Xu, Yifan Zhang 0001, Guangyu Zhu 0002, Yong Rui, Hanqing Lu, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2007 | Semantic Event Extraction from Basketball Games using Multi-Modal AnalysisabstractIn this paper, we present a novel multi-modal framework for semantic event extraction from basketball games based on Webcasting text and broadcast video. We propose novel approaches to text analysis for event detection and semantics extraction, video analysis for event structure modeling and event moment detection, and text/video alignment for event boundary detection in the video. Compared with existing approaches to event detection in sports video which rely heavily on low-level features directly extracted from video itself, our approach aims to bridge the semantic gap between low-level features and high-level events and facilitates personalization of the sports video. Promising results are reported on real-world video clips by using text analysis, video analysis and text/video alignment. Yifan Zhang 0001, Changsheng Xu, Yong Rui, Jinqiao Wang, Hanqing Lu |
ICME | 1 |
| 2005 | Highlight ranking for sports video browsingabstractSports video has been extensively studied for its wide viewer-ship and tremendous commercial potentials. Many studies focused on highlight extraction for summarizing a lengthy video. In this paper, we present an advanced highlight analysis system for sports video browsing, in which highlight evaluation and ranking are concerned besides highlight detection. First, we use replay detection to efficiently localize the highlights. Then incorporating with domain-specific knowledge, we adopt several significant cues to evaluate the importance degree of the highlights with support vector regression. Finally, the highlights are ranked with descending sort according to their importance value. The ranking results can provide a hierarchical video browsing and customized content delivery scheme. Initial experimental results on soccer videos show an encouraging performance comparing with human subjective evaluation. Xiaofeng Tong, Qingshan Liu 0001, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 3 |