EDBT 2026 Demo / reviewers in the wild / expert
Rui-Ze Han
dblp:205/4022 · also Ruize Han
· DBLP profile ↗
40ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-6587-8936ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BEVTrack: Multi-View Multi-Human Registration and Tracking in the Bird's Eye ViewabstractWe handle a new problem of multi-view multi-human tracking in the bird's eye view (BEV). Different from previous works, we require neither the calibration among the multi-view cameras nor the actually captured BEV video. This makes the studied problem closer to real-world applications, however, more challenging. For this purpose, in this work, we propose a novel BEVTrack scheme. Specifically, given multi-view videos, we first use a virtual BEV transform module to obtain the BEV for each view. Then, we propose a unified BEV alignment module to fuse the respectively generated BEVs, in which we specifically design the self-supervised losses by considering both the spatial consistency and the temporal continuity. During the inference, we design the camera-subject collaborative registration and tracking strategy to make use of the mutual dependence between the multi-view cameras and the multiple targets, to achieve the desired BEV tracking. We also build a new benchmark for training and evaluation, the experimental results on which have verified the rationality of the problem and the effectiveness of our method. Zekun Qian, Wei Feng 0005, Rui-Ze Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | From indoor to outdoor: Unsupervised domain adaptive gait recognition
Likai Wang 0002, Wei Feng 0005, Rui-Ze Han, Xiangqun Zhang 0003, Yanjie Wei, Song Wang 0002 |
Pattern Recognit. | 3 |
| 2026 | Category-agnostic object re-identification
Likai Wang 0002, Rui-Ze Han, Bingliang Jiao, Wei Feng 0005 |
Pattern Recognit. | 2 |
| 2025 | FlexiCell: Deep Learning with Learnable Adaptive Filtering and Dual Attention for Cell SegmentationabstractAccurate cell segmentation remains challenging due to morphological variations, diverse imaging modalities, and unclear cellular boundaries. Existing deep learning (DL) methods struggle to extract features adaptively across heterogeneous cellular environments, thereby limiting generalization capacity. To address these challenges, we propose FlexiCell, a novel adaptive segmentation framework that integrates a learnable adaptive filter with dual attention mechanisms. The core innovation lies in the FlexiFilter approach, which combines standard convolution with adaptive residual learning through learnable mixing parameters. These parameters dynamically balance input preservation and feature enhancement. FlexiCell employs multi-scale FlexiFilter blocks with varying kernel sizes, channel and spatial attention networks, and a dedicated boundary extractor for precise edge detection. Extensive experiments demonstrate superior performance compared to benchmark models, achieving 3.8% improvement in detection accuracy and 5.5 % in segmentation quality on our newly developed induced pluripotent stem (iPS) cell datasets. Further evaluation on standardized Cell Tracking Challenge (CTC) benchmarks confirms state-of-the-art performance on mesenchymal stem cells and glioblastoma datasets, outperforming established CTC methods. The framework demonstrates robust generalization across fluorescence, phase contrast, and differential interference contrast microscopy, without requiring dataset-specific optimization. Codes are available at https://github.com/jovialniyo93/FlexiCell. Jovial Niyogisubizo, Keliang Zhao, Shengqi Zhou, Rui-Ze Han, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
BIBM | 4 |
| 2025 | VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object Tracking
Zekun Qian, Rui-Ze Han, Junhui Hou, Linqi Song, Wei Feng 0005 |
ICCV | 2 |
| 2025 | COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue Fusion
Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005 |
ICCV | 2 |
| 2025 | DOVTrack: Data-Efficient Open-Vocabulary TrackingabstractOpen-Vocabulary Multi-Object Tracking (OVMOT) aims to detect and track multi-category objects including both seen and unseen categories during training.
Currently, a significant challenge in this domain is the lack of large-scale annotated video data for training.
To address this challenge, this work aims to effectively train the OV tracker using only the existing limited and sparsely annotated video data.
We propose a comprehensive training sample space expansion strategy that addresses the fundamental limitation of sparse annotations in OVMOT training. Specifically, for the association task, we develop a diffusion-based feature generation framework that synthesizes intermediate object features between sparsely annotated frames, effectively expanding the training sample space by approximately 3× and enabling robust association learning from temporally continuous features. For the detection task, we introduce a dynamic group contrastive learning approach that generates diverse sample groups through affinity, dispersion, and adversarial grouping strategies, tripling the effective training samples for classification while maintaining sample quality. Additionally, we propose an adaptive localization loss that expands positive sample coverage by lowering IoU thresholds while mitigating noise through confidence-based weighting. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the OVMOT benchmark, surpassing existing methods by 3.8\% in TETA metric, without requiring additional data or annotations. The code will be available at https://github.com/zekunqian/DOVTrack. Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005 |
NeurIPS | 2 |
| 2025 | Weakly Supervised Instance Action RecognitionabstractWe study the novel problem of weakly supervised instance action recognition (WSiAR) in multi-person (crowd) scenes. We specifically aim to recognize the action of each subject in the crowd, for which we propose the use of a weakly supervised method, considering the expense of large-scale annotations for training. This problem is of great practical value for video surveillance and sports scene analysis. To this end, we investigated and designed a series of weak annotations for the supervision of weakly supervised instance action recognition (WSiAR). We propose two categories of weak label settings, bag labels and sparse labels, to significantly reduce the number of labels. Based on the former, we propose a novel sub-block-aware multi-instance learning (MIL) loss to obtain more effective information from weak labels during training. With respect to the latter, we propose a pseudo label generation strategy for extending sparse labels. This enables our method to achieve results comparable to those of fully supervised methods but with significantly fewer annotations. The experimental results on two benchmarks verified the rationality of the problem definition and effectiveness of the proposed weakly supervised training method in solving our problem. Haomin Yan, Rui-Ze Han, Wei Feng 0005, Jiewen Zhao, Songmiao Wang |
Comput. Vis. Media | 2 |
| 2025 | Concept-Guided Open-Vocabulary Temporal Action Detection
Songmiao Wang, Rui-Ze Han, Wei Feng 0005 |
J. Comput. Sci. Technol. | 2 |
| 2025 | A DDoS attack detection method based on IQR and DFFCNN in SDN
Meng Yue 0002, Huayang Yan, Rui-Ze Han, Zhijun Wu 0001 |
J. Netw. Comput. Appl. | 3 |
| 2025 | Unveiling the Power of Self-Supervision for Multi-View Multi-Human Association and TrackingabstractMulti-view multi-human association and tracking (MvMHAT), is an emerging yet important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self-learning. In this work, we tackle this problem with an end-to-end neural network in a self-supervised learning manner. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry, and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public. Wei Feng 0005, Fei Wang 0032, Rui-Ze Han, Yiyang Gan, Zekun Qian, Junhui Hou, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | A large-scale combinatorial benchmark for sign language recognition
Liqing Gao, Liang Wang 0001, Lianyu Hu 0003, Rui-Ze Han, Zekang Liu, Fanhua Shang, Wei Feng 0005 |
Pattern Recognit. | 4 |
| 2025 | A New Benchmark and Algorithm for Clothes-Changing Video Person Re-IdentificationabstractPerson re-identification (Re-ID) is a classical computer vision task and has significant applications for public security and information forensics. Recently, long-term Re-ID with clothes-changing has attracted increasing attention. However, existing methods mainly focus on image-based setting, where richer temporal information is overlooked. In this paper, we focus on the relatively new yet practical problem of Clothes-Changing Video-based Re-ID (CCVReID), which is less studied. First, given the dataset shortage, we build two new benchmark datasets for CCVReID problem, including a large-scale synthetic video dataset and a real-world one, both containing human sequences with various clothing changes. Moreover, we systematically study this problem by simultaneously considering the classical appearance feature and temporal feature contained in the video. We develop a dual-branch fusion framework that makes use of the information from both clothes-aware appearance feature and clothes-free gait feature. For better information fusion, a confidence-guided re-ranking strategy is proposed to adaptively balance the weight of these two categories of features. We have released the benchmark and code proposed in this work to the public athttps://github.com/kkw98/CCVReID. Likai Wang 0002, Xiangqun Zhang 0003, Rui-Ze Han, Yanjie Wei, Song Wang 0002, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Synthetic-to-Real Video Person Re-IDabstractPerson re-identification (Re-ID) is an important task and has significant applications for public security and information forensics, which has progressed rapidly with the development of deep learning. In this work, we investigate a novel and challenging setting of Re-ID, i.e., cross-domain video-based person Re-ID. Specifically, we utilize synthetic video datasets as the source domain for training and real-world videos for testing, notably reducing the reliance on expensive real data acquisition and annotation. To harness the potential of synthetic data, we first propose a self-supervised domain-invariant feature learning strategy for both static and dynamic (temporal) features. Additionally, to enhance person identification accuracy in the target domain, we propose a mean-teacher scheme incorporating a self-supervised ID consistency loss. Experimental results across five real datasets validate the rationale behind cross-synthetic-real domain adaptation and demonstrate the efficacy of our method. Notably, the discovery that synthetic data outperforms real data in the cross-domain scenario is a surprising outcome. The code and data are publicly available at https://github.com/XiangqunZhang/UDA_Video_ReID. Xiangqun Zhang 0003, Rui-Ze Han, Likai Wang 0002, Linqi Song, Junhui Hou, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera CalibrationabstractWe tackle a new problem of multi-view camera and sub-ject registration in the bird’ s eye view (BEV) without pregiven camera calibration, which promotes the multi-view subject registration problem to a new calibration-free stage. This greatly alleviates the limitation in many practical applications. However, this is a very challenging problem since its only input is several RGB images from different first-person views (FPVs), without the BEV image and the calibration of the FPVs, while the output is a unified plane aggregated from all views with the positions and orientations of both the subjects and cameras in a BEV. For this purpose, we propose an end-to-end framework solving cam-era and subject registration together by taking advantage of their mutual dependence, whose main idea is as below: i) creating a subject view-transform module (VTM) to project each pedestrian from FPV to a virtual BEV, ii) deriving a multi-view geometry-based spatial alignment module (SAM) to estimate the relative camera pose in a unified BEV, iii) selecting and refining the subject and camera registration results within the unified BEV. We collect a new large-scale synthetic dataset with rich annotations for training and evaluation. Additionally, we also collect a real dataset for cross-domain evaluation. The experimental results show the remarkable effectiveness of our method. The code and proposed datasets are available at BEVSee. Zekun Qian, Rui-Ze Han, Wei Feng 0005, Song Wang 0002 |
CVPR | 2 |
| 2024 | Robust Collaborative Perception without External Localization and Clock DevicesabstractA consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal alignment, traditional methods depend on external devices to provide localization and clock signals. However, hardware-generated signals could be vulnerable to noise and potentially malicious attack, jeopardizing the precision of spatial-temporal alignment. Rather than relying on external hardwares, this work proposes a novel approach: aligning by recognizing the inherent geometric patterns within the perceptual data of various agents. Following this spirit, we propose a robust collaborative perception system that operates independently of external localization and clock devices. The key module of our system, FreeAlign, constructs a salient object graph for each agent based on its detected boxes and uses a graph neural network to identify common subgraphs between agents, leading to accurate relative pose and time. We validate FreeAlign on both real-world and simulated datasets. The results show that, the FreeAlign empowered robust collaborative perception system perform comparably to systems relying on precise localization and clock devices. ${\mathbf{Code}}$ will be released. Zixing Lei, Zhenyang Ni, Rui-Ze Han, Chen Feng 0002, Siheng Chen, Yanfeng Wang 0001 |
ICRA | 3 |
| 2024 | Rethinking the One-shot Object Detection: Cross-Domain Object SearchabstractOne-shot object detection (OSOD) uses a query patch to identify the same category of object in a target image. As the OSOD setting, the target images are required to contain the object category of the query patch, and the image styles (domains) of the query patch and target images are always similar. However, in practical application, the above requirements are not commonly satisfied. Therefore, we propose a new problem namely Cross-Domain Object Search (CDOS), where the object categories of the query patch and target image are decoupled, and the image styles between them may also be significantly different. For this problem, we develop a new method, which incorporates both foreground-background contrastive learning heads and a domain-generalized feature augmentation technique. This makes our method effectively handle the object category gap and domain distribution gap, between the query patch and target image in the training and testing datasets. We further build a new benchmark for the proposed CDOS problem, on which our method shows significant performance improvements over the comparison methods. Shuqi Zheng, Rui-Ze Han, Yuzhong Feng, Junhui Hou, Linqi Song, Wei Feng 0005 |
ACM Multimedia | 3 |
| 2024 | OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object TrackingabstractOpen-vocabulary object perception has become an important topic in artificial intelligence, which aims to identify objects with novel classes that have not been seen during training. Under this setting, open-vocabulary object detection (OVD) in a single image has been studied in many literature. However, open-vocabulary object tracking (OVT) from a video has been studied less, and one reason is the shortage of benchmarks. In this work, we have built a new large-scale benchmark for open-vocabulary multi-object tracking namely OVT-B. OVT-B contains 1,048 categories of objects and 1,973 videos with 637,608 bounding box annotations, which is much larger than the sole open-vocabulary tracking dataset, i.e., OVTAO-val dataset (200+ categories, 900+ videos). The proposed OVT-B can be used as a new benchmark to pave the way for OVT research. We also develop a simple yet effective baseline method for OVT. It integrates the motion features for object tracking, which is an important feature for MOT but is ignored in previous OVT methods. Experimental results have verified the usefulness of the proposed benchmark and the effectiveness of our method. We have released the benchmark to the public at https://github.com/Coo1Sea/OVT-B-Dataset. Haiji Liang, Rui-Ze Han |
NeurIPS | 2 |
| 2024 | Contactless interaction recognition and interactor detection in multi-person scenes
Rui-Ze Han, Wei Feng 0005, Haomin Yan, Song Wang 0002 |
Frontiers Comput. Sci. | 2 |
| 2024 | Benchmarking the Complementary-View Multi-human Association and Tracking
Rui-Ze Han, Wei Feng 0005, Zekun Qian, Haomin Yan, Song Wang 0002 |
Int. J. Comput. Vis. | 1 |
| 2024 | Sign language translation with hierarchical memorized context in question answering scenarios
Liqing Gao, Wei Feng 0005, Rui-Ze Han, Di Lin 0002, Liang Wang 0001 |
Neural Comput. Appl. | 4 |
| 2023 | Combining the Silhouette and Skeleton Data for Gait RecognitionabstractGait recognition, a long-distance biometric technology, has aroused intense interest recently. Currently, the two dominant gait recognition works are appearance-based and model-based, which extract features from silhouettes and skeletons, respectively. However, appearance-based methods are greatly affected by clothes-changing and carrying conditions, while model-based methods are limited by the accuracy of pose estimation. To tackle this challenge, a simple yet effective two-branch network is proposed in this paper, which contains a CNN-based branch taking silhouettes as input and a GCN-based branch taking skeletons as input. In addition, for better gait representation in the GCN-based branch, we present a fully connected graph convolution operator to integrate multi-scale graph convolutions and alleviate the dependence on natural joint connections. Also, we deploy a multi-dimension attention module named STC-Att to learn spatial, temporal and channel-wise attention simultaneously. The experimental results on CASIA-B and OUMVLP show that our method achieves state-of-the-art performance in various conditions. Likai Wang 0002, Rui-Ze Han, Wei Feng 0005 |
ICASSP | 2 |
| 2023 | Relating View Directions of Complementary-View Mobile Cameras via the Human Shadow
Rui-Ze Han, Yiyang Gan, Likai Wang 0002, Nan Li 0048, Wei Feng 0005, Song Wang 0002 |
Int. J. Comput. Vis. | 1 |
| 2022 | Connecting the Complementary-view Videos: Joint Camera Identification and Subject AssociationabstractWe attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds of applications. This is a very challenging problem due to the large view difference between top and side views. In this paper, we develop a new approach that can simultaneously handle three tasks: i) localizing the side-view camera in the top view; ii) estimating the view direction of the side-view camera; iii) detecting and associating the same subjects on the ground across the complementary views. Our main idea is to explore the spatial position layout of the subjects in two views. In particular, we propose a spatial-aware position representation method to embed the spatial-position distribution of the subjects in different views. We further design a cross-view video collaboration framework composed of a camera identification module and a subject association module to simultaneously perform the above three tasks. We collect a new synthetic dataset consisting of top-view and side-view video sequence pairs for performance evaluation and the experimental results show the effectiveness of the proposed method. Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002 |
CVPR | 1 |
| 2022 | Panoramic Human Activity Recognition
Rui-Ze Han, Haomin Yan, Song Wang 0002, Wei Feng 0005 |
ECCV (4) | 1 |
| 2022 | Self-supervised Social Relation Representation for Human Group Detection
Rui-Ze Han, Haomin Yan, Zekun Qian, Wei Feng 0005, Song Wang 0002 |
ECCV (35) | 2 |
| 2022 | Self-Supervised Human Pose based Multi-Camera Video SynchronizationabstractMulti-view video collaborative analysis is an important task and has many applications in multimedia community. However, it always requires the given multiple videos to be temporally synchronized. Existing methods commonly synchronize the videos by the wired communication, which may hinder the practical application in real world, especially for moving cameras. In this paper, we focus on the human-centric video analysis and propose a self-supervised framework for the automatic multi-camera video synchronization. Specifically, we develop SeSyn-Net with the 2D human pose as input for feature embedding and design a series of self-supervised losses to effectively extract the view-invariant but time-discriminative representation for video synchronization. We also build two new datasets for the performance evaluation. Extensive experimental results verify the effectiveness of our method, which achieves the superior performance compared to both the classical and state-of-the-art methods. Liqiang Yin, Rui-Ze Han, Wei Feng 0005, Song Wang 0002 |
ACM Multimedia | 2 |
| 2022 | Multiple Human Association and Tracking From Egocentric and Complementary Top ViewsabstractCrowded scene surveillance can significantly benefit from combining egocentric-view and its complementary top-view cameras. A typical setting is an egocentric-view camera, e.g., a wearable camera on the ground capturing rich local details, and a top-view camera, e.g., a drone-mounted one from high altitude providing a global picture of the scene. To collaboratively analyze such complementary-view videos, an important task is to associate and track multiple people across views and over time, which is challenging and differs from classical human tracking, since we need to not only track multiple subjects in each video, but also identify the same subjects across the two complementary views. This paper formulates it as a constrained mixed integer programming problem, wherein a major challenge is how to effectively measure subjects similarity over time in each video and across two views. Although appearance and motion consistencies well apply to over-time association, they are not good at connecting two highly different complementary views. To this end, we present a spatial distribution based approach to reliable cross-view subject association. We also build a dataset to benchmark this new challenging task. Extensive experiments verify the effectiveness of our method. Rui-Ze Han, Wei Feng 0005, Yujun Zhang 0002, Jiewen Zhao, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Multi-View Multi-Human Association With Deep Assignment NetworkabstractIdentifying the same persons across different views plays an important role in many vision applications. In this paper, we study this important problem, denoted as Multi-view Multi-Human Association (MvMHA), on multi-view images that are taken by different cameras at the same time. Different from previous works on human association across two views, this paper is focused on more general and challenging scenarios of more than two views, and none of these views are fixed or priorly known. In addition, each involved person may be present in all the views or only a subset of views, which are also not priorly known. We develop a new end-to-end deep-network based framework to address this problem. First, we use an appearance-based deep network to extract the feature of each detected subject on each image. We then compute pairwise-similarity scores between all the detected subjects and construct a comprehensive affinity matrix. Finally, we propose a Deep Assignment Network (DAN) to transform the affinity matrix into an assignment matrix, which provides a binary assignment result for MvMHA. We build both a synthetic dataset and a real image dataset to verify the effectiveness of the proposed method. We also test the trained network on other three public datasets, resulting in very good cross-domain performance. Rui-Ze Han, Haomin Yan, Wei Feng 0005, Song Wang 0002 |
IEEE Trans. Image Process. | 1 |
| 2021 | Multiple Human Tracking in Non-Specific Coverage with Wearable CamerasabstractCompared to fixed cameras, wearable cameras have time-varying non-specific view coverage and can be used to alternately observe people at different sites by varying the camera views. However, such view change of wearable cameras may introduce intervals of transitional frames without useful information, which brings new challenge for the important multiple object tracking (MOT) task – existing MOT methods can not handle well frequent disappearing/reappearing targets in the field of view, especially in the presence of informationless transitional sequences of frames. To address this problem, in this paper we propose a Markov Decision Process with jump state (JMDP) to model the target’s lifetime in tracking, and use optical flow of the camera motion and the statistical information of the targets to model the camera state transition. We further develop a frame-level classification algorithm to locate the transitional sequence. By combining all of them, we formulate the proposed non-specific-coverage MOT problem as a joint state transition problem, which can be solved by the state transfer mechanism of the targets and the camera. We collect a new dataset for performance evaluation and the experimental results show the effectiveness of the proposed method. Sibo Wang 0008, Rui-Ze Han, Wei Feng 0005, Song Wang 0002 |
ICASSP | 2 |
| 2021 | Self-supervised Multi-view Multi-Human Association and TrackingabstractMulti-view Multi-human association and tracking (MvMHAT) aims to track a group of people over time in each view, as well as to identify the same person across different views at the same time. This is a relatively new problem but is very important for multi-person scene video surveillance. Different from previous multiple object tracking (MOT) and multi-target multi-camera tracking (MTMCT) tasks, which only consider the over-time human association, MvMHAT requires to jointly achieve both cross-view and over-time data association. In this paper, we model this problem with a self-supervised learning framework and leverage an end-to-end network to tackle it. Specifically, we propose a spatial-temporal association network with two designed self-supervised learning losses, including a symmetric-similarity loss and a transitive-similarity loss, at each time to associate the multiple humans over time and across views. Besides, to promote the research on MvMHAT, we build a new large-scale benchmark for the training and testing of different algorithms. Extensive experiments on the proposed benchmark verify the effectiveness of our method. We have released the benchmark and code to the public. Yiyang Gan, Rui-Ze Han, Liqiang Yin, Wei Feng 0005, Song Wang 0002 |
ACM Multimedia | 2 |
| 2020 | Complementary-View Multiple Human TrackingabstractThe global trajectories of targets on ground can be well captured from a top view in a high altitude, e.g., by a drone-mounted camera, while their local detailed appearances can be better recorded from horizontal views, e.g., by a helmet camera worn by a person. This paper studies a new problem of multiple human tracking from a pair of top- and horizontal-view videos taken at the same time. Our goal is to track the humans in both views and identify the same person across the two complementary views frame by frame, which is very challenging due to very large field of view difference. In this paper, we model the data similarity in each view using appearance and motion reasoning and across views using appearance and spatial reasoning. Combing them, we formulate the proposed multiple human tracking as a joint optimization problem, which can be solved by constrained integer programming. We collect a new dataset consisting of top- and horizontal-view video pairs for performance evaluation and the experimental results show the effectiveness of the proposed method. Rui-Ze Han, Wei Feng 0005, Jiewen Zhao, Zicheng Niu, Yujun Zhang 0002, Song Wang 0002 |
AAAI | 1 |
| 2020 | Key Action and Joint CTC-Attention based Sign Language RecognitionabstractSign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as continuous and dense action sequence, which is difficult to capture key actions corresponding to meaningful sentence. In this paper, we propose to hierarchically search key actions by a pyramid BiLSTM. Specifically, we first construct three BiL-STMs to produce temporal relationships among input video sequence. Then, we associate these BiLSTMs by searching the salient responses in two groups of fixed-scale sliding window and capture key actions. Additionally, in order to balance the sequence alignment and dependency, we propose to jointly train Connectionist Temporal Classification (CTC) and Long Short-Term Memory (LSTM). Experimental results demonstrate the effectiveness of the proposed method. Liqing Gao, Rui-Ze Han, Wei Feng 0005 |
ICASSP | 3 |
| 2020 | Complementary-View Co-Interest Person DetectionabstractFast and accurate identification of the co-interest persons, who draw joint interest of the surrounding people, plays an important role in social scene understanding and surveillance. Previous study mainly focuses on detecting co-interest persons from a single-view video. In this paper, we study a much more realistic and challenging problem, namely co-interest person~(CIP) detection from multiple temporally-synchronized videos taken by the complementary and time-varying views. Specifically, we use a top-view camera, mounted on a flying drone at a high altitude to obtain a global view of the whole scene and all subjects on the ground, and multiple horizontal-view cameras, worn by selected subjects, to obtain a local view of their nearby persons and environment details. We present an efficient top- and horizontal-view data fusion strategy to map multiple horizontal views into the global top view. We then propose a spatial-temporal CIP potential energy function that jointly considers both intra-frame confidence and inter-frame consistency, thus leading to an effective Conditional Random Field~(CRF) formulation. We also construct a complementary-view video dataset, which provides a benchmark for the study of multi-view co-interest person detection. Extensive experiments validate the effectiveness and superiority of the proposed method. Rui-Ze Han, Jiewen Zhao, Wei Feng 0005, Yiyang Gan, Song Wang 0002 |
ACM Multimedia | 1 |
| 2020 | Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable CamerasabstractCompared to a single fixed camera, multiple moving cameras, e.g., those worn by people, can better capture the human interactive and group activities in a scene, by providing multiple, flexible and possibly complementary views of the involved people. In this setting the actual promotion of activity detection is highly dependent on the effective correlation and collaborative analysis of multiple videos taken by different wearable cameras, which is highly challenging given the time-varying view differences across different cameras and mutual occlusion of people in each video. By focusing on two wearable cameras and the interactive activities that involve only two people, in this paper we develop a new approach that can simultaneously: (i) identify the same persons across the two videos, (ii) detect the interactive activities of interest, including their occurrence intervals and involved people, and (iii) recognize the category of each interactive activity. Specifically, we represent each video by a graph, with detected persons as nodes, and propose a unified Graph Neural Network (GNN) based framework to jointly solve the above three problems. A graph matching network is developed for identifying the same persons across the two videos and a graph inference network is then used for detecting the human interactions. We also build a new video dataset, which provides a benchmark for this study, and conduct extensive experiments to validate the effectiveness and superiority of the proposed method. Jiewen Zhao, Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002 |
ACM Multimedia | 2 |
| 2020 | Selective Spatial Regularization by Reinforcement Learned Decision Making for Object TrackingabstractSpatial regularization (SR) is known as an effective tool to alleviate the boundary effect of correlation filter (CF), a successful visual object tracking scheme, from which a number of state-of-the-art visual object trackers can be stemmed. Nevertheless, SR highly increases the optimization complexity of CF and its target-driven nature makes spatially-regularized CF trackers may easily lose the occluded targets or the targets surrounded by other similar objects. In this paper, we propose selective spatial regularization (SSR) for CF-tracking scheme. It can achieve not only higher accuracy and robustness, but also higher speed compared with spatially-regularized CF trackers. Specifically, rather than simply relying on foreground information, we extend the objective function of CF tracking scheme to learn the target-context-regularized filters using target-context-driven weight maps. We then formulate the online selection of these weight maps as a decision making problem by a Markov Decision Process (MDP), where the learning of weight map selection is equivalent to policy learning of the MDP that is solved by a reinforcement learning strategy. Moreover, by adding a special state, representing not-updating filters, in the MDP, we can learn when to skip unnecessary or erroneous filter updating, thus accelerating the online tracking. Finally, the proposed SSR is used to equip three popular spatially-regularized CF trackers to significantly boost their tracking accuracy, while achieving much faster online tracking speed. Besides, extensive experiments on five benchmarks validate the effectiveness of SSR. Qing Guo 0005, Rui-Ze Han, Wei Feng 0005, Zhihao Chen 0004 |
IEEE Trans. Image Process. | 2 |
| 2020 | Fast Learning of Spatially Regularized and Content Aware Correlation Filter for Visual TrackingabstractWith a good balance between accuracy and speed, correlation filter (CF) has become a popular and dominant visual object tracking scheme. It implicitly extends the training samples by circular shifts of a given target patch, which serve as negative samples for fast online learning of the filters. Since all these shifted patches are not real negative samples of the target, CF tracking scheme suffers from the annoying boundary effects that can greatly harm the tracking performance, especially under challenging situations, like occlusion and fast temporal variation. Spatial regularization is known as a potent way to alleviate such boundary effects, but with the cost of highly increased time complexity, caused by complex optimization imported by spatial regularization. In this paper, we propose a new fast learning approach to content-aware spatial regularization, namely weighted sample based CF tracking (WSCF). In WSCF, specifically, we present a simple yet effective energy function that implicitly weighs different training samples by spatial deviations. With the energy function, the learning of correlation filters is composed of two subproblems with closed-form solution and can be efficiently solved in an alternate way. We further develop a content-aware updating strategy to dynamically refine the weight distribution to well adapt to the temporal variations of the target and background. Finally, the proposed WSCF is used to enhance two state-of-the-art CF trackers to significantly boost their tracking accuracy, with little sacrifice on the tracking speed. Extensive experiments on five benchmarks validate the effectiveness of the proposed approach. Rui-Ze Han, Wei Feng 0005, Song Wang 0002 |
IEEE Trans. Image Process. | 1 |
| 2019 | Dynamic Saliency-Aware Regularization for Correlation Filter-Based Object TrackingabstractWith a good balance between tracking accuracy and speed, correlation filter (CF) has become one of the best object tracking frameworks, based on which many successful trackers have been developed. Recently, spatially regularized CF tracking (SRDCF) has been developed to remedy the annoying boundary effects of CF tracking, thus further boosting the tracking performance. However, SRDCF uses a fixed spatial regularization map constructed from a loose bounding box and its performance inevitably degrades when the target or background show significant variations, such as object deformation or occlusion. To address this problem, we propose a new dynamic saliency-aware regularized CF tracking (DSAR-CF) scheme. In DSAR-CF, a simple yet effective energy function, which reflects the object saliency and tracking reliability in the spatial-temporal domain, is defined to guide the online updating of the regularization weight map using an efficient level-set algorithm. Extensive experiments validate that the proposed DSAR-CF leads to better performance in terms of accuracy and speed than the original SRDCF. Wei Feng 0005, Rui-Ze Han, Qing Guo 0005, Jianke Zhu, Song Wang 0002 |
IEEE Trans. Image Process. | 2 |
| 2018 | Content-Related Spatial Regularization for Visual Object TrackingabstractSpatial regularization (SR), being an effective tool to alleviate the boundary effects, can significantly improve the accuracy and robustness of correlation filters (CF) based visual object tracking. The core of SR is a spatially variant weight map that is used to regularize the online learned correlation filters by selecting more meaningful samples. However, most existing trackers apply a data-independent SR weight map. In this paper, we show that a content-related spatial regularization (CRSR) can help to further boost both the tracking accuracy and robustness. Specifically, we present to consider both frame saliency and spatial preference to online generate the CRSR weight map and propose a simple yet effective saliency-embedded CF objective function to simultaneously optimize the filters and CRSR weight map in spatial-temporal domain. Extensive experiments validate that our content-related SR outperforms the classical SR, with higher tracking accuracy and almost two times faster speed. Rui-Ze Han, Qing Guo 0005, Wei Feng 0005 |
ICME | 1 |
| 2017 | Near-surface lighting estimation and reconstructionabstractIn this paper, we propose an effective approach to estimating a near-surface lighting function from a limited number of images captured under different illuminations. Unlike classical methods relying on simplified parallel lighting model or near-point lighting model, our approach directly focuses on the much more realistic near-surface light source and formulates it as a regular grid of near-point light sources. We present an iterative joint optimization strategy to solve the scene normal, reflectance and near-point light source positions. Based on such new model, reliable relighting under arbitrary new illuminations can be faithfully reconstructed by applying the given lighting condition to the same scene. Experiments show that the proposed approach can generate more accurate re-lighting results than state-of-the-art competitors. Qian Zhang 0051, Fei-Peng Tian, Rui-Ze Han, Wei Feng 0005 |
ICME | 3 |