EDBT 2026 Demo / reviewers in the wild / expert
Shuai Wang 0027
dblp:42/1503-27
· DBLP profile ↗
29ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-1570-6570ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Gap: More Powerful Residual Fusion for Deep BNNs
Chengshuo Bai, Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Hailong Zhao, Guanqun Su |
KSEM (4) | 2 |
| 2026 | Cross-Modal Attention Guided Enhanced Fusion Network for RGB-T TrackingabstractVisual tracking that combines RGB and thermal infrared modalities (RGB-T) aims to utilize the useful information of each modality to achieve more robust object localization. Most existing tracking methods based on convolutional neural networks (CNNs) and Transformers emphasize integrating multi-modal features through cross-modal attention, but ignore the potential exploitability of complementary information learned by cross-modal attention for enhancing modal features. In this paper, we propose a novel hierarchical progressive fusion network based on cross-modal attention guided enhancement for RGB-T tracking. Specifically, the complementary information generated by cross-modal attention implicitly reflects the consistent regions of interest of important information between different modalities, which is used to enhance modal features in a targeted manner. In addition, a modal feature refinement module and a fusion module are designed based on dynamic routing to perform noise suppression and adaptive integration on the enhanced multi-modal features. Extensive experiments on GTOT, RGBT234, LasHeR and VTUAV show that our method has competitive performance compared with recent state-of-the-art methods. Jun Liu 0053, Wei Ke 0001, Shuai Wang 0027, Da Yang 0001, Hao Sheng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | InstructTrack: Language-Guided Multi-Object Tracking with Semantic-Aware AssociationabstractLocating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision–language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches. Zishun Zhou, Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Sentan Li, Da Yang 0001, Zhenglong Cui |
MMAsia | 2 |
| 2025 | Uncertainty-Specialized Tracking with Weighted Entropy based Probabilistic GraphabstractMultiple-Object Tracking (MOT) has been an attracting area in these years with excellent progresses. However, there are still complicated uncertainties caused by movement of pedestrians due to environmental disturbance and mental state. To handle this, we characterize pedestrian moving patterns through a dual-phase paradigm, and Probabilistic Graphical Model (PGM) is introduced into our work to build Hidden Markov Model (HMM) to better understand the uncertainties. Moreover, the weighted entropy mechanism is utilized in feature fusion for balanced importance between appearance and motion information. Finally, our method achieves state-of-the-art results on MOT17 and MOT20 datasets, especially in the category of graphical methods. Hao Sheng 0001, Shuai Wang 0027, Da Yang 0001, Guanqun Su |
SMC | 3 |
| 2025 | Semantic Understanding-based Open-Scene Re-IdentificationabstractAlthough current ReID (Re-Identification) methods have become relatively mature, they still require manual extraction of pedestrian images and annotation of features. They lack semantic understanding capabilities in open scenes. While some ReID models integrated with LLMs (Large Language Models) offer more comprehensive functions and better performance, they still fall short in terms of semantic understanding and cross-modal retrieval. To solve these problems, we introduce SUO-ReID, a semantic understanding-based approach for ReID in open scenes. SUO-ReID combines LVLM (Large Vision-Language Model) with ResNet (Residual Network) to extract high-level semantic features of targets in open scenes. It can also perform more flexible and complex functions, such as searching for or comparing targets with specified features, through instruction inputs. Experimental results show that SUO-ReID achieves an accuracy rate of 95.73% on datasets such as Market-1501 and DukeMTMC and exhibits excellent semantic understanding capabilities in open scenes, supporting cross-modal retrieval. It can also provide a detailed description of the features of the identified object and its surrounding scene. This study provides new insights into the application of large vision-language models in the field of ReID. Zhengrui Zhang, Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Guanqun Su |
SMC | 2 |
| 2025 | A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time RenderingabstractReal-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering. Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001 |
IEEE Trans. Computers | 8 |
| 2025 | View-Guided Cost Volume for Light Field Arbitrary-View Disparity EstimationabstractPer-view disparity estimation for light field (LF) is critical for various applications such as light field editing, but previous work mostly focuses on estimating disparity for the center view. In this paper, we propose a view-guided cost volume (VGCV), which successfully generates high-quality disparity maps for LF arbitrary view. Unlike previous methods that construct a static cost for center view only, VGCV is designed with view information and can be applicable to arbitrary-view estimation. In particular, since the key to achieving it is to condition cost on view, we extend previous static cost to a conditional one by introducing the spatial and angular information of target view into cost construction and aggregation, experiments show that this way can effectively adapt VGCV to arbitrary-view task. For construction, previous stereo-matching methods usually adopt correlation (e.g., variance) for dynamic estimation, but just using correlation can lose image structure information, which is essential for scene detail recovery, therefore we design an image-guided construction module and use cross-view attention to adapt cost for conditional construction while keeping its spatial information. Then for aggregation, we present a coordinate-guided aggregation module for VGCV regularization, which is specially designed to solve the problem of LF view deviation. Finally, we implement a Light Field Arbitrary-View Disparity Estimation Network (LFAVNet), then perform it on both synthetic and real LFs. Experiments demonstrate that LFAVNet can generate a higher-quality disparity map for arbitrary view in LF. We also extend our method to center-view estimation and light field editing tasks, which all achieve advanced performance. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Shuai Wang 0027 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | A survey for light field super-resolutionabstractCompared to 2D imaging data, the 4D light field (LF) data retains richer scene’s structure information, which can significantly improve the computer’s perception capability, including depth estimation, semantic segmentation, and LF rendering. However, there is a contradiction between spatial and angular resolution during the LF image acquisition period. To overcome the above problem, researchers have gradually focused on the light field super-resolution (LFSR). In the traditional solutions, researchers achieved the LFSR based on various optimization frameworks, such as Bayesian and Gaussian models. Deep learning-based methods are more popular than conventional methods because they have better performance and more robust generalization capabilities. In this paper, the present approach can mainly divided into conventional methods and deep learning-based methods. We discuss these two branches in light field spatial super-resolution (LFSSR), light field angular super-resolution (LFASR), and light field spatial and angular super-resolution (LFSASR) , respectively. Subsequently, this paper also introduces the primary public datasets and analyzes the performance of the prevalent approaches on these datasets. Finally, we discuss the potential innovations of the LFSR to propose the progress of our research field. Mingyuan Zhao 0001, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen, Tun Wang, Shuai Wang 0027 |
High Confid. Comput. | 9 |
| 2024 | Blockchain-Based Distributed Multiagent Reinforcement Learning for Collaborative Multiobject Tracking FrameworkabstractWith the development of smart cities, video surveillance has become more prevalent in urban areas. The rapid growth of data brings challenges to video processing and analysis. Multi-object tracking (MOT), one of the most fundamental tasks in computer vision, has a wide range of applications and development prospects. MOT aims to locate multiple objects and maintain their unique identities by analyzing the video frame by frame. Most existing MOT frameworks are deployed in centralized systems, which are convenient for management but have problems such as weak algorithm adaptability, limited system scalability, and poor data security. In this paper, we propose a distributed MOT algorithm based on multi-agent reinforcement learning (DMARL-Tracker), which formulates MOT as a Markov decision process (MDP). Each object adjusts its tracking strategy during interactions with the environment. The benchmark results on MOT17 and MOT20 prove that our proposed algorithm achieves state-of-the-art (SOTA) performance. Based on this, we further integrate DMARL-Tracker into the blockchain and propose a blockchain-based collaborative MOT framework. All nodes collaborate and share information through the blockchain, achieving adaptation in different complex scenarios while ensuring data security. The simulation results show that our framework achieves good performance in terms of tracking and resource consumption. Hao Sheng 0001, Shuai Wang 0027, Ruixuan Cong, Da Yang 0001, Yang Zhang 0032 |
IEEE Trans. Computers | 3 |
| 2024 | An Occlusion and Noise-Aware Stereo Framework Based on Light Field Imaging for Robust Disparity EstimationabstractStereo vision is widely studied for depth information extraction. However, occlusion and noise pose significant challenges to traditional methods due to failure in photo consistency. In this paper, an occlusion and noise-aware stereo framework named ONAF is proposed to get a robust depth estimation by integrating the advantages of correspondence cues and refocusing cues from light field(LF). ONAF consists of two special depth cue extractors: correspondence depth cue extractor (CCE) and refocusing depth cue extractor (RCE). CCE extracts accurate correspondence depth cues in occlusion areas based on multi-direction Ray-Epipolar Plane Images(Ray-EPIs) from LF, which are more robust than traditional multi-direction EPIs. RCE generates accurate refocusing depth cues in noise areas, benefitting from the many-to-one integration strategy and the directional perception of texture and occlusion based on multi-direction focal stacks from LF. Attention mechanism is introduced to complementarily fuse CCE and RCE to generate optimum depth maps. The experimental results prove the effectiveness of ONAF, which outperforms state-of-the-art disparity estimation methods, especially in occlusion and noise areas. Da Yang 0001, Zhenglong Cui, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Shuai Wang 0027, Zhang Xiong 0001 |
IEEE Trans. Computers | 6 |
| 2024 | Discriminative Feature Learning With Co-Occurrence Attention Network for Vehicle ReIDabstractVehicle Re-Identification (ReID) aims to find images of the same vehicle from different videos. It remains a challenging task in the video analysis field due to the huge appearance discrepancy of the same vehicle in cross-view matching and the subtle difference of different similar vehicles in same-view matching. In this paper, we propose a Co-occurrence Attention Net (CAN) to deal with these two challenges. Specifically, CAN consists of two branches, a main branch and an aware branch. The main branch is in charge of extracting global features that are consistent in most views. This feature encodes holistic information such as color and pose, however, it can not handle cross/same-view hard cases, as shown in Fig.1. Therefore, the aware branch is designed to focus on the local details and viewpoint information, which can become an important complement for those hard cases. Considering that the positions of local areas such as wheels and logos change with the viewpoint, Aware Attention Module is introduced to find the hidden relationship among local areas and seamlessly combine the viewpoint information simultaneously. Then, CAN is trained by a partition-and-reunion-based loss, which can narrow the intra-class distance and increase the inter-class distance. Further, an adaptive co-occurrence view emphasize strategy is adopted to fully utilize the learned features. Experimental results on three widely used datasets including VeRi-776, VehicleID and VERI-Wild demonstrate the effectiveness of our method and competitive performance with other state-of-the-art methods. Hao Sheng 0001, Shuai Wang 0027, Haobo Chen, Da Yang 0001, Wei Ke 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Distributed Collaborative Object Retrieval With Blockchain-Based Edge ComputingabstractIn the current industrial informatics society, the numerous cameras deployed in the modern city promote the development of various video services, such as security monitoring and object retrieval. However, traditional methods encounter data leakage risks. Some camera owners are reluctant to share their data since the video contains confidential information. Meanwhile, domain diversities between cameras bring obstacles to practical object retrieval applications. To deal with these dilemmas, we propose a blockchain-based collaborative object retrieval (BCOR) system that can protect privacy as much as possible. BCOR includes two core components: multicamera reidentification framework (MC-ReF) and multicamera collaborative chain (M2C-Chain). Specifically, MC-ReF leverages visual relevance attention net (VRANet) to distinguish object identities in edge nodes. Through domain adaptation gradient optimization, VRANet can adapt to different cameras without the need for private camera data. M2C-Chain is responsible for maintaining the security and trust of the system. Through M2C-Chain, the collaboration among different nodes is transferred into a transaction-based manner, which is validated by a deeply integrated consensus. Finally, we implement a prototype system and deploy it into a real-world outdoor scene. The experiments indicate that BCOR achieves 30%–35% average improvement in domain adaptation on mean average precision and Rank-1 indicators. The performance analysis and security experiments also prove the efficiency and stability of BCOR. Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Yang Zhang 0032, Wei Ke 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Blockchain-Empowered Distributed Multicamera Multitarget Tracking in Edge ComputingabstractThe rapid increase in the volume of video data generated from edges in the Industrial Internet of Things, opens up new possibilities for enhancing the application of video service. Multicamera multiobject tracking (MCMT) has always been a fundamental task in video surveillance or traffic control. However, the traditional MCMT methods are limited by the communication bottleneck and computation resources of the centralized curator, and suffer from security and privacy issues. In this article, we first design multicamera multihypothesis tracking (MC-MHT) framework to achieve real-time tracking performance among edge cameras. The complex association of objects is described by multiskip trees. The tracking task is well distributed to each camera. Then, we integrate multicamera tracking chain into MC-MHT to ensure security and trust. The state transition of targets in multicamera is illustrated from the perspective of blockchain transactions. The transactions are validated by an integrated tracking consensus to counter Byzantine behavior. Numerical results derived from real-world scenarios and CAMPUS dataset show that the proposed method achieves real-time performance (24–36 FPs) and 79.0–82.4 MOTA indicator, as well as reduces identity switch errors about 71% under Byzantine attack. Shuai Wang 0027, Hao Sheng 0001, Yang Zhang 0032, Da Yang 0001, Rongshan Chen |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Reinforce Model Tracklet for Multi-Object Tracking
Jianhong Ouyang, Shuai Wang 0027, Yang Zhang 0032, Yubin Wu, Hao Sheng 0001 |
CGI (3) | 2 |
| 2023 | Group Perception Based Self-adaptive Fusion Tracking
Yiyang Xing, Shuai Wang 0027, Yang Zhang 0032, Shuangye Zhao, Yubin Wu, Hao Sheng 0001 |
CGI (4) | 2 |
| 2023 | Direct Inter-Intra View Association for Light Field Super-Resolution
Da Yang 0001, Hao Sheng 0001, Shuai Wang 0027, Rongshan Chen, Zhang Xiong 0001 |
ICONIP (5) | 3 |
| 2023 | CoGAN: Cooperatively trained conditional and unconditional GAN for person image generationabstractAbstract Person image generation aims to synthesize realistic person images that follow the same distribution as the given dataset. Previous attempts can be generally categorized into two classes: conditional GAN and unconditional GAN. The former usually uses pose information as condition to make pose transfer using GAN. The generated person have the same identity as the source person. The latter generates person images from scratch, and the real person images are only used as references for the discriminator. While conditional GAN is widely studied, unconditional GAN is also worth exploring because it can synthesize person image with new identity, which is a useful manner of data augmentation. These two types of generating methods have their different advantages and disadvantages, and sometimes they are complementary. This paper proposes a CoGAN to cooperatively train two types of GANs in an end‐to‐end framework. The two GANs serve different purposes, and can learn from each other during the cooperative learning procedure. The experimental results on public datasets show that the proposed CoGAN improves the performance of both baseline methods, and achieves competitive results compared with state‐of‐the‐art methods. Yang Liu 0088, Hao Sheng 0001, Shuai Wang 0027, Yubin Wu, Zhang Xiong 0001 |
IET Image Process. | 3 |
| 2023 | Hybrid Motion Model for Multiple Object Tracking in Mobile DevicesabstractFor an intelligent transportation system, multiple object tracking (MOT) is more challenging from the traditional static surveillance camera to mobile devices of the Internet of Things (IoT). To cope with this problem, previous works always rely on additional information from multivision, various sensors, or precalibration. Only based on a monocular camera, we propose a hybrid motion model to improve the tracking accuracy in mobile devices. First, the model evaluates camera motion hypotheses by measuring optical flow similarity and transition smoothness to perform robust camera trajectory estimation. Second, along the camera trajectory, smooth dynamic projection is used to map objects from image to world coordinate. Third, to deal with trajectory motion inconsistency, which is caused by occlusion and interaction of long time interval, tracklet motion is described by the multimode motion filter for adaptive modeling. Fourth, in tracklets association, we propose a spatiotemporal evaluation mechanism, which achieves higher discriminability in motion measurement. Experiments on MOT15, MOT17, and KITTI benchmarks show that our proposed method improves the trajectory accuracy, especially in mobile devices and our method achieves competitive results over other state-of-the-art methods. Yubin Wu, Hao Sheng 0001, Yang Zhang 0032, Shuai Wang 0027, Zhang Xiong 0001, Wei Ke 0001 |
IEEE Internet Things J. | 4 |
| 2022 | Group Guided Data Association for Multiple Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Zhang Xiong 0001, Wei Ke 0001 |
ACCV (7) | 3 |
| 2022 | Mask Guided Spatial-Temporal Fusion Network for Multiple Object TrackingabstractMulti-object trackers make the association almost perfectly when no occlusion occurred between two or more targets. However, it is hard to extract reliable features on account of partial occlusion caused by a nearby object, which often leads to tracking failure. In this paper, we utilize mask to guide attention of the neural network in order to focus on the visible part of the target and design a tracklet-level feature extraction method. Then, a tracking framework is proposed based on a mask guided fusion network and multi-hypothesis tracking algorithm. Comprehensive evaluation on the MOT17 dataset shows that our approach achieves competitive results. Shuangye Zhao, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Hao Sheng 0001 |
ICIP | 3 |
| 2022 | Data Association with Graph Network for Multi-Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Wei Ke 0001, Zhang Xiong 0001 |
KSEM (1) | 3 |
| 2022 | Tracking Game: Self-adaptative Agent based Multi-object TrackingabstractMulti-object tracking (MOT) has become a hot task in multi-media analysis. It not only locates the objects but also maintains their unique identities. However, previous methods encounter tracking failures in complex scenes, since they lose most of the unique attributes of each target. In this paper, we formulate the MOT problem as Tracking Game and propose a Self-adaptative Agent Tracker (SAT) framework to solve this problem. The roles in Tracking Game are divided into two classes including the agent player and the game organizer. The organizer controls the game and optimizes the agents' actions from a global perspective. The agent encodes the attributes of targets and selects action dynamically. For these purposes, we design the State Transition Net to update the agent state and the Action Decision Net to implement the flexible tracking strategy for each agent. Finally, we present the organizer-agent coordination tracking algorithm to leverage both global and individual information. The experiments show that the proposed SAT achieves the state-of-the-art performance on both MOT17 and MOT20 benchmarks. Shuai Wang 0027, Da Yang 0001, Yubin Wu, Yang Liu 0088, Hao Sheng 0001 |
ACM Multimedia | 1 |
| 2022 | Extendable Multiple Nodes Recurrent Tracking Framework With RTU++abstractRecently, tracking-by-detection has become a popular paradigm in Multiple-object tracking (MOT) for its concise pipeline. Many current works first associate the detections to form track proposals and then score proposalns by manual functions to select the best. However, long-term tracking information is lost in this way due to detection failure or heavy occlusion. In this paper, the Extendable Multiple Nodes Tracking framework (EMNT) is introduced to model the association. Instead of detections, EMNT creates four basic types of nodes including correct, false, dummy and termination to generally model the tracking procedure. Further, we propose a General Recurrent Tracking Unit (RTU++) to score track proposals by capturing long-term information. In addition, we present an efficient generation method of simulated tracking data to overcome the dilemma of limited available data in MOT. The experiments show that our methods achieve state-of-the-art performance on MOT17, MOT20 and HiEve benchmarks. Meanwhile, RTU++ can be flexibly plugged into other trackers such as MHT, and bring significant improvements. The additional experiments on MOTS20 and CTMC-v1 also demonstrate the generalization ability of RTU++ trained by simulated data in various scenarios. Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Yang Zhang 0032, Yubin Wu |
IEEE Trans. Image Process. | 1 |
| 2021 | A General Recurrent Tracking Framework without Real DataabstractRecent progress in multi-object tracking (MOT) has shown great significance of a robust scoring mechanism for potential tracks. However, the lack of available data in MOT makes it difficult to learn a general scoring mechanism. Multiple cues including appearance, motion and etc., are limitedly utilized in current manual scoring functions. In this paper, we propose a Multiple Nodes Tracking (MNT) framework that adapts to most trackers. Based on this framework, a Recurrent Tracking Unit (RTU) is designed to score potential tracks through long-term information. In addition, we present a method of generating simulated tracking data without real data to overcome the defect of limited available data in MOT. The experiments demonstrate that our simulated tracking data is effective for training RTU and achieves state-of-the-art performance on both MOT17 and MOT16 benchmarks. Meanwhile, RTU can be flexibly plugged into classic trackers such as DeepSORT and MHT, and makes remarkable improvements as well. Shuai Wang 0027, Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Zhang Xiong 0001 |
ICCV | 1 |
| 2021 | Near-Online Tracking With Co-Occurrence Constraints in Blockchain-Based Edge ComputingabstractMultiobject tracking is a basic task in video analysis. Due to the strict requirements on efficiency and resource consumption, most of the applications on edge devices are online or near-online methods. Besides motion modeling, appearance information is also widely used for tracking. However, the influence of occlusion is usually ignored. In this article, spatial-temporal co-occurrence constraints (STCCs) features are introduced to resist occlusions by exploring the rich spatial and temporal information of tracklets. In addition, a novel blockchain-based near-online framework called co-occurrence constraints tracklet tracker (CoCTs) is proposed for cross-camera tracking. It inherits the advantages of the blockchain technology in sharing information. Based on blockchain, an efficient association mechanism and a reliable information sharing method are introduced. Experimental results show that CoCT performs high computational efficiency and low resource consumption. In the edge computing environment, it achieves real-time performance on cross-camera tracking. On the MOT17 benchmark, our method shows the state-of-the-art results compared with other online trackers. Hao Sheng 0001, Shuai Wang 0027, Yang Zhang 0032, Dongxiao Yu, Xiuzhen Cheng, Weifeng Lyu, Zhang Xiong 0001 |
IEEE Internet Things J. | 2 |
| 2020 | A Dual Scale Matching Model for Long-Term Association
Yubin Wu, Shuai Wang 0027, Yang Zhang 0032, Yanbing Chen, Wei Ke 0001, Hao Sheng 0001 |
WASA (1) | 3 |
| 2020 | Multiplex Labeling Graph for Near-Online Tracking in Crowded ScenesabstractIn recent years, the demand for intelligent devices related to the Internet of Things (IoT) is rapidly increasing. In the field of computer vision, many algorithms have been preinstalled in IoT devices to achieve higher efficiency, such as face recognition, area detection, target tracking, etc. Tracking is an important but complex task that needs high efficiency solutions in real applications. There is a common assumption that detection can only represent one pedestrian to describe nonoverlapping in physical space. In fact, the pixels of the image do not exactly correspond to the positions in the real world. In order to overcome the limitation of this assumption, we remove this unreasonable assumption and present a novel idea that each detector response can have multiple labels to describe different targets at the same time. Therefore, we propose a graph-based method for near-online tracking in this article. We introduce a detection multiplexing method for tracking in the monocular image and propose a multiplex labeling graph (MLG) model. Each node in MLG has the ability to represent multiple targets. In addition, we improve the shortage of graph-based trackers in using temporal features. We construct long short-term memory networks to model motion and appearance features for MLG optimization. On the public multiobject tracking challenge benchmark, our near-online method gains satisfactory efficiency and achieves state-of-the-art results without additional private detection as well. Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Hypothesis Testing Based Tracking With Spatio-Temporal Joint Interaction ModelingabstractData association is one of the key research in tracking-by-detection framework. Due to frequent interactions among targets, there are various relationships among trajectories in crowded scenes which leads to problems in data association, such as association ambiguity, association omission, etc. To handle these problems, we propose hypothesis-testing based tracking (HTBT) framework to build potential associations between target by constructing and testing hypotheses. In addition, a spatio-temporal interaction graph (STIG) model is introduced to describe the basic interaction patterns of trajectories and test the potential hypotheses. Based on network flow optimization, we formulate offline tracking as a MAP problem. Experimental results show that our tracking framework improves the robustness of tracklet association when detection failure occurs during tracking. On the public MOT16, MOT17 and MOT20 benchmark, our method achieves competitive results compared with other state-of-the-art methods. Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Long-Term Tracking With Deep Tracklet AssociationabstractRecently, most multiple object tracking (MOT) algorithms adopt the idea of tracking-by-detection. Relevant research shows that the performance of the detector obviously affects the tracker, while the improvement of detector is gradually slowing down in recent years. Therefore, trackers using tracklet (short trajectory) are proposed to generate more complete trajectories. Although there are various tracklet generation algorithms, the fragmentation problem still often occurs in crowded scenes. In this paper, we introduce an iterative clustering method that generates more tracklets while maintaining high confidence. Our method shows robust performance on avoiding internal identity switch. Then we propose a deep association method for tracklet association. In terms of motion and appearance, we construct motion evaluation network (MEN) and appearance evaluation network (AEN) to learn long-term features of tracklets for association. In order to explore more robust features of tracklets, a tracklet-based training mechanism is also introduced. Tracklet groups are used as the input of the networks instead of discrete detections. Experimental results show that our training method enhances the performance of the networks. In addition, our tracking framework generates more complete trajectories while maintaining the unique identity of each target as the same time. On the latest MOT 2017 benchmark, we achieve state-of-the-art results. Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 4 |