EDBT 2026 Demo / reviewers in the wild / expert
Renshu Gu
dblp:11/10145
· DBLP profile ↗
31ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-3900-2148ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-quality pattern decoding of large-scale color fabricabstractThis work develops a novel solution for generating binary patterns of large-scale fabrics using deep neural networks. It contributes to textile engineering by enabling the analysis of ancient and modern textile products. There are only two possible over-under relationships between warp and weft yarns at each crossing point, which can be simulated by a binary matrix. Generating a binary pattern from an observed fabric pattern can help designers save time and effort in reproducing fabrics. Deep neural networks have recently been applied in this field and can generate accurate binary patterns that match fabrics; however, these methods still require improvements. This paper introduces a feature point matching-based image stitching method to address the mismatch between the resolution of large fabric images and the network input requirements. Then, we preserve the contrast of color fabric patterns using principal component analysis for grayscale conversion. Finally, we propose a method for deriving pixel-wise confidence values of the label image based on receptive field size and stitching label images by accumulating weights. We show results on Jacquard fabric samples with an average of 266 thousand intersections. The ablation study showed that incorporating the two newly proposed methods achieved the highest accuracy for the textile binary pattern, with an average of 0.952 across samples. Masahiro Toyoura, Qingqi Huang, Renshu Gu, Gang Xu 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Unsupervised domain adaptation for cross-modal volumetric medical image segmentation by synergistic alignment and decoupled learning
Renshu Gu, Masahiro Toyoura, Gang Xu 0001 |
Pattern Recognit. | 1 |
| 2026 | LGD: Leveraging generative descriptions for zero-shot referring image segmentationabstractZero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of aligning and matching semantics across visual and textual modalities without training. Previous works address this challenge by utilizing Vision-Language Models and mask proposal networks for region-text matching. However, this paradigm may lead to incorrect target localization due to the inherent ambiguity and diversity of free-form referring expressions. To alleviate this issue, we present LGD (Leveraging Generative Descriptions), a framework that utilizes the advanced language generation capabilities of Multi-Modal Large Language Models to enhance region-text matching performance in Vision-Language Models. Specifically, we first design two kinds of prompts, the attribute prompt and the surrounding prompt, to guide the Multi-Modal Large Language Models in generating descriptions related to the crucial attributes of the referent object and the details of surrounding objects, referred to as attribute description and surrounding description, respectively. Secondly, three visual-text matching scores are introduced to evaluate the similarity between instance-level visual features and textual features, which determines the mask most associated with the referring expression. The proposed method achieves new state-of-the-art performance on three public datasets RefCOCO, RefCOCO+ and RefCOCOg, with maximum improvements of 9.97 % in oIoU and 11.29 % in mIoU compared to previous methods Jiachen Li 0002, Qing Xie 0002, Renshu Gu, Jinyu Xu 0001, Yongjian Liu, Xiaohan Yu 0001 |
Pattern Recognit. | 3 |
| 2026 | Feature-Preserving Offset MeshingabstractWe introduce a new offset meshing method that handles clean 3D surface meshes of arbitrary geometry and topology—where “clean” refers to meshes that are watertight, manifold, and free of self-intersections. Our approach also extends to imperfect, or “dirty,” meshes that violate these conditions, although the problem becomes significantly more difficult in such scenarios, and faithful feature preservation near defective areas cannot always be assured. In contrast to prior techniques, which have largely focused on constant-radius offsets, our method is, to our knowledge, the first to support mitered offsets while effectively preserving sharp features. Our method is designed based on several core principles: (1) explicitly generating the offset vertices and triangles with feature-capturing energy and constraints; (2) prioritizing the generation of the offset geometry before establishing its connectivity, (3) employing exact algorithms in critical pipeline steps for robustness, balancing the use of floating-point computations for efficiency, (4) applying various conservative speed up strategies including early reject non-contributing computations to the final output. Our approach further uniquely supports variable offset distances on input surface elements, offering a wider range of practical applications compared to conventional methods. For benchmarking purposes, we performed an extensive comparison against state-of-the-art offset methods using a curated subset of the Thingi10K dataset. Our results demonstrate the superiority of our approach over current state-of-the-art methods in terms of element count, feature preservation, and non-uniform offset distances of the resulting offset mesh surfaces, marking a significant advancement in the field. Hongyi Cao, Gang Xu 0001, Renshu Gu, Jinlan Xu, Timon Rabczuk, Yuzhe Luo, Xifeng Gao |
ACM Trans. Graph. | 3 |
| 2025 | OmniSR: Shadow Removal Under Direct and Indirect LightingabstractShadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in removing shadows from indirect illumination is obtaining shadow-free images to train the shadow removal network. To overcome this challenge, we propose a novel rendering pipeline for generating shadowed and shadow-free images under direct and indirect illumination, and create a comprehensive synthetic dataset that contains over 30,000 image pairs, covering various object types and lighting conditions. We also propose an innovative shadow removal network that explicitly integrates semantic and geometric priors through concatenation and attention mechanisms. The experiments show that our method outperforms state-of-the-art shadow removal techniques and can effectively generalize to indoor and outdoor scenes under various lighting conditions, enhancing the overall effectiveness and applicability of shadow removal methods. Jiamin Xu, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001 |
AAAI | 5 |
| 2025 | Detail-Preserving Latent Diffusion for Stable Shadow RemovalabstractAchieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To address this, we leverage the rich visual priors of a pre-trained Stable Diffusion (SD) model and propose a two-stage fine-tuning pipeline to adapt the SD model for stable and efficient shadow removal. In the first stage, we fix the VAE and fine-tune the denoiser in latent space, which yields substantial shadow removal but may lose some high-frequency details. To resolve this, we introduce a second stage, called the detail injection stage. This stage selectively extracts features from the VAE encoder to modulate the decoder, injecting fine details into the final results. Experimental results show that our method outperforms state-of-the-art shadow removal techniques. The cross-dataset evaluation further demonstrates that our method generalizes effectively to unseen data, enhancing the applicability of shadow removal methods. Jiamin Xu, Chi Wang 0004, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001 |
CVPR | 5 |
| 2025 | Uncertainty Estimation with Self-Distillation for semi-supervised few-shot classification
Ping Li 0006, Renshu Gu |
Knowl. Based Syst. | 3 |
| 2025 | Unlabeled data augmentation with diffusion model for semi-supervised object detection
Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008 |
Vis. Comput. | 2 |
| 2024 | Diffusers Generated Unlabeled Images Improves Semi-supervised Object DetectionabstractIn the field of object detection, particularly in medical imaging, the scarcity of data often poses a significant challenge to model performance. To address this issue, this study proposes a semi-supervised learning approach based on a generative model. We begin by fine-tuning a pre-trained generative model using our dataset to better adapt the generative model to our specific data distribution. The fine-tuned generative model is then used to generate additional unlabeled data. These generated unlabeled data, combined with the original dataset, are employed in a semi-supervised training process. Experimental results demonstrate that our method significantly enhances the performance of the object detection model, especially in scenarios with limited labeled data, such as medical imaging. By incorporating the generated unlabeled training data into the semi-supervised framework, we observed a notable improvement in model accuracy. Specifically, our experiments showed an increase of up to $6.92 \%$ after adding the generated iamges. Moreover, it is foreseeable that incorporating a higher proportion of generated unlabeled data could lead to even more significant improvements in performance. Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008 |
CW | 2 |
| 2024 | Gland Segmentation in Colon Histology Images via Attention-based Multimodal Information FusionabstractIntegrating pathological image and multimodal information such as patient metadata can help improve the segmentation performance. Most existing segmentation methods overlook the importance to exploit patient metadata, and suffer from suboptimal segmentation performance. To tackle these challenges, a novel multimodal feature fusion module (MFFM) is proposed. By incorporating a cross-modal attention mechanism and a self-attention mechanism, MFFM effectively and efficiently integrates information from different modalities. Experiments are conducted on GlaS, a glandular segmentation dataset, and the experimental results demonstrate that the method outperforms the state-of-the-art segmentation networks. Renshu Gu, Xiangyang Wu 0001, Gang Xu 0001 |
CW | 2 |
| 2024 | Chinese Ink Cartoons Generation With Multiscale Semantic ConsistencyabstractChinese Ink Cartoon (CIC) generation is challenging, due to the lack of paired photo-cartoon data and the severe geometric deformations between photos and cartoons. To tackle this challenge, in this paper, we build a high-quality CIC dataset with rich annotations, and propose a novel CIC generation method based on Multiscale Semantic Consistency (MuSeC). The CIC dataset consists of $\mathbf{1, 2 0 0}$ high-resolution CIC paintings with nearly 2,000 annotations of human faces and bodies, respectively. The generation task is formulated as an unsupervised image-to-image translation problem. First, we constrain the generated cartoon to pixel-wisely convey the semantic structure of the input photo through a learned CIC face parsing network. Additionally, we use patch-wise contrastive learning and global semantic consistency loss. In this way, the generated cartoon is optimized to precisely present the identity, attributes and structure of the input photo. Experimental results show that the proposed method can generate high-quality CICs, and outperforms previous methods both quantitatively and qualitatively. Our dataset and code have been released at github.com/AiArt-Gao/MuSeC Shuran Su, Mei Du, Jintao Mao, Renshu Gu, Fei Gao 0006 |
CW | 4 |
| 2024 | Diffusion-driven Cycle-consistent Domain Adaptation for Cross-modality Medical Image SegmentationabstractMedical image segmentation often suffers from performance degradation when applied to images from different domains. To address this, we propose DiMA-Seg (Diffusion Model Adaptation for Segmentation), a novel framework for unsupervised domain adaptation in medical image segmentation. DiMA-Seg combines GAN-based image translation with diffusion model-based feature extraction, leveraging the strengths of both approaches. Our method utilizes the hierarchical nature of diffusion models to extract multi-scale features for accurate segmentation in the target domain. Experiment on MMWHS dataset demonstrates that DiMA-Seg outperforms existing methods in segmentation accuracy. Renshu Gu, Xiangyang Wu 0001, Masahiro Toyoura, Gang Xu 0001 |
CW | 2 |
| 2024 | Micro-Action Recognition via Hierarchical Fusion and InferenceabstractMicro-actions are spontaneous body movements that indicate a person's true feelings and potential intentions, and micro-action recognition is important in human behavior analysis. Yet, recognizing micro-actions is challenging because they are subtle and appear for a very short time compared to normal actions. In this paper, we propose a micro-action recognition framework based on Hierarchical Fusion and Inference (HiFI) to capture subtle multimodal information. Specifically, we first hierarchically integrate multimodal local and global information, including the 2D key-points of faces, hands and bodies, the depth information, and the RGB image sequences. Afterward, both 3D-CNNs and Transformers are used to effectively capture local and long-range dependence. Finally, we propose a novel from-fine-to-coarse (F2C) inference strategy, based on hybrid ensemble of multi-branches, to boost the accuracy and credibility of coarse action recognition. Our solution ranked 4th in the MAC Challenge Track 1. Fan Gong, Qijian Bao, Fei Gao 0006, Renshu Gu, Gang Xu 0001 |
ACM Multimedia | 6 |
| 2024 | 3D Human Pose Estimation from Multiple Dynamic Views via Single-view Pretraining with Procrustes Alignmentabstract3D Human pose estimation from multiple cameras with unknown calibration has received less attention than it should. The few existing data-driven solutions do not fully exploit 3D training data that are available on the market, and typically train from scratch for every novel multi-view scene, which impedes both accuracy and efficiency. We show how to exploit 3D training data to the fullest and associate multiple dynamic views efficiently to achieve high precision on novel scenes using a simple yet effective framework, dubbed Multiple Dynamic View Pose estimation (MDVPose). MDVPose utilizes novel scenarios data to finetune a single-view pretrained motion encoder in multi-view setting, aligns arbitrary number of views in a unified coordinate via Procruste alignment, and imposes multi-view consistency. The proposed method achieves 22.1 mm P-MPJPE or 34.2 mm MPJPE on the challenging in-the-wild Ski-Pose PTZ dataset, which outperforms the state-of-the-art method by 24.8% P-MPJPE (-7.3 mm) and 19.0% MPJPE (-8.0 mm). It also outperforms the state-of-the-art methods by a large margin (-18.2mm P-MPJPE and -28.3mm MPJPE) on the EgoBody dataset. In addition, MDVPose achieves robust performance on the Human3.6M datasets featuring multiple static cameras. Code is available at https://github.com/iGame-Lab/MDVPose. Renshu Gu, Yixuan Si, Fei Gao 0006, Jiamin Xu, Gang Xu 0001 |
ACM Multimedia | 1 |
| 2024 | Feature-preserving quadrilateral mesh Boolean operation with cross-field guided layout blending
Haiyan Wu, Gang Xu 0001, Ran Ling, Renshu Gu |
Comput. Aided Geom. Des. | 5 |
| 2024 | Mmy-net: a multimodal network exploiting image and patient metadata for simultaneous segmentation and diagnosis
Renshu Gu, Yueyu Zhang, Lisha Wang, Dechao Chen, Yaqi Wang 0002, Ruiquan Ge, Zicheng Jiao, Juan Ye, Gangyong Jia, Linyan Wang |
Multim. Syst. | 1 |
| 2024 | Boosting Research for Carbon Neutral on Edge UWB Nodes Integration Communication Localization Technology of IoV
Ouhan Huang, Huanle Rao, Renshu Gu, Hong Xu 0014, Gangyong Jia |
IEEE Trans. Sustain. Comput. | 4 |
| 2023 | VTP: volumetric transformer for multi-view multi-person 3D pose estimation
Renshu Gu, Ouhan Huang, Gangyong Jia |
Appl. Intell. | 2 |
| 2023 | Split and Connect: A Universal Tracklet Booster for Multi-Object TrackingabstractMulti-object tracking (MOT) is an essential task in the computer vision field. With the fast development of deep learning technology in recent years, MOT has achieved great improvement. However, some challenges still remain, such as sensitiveness to occlusion, instability under different lighting conditions, and non-robustness to deformable objects, causing incorrect temporal associations. To address such common challenges in most of the existing trackers, in this paper, a tracklet booster (TBooster) algorithm is proposed to correct the association errors resulting from existing trackers. The correction of the association error from TBooster has two folds: split tracklets on potential ID-change positions and then connect multiple tracklets into one if they are from the same object. To achieve this goal, the TBooster consists of two components,i.e., Splitter and Connector. In Splitter, an architecture with stacked temporal dilated convolution blocks is employed for the splitting position prediction via label smoothing strategy with adaptive Gaussian kernels. In Connector, a multi-head self-attention-based encoder is exploited for the tracklet embedding, which is further used to connect tracklets into full tracks. We conduct sufficient experiments on MOT17 and MOT20 benchmark datasets and achieve promising results. Combined with the proposed tracklet booster, existing trackers can achieve large improvements on the IDF1 score, which shows the effectiveness of the proposed TBooster. Gaoang Wang, Yizhou Wang 0005, Renshu Gu, Weijie Hu, Jenq-Neng Hwang |
IEEE Trans. Multim. | 3 |
| 2023 | Textile image recoloring by polarization observation
Haipeng Luan, Masahiro Toyoura, Renshu Gu, Takamasa Terada, Haiyan Wu, Takuya Funatomi, Gang Xu 0001 |
Vis. Comput. | 3 |
| 2022 | IGA-Reuse-NET: A deep-learning-based isogeometric analysis-reuse approach with topology-consistent parameterizationabstractIn this paper, a deep learning framework combined with isogeometric analysis (IGA for short) called IGA-Reuse-Net is proposed for efficient reuse of numerical simulation on a set of topology-consistent models. Compared with previous data-driven numerical simulation methods only for simple computational domains, our method can predict high-accuracy PDE solutions over topology-consistent geometries with complex boundaries. UNet3+ architecture with interlaced sparse self-attention (ISSA) module is used to enhance the performance of the network. In addition, we propose a new loss function that combines a coefficients loss and a numerical solution loss. Several training datasets with topology-consistent models are constructed for the proposed framework. To verify the effectiveness of our approach, two different types of Poisson equations with different source functions are solved on three datasets with different topologies. Our framework can achieve a good trade-off between accuracy and efficiency. It outperforms the physics-informed neural network (PINN for short) model and yields promising results of prediction. Jinlan Xu, Fei Gao 0006, Charlie C. L. Wang, Renshu Gu, Timon Rabczuk, Gang Xu 0001 |
Comput. Aided Geom. Des. | 5 |
| 2022 | Unsupervised universal hierarchical multi-person 3D pose estimation for natural scenes
Renshu Gu, Zhongyu Jiang, Gaoang Wang, Kevin McQuade, Jenq-Neng Hwang |
Multim. Tools Appl. | 1 |
| 2022 | LASOR: Learning Accurate 3D Human Pose and Shape via Synthetic Occlusion-Aware Data and Neural Mesh RenderingabstractA key challenge in the task of human pose and shape estimation is occlusion, including self-occlusions, object-human occlusions, and inter-person occlusions. The lack of diverse and accurate pose and shape training data becomes a major bottleneck, especially for scenes with occlusions in the wild. In this paper, we focus on the estimation of human pose and shape in the case of inter-person occlusions, while also handling object-human occlusions and self-occlusion. We propose a novel framework that synthesizes occlusion-aware silhouette and 2D keypoints data and directly regress to the SMPL pose and shape parameters. A neural 3D mesh renderer is exploited to enable silhouette supervision on the fly, which contributes to great improvements in shape estimation. In addition, keypoints-and-silhouette-driven training data in panoramic viewpoints are synthesized to compensate for the lack of viewpoint diversity in any existing dataset. Experimental results show that we are among the state-of-the-art on the 3DPW and 3DPW-Crowd datasets in terms of pose estimation accuracy. The proposed method evidently outperforms Mesh Transformer, 3DCrowdNet and ROMP in terms of shape estimation. Top performance is also achieved on SSP-3D in terms of shape prediction accuracy. Demo and code will be available at https://igame-lab.github.io/LASOR/. Kaibing Yang, Renshu Gu, Maoyu Wang, Masahiro Toyoura, Gang Xu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Track without Appearance: Learn Box and Tracklet Embedding with Local and Global Motion Patterns for Vehicle TrackingabstractVehicle tracking is an essential task in the multi-object tracking (MOT) field. A distinct characteristic in vehicle tracking is that the trajectories of vehicles are fairly smooth in both the world coordinate and the image coordinate. Hence, models that capture motion consistencies are of high necessity. However, tracking with the standalone motion-based trackers is quite challenging because targets could get lost easily due to limited information, detection error and occlusion. Leveraging appearance information to assist object re-identification could resolve this challenge to some extent. However, doing so requires extra computation while appearance information is sensitive to occlusion as well. In this paper, we try to explore the significance of motion patterns for vehicle tracking without appearance information. We propose a novel approach that tackles the association issue for long-term tracking with the exclusive fully-exploited motion information. We address the tracklet embedding issue with the proposed reconstruct-to-embed strategy based on deep graph convolutional neural networks (GCN). Comprehensive experiments on the KITTI-car tracking dataset and UA-Detrac dataset show that the proposed method, though without appearance information, could achieve competitive performance with the state-of-the-art (SOTA) trackers. The source code will be available at https://github.com/GaoangW/LGMTracker. Gaoang Wang, Renshu Gu, Zuozhu Liu, Weijie Hu, Mingli Song, Jenq-Neng Hwang |
ICCV | 2 |
| 2021 | ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving ApplicationsabstractThe Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community. Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu |
ICMR | 10 |
| 2021 | Weakly supervised instance segmentation using multi-prior fusion
Shengyu Hao, Gaoang Wang, Renshu Gu |
Comput. Vis. Image Underst. | 3 |
| 2020 | Exploring Severe Occlusion: Multi-Person 3D Pose Estimation with Gated Convolutionabstract3D human pose estimation (HPE) is crucial in many fields, such as human behavior analysis, augmented reality/virtual reality (AR/VR) applications, and self-driving industry. Videos that contain multiple potentially occluded people captured from freely moving monocular cameras are very common in realworld scenarios, while 3D HPE for such scenarios is quite challenging, partially because there is a lack of such data with accurate 3D ground truth labels in existing datasets. In this paper, we propose a temporal regression network with a gated convolution module to transform 2D joints to 3D and recover the missing occluded joints in the meantime. A simple yet effective localization approach is further conducted to transform the normalized pose to the global trajectory. To verify the effectiveness of our approach, we also collect a new moving camera multi-human (MMHuman) dataset that includes multiple people with heavy occlusion captured by moving cameras. The 3D ground truth joints are provided by accurate motion capture (MoCap) system. From the experiments on static-camera based Human3.6M data and our own collected moving-camera based data, we show that our proposed method outperforms most state-of-the-art 2D-to-3D pose estimation methods, especially for the scenarios with heavy occlusions. Renshu Gu, Gaoang Wang, Jenq-Neng Hwang |
ICPR | 1 |
| 2020 | Multi-Person Hierarchical 3D Pose Estimation in Natural VideosabstractDespite the increasing need of analyzing human poses on the street and in the wild, multi-person 3D pose estimation using monocular static or moving camera in real-world scenarios remains a challenge, either requiring large-scale training data or high computation complexity due to the high degrees of freedom in 3D human poses. We propose a novel scheme to effectively track and hierarchically estimate 3D human poses in natural videos in an efficient fashion. Without the need of using labelled 3D training data, we formulate torso estimation as a Perspective-N-Point (PNP) problem, and limb pose estimation as an optimization problem, and hierarchically structure the high dimensional poses to efficiently address the challenge. Experiments show good performance and high efficiency of multi-person 3D pose estimation on real-world videos, including street scenarios and various human daily activities from fixed and moving cameras, resulting in great new opportunities to understand and predict human behaviors. Renshu Gu, Gaoang Wang, Zhongyu Jiang, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Exploit the Connectivity: Multi-Object Tracking with TrackletNetabstractMulti-object tracking (MOT) is an important topic and critical task related to both static and moving camera applications, such as traffic flow analysis, autonomous driving and robotic vision. However, due to unreliable detection, occlusion and fast camera motion, tracked targets can be easily lost, which makes MOT very challenging. Most recent works exploit spatial and temporal information for MOT, but how to combine appearance and temporal features is still not well addressed. In this paper, we propose an innovative and effective tracking method called TrackletNet Tracker (TNT) that combines temporal and appearance information together as a unified framework. First, we define a graph model which treats each tracklet as a vertex. The tracklets are generated by associating detection results frame by frame with the help of the appearance similarity and the spatial consistency. To compensate camera movement, epipolar constraints are taken into consideration in the association. Then, for every pair of two tracklets, the similarity, called the connectivity in the paper, is measured by our designed multi-scale TrackletNet. Afterwards, the tracklets are clustered into groups and each group represents a unique object ID. Our proposed TNT has the ability to handle most of the challenges in MOT, and achieves promising results on MOT16 and MOT17 benchmark datasets compared with other state-of-the-art methods. Gaoang Wang, Yizhou Wang 0005, Haotian Zhang 0005, Renshu Gu, Jenq-Neng Hwang |
ACM Multimedia | 4 |
| 2018 | Joint Multi-View People Tracking and Pose Estimation for 3D Scene ReconstructionabstractThe goal of data analytics in surveillance videos is to fully understand and reconstruct the 3D scene, i.e., to recover the trajectory and action of each object. In a surveillance system with camera arrays of overlapping views, we propose a novel video scene reconstruction framework to collaboratively track multiple human objects and estimate their 3D poses. First, tracklets are extracted from each single view following the tracking-by-detection paradigm. We propose an effective integration of visual and semantic object attributes, i.e., appearance models, geometry information and poses/actions, to associate tracklets across different views. Based on the optimum viewing perspectives derived from tracking, a hierarchical estimation of human poses is introduced to generate the 3D skeleton of each object. The estimated body joint points are fed back to the tracking stage to enhance tracklet association. Experiments on benchmarks of multiview tracking and 3D pose estimation validate the effectiveness of the proposed method. Renshu Gu, Jenq-Neng Hwang |
ICME | 2 |
| 2017 | Vehicle detection and recognition for intelligent traffic surveillance system
Congzhe Zhang, Renshu Gu |
Multim. Tools Appl. | 3 |