EDBT 2026 Demo / reviewers in the wild / expert
Zhong Zhou
dblp:95/4490
· DBLP profile ↗
117ranked-venue papers
10as first author
58since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 7 first-author · 36 since 2021Artificial intelligence and machine learning · 31 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 6 since 2021Computer networks · 12 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian UpdatingabstractImage geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainable abilities, they often struggle with confirmation bias, i.e., committing to early, potentially incorrect guesses driven by visual clues with varied geographic likelihoods. In this paper, we propose GeoBayes, a novel training-free framework that formulates geolocalization as a Maximum a Posteriori (MAP) estimation task over multiple geographic hypotheses and performs probabilistic thought via sequential Bayesian reasoning. GeoBayes treats each visual object and its associated geographic clues as probabilistic evidence, integrating them iteratively through a Hypothesize–Verify–Update loop. At each step, it evaluates how new evidence supports existing hypotheses and updates their posterior probabilities, gradually converging on the most probable location. This allows GeoBayes to explicitly quantify and fuse the varied geographic probabilities implied by various visual elements, reducing the risk of overcommitting to misleading clues. Furthermore, considering the natural hierarchy of geographic labels (e.g., country, city), GeoBayes introduces a state memory mechanism that stores hypotheses, inference context, and evidence scores across levels. This design enables the framework to propagate prior knowledge across levels of the geographic hierarchy and incorporate geographic structural constraints into the Bayesian update process, achieving a coarse-to-fine geo-localization. Experiments on IM2GPS3k and YFCC4K show that GeoBayes improves MLLM-based geo-localization accuracy without extra training. This demonstrates the effectiveness of probabilistic reasoning for robust and interpretable geo-localization. Kaige Li, Junhao Fang, Qichuan Geng, Zhong Zhou |
AAAI | 7 |
| 2026 | Pedestrian Scene Coverage Control Using Perceptive Quality-Based Virtual Potential FieldabstractABSTRACT Comprehensive observation of target area with pedestrians through pan‐tilt‐zoom (PTZ) camera networks is crucial in various surveillance applications. However, the dynamic configuration of PTZ cameras increases the difficulty of coordinating multiple cameras to monitor large‐scale scenes. Since coverage control in PTZ camera networks has been proven to be an NP‐hard problem, many studies have adopted virtual potential field (VPF) algorithms to efficiently obtain approximate solutions. The VPF methods treat camera viewpoints as charged particles. Through repulsive forces between these particles, PTZ camera networks expand scene coverage and reduce overlap between camera fields of view (FoVs). However, VPF‐based methods cannot leverage the scene layout and target priority information, failing to cover pedestrians and other critical areas. In this work, we introduce a unified perception quality measure framework that quantifies surveillance importance for scenes, cameras, and pedestrians. Building on this framework, we design a coverage control scheme using a perceptive quality‐based virtual potential field. This scheme models target regions and pedestrian priorities as virtual gravitational and attractive forces. It maximizes coverage of key regions, minimizes camera overlap, and supports high‐resolution monitoring and tracking of pedestrians. Extensive experiments show that our approach outperforms state‐of‐the‐art methods, achieving superior scene and pedestrian coverage performance. Liangliang Cai, Zhuocheng Liu, Zhong Zhou |
Comput. Animat. Virtual Worlds | 3 |
| 2026 | Out-of-Distribution Semantic Segmentation With Disentangled and Calibrated RepresentationabstractOut-of-distribution (OoD) semantic segmentation aims to recognize pixels of classes undefined in the training dataset. Existing methods mostly focus on training the model to fit real OoD data samples to identify OoD pixels, which requires extra data collection and annotation efforts. By contrast, synthesizing OoD data with training data provides a more resource-efficient alternative. However, synthetic data generated from controlled settings lacks diversity, causing the model to suffer from overfitting. To this end, we propose a disentangled representation learning (DRL) method to guide the model to disentangle semantic-related and semantic-unrelated features from synthetic OoD data. DRL encourages the model to utilize the former to identify semantic categories, rather than overfitting to such semantic-unrelated features as synthetic artificiality. Specifically, DRL first incorporates two disentanglers to extract the semantic-related and -unrelated features and then applies a shuffle and reconstruction mechanism to regularize the disentangled features. Furthermore, to facilitate disentangling, we propose a pixel-wise feature similarity calibration (PSC) module, which utilizes more accurate ID-OoD similarity to calibrate inaccurate ID-OoD similarity learned exclusively from ID data. Thus, PSC delivers accurate and stable pixel-wise features for effective disentangling. Extensive experiments illustrate that the proposed method exhibits strong generalization ability. It attains 74.04% AuPRC and 20.82% FPR on Road Anomaly, 69.85% AuPRC and 5.78% FPR on Fishyscapes LostAndFound Validation Set, using SegFormer with the MiT-B5 backbone. Source code is available at https://github.com/WanMotion/DisentangledOoDSeg. Maoxian Wan, Kaige Li, Qichuan Geng, Binyi Su, Zhong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MaskScene: Hierarchical Conditional Masked Models for Real-Time 3D Indoor Scene SynthesisabstractIndoor scene synthesis is essential for creative industries. Driven by this demand, recent advances in scene synthesis using diffusion and autoregressive models have shown promising results. However, existing models struggle to achieve real-time performance, high visual fidelity, and flexible scene editing simultaneously. To tackle this challenge, we propose MaskScene, a novel hierarchical conditional masked model for real-time 3D indoor scene synthesis and editing. Specifically, MaskScene introduces a hierarchical scene representation that explicitly encodes scene relationships, semantics, and tokens. Based on this representation, we design a hierarchical conditional masked modeling architecture that enables parallel and iterative decoding, conditioned on both semantics and relationships. MaskScene leverages local object masking and hierarchical scene structures to infer occluded regions from partial observations, enabling rapid construction of 3D indoor environments that accurately reflect real-world scenes. Compared to state-of-the-art methods, MaskScene achieves 80× faster generation speed and improves scene quality by 10%, while also supporting zero-shot editing, such as scene completion and rearrangement, without extra fine-tuning. Qichuan Geng, Zhong Zhou, Wenfeng Song |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Enhancing open-vocabulary scene understanding via push-pull alignment in gaussian splatting
Shengjia Liang, Yuan Xiong, Qichuan Geng, Zhong Zhou |
Vis. Comput. | 6 |
| 2026 | CleanSplat: curriculum structural Gaussian splatting for spot-free novel view synthesisabstract3D Gaussian splatting (3DGS) has gained significant attention for its real-time, photorealistic rendering in novel view synthesis. However, its performance degrades severely when applied to real-world scenes with transients that break cross-view consistency. Existing methods typically attempt to identify and mask out these transients, but inherent masking errors often leave behind incoherent floating Gaussians, resulting in spotted artifacts that degrade the final scene quality. To address these limitations, we propose CleanSplat, a curriculum structural Gaussian splatting method that purifies the scene Gaussians through curriculum 3DGS optimization for transient removal and structural pruning of spotted artifacts. We introduce a curriculum-guided masking paradigm that generates coarse-to-fine transient masks from multi-scale features. The progressive optimization is driven by modulating the masking supervision based on current training state. To clear the spots, we propose a structure-aware handling strategy that employs a superpoint graph (SPG) partitioning of the Gaussians to perform principled identification and hierarchical pruning. This allows for the filtering of both intra-superpoint outliers and entire spurious superpoints based on local 3D coherence instead of only simple photometric consistency. By integrating curriculum 3DGS optimization and structural pruning, our method effectively separates the transients and purifies the static scene Gaussians. Extensive experiments on challenging datasets demonstrate that CleanSplat significantly outperforms state-of-the-art methods, delivering more detailed and cleaner novel view synthesis. Shengjia Liang, Qichuan Geng, Yuan Xiong, Zhong Zhou |
Virtual Real. Intell. Hardw. | 5 |
| 2025 | Incremental Few-Shot Semantic Segmentation Via Multi-Level Switchable Visual Prompts
Maoxian Wan, Kaige Li, Qichuan Geng, Zhong Zhou |
ICCV | 5 |
| 2025 | Adaptive Prompt Learning via Gaussian Outlier Synthesis for Out-Of-Distribution Detection
Dongyu She, Zhong Zhou |
ICCV | 3 |
| 2025 | 3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Contextabstract3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due to the misalignment between these reference points and the lanes, it is difficult to obtain complete and discriminative context for complex instances. In this paper, we are devoted to introducing 3D priors adaptive to lane appearances, which serve as references to aggregate the lane context. Specifically, we propose a projection-consistent reference generation strategy to keep the projected 3D reference points geometrically consistent with the corresponding lanes in images. In addition, a segmentation-lifting denoising strategy is designed to improve the ability of the model to map the lane segmentation into 3D space. To leverage more lane-related information, we propose a decoupled lane-context aggregation module by considering the perspectives of individual geometries and integrated layout, namely intra-lane and inter-lane context. Extensive experiments on the OpenLane dataset show that our approach outperforms previous methods and achieves the state-of-the-art performance. The code will be made publicly available. Yiqiu Bing, Huilin Niu, Hong Zhang 0009, Zhong Zhou, Qichuan Geng |
ICRA | 5 |
| 2025 | Aligned and Detail Guided Retrieval: Multi-scale Fine-Grained Features Enhancement for End-to-End Person Search
Feihu Yan, Kunlin Zou, Zhong Zhou, Haiyong Chen |
ICXR | 4 |
| 2025 | Understanding Matters: Semantic-Structural Determined Visual Relocalization for Large ScenesabstractScene Coordinate Regression (SCR) estimates 3D scene coordinates from 2D images, and has become an important approach in visual relocalization. Existing methods exhibit high localization accuracy in small scenes, but still face substantial challenges in large-scale scenes, which usually have significant variations in depth, scale, and occlusion. Although structure-guided scene partitioning is commonly adopted, the over-partitioned elements and large feature variances within subscenes impede the estimation of the 3D coordinates, introducing misleading information for subsequent processing. To address the above-mentioned issues, we propose the Semantic-Structural Determined Visual Relocalization method for SCR, which leverages semantic-structural partition learning and partition-determined pose refinement to better understand the semantic and structural information on large scenes. Firstly, we partition the scene into small subscenes with label assignments, ensuring semantic consistency and structural continuity within each subscene. A classifier is then trained with sampling-based learning to predict these labels. Secondly, the partition predictions are encoded into embeddings and integrated with local features for intra-class compactness and inter-class separation, producing partition-aware features. To further decrease feature variances, we employ a discriminability metric and suppress ambiguous points, improving subsequent computations. Experimental results on the Cambridge Landmarks dataset demonstrate that the proposed method achieves significant improvements with fewer training costs on large-scale scenes, reducing the median error by 38% compared to the state-of-the-art SCR method DSAC*. Code is available: https://gitee.com/VR_NAVE/ss-dvr. Jingyi Nie, Liangliang Cai, Qichuan Geng, Zhong Zhou |
IJCAI | 4 |
| 2025 | Boosting Few-Shot Open-Set Object Detection via Prompt Learning and Robust Decision BoundaryabstractFew-shot Open-set Object Detection (FOOD) poses a challenge in many open-world scenarios. It aims to train an open-set detector to detect known objects while rejecting unknowns with scarce training samples. Existing FOOD methods are subject to limited visual information, and often exhibit an ambiguous decision boundary between known and unknown classes. To address these limitations, we propose the first prompt-based few-shot open-set object detection framework, which exploits additional textual information and delves into constructing a robust decision boundary for unknown rejection. Specifically, as no available training data for unknown classes, we select pseudo-unknown samples with Attribution-Gradient based Pseudo-unknown Mining (AGPM), which leverages the discrepancy in attribution gradients to quantify uncertainty. Subsequently, we propose Conditional Evidence Decoupling (CED) to decouple and extract distinct knowledge from selected pseudo-unknown samples by eliminating opposing evidence. This optimization process can enhance the discrimination between known and unknown classes. To further regularize the model and form a robust decision boundary for unknown rejection, we introduce Abnormal Distribution Calibration (ADC) to calibrate the output probability distribution of local abnormal features in pseudo-unknown samples. Our method achieves superior performance over previous state-of-the-art approaches, improving the average recall of unknown class by 7.24% across all shots in VOC10-5-5 dataset settings and 1.38% in VOC-COCO dataset settings. Our source code is available at https://gitee.com/VR_NAVE/ced-food. Zhaowei Wu, Binyi Su, Qichuan Geng, Zhong Zhou |
IJCAI | 5 |
| 2025 | DGDiff: Immersive 3D Indoor Scene Synthesis via Dialog-Graph Conditioned DiffusionabstractImmersive 3D indoor scene synthesis is essential for applications such as AR/VR and 3D content creation. However, existing approaches fail to meet the immersive AR/VR requirements for fidelity, user-system interaction, and production speed simultaneously. Traditional scene synthesis methods are overly rigid, limiting user interactivity, whereas large language model (LLM)-based approaches suffer from slow response times and imprecise spatial structuring. To address these issues, we propose DGDiff, a novel dialog-graph conditioned diffusion framework for immersive, controllable, continuous synthesis and editing of 3D indoor scenes. This framework combines a conversational module powered by LLMs with a multimodal diffusion model. The conversational module translates user dialogue into structured semantic graphs, while the diffusion model integrates textual and graph-based conditions to synthesize realistic, editable indoor scenes. Experimental results demonstrate that DGDiff outperforms single-modality baselines, achieving an improvement of over 10 % in FID and a reduction of approximately 30 % in response time for dynamic scene interactive editing, offering an immersive and user-friendly synthesis experience. Project page: https://gitee.com/VR_NAVE/dgdiff.git Qichuan Geng, Zhong Zhou |
ISMAR | 4 |
| 2025 | A fast and accurate detection model of internal defects in tunnel lining for ground penetrating radar image data
Shirong Zhou, Zhong Zhou |
Adv. Eng. Informatics | 4 |
| 2025 | A deep learning-based algorithm for fast identification of multiple defects in tunnels
Zhong Zhou, Hongchang Li, Shirong Zhou, Longbin Yan |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Visual Expansion and Real-Time Calibration for Pan-Tilt-Zoom Cameras Assisted by Panoramic ModelsabstractABSTRACT Pan‐tilt‐zoom (PTZ) cameras, which dynamically adjust their field of view (FOV), are pervasive in large‐scale scenes, such as train stations, squares, and airports. In real scenarios, PTZ cameras are required to quickly judge their directions using contextual clues from the surrounding environment. To achieve this goal, some research projects camera videos into three‐dimensional (3D) models or panoramas and allows operators to establish spatial relationships. However, these works face several challenges in terms of real‐time processing, localization accuracy, and realistic reference. To address this problem, a visual expansion and real‐time calibration for PTZ cameras assisted by panoramic models is proposed. The calibration method consists of three parts: Providing a real environment background by building a panoramic model, meeting the needs of real‐time processing by establishing a PTZ camera motion estimation model and achieving high‐precision alignment between PTZ images and panoramic models using only two feature point pairs. Our methods were validated using both the public and our Scene dataset. The experimental results indicate that our method outperforms other state‐of‐the‐art methods in terms of real‐time processing, accuracy, and robustness. Liangliang Cai, Zhong Zhou |
Comput. Animat. Virtual Worlds | 2 |
| 2025 | RMGNet: The Progressive Relationship-Mining Graph Neural Network for Text-to-Image Person Re-IdentificationabstractThe Text-to-Image Person Re-identification (TI-ReID) task objective is to precisely identify the person’s images with the textual description of the person. The mainstream research methods focus on cross-modal aligning local features, and overlook the learning of intra-modal and cross-modal relationships between different features. This renders the person features lacking in high-level semantic information. To resolve such issues, we propose the Progressive Relationship-Mining Graph Network (RMGNet), including the Intra-Modal Relationship-Mining (IMRM) and the Cross-Modal Relationship-Mining (CMRM) module. These modules are employed to model and mine semantic relationship information among different features. Specifically, the IMRM module models and mines the high-level semantic interrelationships inherent in the image and text features. The CMRM module introduces the nearest neighbor method to model cross-modal semantic relationships to enhance the cross-modal semantic correspondence capabilities of person features. On this basis, we design the Adaptive Corner Center (Acc) loss and the Coarse-to-Fine Learning (C2FL) strategy. These ensure the network receives consistent and effective metric learning supervision throughout the entirety of the training process. To validate the efficacy of the proposed method, extensive experiments are conducted on three prevalent datasets: CHUK-PEDES, ICFC-PEDES, and RSTPReid. The achieved mAP of 70.59%, 41.62%, and 49.58% surpassed those current state-of-the-art methods. Xin Zhang 0116, Kun Liu 0009, Xinwang Wang, Zhong Zhou, Haiyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MAP: Masked Adversarial Perturbation for Boosting Black-Box Attack TransferabilityabstractThe transferability of adversarial examples is vital for black-box attacks, as it enables the adversary to deceive the target model without knowing its internals. Despite numerous methods focusing on transferability, they still struggle with transferring across models with distinct architectural components (e.g., CNNs and ViTs). In this work, we argue that the limited adversarial perturbation diversity leads to overfitting of the surrogate model, which acts as a key factor in reducing transferability. To this end, we propose a Masked Adversarial Perturbation (MAP) method to boost adversarial transferability across various architectures from a novel perspective of diversifying perturbation. Specifically, MAP randomly masks perturbation patches during iterations and compels the remaining ones to retain the attack effect, which diversifies perturbations to mitigate their overfitting to the surrogate model. Naturally, MAP spreads perturbation over local patches to alleviate their co-adaptation and prevent perturbations from overly relying on specific patterns. Consequently, it can deceive convolution operation and self-attention mechanism indiscriminately by attacking their basic input units, i.e., a single patch, showing superior transferability over previous methods. Extensive experiments illustrate that MAP consistently and significantly boosts diverse black-box attacks to achieve state-of-the-art performance. Kaige Li, Maoxian Wan, Qichuan Geng, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 6 |
| 2025 | LangLoc: Language-Driven Localization via Formatted Spatial Description GenerationabstractExisting localization methods commonly employ vision to perceive scene and achieve localization in GNSS-denied areas, yet they often struggle in environments with complex lighting conditions, dynamic objects or privacy-preserving areas. Humans possess the ability to describe various scenes using natural language, effectively inferring their location by leveraging the rich semantic information in these descriptions. Harnessing language presents a potential solution for robust localization. Thus, this study introduces a new task, Language-driven Localization, and proposes a novel localization framework, LangLoc, which determines the user's position and orientation through textual descriptions. Given the diversity of natural language descriptions, we first design a Spatial Description Generator (SDG), foundational to LangLoc, which extracts and combines the position and attribute information of objects within a scene to generate uniformly formatted textual descriptions. SDG eliminates the ambiguity of language, detailing the spatial layout and object relations of the scene, providing a reliable basis for localization. With generated descriptions, LangLoc effortlessly achieves language-only localization using text encoder and pose regressor. Furthermore, LangLoc can add one image to text input, achieving mutual optimization and feature adaptive fusion across modalities through two modality-specific encoders, cross-modal fusion, and multimodal joint learning strategies. This enhances the framework's capability to handle complex scenes, achieving more accurate localization. Extensive experiments on the Oxford RobotCar, 4-Seasons, and Virtual Gallery datasets demonstrate LangLoc's effectiveness in both language-only and visual-language localization across various outdoor and indoor scenarios. Notably, LangLoc achieves noticeable performance gains when using both text and image inputs in challenging conditions such as overexposure, low lighting, and occlusions, showcasing its superior robustness. Changhao Chen, Kaige Li, Yuan Xiong, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 6 |
| 2025 | FDCPNet:feature discrimination and context propagation network for 3D shape representationabstractThree-dimensional (3D) shape representation using mesh data is essential in various applications, such as virtual reality and simulation technologies. Current methods for extracting features from mesh edges or faces struggle with complex 3D models because edge-based approaches miss global contexts and face-based methods overlook variations in adjacent areas, which affects the overall precision. To address these issues, we propose the Feature Discrimination and Context Propagation Network (FDCPNet), which is a novel approach that synergistically integrates local and global features in mesh datasets. FDCPNet is composed of two modules: (1) the Feature Discrimination Module, which employs an attention mechanism to enhance the identification of key local features, and (2) the Context Propagation Module, which enriches key local features by integrating global contextual information, thereby facilitating a more detailed and comprehensive representation of crucial areas within the mesh model. Experiments on popular datasets validated the effectiveness of FDCPNet, showing an improvement in the classification accuracy over the baseline MeshNet. Furthermore, even with reduced mesh face numbers and limited training data, FDCPNet achieved promising results, demonstrating its robustness in scenarios of variable complexity. Yuan Xiong, Zhong Zhou |
Virtual Real. Intell. Hardw. | 5 |
| 2024 | LRGAN: Learnable Weighted Recurrent Generative Adversarial Network for End-to-End Shadow GenerationabstractIn augmented reality(AR) applications, it is a challenging task to generate virtual object shadows while maintaining the precision and consistency of virtual and real areas. To achieve the above target, we propose a learnable weighted recurrent generative adversarial network(LRGAN) for end-to-end shadow generation. Without any additional computational overhead, LRGAN only needs to analyze the background context to create a bridge between the target shadows and the background. Our model incorporates multiple progressive steps to recurrently compute the precise reference masks, based on which a fine-grained shadow generation module generates the shadows. A learnable weighted fusion module, which can normalize pixel values to deal with pixel overflow, fuses the generated shadows with the original image. In addition, we adopt the combined method of module training and the whole model training. Experimental results show that our proposed LRGAN not only improves the plausibility of shadow location and shape but also achieves color harmony in the shadow areas. In the absence of other prior knowledge or post-processing, it outperforms the State-of-the-Art end-to-end methods. Junsheng Xue, Hai Huang 0001, Zhong Zhou, Shibiao Xu, Aoran Chen |
IJCNN | 3 |
| 2024 | SADNet: Generating immersive virtual reality avatars by real-time monocular pose estimationabstractSummary Generating immersive virtual reality avatars is a challenging task in VR/AR applications, which maps physical human body poses to avatars in virtual scenes for an immersive user experience. However, most existing work is time‐consuming and limited by datasets, which does not satisfy immersive and real‐time requirements of VR systems. In this paper, we aim to generate 3D real‐time virtual reality avatars based on a monocular camera to solve these problems. Specifically, we first design a self‐attention distillation network (SADNet) for effective human pose estimation, which is guided by a pre‐trained teacher. Secondly, we propose a lightweight pose mapping method for human avatars that utilizes the camera model to map 2D poses to 3D avatar keypoints, generating real‐time human avatars with pose consistency. Finally, we integrate our framework into a VR system, displaying generated 3D pose‐driven avatars on Helmet‐Mounted Display devices for an immersive user experience. We evaluate SADNet on two publicly available datasets. Experimental results show that SADNet achieves a state‐of‐the‐art trade‐off between speed and accuracy. In addition, we conducted a user experience study on the performance and immersion of virtual reality avatars. Results show that pose‐driven 3D human avatars generated by our method are smooth and attractive. Ling Jiang 0001, Yuan Xiong, Qianqian Wang 0011, Wei Wu 0008, Zhong Zhou |
Comput. Animat. Virtual Worlds | 6 |
| 2024 | DreamWalk: Dynamic remapping and multiperspectivity for large-scale redirected walkingabstractSummary Redirected walking (RDW) provides an immersive user experience in virtual reality applications. In RDW, the size of the physical play area is limited, which makes it challenging to design the virtual path in a larger virtual space. Mainstream RDW approaches rigidly manipulate gains to guide the user to follow predetermined rules. However, these methods may cause simulator sickness, boundary collision, and reset. Static mapping approaches warp the virtual path through expensive vertex replacement in the stage of model pre‐processing. They are restricted to narrow spaces with non‐looping pathways, partition walls, and planar surfaces. These methods fail to provide a smooth walking experience for large‐scale open scenes. To tackle these problems, we propose a novel approach that dynamically redirects the user to walk in a non‐linear virtual space. More specifically, we propose a Bezier‐curve‐based mapping algorithm to warp the virtual space dynamically and apply multiperspective fusion for visualization augmentation. We conduct comparable experiments to show its superiority over state‐of‐the‐art large‐scale redirected walking approaches on our self‐collected photogrammetry dataset. Yuan Xiong, Tianjing Li, Zhong Zhou |
Comput. Animat. Virtual Worlds | 4 |
| 2024 | Toward Generalized Few-Shot Open-Set Object DetectionabstractOpen-set object detection (OSOD) aims to detect the known categories and reject unknown objects in a dynamic world, which has achieved significant attention. However, previous approaches only consider this problem in data-abundant conditions, while neglecting the few-shot scenes. In this paper, we seek a solution for the generalized few-shot open-set object detection (G-FOOD), which aims to avoid detecting unknown classes as known classes with a high confidence score while maintaining the performance of few-shot detection. The main challenge for this task is that few training samples induce the model to overfit on the known classes, resulting in a poor open-set performance. We propose a new G-FOOD algorithm to tackle this issue, named Few-shOt Open-set Detector (FOOD), which contains a novel class weight sparsification classifier (CWSC) and a novel unknown decoupling learner (UDL). To prevent over-fitting, CWSC randomly sparses parts of the normalized weights for the logit prediction of all classes, and then decreases the co-adaptability between the class and its neighbors. Alongside, UDL decouples training the unknown class and enables the model to form a compact unknown decision boundary. Thus, the unknown objects can be identified with a confidence probability without any threshold, prototype, or generation. We compare our method with several state-of-the-art OSOD methods in few-shot scenes and observe that our method improves the F-score of unknown classes by 4.80%-9.08% across all shots in VOC-COCO dataset settings. Binyi Su, Hua Zhang 0008, Jingzhi Li 0002, Zhong Zhou |
IEEE Trans. Image Process. | 4 |
| 2024 | Exploring Scale-Aware Features for Real-Time Semantic Segmentation of Street ScenesabstractReal-time semantic segmentation of street scenes is an essential and challenging task for autonomous driving systems, which needs to achieve both high accuracy and efficiency. Moreover, numerous objects and stuff at different scales in street scenes further increase the difficulty of this task. To address this challenge, we develop a lightweight and high-accuracy network termed Scale-Aware Network (SANet), which aims to selectively aggregate multi-scale features while maintaining high efficiency. In SANet, we first design a Selective Context Encoding (SCE) module, which considers the intrinsic differences of various pixels to selectively encode private contexts for each pixel, thus learning more desirable contextual features while reducing redundancy. With the context embedding in hand, we then design a Selective Feature Fusion (SFF) module to recursively fuses them with multiple features at different levels or scales to generate scale-aware features, where each feature map contains scale-specific information. Extensive experiments on challenging street scene datasets, i.e., Cityscapes and CamVid, illustrate that our SANet achieves a leading trade-off between segmentation accuracy and speed. Concretely, our method yields$78.1\%$mIoU at$109.0$FPS on the Cityscapes test set and$77.2\%$mIoU at$250.4$FPS on the CamVid test set. Code will be available at https://github.com/kaigelee/SANet. Kaige Li, Qichuan Geng, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Multimodality Adaptive Transformer and Mutual Learning for Unsupervised Domain Adaptation Vehicle Re-IdentificationabstractUnsupervised Domain Adaptation Vehicle Re-Identification (UDA vehicle re-ID) aims to enable the model trained in the source domain dataset to adapt to the target domain data and obtain accurate re-identification results, which has received widespread attention due to its practicality in the field of intelligent transportation systems. Most current UDA vehicle re-ID research ignores the mining and utilization of attribute information. Meanwhile, the Convolutional Neural Networks-based (CNN-based) network will cause the loss of fine-grained information, reducing the expression and generalization ability of vehicle features. To alleviate such issues, we are motivated by the Transformer, which can exploit distinguishable attribute information and fuse multimodal features effectively. Therefore, this paper proposes a Multimodality Adaptive Transformer Network (MATNet) to intensify the ability to learn vehicle fine-grained features related to attributes. Moreover, the noise contained in pseudo-labels assigned by cluster algorithms interferes with the performance of the UDA vehicle re-ID method. We also design the Dual Mutual Dynamic Update Pseudo-Label generation strategy (DMDU) to improve the accuracy of pseudo-labels and alleviate error accumulation. The strategy is based on mutual learning, which can effectively utilize the congruous and particular knowledge of the two models to generate pseudo-labels. Extensive experiments on two large-scale public datasets, including VeRi-776 and VehicleID, illustrate that our method outperforms the state-of-the-art methods. Xin Zhang 0116, Yunan Ling, Kaige Li, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | QR-CLIP: Introducing Explicit Knowledge for Location and Time ReasoningabstractThis article focuses on reasoning about the location and time behind images. Given that pre-trained vision-language models (VLMs) exhibit excellent image and text understanding capabilities, most existing methods leverage them to match visual cues with location and time-related descriptions. However, these methods cannot look beyond the actual content of an image, failing to produce satisfactory reasoning results, as such reasoning requires connecting visual details with rich external cues (e.g., relevant event contexts). To this end, we propose a novel reasoning method, QR-CLIP , that aims at enhancing the model’s ability to reason about location and time through interaction with external explicit knowledge such as Wikipedia. Specifically, QR-CLIP consists of two modules: (1) The Quantity module abstracts the image into multiple distinct representations and uses them to search and gather external knowledge from different perspectives that are beneficial to model reasoning. (2) The Relevance module filters the visual features and the searched explicit knowledge and dynamically integrates them to form a comprehensive reasoning result. Extensive experiments demonstrate the effectiveness and generalizability of QR-CLIP . On the WikiTiLo dataset, QR-CLIP boosts the accuracy of location (country) and time reasoning by 7.03% and 2.22%, respectively, over previous SOTA methods. On the more challenging TARA dataset, it improves the accuracy for location and time reasoning by 3.05% and 2.45%, respectively. The source code is at https://github.com/Shi-Wm/QR-CLIP . Dehong Gao, Yuan Xiong, Zhong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | VirtualLoc: Large-scale Visual Localization Using Virtual ImagesabstractRobust and accurate camera pose estimation is fundamental in computer vision. Learning-based regression approaches acquire six-degree-of-freedom camera parameters accurately from visual cues of an input image. However, most are trained on street-view and landmark datasets. These approaches can hardly be generalized to overlooking use cases, such as the calibration of the surveillance camera and unmanned aerial vehicle. Besides, reference images captured from the real world are rare and expensive, and their diversity is not guaranteed. In this article, we address the problem of using alternative virtual images for visual localization training. This work has the following principle contributions: First, we present a new challenging localization dataset containing six reconstructed large-scale three-dimensional scenes, 10,594 calibrated photographs with condition changes, and 300k virtual images with pixelwise labeled depth, relative surface normal, and semantic segmentation. Second, we present a flexible multi-feature fusion network trained on virtual image datasets for robust image retrieval. Third, we propose an end-to-end confidence map prediction network for feature filtering and pose estimation. We demonstrate that large-scale rendered virtual images are beneficial to visual localization. Using virtual images can solve the diversity problem of real images and leverage labeled multi-feature data for deep learning. Experimental results show that our method achieves remarkable performance surpassing state-of-the-art approaches. To foster research on improvement for visual localization using synthetic images, we release our benchmark at https://github.com/YuanXiong/contributions . Yuan Xiong, Zhong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Mirror world: creating digital twins of the space and persons from video streamings
Ling Jiang 0001, Liangliang Cai, Wei Wu 0008, Zhong Zhou |
Vis. Comput. | 4 |
| 2023 | HSIC-based Moving Weight Averaging for Few-Shot Open-Set Object DetectionabstractWe study the problem of few-shot open-set object detection (FOOD), whose goal is to quickly adapt a model to a small set of labeled samples and reject unknown class samples. Recent works usually use the weight sparsification for unknown rejection, but due to the lack of tailored considerations for data-scarce scenarios, the performance is not satisfactory. In this work, we solve the challenging few-shot open-set object detection problems from three aspects. First, different from previous pseudo-unknown sample mining methods, we employ the evidential uncertainty estimated by the Dirichlet distribution of probability to mine the pseudo-unknown samples from the foreground and background proposal space. Second, based on the statistical analysis between the number of pseudo-unknown samples and the Intersection over Union (IoU), we propose an IoU-aware unknown objective, which sharps the unknown decision boundary by considering the localization quality. Third, to suppress the over-fitting problem and improve the model's generalization ability for unknown rejection, we propose the HSIC-based (Hilbert-Schmidt Independence Criterion) moving weight averaging to update the weights of classification and regression heads, which considers the degree of independence between the current weights and previous weights stored in the long-term memory banks. We compare our method with several state-of-the-art methods and observe that our method improves the mean recall of unknown classes by 12.87% across all shots in the VOC-COCO dataset settings. Our code is available at https://github.com/binyisu/food. Binyi Su, Hua Zhang 0008, Zhong Zhou |
ACM Multimedia | 3 |
| 2023 | Recommending Analogical APIs via Knowledge Graph EmbeddingabstractLibrary migration, which replaces the current library with a different one to retain the same software behavior, is common in software evolution. An essential part of this is finding an analogous API for the desired functionality. However, due to the multitude of libraries/APIs, manually finding such an API is time-consuming and error-prone. Researchers created automated analogical API recommendation techniques, notably documentation-based methods. Despite potential, these methods have limitations, e.g., incomplete semantic understanding in documentation and scalability issues. In this study, we present KGE4AR, a novel documentation-based approach using knowledge graph (KG) embedding for recommending analogical APIs during library migration. KGE4AR introduces a unified API KG to comprehensively represent documentation knowledge, capturing high-level semantics. It further embeds this unified API KG into vectors for efficient, scalable similarity calculation. We assess KGE4AR with 35,773 Java libraries in two scenarios, with and without target libraries. KGE4AR notably outperforms state-of-the-art techniques (e.g., 47.1%-143.0% and 11.7%-80.6% MRR improvements), showcasing scalability with growing library counts. Mingwei Liu 0002, Yiling Lou, Xin Peng 0001, Zhong Zhou, Xueying Du, Tianyong Yang |
ESEC/SIGSOFT FSE | 5 |
| 2023 | Model-guided 3D stitching for augmented virtual environment
Zhong Zhou, Zhe Zhu, Jingdi You |
Sci. China Inf. Sci. | 1 |
| 2023 | Center-point-pair detection and context-aware re-identification for end-to-end multi-object tracking
Yunan Ling, Yuanzhe Yang, Chengxiang Chu, Zhong Zhou |
Neurocomputing | 5 |
| 2023 | Geometric-driven structure recovery from a single omnidirectional image based on planar depth map learning
Likai Xiao, Zhong Zhou |
Neural Comput. Appl. | 3 |
| 2023 | PVEL-AD: A Large-Scale Open-World Dataset for Photovoltaic Cell Anomaly DetectionabstractThe anomaly detection in photovoltaic (PV) cell electroluminescence (EL) image is of great significance for the vision-based fault diagnosis. Many researchers are committed to solving this problem, but a large-scale open-world dataset is required to validate their novel ideas. We build a PV EL Anomaly Detection (PVEL-AD1, 2, 3) dataset for polycrystalline solar cell, which contains 36 543 near-infrared images with various internal defects and heterogeneous background. This dataset contains anomaly free images and anomalous images with ten different categories. Moreover, 37 380 ground truth bounding boxes are provided for eight types of defects. We also carry out a comprehensive evaluation of the state-of-the-art object detection methods based on deep learning. The evaluation results on this dataset provide the initial benchmark, which is convenient for follow-up researchers to conduct experimental comparisons. To the best of our knowledge, this is the first public dataset for PV solar cell anomaly detection that provides box-wise ground truth. Furthermore, this dataset can also be used for the evaluation of many computer vision tasks such as few-shot detection, one-class classification, and anomaly generation. Binyi Su, Zhong Zhou, Haiyong Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Context and Spatial Feature Calibration for Real-Time Semantic SegmentationabstractContext modeling or multi-level feature fusion methods have been proved to be effective in improving semantic segmentation performance. However, they are not specialized to deal with the problems of pixel-context mismatch and spatial feature misalignment, and the high computational complexity hinders their widespread application in real-time scenarios. In this work, we propose a lightweight Context and Spatial Feature Calibration Network (CSFCN) to address the above issues with pooling-based and sampling-based attention mechanisms. CSFCN contains two core modules: Context Feature Calibration (CFC) module and Spatial Feature Calibration (SFC) module. CFC adopts a cascaded pyramid pooling module to efficiently capture nested contexts, and then aggregates private contexts for each pixel based on pixel-context similarity to realize context feature calibration. SFC splits features into multiple groups of sub-features along the channel dimension and propagates sub-features therein by the learnable sampling to achieve spatial feature calibration. Extensive experiments on the Cityscapes and CamVid datasets illustrate that our method achieves a state-of-the-art trade-off between speed and accuracy. Concretely, our method achieves 78.7% mIoU with 70.0 FPS and 77.8% mIoU with 179.2 FPS on the Cityscapes and CamVid test sets, respectively. The code is available at https://nave.vr3i.com/ and https://github.com/kaigelee/CSFCN. Kaige Li, Qichuan Geng, Maoxian Wan, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 5 |
| 2023 | Web-based mixed reality video fusion with remote renderingabstractMixed Reality (MR) video fusion system fuses video imagery with 3D scenes. It makes the scene much more realistic and helps the users understand the video contents and temporalspatial correlation between them, thus reducing the user’s cognitive load. Nowadays, MR video fusion has been used in various applications. However, video fusion systems require powerful client machines because video streaming delivery, stitching, and rendering are computation-intensive. Moreover, huge bandwidth usage is also another critical factor that affects the scalability of video fusion systems. The framework proposed in this paper overcomes this client limitation by utilizing remote rendering. Furthermore, the framework we built is based on browsers. Therefore, the user could try the MR video fusion system with a laptop or even pad, no extra plug-ins or application programs need to be installed. Several experiments on diverse metrics demonstrate the effectiveness of the proposed framework. Zhong Zhou |
Virtual Real. Intell. Hardw. | 2 |
| 2022 | RepF-Net: Distortion-Aware Re-projection Fusion Network for Object Detection in Panorama Image
Zhong Zhou |
ACCV (3) | 3 |
| 2022 | Coverage Control for PTZ Camera Networks Using Scene Potential MapabstractPan-Tilt-Zoom (PTZ) camera networks are pervasive in many applications, such as video surveillance, sports analysis and epidemic prevention. However, it is difficult to control several PTZ cameras observing a large scene due to the high complexity of PTZ camera networks. This problem, called the coverage control for PTZ cameras, draws broad attention of researchers. Virtual potential field (VPF) can enhance the coverage performance and reduce the overlap region of cameras, but it does not fully take the actual scenario into account. In this work, we introduce scene potential map (SPM) to the VPF method, investigating both the scene and the task in the optimization. The proposed scene potential map characterizes the importance of target region and evaluates the perception quality of PTZ camera. Then we propose a novel virtual force analysis method to optimalize poses of the PTZ camera network. We also develop a region partition method based on the perception quality measure to divide the target region and achieve better zoom level, realizing an excellent coverage performance of PTZ camera networks. Finally, the evaluation experiments clearly demonstrated that our proposed coverage control scheme can achieve remarkably better performance than state-of-the-arts. Liangliang Cai, Hanyuan Ma, Zhuocheng Liu, Zhaoxin Li, Zhong Zhou |
ICME | 5 |
| 2022 | Multi-Task Feature Decomposition Based Marginal Distribution for Person SearchabstractPerson search is a composite task, aiming at locating and identifying a query person from uncropped images. It requires jointly solving Pedestrian Detection and Person Re-identification. One major challenge in person search is the contradictory goals of detection and re-identification. The model has to simultaneously model the universality and specificity of persons. In this paper, we propose a novel parameter-free approach called Feature Decomposition Person Search (FDPS) to separate various tasks. FDPS decomposes the ROI feature map to extract sub-features based on the marginal distribution for different tasks. Also, we find that the Online Instance Match loss pays imbalanced attention to positive and negative categories. We present a Balance Online Instance Match (BOIM) loss to enhance the contribution of negative categories during training. Our method achieves the state-of-the-art performance in one-step methods on two prevailing benchmarks, with high efficiency. Yuanzhe Yang, Qichuan Geng, Chengxiang Chu, Zhong Zhou |
ICME | 5 |
| 2022 | Head Point Positioning and Spatial-Channel Self-Attention Network for Multi-Object TrackingabstractMulti-Object Tracking (MOT) aims to generate trajectories for multiple objects in the surveillance scene. This is a challenging task because the pedestrians in tracking video often gather together and occlude each other. Consequently, the two main problems in the popular tracking-by-detection framework are how to alleviate unreliable detection and extract robust object appearance features. In this paper, we propose a new tracking method that is composed of two novel types of modules - an object detection strategy based on pedestrian head point positioning and a Spatial-Channel Self-Attention feature extraction network (SCSAN). Specifically, the proposed detection strategy generates more accurate tracking object bounding boxes with Soft-Head-NMS, which combines the advantages of object detection and head point positioning. The head point location information is used as a guidance to screen unreliable detection. The SCSAN utilizes the Spatial-Channel Self-Attention mechanism to lead and determine the optimal attention value for each area and channel. Extensive experiments are carried out to demonstrate the proposed tracker achieves competitive results and is state-of-the-art in half metrics. Yuanzhe Yang, Chengxiang Chu, Zhong Zhou |
ICPR | 5 |
| 2022 | Robust and efficient edge-based visual odometryabstractVisual odometry, which aims to estimate relative camera motion between sequential video frames, has been widely used in the fields of augmented reality, virtual reality, and autonomous driving. However, it is still quite challenging for state-of-the-art approaches to handle low-texture scenes. In this paper, we propose a robust and efficient visual odometry algorithm that directly utilizes edge pixels to track camera pose. In contrast to direct methods, we choose reprojection error to construct the optimization energy, which can effectively cope with illumination changes. The distance transform map built upon edge detection for each frame is used to improve tracking efficiency. A novel weighted edge alignment method together with sliding window optimization is proposed to further improve the accuracy. Experiments on public datasets show that the method is comparable to state-of-the-art methods in terms of tracking accuracy, while being faster and more robust. Feihu Yan, Zhaoxin Li, Zhong Zhou |
Comput. Vis. Media | 3 |
| 2022 | Mutual purification for unsupervised domain adaptation in person re-identification
Lei Zhang 0195, Qishuai Diao, Zhong Zhou, Wei Wu 0008 |
Neural Comput. Appl. | 4 |
| 2022 | Person Re-identification with pose variation aware data augmentation
Lei Zhang 0195, Qishuai Diao, Zhong Zhou, Wei Wu 0008 |
Neural Comput. Appl. | 4 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Attention-Guided Collaborative CountingabstractExisting crowd counting designs usually exploit multi-branch structures to address the scale diversity problem. However, branches in these structures work in a competitive rather than collaborative way. In this paper, we focus on promoting collaboration between branches. Specifically, we propose an attention-guided collaborative counting module (AGCCM) comprising an attention-guided module (AGM) and a collaborative counting module (CCM). The CCM promotes collaboration among branches by recombining each branch's output into an independent count and joint counts with other branches. The AGM capturing the global attention map through a transformer structure with a pair of foreground-background related loss functions can distinguish the advantages of different branches. The loss functions do not require additional labels and crowd division. In addition, we design two kinds of bidirectional transformers (Bi-Transformers) to decouple the global attention to row attention and column attention. The proposed Bi-Transformers are able to reduce the computational complexity and handle images in any resolution without cropping the image into small patches. Extensive experiments on several public datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting methods. Hong Mo, Wenqi Ren, Feihu Yan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 5 |
| 2022 | FSRDD: An Efficient Few-Shot Detector for Rare City Road Damage DetectionabstractRoad damage detection (RDD) is indispensable for safe autonomous driving. Existing RDD models focus on designing feature representations following expert knowledge. However, collecting and labeling all types of samples is time-consuming and leads to insufficient training data. To alleviate the adverse effect of few training samples, a novel few-shot road damage detector (FSRDD) is proposed in this paper to detect rare road damages. The proposed FSRDD includes three stages. First, fully annotated abundant base classes are leveraged to train a base detector, where ghost attention (GA) and proposal feature metric (PFM) modules are developed to eliminate the redundant information and measure the proposal features, respectively. Second, the recognition branch of the detector is fine-tuned using a few samples of all classes. Finally, the test set is inferred with the help of an offline scale-aware prototypical calibration block (SPCB). Extensive experiments show that our FSRDD achieves 10-shot rare road damage detection with 33.4% and 12.9% mAP50 on RDD and CNRDD datasets, respectively, significantly outperforming state-of-the-art methods. Binyi Su, Hua Zhang 0008, Zhaohui Wu 0005, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | 3D scene graph prediction from point cloudsabstractIn this study, we propose a novel 3D scene graph prediction approach for scene understanding from point clouds. It can automatically organize the entities of a scene in a graph, where objects are nodes and their relationships are modeled as edges. More specifically, we employ the DGCNN to capture the features of objects and their relationships in the scene. A Graph Attention Network (GAT) is introduced to exploit latent features obtained from the initial estimation to further refine the object arrangement in the graph structure. A one loss function modified from cross entropy with a variable weight is proposed to solve the multi-category problem in the prediction of object and predicate. Experiments reveal that the proposed approach performs favorably against the state-of-the-art methods in terms of predicate classification and relationship prediction and achieves comparable performance on object classification prediction. The 3D scene graph prediction approach can form an abstract description of the scene space from point clouds. Fanfan Wu, Feihu Yan, Zhong Zhou |
Virtual Real. Intell. Hardw. | 4 |
| 2021 | Monocular Dense SLAM with Consistent Deep Depth Prediction
Feihu Yan, Jiawei Wen, Zhaoxin Li, Zhong Zhou |
CGI | 4 |
| 2021 | Object Quality Guided Feature Fusion for Person Re-identificationabstractPerson re-identification (Re-ID) is an essential task in computer vision, which aims to match a person of interest across multiple non-overlapping camera views. It is a fundamental challenging task because of the conflicts between large variations of samples and the limited scale of training sets. Data augmentation method based on generative adversarial network (GAN) is an efficient way to relieve this dilemma. However, existing methods do not consider how to keep identity information and filter the noise of the generated auxiliary samples during Re-ID training. In this paper, we propose object quality guided feature fusion network for person re-identification, which consists of a self-supervised object quality estimation module and a feature fusion module. Specifically, the former evaluates the quality of the auxiliary data to filter the noise and the disturbing features, while the later accomplishes the feature fusion based on object quality estimation in the collection-to-collection recognition manner to make full use of auxiliary data. Extensive performance analysis and experiments are conducted on two benchmark datasets (Market-1501 and DukeMTMC-reID) to show that our proposed approach outperforms or shows comparable results to the existing best performed methods. Lei Zhang 0195, Qishuai Diao, Danyang Huang, Zhong Zhou, Wei Wu 0008 |
ICTAI | 5 |
| 2021 | Diabetic Retinopathy Grading Base on Contrastive Learning and Semi-supervised Learning
Yunchao Gu, JunJun Pan, Zhong Zhou |
ISBRA | 4 |
| 2021 | Distortion-Aware Room Layout Estimation from A Single Fisheye ImageabstractOmnidirectional images of 180° or 360° field of view provide the entire visual content around the capture cameras, giving rise to more sophisticated scene understanding and reasoning and bringing broad application prospects for VR/AR/MR. As a result, researches on omnidirectional image layout estimation have sprung up in recent years. However, existing layout estimation methods designed for panorama images cannot perform well on fisheye images, mainly due to lack of public fisheye dataset as well as the significantly differences in the positions and degree of distortions caused by different projection models. To fill theses gaps, in this work we first reuse the released large-scale panorama datasets and reproduce them to fisheye images via projection conversion, thereby circumventing the challenge of obtaining high-quality fisheye datasets with ground truth layout annotations. Then, we propose a distortion-aware module according to the distortion of the orthographic projection (i.e., OrthConv) to perform effective features extraction from fisheye images. Additionally, we exploit bidirectional LSTM with two-dimensional step mode for horizontal and vertical prediction to capture the long-range geometric pattern of the object for the global coherent predictions even with occlusion and cluttered scenes. We extensively evaluate our deformable convolution for room layout estimation task. In comparison with state-of-the-art approaches, our approach produces considerable performance gains in real-world dataset as well as in synthetic dataset. This technology provides high-efficiency and low-cost technical implementations for VR house viewing and MR video surveillance. We present an MR-based building video surveillance scene equipped with nine fisheye lens can achieve an immersive hybrid display experience, which can be used for intelligent building management in the future. Likai Xiao, Zhaoxin Li, Zhong Zhou |
ISMAR | 5 |
| 2021 | Long-Term Visual Localization with Semantic Enhanced Global RetrievalabstractVisual localization under varying conditions such as changes in illumination, season and weather is a fundamental task for applications such as autonomous navigation. In this paper, we present a novel method of using semantic information for global image retrieval. By exploiting the distribution of different classes in a semantic scene, the discriminative features of the scene’s structure layout is embedded into a normalized vector that can be used for retrieval, i.e. semantic retrieval. Color image retrieval is based on low-level visual features extracted by algorithms or Convolutional Neural Networks (CNNs), while semantic retrieval is based on high-level semantic features which are robust in scene appearance variations. By combining semantic retrieval with color image retrieval in the global retrieval step, we show that these two methods can complement with each other and significantly improve the localization performance. Experiments on the challenging CMU Seasons dataset show that our method is robust across large variations of appearance and achieves state-of-the-art localization performance. Hongrui Chen, Yuan Xiong, Zhong Zhou |
MSN | 4 |
| 2021 | Asymmetric Mutual Learning for Unsupervised Cross-Domain Person Re-identification
Danyang Huang, Lei Zhang 0195, Qishuai Diao, Wei Wu 0008, Zhong Zhou |
PRICAI (3) | 5 |
| 2021 | Online Multi-Object Tracking with Pose-Guided Object Location and Dual Self-Attention Network
Yuanzhe Yang, Chengxiang Chu, Zhong Zhou |
PRICAI (3) | 5 |
| 2021 | DSNet: Deep Shadow Network for Illumination EstimationabstractIllumination consistency has applications to modeling and rendering in virtual reality. In 3D reconstruction and Mixed Reality(MR) fusion, the appearance of a large-scale outdoor scene may change in response to lighting and seasons, for example. Since 3D reconstruction from scratch is costly, it is helpful to be able to update existing models with recently captured photographs. However, the illumination conditions of the captured photograph can be arbitrary, making it challenging to fit to the existing model. To tackle this problem, this paper proposes a novel approach that can precisely estimate the illumination of the input image. Our Deep Shadow Network (DSNet) collaboratively utilizes illumination-based data augmentation for sun position estimation, along with a dataset of illumination-based augmented renderings. Our run-time rendering and optimization strategy is also discussed. We show that accurate simulation of illumination can improve the performance of visual applications including place recognition and long-term localization. Experimental results validate the effectiveness of the proposed approach, and show its superiority over the state-of-the-art. Yuan Xiong, Hongrui Chen, Zhe Zhu, Zhong Zhou |
VR | 5 |
| 2021 | Weakly supervised object-aware convolutional neural networks for semantic feature matching
Wei Lyu, Lang Chen, Zhong Zhou, Wei Wu 0008 |
Neurocomputing | 3 |
| 2021 | Gated Path Selection Network for Semantic SegmentationabstractSemantic segmentation is a challenging task that needs to handle large scale variations, deformations, and different viewpoints. In this paper, we develop a novel network named Gated Path Selection Network (GPSNet), which aims to adaptively select receptive fields while maintaining the dense sampling capability. In GPSNet, we first design a two-dimensional SuperNet, which densely incorporates features from growing receptive fields. And then, a Comparative Feature Aggregation (CFA) module is introduced to dynamically aggregate discriminative semantic context. In contrast to previous works that focus on optimizing sparse sampling locations on regular grids, GPSNet can adaptively harvest free form dense semantic context information. The derived adaptive receptive fields and dense sampling locations are data-dependent and flexible which can model various contexts of objects. On two representative semantic segmentation datasets, i.e., Cityscapes and ADE20K, we show that the proposed approach consistently outperforms previous methods without bells and whistles. Qichuan Geng, Hong Zhang 0009, Xiaojuan Qi 0001, Gao Huang 0001, Ruigang Yang, Zhong Zhou |
IEEE Trans. Image Process. | 6 |
| 2020 | Attention Guided Region Division for Crowd CountingabstractCrowd counting has drawn more and more attention in computer vision. There are two mainstream approaches to deal with crowd counting tasks, regression and detection. Regression-based methods usually overestimate the count in sparse areas, while detection-based methods tend to underestimation in dense areas. In this paper, we propose a two-branch network combining regression and detection. We introduce the attention mechanism to make the network adaptively divide dense and sparse areas and employ appropriate methods on them respectively. The regression branch predicts density map in extremely dense areas. An improved detection network is applied to detect multi-scale heads in relatively sparse areas. Our method is able to obtain precise head bounding boxes in sparse areas with ensuring counting accuracy in dense areas. Experimental results show that our method achieves state-of-the-art on challenging public crowd counting datasets. Xiaoqi Pan, Hong Mo, Zhong Zhou, Wei Wu 0008 |
ICASSP | 3 |
| 2020 | Pose Variation Adaptation for Person Re- identificationabstractPerson re-identification (reid) aims at matching pedestrians observed from non-overlapping camera views. It has important applications in surveillance video analysis such as human retrieval, human tracking and activity analysis. Although a large number of effective feature learning and distance metric optimizing approaches have been proposed, it still suffers from pedestrian appearance variations caused by pose changing. Most of the previous methods address this problem by learning a pose-invariant descriptor subspace. In this paper, we propose a pose variation adaptation method for person reid. It can reduce the probability of deep learning network over-fitting. Specifically, we introduce a pose transfer generative adversarial network with a similarity measurement module. With the learned pose transfer model, training images can be transferred to any given poses, and with the original images, forming an augmented training dataset. It increases data diversity against over-fitting. In contrast to previous GAN-based methods, we consider the influence of pose variations on similarity measures to generate shaper and more realistic samples for person reid. Besides, we optimize hard example mining to introduce a novel manner of samples used with the learned pose transfer model. It focuses on the inferior samples which are caused by pose variations to increase the number of effective hard examples for learning discriminative features and improving the generalization ability. We extensively conduct comparative evaluations to demonstrate the advantages and superiorities of the proposed method over the state-of-the-art person reid approaches on Market-1501 and DukeMTM C- reID. Lei Zhang 0195, Qishuai Diao, Zhong Zhou, Wei Wu 0008 |
ICPR | 5 |
| 2020 | Automatic façade recovery from single nighttime image
Qichuan Geng, Zhong Zhou, Wei Wu 0008 |
Frontiers Comput. Sci. | 3 |
| 2020 | Background Noise Filtering and Distribution Dividing for Crowd CountingabstractCrowd counting is a challenging problem due to the diverse crowd distribution and background interference. In this paper, we propose a new approach for head size estimation to reduce the impact of different crowd scale and background noise. Different from just using local information of distance between human heads, the global information of the people distribution in the whole image is also under consideration. We obey the order of far- to near-region (small to large) to spread head size, and ensure that the propagation is uninterrupted by inserting dummy head points. The estimated head size is further exploited, such as dividing the crowd into parts of different densities and generating a high-fidelity head mask. On the other hand, we design three different head mask usage mechanisms and the corresponding head masks to analyze where and which mask could lead to better background filtering1. Based on the learned masks, two competitive models are proposed which can perform robust crowd estimation against background noise and diverse crowd scale. We evaluate the proposed method on three public crowd counting datasets of ShanghaiTech [2], UCFQNRF [3] and UCFCC_50 [4]. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting approaches. Hong Mo, Wenqi Ren, Yuan Xiong, Xiaoqi Pan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 5 |
| 2019 | Multi-scale Vehicle Re-identification Using Self-adapting Label Smoothing RegularizationabstractVehicle re-identification (re-id) plays an important role in intelligent surveillance. Since difference vehicle models may have similar appearances, together with the problem of image scale variations, the vehicle re-id remains long-term challenging. We present a novel multi-scale vehicle re-id framework using self-adapting label smoothing regularization (SLSR). It integrates the appearance information from multi-scale images to alleviate the influence of scale changes caused by perspectives. To enhance the generalization ability in feature representations, we design the self-adapting label smoothing regulation in semi-supervised training process. It dynamically assigns labels to fake images to realize data augmentation. We validate the effectiveness of our proposed framework on popular VeRi and VehicleID datasets. Extensive experimental results demonstrate that our method outperforms most state-of-the-art methods on both datasets. Especially, we exceeds the latest method by 3.81% in mAP and 5.32% in rank-1 on VeRi dataset. Lei Zhang 0195, Zhong Zhou, Wei Wu 0008 |
ICASSP | 4 |
| 2019 | 3D Room Reconstruction from A Single Fisheye ImageabstractWe propose a rapid and accurate approach to recover the layout of a room automatically from a single fisheye image. It decomposes the fisheye image to a set of perspective images and jointly extract line images from the fisheye image and perspective images for geometric information. The semantic information gained from semantic segmentation on a cylinder expansion of the fisheye image are then used for structure line determination. By considering distinct features contained in the perspective images, the invalid hypotheses are filtered effectively and the most accurate structure lines are selected to minimize computational cost. To evaluate the effectiveness of the proposed approach, we construct an annotated fisheye image dataset. Comprehensive experimental evaluation on the dataset illustrate that our proposed approach produces higher quality layout estimations than existing layout reconstruction approaches and being 6 times faster in the reconstruction time. Yuehua Wang, Zhong Zhou |
IJCNN | 5 |
| 2019 | SANTLR: Speech Annotation Toolkit for Low Resource Languages
Zhong Zhou, Siddharth Dalmia, Alan W. Black, Florian Metze |
INTERSPEECH | 2 |
| 2019 | Handling pure camera rotation in semi-dense monocular SLAM
Feihu Yan, Zhong Zhou |
Vis. Comput. | 3 |
| 2019 | A survey on image and video stitchingabstractImage/video stitching is a technology for solving the field of view (FOV) limitation of images/ videos. It stitches multiple overlapping images/videos to generate a wide-FOV image/video, and has been used in various fields such as sports broadcasting, video surveillance, street view, and entertainment. This survey reviews image/video stitching algorithms, with a particular focus on those developed in recent years. Image stitching first calculates the corresponding relationships between multiple overlapping images, deforms and aligns the matched images, and then blends the aligned images to generate a wide-FOV image. A seamless method is always adopted to eliminate such potential flaws as ghosting and blurring caused by parallax or objects moving across the overlapping regions. Video stitching is the further extension of image stitching. It usually stitches selected frames of original videos to generate a stitching template by performing image stitching algorithms, and the subsequent frames can then be stitched according to the template. Video stitching is more complicated with moving objects or violent camera movement, because these factors introduce jitter, shakiness, ghosting, and blurring. Foreground detection technique is usually combined into stitching to eliminate ghosting and blurring, while video stabilization algorithms are adopted to solve the jitter and shakiness. This paper further discusses panoramic stitching as a special-extension of image / video stitching. Panoramic stitching is currently the most widely used application in stitching. This survey reviews the latest image/video stitching methods, and introduces the fundamental principles/advantages/weaknesses of image/video stitching algorithms. Image/video stitching faces long-term challenges such as wide baseline, large parallax, and low-texture problem in the overlapping region. New technologies may present new opportunities to address these issues, such as deep learning-based semantic correspondence, and 3D image stitching. Finally, this survey discusses the challenges of image/video stitching and proposes potential solutions. Wei Lyu, Zhong Zhou, Lang Chen |
Virtual Real. Intell. Hardw. | 2 |
| 2018 | Unified Framework for Joint Attribute Classification and Person Re-identification
Chenxin Sun, Lei Zhang 0195, Yuehua Wang, Wei Wu 0008, Zhong Zhou |
ICANN (1) | 6 |
| 2018 | Multi-Attribute Driven Vehicle Re-Identification with Spatial-Temporal Re-RankingabstractVehicle re-identification (re-id) is a promising topic, which focuses on retrieving the same vehicles across different cameras. It is challenging due to the variations of illumination and camera viewpoints. To solve these problems, we present a multi -attribute driven vehicle re-id approach to learn discriminative representations. The proposed approach consists of a multi-branch architecture and a re-ranking strategy. The multi-branch architecture extracts color, model, and appearance features, which explicitly leverages the vehicle attribute cues to enhance the generalization ability, especially for the different vehicles with similar appearance and the same vehicles with different orientations. The re-ranking strategy introduces the spatial-temporal relationship among vehicles from multiple cameras to construct the similar appearance sets and utilizes Jaccard distance between these similar appearance sets to re-rank. Extensive experimental results demonstrate that our proposed approach significantly outperforms state-of-the-art re-id methods on the popular VeRi-776 dataset and VehiclelD dataset. Zhong Zhou, Wei Wu 0008 |
ICIP | 3 |
| 2018 | Orientation-Guided Similarity Learning for Person Re-identificationabstractPerson re-identification (re-id) is a promising topic in computer vision, which concentrates on similarity learning of individuals across different camera views. It remains challenging due to the unpredictable orientation variations, the partial occlusions, and the inaccurate detections. To solve these problems, we present an orientation-guided similarity learning architecture to learn discriminative feature representations and define similarity metric for person re-id. Our proposed architecture explicitly leverages pedestrian orientation and body part cues to enhance the generalization ability. In the architecture, an orientation-guided loss function that pulls the positive samples with the same orientations closer is designed to alleviate the orientation variations. Meanwhile, an aligned dense network with pose estimation is presented to extract robust global-local fusion representations, which effectively exploits local features to overcome partial occlusions. In the end, we introduce a two-stage Top-k reranking strategy to optimize initial re-id results by min-hash and weighted distance. Extensive experimental results demonstrate that our proposed approach significantly outperforms state-of-the-art re-id methods on the popular CUHK03, Market1501, and DukeMTMC-reID datasets. Chenxin Sun, Yuehua Wang, Zhong Zhou, Wei Wu 0008 |
ICPR | 5 |
| 2018 | Online Inter-Camera Trajectory Association Exploiting Person Re-Identification and Camera TopologyabstractOnline inter-camera trajectory association is a promising topic in intelligent video surveillance, which concentrates on associating trajectories belong to the same individual across different cameras according to time. It remains challenging due to the inconsistent appearance of a person in different cameras and the lack of spatio-temporal constraints between cameras. Besides, the orientation variations and the partial occlusions significantly increase the difficulty of inter-camera trajectory association. Targeting to solve these problems, this work proposes an orientation-driven person re-identification (ODPR) and an effective camera topology estimation based on appearance features for online inter-camera trajectory association. ODPR explicitly leverages the orientation cues and stable torso features to learn discriminative feature representations for identifying trajectories across cameras, which alleviates the pedestrian orientation variations by the designed orientation-driven loss function and orientation aware weights. The effective camera topology estimation introduces appearance features to generate the correct spatio-temporal constraints for narrowing the retrieval range, which improves the time efficiency and provides the possibility for intelligent inter-camera trajectory association in large-scale surveillance environments. Extensive experimental results demonstrate that our proposed approach significantly outperforms most state-of-the-art methods on the popular person re-identification datasets and the public multi-target, multi-camera tracking benchmark. Sichen Bai, Chang Xing, Zhong Zhou, Wei Wu 0008 |
ACM Multimedia | 5 |
| 2018 | MR video fusion: interactive 3D modeling and stitching on wide-baseline videosabstractA major challenge facing camera networks today is how to effectively organizing and visualizing videos in the presence of complicated network connection and overwhelming and even increasing amount of data. Previous works focus on 2D stitching or dynamic projection to 3D models, such as panorama and Augmented Virtual Environment (AVE), and haven't given an ideal solution. We present a novel method of multiple video fusion in 3D environment, which produces a highly comprehensive imagery and yields a spatio-temporal consistent scene. User initially interact with a newly designed background model named video model to register and stitch videos' background frames offline. The method then fuses the offline results to render videos in a real time manner. We demonstrate our system on 3 real scenes, each of which contains dozens of wide-baseline videos. The experimental results show that, our 3D modeling interface developed with the our presented model and method can efficiently assist the users to seamlessly integrate videos by comparing to commercial-off-the-shelf software with less operating complexity and more accurate 3D environment. The stitching method proposed by us is much more robust against the position, orientation, attribute differences among videos than the start-of-the-art methods. More importantly, this study sheds light on how to use the 3D techniques to solve 2D problems in realistic and we validate its feasibility. Mingjun Cao 0002, Jingdi You, Yuehua Wang, Zhong Zhou |
VRST | 6 |
| 2018 | Survey of recent progress in semantic image segmentation with CNNs
Qichuan Geng, Zhong Zhou, Xiaochun Cao |
Sci. China Inf. Sci. | 2 |
| 2017 | MR sand table: Mixing real-time video streaming in physical modelsabstractA novel prototype of MR (Mixed Reality) Sand Table is presented in this paper, that fuses multiple real-time video streaming into a physically united view. The main processes include geometric calibration and alignment, image blending and the final projection. Firstly we proposed a two-step MR alignment scheme which estimates the transform matrix between input video streaming and the sand table for coarse alignment, and deforms the input frame using moving least squares for accurate alignment. To overcome the video border distinction problem, we make a border-adaptive image stitching with brightness diffusion to merge the overlapping area. With the projection, the video area can be mixed into the sand table in real-time to provide a live physical mixed reality model. We build a prototype to demonstrate the effectiveness of the proposed method. This design could also be easily extended to large size with help of multiple projectors. The system proposed in this paper supports multiple user interaction in a broad area of applications such as surveillance, demonstration, action preview and discussion assistances. Zhong Zhou, Zhiyi Bian, Zheng Zhuo |
VR | 1 |
| 2017 | Automatic Mesh Animation Preview With User Voting-Based RefinementabstractWith the rapid growth in the number and quality of mesh animations, browsing mesh animations wastes considerable bandwidth and requires tremendous rendering resources. Accordingly, the preview technique, which provides users a rapid understanding of a mesh animation before downloading, has received increasing attention. In this paper, we propose an automatic mesh animation preview method that incorporates the interframe motion saliency, intraframe surface saliency, user preference, and camera smoothness constraint to formulate the viewpoint selection as a minimization problem. Then, the minimization is solved by finding the shortest path, and the viewpoints for mesh animation preview are generated accordingly. A voting mechanism is introduced into this process to collect user feedbacks and periodically use user voting feedbacks to refine the preview camera path. The experiment results show that our mesh preview method helps users acquire a good understanding of animation contents. A user study demonstrates that our preview results are superior to those generated by typical preview methods in terms of the subjective visual quality. Zhong Zhou, Jingchang Zhang |
IEEE Trans. Multim. | 1 |
| 2016 | Game theory-based model for maximizing SSP utility in cognitive radio networks
Zhong Zhou, Wei Wu 0008 |
Comput. Commun. | 2 |
| 2015 | Light field projection for lighting reproductionabstractWe propose a novel approach to generate 4D light field in the physical world for lighting reproduction. The light field is generated by projecting lighting images on a lens array. The lens array turns the projected images into a controlled anisotropic point light source array which can simulate the light field of a real scene. In terms of acquisition, we capture an array of light probe images from a real scene, based on which an incident light field is generated. The lens array and the projectors are geometric and photometrically calibrated, and an efficient resampling algorithm is developed to turn the incident light field into the images projected onto the lens array. The reproduced illumination, which allows per-ray lighting control, can produce realistic lighting result on real objects, avoiding the complex process of geometric and material modeling. We demonstrate the effectiveness of our approach with a prototype setup. Zhong Zhou, Xiaofeng Qiu, Ruigang Yang, Qinping Zhao |
VR | 1 |
| 2015 | Video driven pedestrian visualization with characteristic appearancesabstractAugmented virtual environment (AVE) could visualize plausible live views from videos by projecting dynamic imagery to the 3D environment. Static objects in a video can be rendered in new views since they are easily modeled beforehand, while moving ones that don't have exact online models will be distorted from different views without proper depth. To cope with the problem, we introduce a novel method to visualize pedestrians, which are common in outdoor surveillance. Our method detects pedestrians and produces their trajectories. Then such pedestrian characteristic appearances as geometric information, texture and walking animation are transferred to a stand-in 3D animation model in the virtual environment. Experiments show our visualization can reveal person's characteristic appearances. Wei Wu 0008, Zhong Zhou |
VRST | 3 |
| 2015 | An efficient MAC protocol for underwater multi-user uplink communication networks
Yu Luo 0001, Lina Pu, Zheng Peng 0001, Zhong Zhou, Jun-Hong Cui |
Ad Hoc Networks | 4 |
| 2015 | Stable and Fast Fluid-Solid Coupling for Incompressible SPHabstractAbstract The solid boundary handling has been a research focus in physically based fluid animation. In this paper, we propose a novel stable and fast particle method to couple predictive–corrective incompressible smoothed particle hydrodynamics and geometric lattice shape matching (LSM), which animates the visually realistic interaction of fluids and deformable solids allowing larger time steps or velocity differences. By combining the boundary particles sampled from solids with a momentum‐conserving velocity‐position correction scheme, our approach can alleviate the particle deficiency issues and prevent the penetration artefacts at the fluid–solid interfaces simultaneously. We further simulate the stable deformation and melting of solid objects coupled to smoothed particle hydrodynamics fluids based on a highly extended LSM model. In order to improve the time performance of each time step, we entirely implement the unified particle framework on GPUs using compute unified device architecture. The advantages of our two‐way fluid–solid coupling method in computer animation are demonstrated via several virtual scenarios. Xuqiang Shao, Zhong Zhou, Nadia Magnenat-Thalmann |
Comput. Graph. Forum | 2 |
| 2015 | Realistic and stable simulation of turbulent details behind objects in smoothed-particle hydrodynamics fluidsabstractAbstract This paper presents a novel realistic and stable turbulence synthesis method to simulate the turbulent details generated behind objects in smoothed particle hydrodynamics (SPH) fluids. Firstly, by approximating the boundary layer theory on the fly in SPH fluids, we propose a vorticity production model to identify which fluid particles shed from object surfaces and which are seeded as vortex particles. Then, we employ an SPH‐like summation interpolant formulation of the Biot–Savart law to calculate the fluctuating velocities stemming from the generated vorticity field. Finally, the stable evolution of the vorticity field is achieved by combining an implicit vorticity diffusion technique and an artificial dissipation term. Moreover, in order to efficiently catch turbulent details for rendering, we propose an octree‐based adaptive surface reconstruction method for particle‐based fluids. The experiment results demonstrate that our turbulence synthesis method provides an effect way to model the obstacle‐induced turbulent details in SPH fluids and can be easily added to existing particle‐based fluid–solid coupling pipelines. Copyright © 2014 John Wiley & Sons, Ltd. Xuqiang Shao, Zhong Zhou, Wei Wu 0008 |
Comput. Animat. Virtual Worlds | 2 |
| 2015 | Streaming 3D deforming surfaces with dynamic resolution controlabstractAbstract Real‐time streaming of shape deformations in a shared distributed virtual environment is a challenging task due to the difficulty of transmitting large amounts of 3D animation data to multiple receiving parties at a high frame rate. In this paper, we present a framework for streaming 3D shape deformations, which allows shapes with multi‐resolutions to share the same deformations simultaneously in real time. The geometry and motion of deformingmeshorpoint‐sampledsurfaces are compactly encoded, transmitted, and reconstructed using the spectra of the manifold harmonics. A receiver‐based multi‐resolution surface reconstruction approach is introduced, which allows deforming shapes to switch smoothly between continuous multi‐resolutions. On the basis of this dynamic reconstruction scheme, a frame rate control algorithm is further proposed to achieve rendering at interactive rates. We also demonstrate an efficient interpolation‐based strategy to reduce computing of deformation. The experiments conducted on bothmeshandpoint‐sampledsurfaces show that our approach achieves efficient performance even if deformations of complex 3D surfaces are streamed. Copyright © 2013 John Wiley & Sons, Ltd. Lin Zhang 0009, Fei Dou, Zhong Zhou, Wei Wu 0008 |
Comput. Animat. Virtual Worlds | 3 |
| 2015 | Efficient 3-D Scene Prefetching From Learning User Access PatternsabstractRendering large-scale 3-D scenes on a thin client is attracting increasing attention with the development of the mobile Internet. Efficient scene prefetching to provide timely data with a limited cache is one of the most critical issues for remote 3-D data scheduling in networked virtual environment applications. Existing prefetching schemes predict the future positions of each individual user based on user traces. In this paper, we investigate scene content sequences accessed by various users instead of user viewpoint traces and propose a user access pattern-based 3-D scene prefetching scheme. We make a relationship graph-based clustering to partition history user access sequences into several clusters and choose representative sequences from among these clusters as user access patterns. Then, these user access patterns are prioritized by their popularity and users' personal preference. Based on these access patterns, the proposed prefetching scheme predicts the scene contents that will most likely be visited in the future and delivers them to the client in advance. The experiment results demonstrate that our user access pattern-based prefetching approach achieves a high hit ratio and outperforms the prevailing prefetching schemes in terms of access latency and cache capacity. Zhong Zhou, Jingchang Zhang |
IEEE Trans. Multim. | 1 |
| 2015 | Learning Spatial and Temporal Extents of Human Actions for Action DetectionabstractFor the problem of action detection, most existing methods require that relevant portions of the action of interest in training videos have been manually annotated with bounding boxes. Some recent works tried to avoid tedious manual annotation , and proposed to automatically identify the relevant portions in training videos. However, these methods only concerned the identification in either spatial or temporal domain, and may get irrelevant contents from another domain. These irrelevant contents are usually undesirable in the training phase, which will lead to a degradation of the detection performance. This paper advances prior work by proposing a joint learning framework to simultaneously identify the spatial and temporal extents of the action of interest in training videos. To get pixel-level localization results, our method uses dense trajectories extracted from videos as local features to represent actions. We first present a trajectory split-and-merge algorithm to segment a video into the background and several separated foreground moving objects. In this algorithm, the inherent temporal smoothness of human actions is exploited to facilitate segmentation. Then, with the latent SVM framework on segmentation results, spatial and temporal extents of the action of interest are treated as latent variables that are inferred simultaneously with action recognition. Experiments on two challenging datasets show that action detection with our learned spatial and temporal extents is superior than state-of-the-art methods. Zhong Zhou, Feng Shi 0002, Wei Wu 0008 |
IEEE Trans. Multim. | 1 |
| 2015 | Non-Rigid Structure-From-Motion on Degenerate Deformations With Low-Rank Shape Deformation ModelabstractNon-rigid structure-from-motion (NRSfM) is the process of recovering time-varying 3D structures and poses of a deformable object from an uncalibrated monocular video sequence. Currently, most NRSfM algorithms utilize a non- degenerate assumption for non-rigid object deformations whereby the 3D structures of a non-rigid object can be assumed to be a linear combination of basis shapes with full rank three. Unfortunately, this assumption will produce extra degrees-of-freedom when the non-rigid object has some degenerate deformations with shape bases of rank less than three. These extra degrees-of-freedom will yield spurious shape deformations due to non-negligible noise in real applications, which will cause substantial reconstruction errors. To solve this problem, we propose a low-rank shape deformation model to represent 3D structures of degenerate deformations. When modeling degenerate deformations, the proposed model exploits the rank-deficient nature of degenerate deformations in addition to the low-rank property of non-rigid objects' trajectories, thus providing a more accurate and compact representation compared with existing models. Based on this model, we formulate the NRSfM problem as two coherent optimization problems. These problems are solved with iterative non-linear optimization algorithms. Experiments on synthetic and motion capture data are conducted. The results exhibit the significant advantages of our approach over state-of-the-art NRSfM algorithms for the 3D recovery of non-rigid objects with degenerate deformations. Zhong Zhou, Feng Shi 0002, Jiangjian Xiao, Wei Wu 0008 |
IEEE Trans. Multim. | 1 |
| 2015 | Progressive Motion Vector Clustering for Motion Estimation and Auxiliary TrackingabstractThe motion vector similarity between neighboring blocks is widely used in motion estimation algorithms. However, for nonneighboring blocks, they may also have similar motions due to close depths or belonging to the same object inside the scene. Therefore, the motion vectors usually have several kinds of patterns, which reveal a clustering structure. In this article, we propose a progressive clustering algorithm, which periodically counts the motion vectors of the past blocks to make incremental clustering statistics. These statistics are used as the motion vector predictors for the following blocks. It is proved to be much more efficient for one block to find the best-matching candidate with the predictors. We also design the clustering based search with CUDA for GPU acceleration. Another interesting application of the clustering statistics is persistent static object tracking. Based on the statistics, several auxiliary tracking areas are created to guide the object tracking. Even when the target object has significant changes in appearance or it disappears occasionally, its position still can be predicted. The experiments on Xiph.org Video Test Media dataset illustrate that our clustering based search algorithm outperforms the mainstream and some state-of-the-art motion estimation algorithms. It is 33 times faster on average than the full search algorithm with only slightly higher mean-square error values in the experiments. The tracking results show that the auxiliary tracking areas help to locate the target object effectively. Zhong Zhou, Wei Wu 0008 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Reconstruction of three-dimensional flame with color temperatureabstractThe reconstruction of flame from the captured images is a difficult and computationally expensive problem. Reconstruction from color images will keep the colorful appearance, as is beneficial for visually realistic flame modeling. Most of existing color-image-based methods rebuild three density fields from RGB intensities; however, these methods suffer from the color distortion problem due to the high correlation of RGB intensities. A novel method for 3D flame reconstruction using color temperature is presented in this paper. Color-temperature mapping is calculated to avoid color distortion; this method maps the RGB intensities into the color temperature and its joint intensity. We improve the multiplication reconstruction with visual hull restriction so that the energy distribution is more reasonable, which allows avoidance of the impossible zones. Experimental results indicate that our approach is efficient in the visually plausible 3D flame generation and produces better color restorations. Zhaohui Wu 0005, Zhong Zhou, Delei Tian, Wei Wu 0008 |
Vis. Comput. | 2 |
| 2014 | A two-phase broadcast scheme for underwater acoustic networksabstractReliable broadcast is a critical service for Underwater Acoustic Networks (UANs). In this paper, we propose a Two-phase Broadcast Scheme (TBS) for UANs. TBS includes two phases: Fast Spreading phase and Data Recovery phase. It does not require topology or neighbor information. In the Fast Spreading phase, which is a best effort phase, opportunistic overhearing and network coding are combined to accumulate encoded packets. A probability based forwarder selection scheme is employed to alleviate the broadcast storm problem. A rebroadcast scheduling algorithm is also proposed to reduce collisions. The Data Recovery phase is to guarantee reliability if a node fails to complete data decoding in the Fast Spreading phase. Through delayed request sending, the Data Recovery phase at a node will not interfere with the Fast Spreading phase at other nodes. Through simulations, we demonstrate the advantages of TBS in terms of efficiency and reliability. Haining Mo, Zheng Peng 0001, Zhong Zhou, Jun-Hong Cui |
GLOBECOM | 3 |
| 2014 | Automatic mesh animation previewabstractWith growing number of high quality 3D models published online, the technique of generating efficient and economical 3D model previews has raised increasing concerns. Although several previous work has been done on the preview of static mesh, that of animated mesh is different and more complex due to the difficulty in describing the animation. In this paper, we present a novel method of automatic preview generation for 3D mesh animation. A new measure named inter-frame surface saliency, which evaluates both inter-frame motions and surface saliency in each frame, is introduced. Given an animated mesh, an energy function combining this measure and camera smoothness is constructed for the representative viewpoints selection in key frames, and then an optimal camera path is generated. Finally, a brief but informative preview could be created by moving the camera along this path with frame rate control. Qiaodong Cui, Fei Dou, Lin Zhang 0009, Zhong Zhou |
ICME | 5 |
| 2014 | Maximizing investment income of SSP for spectrum trading in cognitive radio networksabstractMore and more researches have demonstrated the benefits of cognitive radio technology in improving flexibility and efficiency of spectrum utilization. In order to encourage primary users (PUs) to share their idle spectrum resources with secondary users (SUs), spectrum trading frameworks are developed. In this paper, the investment problem of spectrum service provider (SSP) is considered which obtains spectrum from PUs and provides service to multiple SUs. The SUs' actions are estimated according to statistical data. A estimation method for channels number is proposed basing on maximizing the SSP's investment income. A Markov chain model is used to analyze the SSP's state transition and calculate the SU's waiting time and queuing size by queuing theory. The optimal number of channels is deduced with marginal analysis theory. In either spectrum purchase or auction, the SSP could adjust its investment strategy timely and flexibly according to these parameters. Zhong Zhou, Wei Wu 0008 |
WCNC | 2 |
| 2013 | Synthesizing Solid-Induced Turbulence for Particle-Based FluidsabstractSimulating the accompanying turbulent details of fluid-solid coupling is still challenging, as numerical dissipation always plagues current fluid solvers. In this paper, we propose a novel particle-based method to simulate turbulent details generated behind objects in SPH fluids. A turbulence production model, approximating the boundary layer theory on the fly in SPH fluids, is proposed to identify which fluid particles shed from object surfaces and are seeded as vortex particles. Then the fluctuating velocities stemming from the generated vorticity field are calculated using an SPH-like summation interpolant formulation of the Biot-Savart law. And the stable evolution of vorticity field is solved by introducing an artificial dissipation term into the vorticity N-S equation. The advantages of our turbulence synthesis method in computer animation are demonstrated via various virtual scenarios. Xuqiang Shao, Zhong Zhou, Wei Wu 0008 |
CAD/Graphics | 2 |
| 2013 | Robust Trajectory Clustering for Motion SegmentationabstractDue to occlusions and objects' non-rigid deformation in the scene, the obtained motion trajectories from common trackers may contain a number of missing or mis-associated entries. To cluster such corrupted point based trajectories into multiple motions is still a hard problem. In this paper, we present an approach that exploits temporal and spatial characteristics from tracked points to facilitate segmentation of incomplete and corrupted trajectories, thereby obtain highly robust results against severe data missing and noises. Our method first uses the Discrete Cosine Transform (DCT) bases as a temporal smoothness constraint on trajectory projection to ensure the validity of resulting components to repair pathological trajectories. Then, based on an observation that the trajectories of foreground and background in a scene may have different spatial distributions, we propose a two-stage clustering strategy that first performs foreground-background separation then segments remaining foreground trajectories. We show that, with this new clustering strategy, sequences with complex motions can be accurately segmented by even using a simple translational model. Finally, a series of experiments on Hopkins 155 dataset and Berkeley motion segmentation dataset show the advantage of our method over other state-of-the-art motion segmentation algorithms in terms of both effectiveness and robustness. Feng Shi 0002, Zhong Zhou, Jiangjian Xiao, Wei Wu 0008 |
ICCV | 2 |
| 2013 | A new trajectory clustering algorithm using temporal smoothness for motion segmentationabstractIn this paper, a new trajectory clustering algorithm for motion segmentation is proposed. Our key contribution is to use temporal smoothness constraint to facilitate segmentation of incomplete trajectories, which leads to high robustness to missing data. We further show that most motions in foreground of a scene can be approximately represented by a set of translational motion models. Based on this assumption, a new clustering strategy is proposed to separate foreground objects from background. Finally, a series of experiments show that our approach is more effective and outperforms several state-of-the-art methods. Zhong Zhou, Jiangjian Xiao |
ICIP | 2 |
| 2013 | Effective Relay Selection for Underwater Cooperative Acoustic NetworksabstractCooperative communication has been studied extensively as a promising technique for improving the performance of terrestrial wireless networks. However, in underwater cooperative acoustic networks, long propagation delays and complex acoustic channels make the conventional relay selection schemes designed for terrestrial wireless networks inefficient. In this paper, we develop a new best relay selection criterion, called COoperative Best Relay Assessment (COBRA), for underwater cooperative acoustic networks to minimize the one-way packet transmission time. The new criterion takes into account both the spectral efficiency and the underwater long propagation delay to improve the overall throughput performance of the network with energy constraint. A best relay selection algorithm is also proposed based on COBRA criterion. This algorithm only requires the channel statistical information instead of the instantaneous channel state. Our simulation results show a significant decrease on one-way packet transmission time with COBRA. The throughput and delivery ratio performance improvement further verifies the advantages of our proposed criterion over the conventional channel state based algorithms. Yu Luo 0001, Lina Pu, Zheng Peng 0001, Zhong Zhou, Jun-Hong Cui, Zhaoyang Zhang 0001 |
MASS | 4 |
| 2013 | A Unified Framework for Joint Video Pedestrian Segmentation and Pose TrackingabstractPedestrian segmentation and pose tracking are performed to infer human silhouettes and skeletons, respectively. Although the two tasks are complementary in nature, few works have been done on combining them together to improve each other, and some related methods are limited to still images. In this paper, we propose an approach to jointly solving them in monocular videos via a unified framework. Basically, the framework is built on EM-based maximum likelihood estimation, in which pose tracking is fulfilled through Bayesian filtering using body silhouette as an observation cue, and pedestrian segmentation is inferred by guided filtering with constraint of body skeleton. The two sets of parameters are alternatively updated along the video. In the initialization of the framework, we utilize a hierarchical shape matching scheme to obtain the silhouette and skeleton in the first frame. Experiments on challenging pedestrian datasets verify the approach's effectiveness to cluttered backgrounds, moving camera and various articulated bodies, and the performance is improved significantly by solving the two tasks together. Zhong Zhou, Wei Wu 0008 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Replica-aided load balancing in overlay networks
Yuehua Wang, Zhong Zhou, Ling Liu 0001, Wei Wu 0008 |
J. Netw. Comput. Appl. | 2 |
| 2013 | Mobi-Sync: Efficient Time Synchronization for Mobile Underwater Sensor NetworksabstractTime synchronization is an important requirement for many services provided by distributed networks. A lot of time synchronization protocols have been proposed for terrestrial Wireless Sensor Networks (WSNs). However, none of them can be directly applied to Underwater Sensor Networks (UWSNs). A synchronization algorithm for UWSNs must consider additional factors such as long propagation delays from the use of acoustic communication and sensor node mobility. These unique challenges make the accuracy of synchronization procedures for UWSNs even more critical. Time synchronization solutions specifically designed for UWSNs are needed to satisfy these new requirements. This paper proposes Mobi-Sync, a novel time synchronization scheme for mobile underwater sensor networks. Mobi-Sync distinguishes itself from previous approaches for terrestrial WSN by considering spatial correlation among the mobility patterns of neighboring UWSNs nodes. This enables Mobi-Sync to accurately estimate the long dynamic propagation delays. Simulation results show that Mobi-Sync outperforms existing schemes in both accuracy and energy efficiency. Jun Liu 0006, Zhong Zhou, Zheng Peng 0001, Jun-Hong Cui, Michael Zuba, Lance Fiondella |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Practical Coding-based Multi-Hop Reliable Data Transfer for underwater acoustic networksabstractIn this paper, we investigate reliable data transfer for multi-hop underwater acoustic networks. Motivated by experiences from real-world field tests, we propose a Practical Coding-based Multi-hop Reliable Data Transfer (PCMRDT) protocol. For the per-hop reliable data transfer, PCMRDT combines random linear coding and selective repeat to achieve high reliability and efficiency. We analyze the data recovery capability for random linear coding so that we can set an appropriate coding rate. In addition, PCMRDT utilizes a multi-hop coordination mechanism to eliminate collisions and decreases average end-to-end delay over multiple hops. Simulation results show that PCMRDT can significantly reduce the network delay with high energy efficiency. Haining Mo, Zhong Zhou, Michael Zuba, Zheng Peng 0001, Jun-Hong Cui, Yantai Shu |
GLOBECOM | 2 |
| 2012 | Topological Similarity-Based Scheme for Large-Scale Group Communication ServicesabstractGroup communication is essential for multi-user applications. However, due to unpredictable node departures and non-deterministic network partitions, providing reliable and scalable group communication services is challenging when the applications are utilized by the users with heterogeneous capacities on a large scale. To address this challenge, we propose a novel replication scheme to achieve high reliability and low-cost scalability in group communication with following three features. First, it introduces a new concept of replication based on topological similarity, which empowers each node with an ability of measuring similarity between the nodes in topology. By eliminating the topological similarity between the replicas, it intelligently mitigates service interruptions caused by node failures and network partitions. Second, instead of specifying the number of replicas, it provides a technique for nodes to dynamically adapt the replication placement schemes by exploiting functionality importance of the nodes in the group- communication session. It eliminates the bottleneck problem and improves the network resource utilization. Third, the scheme is self-converging and it can stabilize within a few adaptations even facing a high churn rate. Extensive simulations show that it yields significant improvements in reduction of replication overhead and service interruption when comparing to existing approaches. Yuehua Wang, Zhong Zhou, Ling Liu 0001, Liang Cheng 0001, Wei Wu 0008 |
ICCCN | 2 |
| 2012 | Clustering Based Search Algorithm for Motion EstimationabstractMotion estimation is the key part of video compression since it removes the temporal redundancy within frames and significantly affects the encoding quality and efficiency. In this paper, a novel fast motion estimation algorithm named Clustering Based Search algorithm is proposed, which is the first to define the clustering feature of motion vectors in a sequence. The proposed algorithm periodically counts the motion vectors of past blocks to make progressive clustering statistics, and then utilizes the clusters as motion vector predictors for the following blocks. It is found to be much more efficient for one block to find the best-matched candidate with the predictors. Compared with the mainstream search algorithms, this algorithm is almost the most efficient one, 35 times faster in average than the full search algorithm, while its mean-square error is even competitively close to that of the full search algorithm. Zhong Zhou, Wei Wu 0008 |
ICME | 2 |
| 2012 | A new metric for measuring image-based 3D reconstruction
Zhong Zhou, Wei Wu 0008 |
ICPR | 3 |
| 2012 | Iterative pedestrian segmentation and pose tracking under a probabilistic frameworkabstractThis paper presents a unified probabilistic framework to tackle two closely related visual tasks: pedestrian segmentation and pose tracking along monocular videos. Although the two tasks are complementary in nature, most previous approaches focus on them individually. Here, we resolve the two problems simultaneously by building and inferring a single body model. More specifically, pedestrian segmentation is performed by optimizing body region with constraint of body pose in a Markov Random Field (MRF), and pose parameters are reasoned about through a Bayesian filtering, which takes body silhouette as an observation cue. Since the two processes are inter-related, we resort to an Expectation-Maximization (EM) algorithm to refine them alternatively. Additionally, a template matching scheme is utilized for initialization. Experimental results on challenging videos verify the framework's robustness to non-rigid human segmentation, cluttered backgrounds and moving cameras. Zhong Zhou, Wei Wu 0008 |
ICRA | 2 |
| 2012 | Texture and Shadow Insensitive Metric for Image-Based ReconstructionabstractThis paper proposes an accurate metric for image based 3d reconstruction without ground truth. Specially, our metric is insensitive to texture changing and shadows, which are commonly occurred in real world scenes. Based on the interreflected rendering model, we improve the accuracy of previous irradiance based metric. Additionally, we estimate the reflectance of each vertex on the surface to support the case with varying reflectance. We also consider the difference between estimated and observed irradiance in our metric to further eliminate the boundary effect of texture changing or self-shadow. Experiments on both indoor and outdoor datasets illustrate the effectiveness of our metric. Our evaluation results are not only more accurate than the results of previous metrics, but also insensitive to the texture and shadow. Zhong Zhou, Wei Wu 0008 |
ICTAI | 3 |
| 2012 | Robust Frame Registration for Multiple Camera Setups in Dynamic ScenesabstractIn this paper, we propose a novel method to register frames from multiple cameras into a consistent global scale. Assuming a moving object is observed in multiple camera setups, we use initial frames to create a global reference structure where the pose variation of each new frame is estimated using a RANSAC-based registration algorithm. We further combine the registration method with other state-of the-art techniques to build a high quality 3D reconstruction system with a smaller number of cameras than used by more traditional methods. Experimental results show that our method performs better and is more economical than the registration of separate monocular structures from motion methods. 3D reconstruction results on various challenging real world multi-camera video datasets also illustrate the feasibility and robustness of our method. Zhong Zhou, Ye Duan, Wei Wu 0008 |
ICTAI | 2 |
| 2012 | Displacement residual based DDM matching algorithm
Lin Zhang 0009, Zhong Zhou, Lin Liu 0001, Wei Wu 0008 |
Sci. China Inf. Sci. | 2 |
| 2012 | Radiance-based color calibration for image-based modeling with multiple cameras
Zhong Zhou, Wei Wu 0008 |
Sci. China Inf. Sci. | 2 |
| 2012 | Fault Tolerance and Recovery for Group Communication Services in Distributed Networks
Yuehua Wang, Zhong Zhou, Ling Liu 0001, Wei Wu 0008 |
J. Comput. Sci. Technol. | 2 |
| 2012 | Particle-based simulation of bubbles in water-solid interactionabstractABSTRACT In this paper, a particle‐based multiphase method for creating realistic animations of bubbles in water–solid interaction is presented. To generate bubbles from gas dissolved in the water on the fly, we propose an approximate model for the creation of bubbles, which takes into account the influence of gas concentration in the water, the solid material, and water–solid velocity difference. As the air particle on the bubble surface is treated as a virtual nucleation site, the bubble absorbs air from surrounding water and grows. The density and pressure forces of air bubbles are computed separately using smoothed particle hydrodynamics; then, the two‐way coupling of bubbles with water and solid is solved by a new drag force, so the generated bubbles’ flow on the surface of solid and the deformation in the rising process can be simulated. Additionally, touching bubbles merge together under the cohesion forces weighted by the smoothing kernel and velocity difference. The experimental results show that this method is capable of simulating bubbles in water–solid interaction under different physical conditions. Copyright © 2012 John Wiley & Sons, Ltd. Xuqiang Shao, Zhong Zhou, Wei Wu 0008 |
Comput. Animat. Virtual Worlds | 2 |
| 2012 | Detail-feature-preserving surface reconstructionabstractABSTRACT In this paper, we propose a feature‐preserving surface reconstruction method from sparse noisy 3D measurements such as range scanning or passive multiview stereo. In contrast to earlier methods, we define a novel type of explicit 3D filter—regularized weighted least squares filter—to characterize the detail features such as surface wrinkles and sharp features. To account for noise, we rasterize input‐oriented points into a probabilistic volume (base volume) and then create a guidance volume by Gaussian filtering. Both the base volume and the guidance volume are further filtered by regularized weighted least squares filter to detect and recover detail features. After the two‐stage filtering, a global minimal surface is computed by graph cut and meshed as a geometric model. Experimental results on various datasets show that our method is robust to noise, outliers, and missing parts, which makes it more suitable to fit indoor/outdoor multiview stereo data. Unlike other methods, our method can completely recover scene structures and preserve detail features from noisy point samples. Copyright © 2012 John Wiley & Sons, Ltd. Zhong Zhou, Ye Duan, Wei Wu 0008 |
Comput. Animat. Virtual Worlds | 2 |
| 2011 | Realistic Fire Simulation: A SurveyabstractAs fire is one of the basic elements of nature, realistic fire simulation plays a significant role in fire control, film special effects, military emulation, and virtual reality. Fire simulation has been being an eternal theme in the field of computer graphics, because of the anisotropic shape and complex physical mechanism of the fire. This paper presents a survey on the development of flame simulation, with a detailed introduction to the different kinds of methods applied in the field. The methods can be classified into several different kind, mainly include texture mapping method, particle system method, mathematical physics-based method, cellular automation method, image based tomographic reconstruction, and other methods. This paper analyzes the performance of different methods from real-time, reality, spatio-temporal complexity, edit ability, and interactivity. Finally, in connection with issues in the present research and possible future direction of development, the paper puts forth a number of theoretical and technical problems, hoping they can be resolved in the future. Zhaohui Wu 0005, Zhong Zhou, Wei Wu 0008 |
CAD/Graphics | 2 |
| 2011 | Combining Shape and Appearance for Automatic Pedestrian SegmentationabstractIn this paper we present an approach to automatically segmenting non-rigid pedestrians in still images. Inspired by global shape matching as well as interactive figure-ground separation methods, this approach fulfills the task combining shape and appearance cues in a unified framework. The main idea is to initially extract pedestrian silhouette and skeleton via hierarchical shape matching, and then generate an appearance trimap to refine segmentation. The major contributions of this paper include: 1) a novel shape matching scheme, which is proposed to replace the commonly used Chamfer matching in the shape matching stage, 2) a head-torso parsing method, which is developed for localizing pedestrian to reduce the search space, 3) an automatic trimap generation method used to refine segmentation. Experiments on public datasets demonstrate that the approach improves pedestrian segmentation efficiently and effectively. Zhong Zhou, Wei Wu 0008 |
ICTAI | 2 |
| 2010 | Real-time stereo-vision system for 3D teleimmersive collaborationabstractThough the variety of desktop real time stereo vision systems has grown considerably in the past several years, few make any verifiable claims about the accuracy of the algorithms used to construct 3D data or describe how the data generated by such systems, which is large in size, can be effectively distributed. In this paper, we describe a system that creates an accurate (on the order of a centimeter), 3D reconstruction of an environment in real time (under 30 ms) that also allows for remote interaction between users. This paper addresses how to reconstruct, compress, and visualize the 3D environment. In contrast to most commercial desktop real time stereo vision systems our algorithm produces 3D meshes instead of dense point clouds, which we show allows for better quality visualizations. The chosen representation of the data also allows for high compression ratios for transfer to remote sites. We demonstrate the accuracy and speed of our results on a variety of benchmarks. Ramanarayan Vasudevan, Zhong Zhou, Gregorij Kurillo, Edgar J. Lobaton, Ruzena Bajcsy, Klara Nahrstedt |
ICME | 2 |
| 2010 | A Hybrid Deformation Model for Virtual CuttingabstractIn this paper, we present a novel hybrid deformation model, which can simulate large deformations and non-linear behaviors of soft tissues in real time. The model partitions a soft tissue into operational and non-operational regions, and simulates the large deformations by coupling the geometry constrained mesh less method FLSM (fast lattice shape matching) with finite element algorithm TLED (total Lagrangian explicit dynamics). With the virtual node cutting method, the soft tissues can be cut realistically and robustly. The virtual cutting simulation features real time performance with the GPU-based acceleration. In combination with our multi-camera based real-time modeling framework, the deformable model and its interaction of virtual cutting can be extended to support remote virtual surgery. Xuqiang Shao, Zhong Zhou, Wei Wu 0008 |
ISM | 2 |
| 2010 | Static Object Tracking in Road Panoramic VideosabstractIn panoramic videos, the object movement between adjacent side images leads to deformation and discontinuity, which makes the traditional video tracking approaches insufficient. An effective static object tracking algorithm is proposed in this paper to resolve the tracking problems from the deformation and discontinuity in cubic panorama. The algorithm extends the relevant side images with boundary consistency, and then conducts a background eliminated mean-shift algorithm to track objects on the extended images. Experiment results show that the algorithm can track static objects correctly in reasonable situations in real-time. Zhong Zhou, Chen Ke, Wei Wu 0008 |
ISM | 1 |
| 2009 | Expanding Line Search for Panorama Motion EstimationabstractThis paper describes an effective motion estimation algorithm for panoramic video. According to the characteristics of the block motion in panoramic video, the proposed algorithm extends the reference frames and constructs search lines for the line search. With the constructed search lines, the line search estimates the motion of the corresponding macro block. It starts at the block matched by the line search and estimates the motion of adjacent blocks. Experiment results show that the algorithm can estimate the motion of macro blocks in cubic panoramic video effectively. It is also shown that the algorithm improves the search speed with effective motion estimation. Zhong Zhou, Jingxiang Chen, Wei Wu 0008 |
ISM | 2 |
| 2009 | LoI: Efficient relevance evaluation and filtering for distributed simulation
Zhong Zhou, Wei Wu 0008 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | Algorithm of simulation time synchronization over large-scale nodes
Qinping Zhao, Zhong Zhou, Fang Lü |
Sci. China Ser. F Inf. Sci. | 2 |