Guile Wu

dblp:180/0097 · DBLP profile ↗
← Back
24ranked-venue papers
14as first author
15since 2021 · last 2026
0000-0001-7319-473XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 12 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 11 first-author · 13 since 2021
YearPublicationVenuePosition
2026 Unigaussian: Driving Scene Reconstruction From Multiple Camera Models Via Unified Gaussian Representations
abstract
Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras. In fact, how to effectively simulate fisheye cameras in driving scene remains an unsolved problem. In this work, we propose UniGaussian, a novel approach that learns a unified 3D Gaussian representation from multiple camera models for urban scene reconstruction in autonomous driving. Our contributions are two-fold. First, we propose a new differentiable rendering method that distorts 3D Gaussians using a series of affine transformations tailored to fisheye camera models. This addresses the compatibility issue of 3D Gaussian splatting with fisheye cameras, which is hindered by light ray distortion caused by lenses or mirrors. Besides, our method maintains real-time rendering while ensuring differentiability. Second, built on the differentiable rendering method, we design a new framework that learns a unified Gaussian representation from multiple camera models. By applying affine transformations to adapt different camera models and regularizing the shared Gaussians with supervision from different modalities, our framework learns a unified 3D Gaussian representation with input data from multiple sources and achieves holistic driving scene understanding. As a result, our approach models multiple sensors (pinhole and fisheye cameras) and modalities (depth, semantic, normal and LiDAR point clouds). Our experiments show that our method achieves superior rendering quality and fast rendering speed for driving scene simulation.
Guile Wu, Runhao Li, Zheyuan Yang, Tongtong Cao, Xingxin Chen
3DV2
2024 VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
Zheyuan Yang, Guile Wu, Kejian Lin, Jinjun Shan
ECCV (66)3
2024 Federated zero-shot learning with mid-level semantic knowledge transfer
abstract
Conventional centralized deep learning paradigms are not feasible when data from different sources cannot be shared due to data privacy or transmission limitation. To resolve this problem, federated learning has been introduced to transfer knowledge across multiple sources (clients) with non-shared data while optimizing a globally generalized central model (server). Existing federated learning paradigms mostly focus on transmitting image encoders that take instance-sensitive images as input, making them less generalizable and vulnerable to privacy inference attacks. In contrast, in this work, we consider transferring mid-level semantic knowledge (such as attribute) which is not sensitive to specific objects of interest and therefore is more privacy-preserving and general. To this end, we formulate a new Federated Zero-Shot Learning (FZSL) paradigm to learn mid-level semantic knowledge at multiple local clients with non-shared local data and cumulatively aggregate a globally generalized central model for deployment. To improve model discriminative ability, we explore semantic knowledge available from either a language or a vision-language foundation model in order to enrich the mid-level semantic space in FZSL. Extensive experiments on five zero-shot learning benchmark datasets validate the effectiveness of our approach for optimizing a generalizable federated learning model with mid-level semantic knowledge transfer.
Shitong Sun, Chenyang Si, Guile Wu, Shaogang Gong
Pattern Recognit.3
2023 Hierarchical Quantization Consistency for Fully Unsupervised Image Retrieval
Guile Wu, Chao Zhang 0023, Stephan Liwicki
BMVC1
2023 MV-DeepSDF: Implicit Modeling with Multi-Sweep Point Clouds for 3D Vehicle Reconstruction in Autonomous Driving
abstract
Reconstructing 3D vehicles from noisy and sparse partial point clouds is of great significance to autonomous driving. Most existing 3D reconstruction methods cannot be directly applied to this problem because they are elaborately designed to deal with dense inputs with trivial noise. In this work, we propose a novel framework, dubbed MV-DeepSDF, which estimates the optimal Signed Distance Function (SDF) shape representation from multi-sweep point clouds to reconstruct vehicles in the wild. Although there have been some SDF-based implicit modeling methods, they only focus on single-view-based reconstruction, resulting in low fidelity. In contrast, we first analyze multi-sweep consistency and complementarity in the latent feature space and propose to transform the implicit space shape estimation problem into an element-to-set feature extraction problem. Then, we devise a new architecture to extract individual element-level representations and aggregate them to generate a set-level predicted latent code. This set-level latent code is an expression of the optimal 3D shape in the implicit space, and can be subsequently decoded to a continuous SDF of the vehicle. In this way, our approach learns consistent and complementary information among multi-sweeps for 3D vehicle reconstruction. We conduct thorough experiments on two real-world autonomous driving datasets (Waymo and KITTI) to demonstrate the superiority of our approach over state-of-the-art alternative methods both qualitatively and quantitatively.
Kelly Zhu, Guile Wu, Jinjun Shan
ICCV3
2023 Towards Universal LiDAR-Based 3D Object Detection by Multi-Domain Knowledge Transfer
abstract
Contemporary LiDAR-based 3D object detection methods mostly focus on single-domain learning or cross-domain adaptive learning. However, for autonomous driving systems, optimizing a specific LiDAR-based 3D object detector for each domain is costly and lacks of scalability in real-world deployment. It is desirable to train a universal LiDAR-based 3D object detector from multiple domains. In this work, we propose the first attempt to explore multi-domain learning and generalization for LiDAR-based 3D object detection. We show that jointly optimizing a 3D object detector from multiple domains achieves better generalization capability compared to the conventional single-domain learning model. To explore informative knowledge across domains towards a universal 3D object detector, we propose a multi-domain knowledge transfer framework with universal feature transformation. This approach leverages spatial-wise and channel-wise knowledge across domains to learn universal feature representations, so it facilitates to optimize a universal 3D object detector for deployment at different domains. Extensive experiments on four benchmark datasets (Waymo, KITTI, NuScenes and ONCE) show the superiority of our approach over the state-of-the-art approaches for multi-domain learning and generalization in LiDAR-based 3D object detection.
Guile Wu, Tongtong Cao, Xingxin Chen
ICCV1
2022 Learning Unbiased Transferability for Domain Adaptation by Uncertainty Modeling
Jian Hu 0002, Haowen Zhong, Shaogang Gong, Guile Wu, Junchi Yan
ECCV (31)5
2022 Learning hybrid ranking representation for person re-identification
Guile Wu, Xiatian Zhu, Shaogang Gong
Pattern Recognit.1
2021 Generalising without Forgetting for Lifelong Person Re-Identification
abstract
Existing person re-identification (Re-ID) methods mostly prepare all training data in advance, while real-world Re-ID data are inherently captured over time or from different locations, which requires a model to be incrementally generalised from sequential learning of piecemeal new data without forgetting what is already learned. In this work, we call this lifelong person Re-ID, characterised by solving a problem of unseen class identification subject to continuous new domain generalisation and adaptation with class imbalanced learning. We formulate a new Generalising without Forgetting method (GwFReID) for lifelong Re-ID and design a comprehensive learning objective that accounts for classification coherence, distribution coherence and representation coherence in a unified framework. This design helps to simultaneously learn new information, distil old knowledge and solve class imbalance, which enables GwFReID to incrementally improve model generalisation without catastrophic forgetting of what is already learned. Extensive experiments on eight Re-ID benchmarks, CIFAR-100 and ImageNet show the superiority of GwFReID over the state-of-the-art methods.
Guile Wu, Shaogang Gong
AAAI1
2021 Decentralised Learning from Independent Multi-Domain Labels for Person Re-Identification
abstract
Deep learning has been successful for many computer vision tasks due to the availability of shared and centralised large-scale training data. However, increasing awareness of privacy concerns poses new challenges to deep learning, especially for human subject related recognition such as person re-identification (Re-ID). In this work, we solve the Re-ID problem by decentralised learning from non-shared private training data distributed at multiple user sites of independent multi-domain label spaces. We propose a novel paradigm called Federated Person Re-Identification (FedReID) to construct a generalisable global model (a central server) by simultaneously learning with multiple privacy-preserved local models (local clients). Specifically, each local client receives global model updates from the server and trains a local model using its local data independent from all the other clients. Then, the central server aggregates transferrable local model updates to construct a generalisable global feature embedding model without accessing local data so to preserve local privacy. This client-server collaborative learning process is iteratively performed under privacy control, enabling FedReID to realise decentralised learning without sharing distributed data nor collecting any centralised data. Extensive experiments on ten Re-ID benchmarks show that FedReID achieves compelling generalisation performance beyond any locally trained models without using shared training data, whilst inherently protects the privacy of each local client. This is uniquely advantageous over contemporary Re-ID methods.
Guile Wu, Shaogang Gong
AAAI1
2021 Peer Collaborative Learning for Online Knowledge Distillation
abstract
Traditional knowledge distillation uses a two-stage training strategy to transfer knowledge from a high-capacity teacher model to a compact student model, which relies heavily on the pre-trained teacher. Recent online knowledge distillation alleviates this limitation by collaborative learning, mutual learning and online ensembling, following a one-stage end-to-end training fashion. However, collaborative learning and mutual learning fail to construct an online high-capacity teacher, whilst online ensembling ignores the collaboration among branches and its logit summation impedes the further optimisation of the ensemble teacher. In this work, we propose a novel Peer Collaborative Learning method for online knowledge distillation, which integrates online ensembling and network collaboration into a unified framework. Specifically, given a target network, we construct a multi-branch network for training, in which each branch is called a peer. We perform random augmentation multiple times on the inputs to peers and assemble feature representations outputted from peers with an additional classifier as the peer ensemble teacher. This helps to transfer knowledge from a high-capacity teacher to peers, and in turn further optimises the ensemble teacher. Meanwhile, we employ the temporal mean model of each peer as the peer mean teacher to collaboratively transfer knowledge among peers, which helps each peer to learn richer knowledge and facilitates to optimise a more stable model with better generalisation. Extensive experiments on CIFAR-10, CIFAR-100 and ImageNet show that the proposed method significantly improves the generalisation of various backbone networks and outperforms the state-of-the-art methods.
Guile Wu, Shaogang Gong
AAAI1
2021 Decentralised Person Re-Identification with Selective Knowledge Aggregation
Shitong Sun, Guile Wu, Shaogang Gong
BMVC2
2021 Collaborative Optimization and Aggregation for Decentralized Domain Generalization and Adaptation
abstract
Contemporary domain generalization (DG) and multisource unsupervised domain adaptation (UDA) methods mostly collect data from multiple domains together for joint optimization. However, this centralized training paradigm poses a threat to data privacy and is not applicable when data are non-shared across domains. In this work, we propose a new approach called Collaborative Optimization and Aggregation (COPA), which aims at optimizing a generalized target model for decentralized DG and UDA, where data from different domains are non-shared and private. Our base model consists of a domain-invariant feature extractor and an ensemble of domain-specific classifiers. In an iterative learning process, we optimize a local model for each domain, and then centrally aggregate local feature extractors and assemble domain-specific classifiers to construct a generalized global model, without sharing data from different domains. To improve generalization of feature extractors, we employ hybrid batch-instance normalization and collaboration of frozen classifiers. For better decentralized UDA, we further introduce a prediction agreement mechanism to overcome local disparities towards central model aggregation. Extensive experiments on five DG and UDA benchmark datasets show that COPA is capable of achieving comparable performance against the state-of-the-art DG and UDA methods without the need for centralized data collection in model training.
Guile Wu, Shaogang Gong
ICCV1
2021 Striking a Balance between Stability and Plasticity for Class-Incremental Learning
abstract
Class-incremental learning (CIL) aims at continuously updating a trained model with new classes (plasticity) without forgetting previously learned old ones (stability). Contemporary studies resort to storing representative exemplars for rehearsal or preventing consolidated model parameters from drifting, but the former requires an additional space for storing exemplars at every incremental phase while the latter usually shows poor model generalization. In this paper, we focus on resolving the stability-plasticity dilemma in class-incremental learning where no exemplars from old classes are stored. To make a trade-off between learning new information and maintaining old knowledge, we reformulate a simple yet effective baseline method based on a cosine classifier framework and reciprocal adaptive weights. With the reformulated baseline, we present two new approaches to CIL by learning class-independent knowledge and multi-perspective knowledge, respectively. The former exploits class-independent knowledge to bridge learning new and old classes, while the latter learns knowledge from different perspectives to facilitate CIL. Extensive experiments on several widely used CIL benchmark datasets show the superiority of our approaches over the state-of-the-art methods.
Guile Wu, Shaogang Gong
ICCV1
2021 Semi-Supervised Few-Shot Learning with Pseudo Label Refinement
abstract
Few-shot classification aims at recognising novel categories with very limited labelled samples. Although substantial achievements have been obtained, few-shot classification remains challenging due to the scarcity of labelled examples. Recent studies resort to leveraging unlabelled data to expand the training set using pseudo labelling, but this strategy often yields significant label noise. In this work, we introduce a new baseline method for semi-supervised few-shot learning by iterative pseudo label refinement to reduce noise. Then, we investigate the label noise propagation problem and improve the baseline with a denoising network to learn distributions of clean and noisy pseudo-labelled examples via a mixture model. This helps to estimate confidence values of pseudo labelled examples and to select the reliable ones with less noise for iteratively refining a few-shot classifier. Extensive experiments on three widely used benchmarks, minilma- genet, tieredImagenet and CIFAR-FS, show the superiority of the proposed methods over the state-of-the-art methods.
Guile Wu, Shaogang Gong, Xu Lan
ICME2
2020 Tracklet Self-Supervised Learning for Unsupervised Person Re-Identification
abstract
Existing unsupervised person re-identification (re-id) methods mainly focus on cross-domain adaptation or one-shot learning. Although they are more scalable than the supervised learning counterparts, relying on a relevant labelled source domain or one labelled tracklet per person initialisation still restricts their scalability in real-world deployments. To alleviate these problems, some recent studies develop unsupervised tracklet association and bottom-up image clustering methods, but they still rely on explicit camera annotation or merely utilise suboptimal global clustering. In this work, we formulate a novel tracklet self-supervised learning (TSSL) method, which is capable of capitalising directly from abundant unlabelled tracklet data, to optimise a feature embedding space for both video and image unsupervised re-id. This is achieved by designing a comprehensive unsupervised learning objective that accounts for tracklet frame coherence, tracklet neighbourhood compactness, and tracklet cluster structure in a unified formulation. As a pure unsupervised learning re-id model, TSSL is end-to-end trainable at the absence of source data annotation, person identity labels, and camera prior knowledge. Extensive experiments demonstrate the superiority of TSSL over a wide variety of the state-of-the-art alternative methods on four large-scale person re-id benchmarks, including Market-1501, DukeMTMC-ReID, MARS and DukeMTMC-VideoReID.
Guile Wu, Xiatian Zhu, Shaogang Gong
AAAI1
2019 Spatio-Temporal Associative Representation for Video Person Re-Identification
Guile Wu, Xiatian Zhu, Shaogang Gong
BMVC1
2019 Person Re-Identification by Ranking Ensemble Representations
abstract
Existing deep learning algorithms for person re-identification (re-id) typically rely on single-sample classification or pairwise matching constraints. This indicates a breach of deployment due to ignoring the probe-specific matching information against the gallery set encoded in ranking lists. In this work, we address this problem by exploring the idea of RANkinG Ensembles (RANGE) that learns such information from the ranking lists. Specifically, given an off-the-self deep re-id feature representation model, we construct per-probe ranking lists and exploit them to learn inter ranking ensemble representation. To mitigate the harm of inevitable false gallery positives, we further introduce a complementary intra ranking ensemble representation. Extensive experiments show that both supervised and unsupervised re-id benefit from the proposed RANGE method on four challenging benchmarks: MSMT17, Market-1501, DukeMTMC-ReID, and CUHK03.
Guile Wu, Xiatian Zhu, Shaogang Gong
ICIP1
2019 Feature covariance matrix-based dynamic hand gesture recognition
Linpu Fang, Guile Wu, Wenxiong Kang, Qiuxia Wu, Zhiyong Wang 0001, David Dagan Feng
Neural Comput. Appl.2
2018 Real-Time Long-Term Tracking With Prediction-Detection-Correction
abstract
Real-time long-term visual tracking is one of the most challenging problems in computer vision due to various factors such as occlusion and motion ambiguity. To achieve robust long-term tracking, most state-of-the-art methods typically construct an online detector in each frame. However, they fail to achieve real-time performance due to high computational complexity. In this paper, we propose a novel real-time long-term tracking algorithm by exploiting a joint Prediction-Detection-Correction Tracking framework (PDCT). We utilize a superpixel optical flow to construct a predictor to estimate the target motion and internal scale variation. To locate the target at a finer level, we develop an improved kernelized correlation detector with an adaptive online learning rate and translation-scale parameters from the predictor. To refine the tracking result and redetect the target in the case of a tracking failure, we devise a corrector utilizing dual online SVMs with dense sampling and reliable history samples. The SVMs are trained with passive-aggressive learning and online retraining strategies. In addition, we employ a selection mechanism for the correlation responses to maintain reliable samples effectively. As a result, our proposed tracker is able to refine tracking results via the corrector and detector and maintains reliable tracking results for subsequent tracking. Extensive experiments on the widely used object tracking benchmark show that the proposed tracker is superior to state-of-the-art trackers in terms of both effectiveness and efficiency, and the integration of each component is effective under the PDCT framework.
Ningxin Liang, Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.2
2017 Visual tracking utilizing robust complementary learner and adaptive refiner
Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing2
2017 Exploiting superpixel and hybrid hash for kernel-based visual tracking
Guile Wu, Wenxiong Kang
Pattern Recognit.1
2017 Vision-Based Fingertip Tracking Utilizing Curvature Points Clustering and Hash Model Representation
abstract
Fingertip tracking plays an increasingly important role in augmented reality and virtual-reality applications. However, existing approaches (either continuous-detection-based or separated-detection-tracking-based methods) cannot effectively learn temporal-spatial information or even separate fingertip tracking into unrelated stages, which causes poor real-time performance and incomplete tracking continuity. Moreover, due to the need for high-cost devices, the high degrees of freedom of the hand, and subtle differences among fingers, fingertip tracking remains a challenging task. To address these problems, we propose a novel tracking-combined-with-detection approach for vision-based fingertip tracking. By adopting clustering and geometric constraint analysis, we develop a curvature points clustering method for fingertip detection. Then, by exploiting the identified fingertip points for motion estimation with bidirectional optical flows and temporal-spatial probability calculation, the tracking stage is effectively integrated with the detection stage. To accurately locate the fingertip, we represent the fingertip model with a perceptual hash sequence and locate the fingertip by searching for the best-matching region. Extensive experimental results show the superiority of the proposed algorithm to commonly used and state-of-the-art methods and demonstrate its effectiveness and practicability.
Guile Wu, Wenxiong Kang
IEEE Trans. Multim.1
2016 Robust Fingertip Detection in a Complex Environment
abstract
Fingertip detection has a broad application in gesture recognition and finger tracking. It is also an important foundation of human-computer interaction systems. However, most algorithms are suitable for simple conditions with low accuracy because the hand is a nonrigid object, and its appearance model is complex. To address the challenging problem of accurately detecting fingertips in a complex environment, we propose a novel and robust fingertip detection algorithm in this paper. Unlike existing methods, our study requires no special device or mark, and users are free to move their hands. Via dense optical flow and a skin filter, we perform complete hand region segmentation in a complex environment. We find the maximum value of the local centroid distance outside the centroid circles and identify fingertips. Our algorithm performs favorably compared with common hand region segmentation and fingertip detection methods. Thorough experimentation proves that our proposed algorithm is effective and robust.
Guile Wu, Wenxiong Kang
IEEE Trans. Multim.1