Guan Luo

dblp:60/6918 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-2151-3011ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Integrating Diverse Assignment Strategies into DETRs
abstract
Label assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich supervisory signals, they often introduce complex, architecture-specific modifications and typically focus on a single auxiliary strategy, lacking a unified and scalable design. In this paper, we first systematically investigate the effects of ``one-to-many'' supervision and reveal a surprising insight that performance gains are driven not by the sheer quantity of supervision, but by the diversity of the assignment strategies employed. This finding suggests that a more elegant, parameter-efficient approach is attainable. Building on this insight, we propose LoRA-DETR, a flexible and lightweight framework that seamlessly integrates diverse assignment strategies into any DETR-style detector. Our method augments the primary network with multiple Low-Rank Adaptation (LoRA) branches during training, each instantiating a different one-to-many assignment rule. These branches act as auxiliary modules that inject rich, varied supervisory gradients into the main model and are discarded during inference, thus incurring no additional computational cost. This design promotes robust joint optimization while maintaining the architectural simplicity of the original detector. Extensive experiments on different baselines validate the effectiveness of our approach. Our work presents a new paradigm for enhancing detectors, demonstrating that diverse ``one-to-many'' supervision can be integrated to achieve state-of-the-art results without compromising model elegance.
Hanshi Wang, Fudong Ge, Guan Luo, Weiming Hu 0004
AAAI5
2025 Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders
abstract
Recent 3D content generation pipelines commonly employ Variational Autoencoders (VAEs) to encode shapes into compact latent representations for diffusion-based generation. However, the widely adopted uniform point sampling strategy in Shape VAE training often leads to a significant loss of geometric details, limiting the quality of shape reconstruction and downstream generation tasks. We present Dora-Vae, a novel approach that enhances VAE reconstruction through our proposed sharp edge sampling strategy and a dual cross-attention mechanism. By identifying and prioritizing regions with high geometric complexity during training, our method significantly improves the preservation of fine-grained shape features. Such sampling strategy and the dual attention mechanism enable the VAE to focus on crucial geometric details that are typically missed by uniform sampling approaches. To systematically evaluate VAE reconstruction quality, we additionally propose Dora-Bench, a benchmark that quantifies shape complexity through the density of sharp edges, introducing a new metric focused on reconstruction accuracy at these salient geometric features. Extensive experiments on the Dora-Bench demonstrate that Dora-Vae achieves comparable reconstruction quality to the state-of-the-art dense XCube-Vae while requiring a latent space at least 8× smaller (1,280 vs. > 10,000 codes). Project page: https://aruichen.github.io/Dora.
Yixun Liang, Guan Luo, Jiarui Liu 0003, Xiu Li 0001, Xiaoxiao Long, Jiashi Feng, Ping Tan 0002
CVPR4
2025 MS3D: High-Quality 3D Generation via Multi-Scale Representation Modeling
Guan Luo
ICCV1
2025 Towards Interpretable User Intent Analysis with Deficient Evidence Fusion for Pseudo-Modalities
abstract
Intent analysis enhances retrieval systems' ability to understand users' motivations and interact with people intelligently. Beyond accurately obtaining intent classification results, explaining the intent classification process remains an essential yet challenging issue in multimedia retrieval. In this study, we propose an interpretable intent analysis method, the Deficient Evidence Network (DENet), to generate interpretations of semantic elements while accurately computing intent labels. Our approach masks parts of the input text to create pseudo-modalities with deficient evidence and integrates these evidence using Dempster-Shafer theory. By evaluating the deficient evidence of different masked semantic elements, our method extracts interpretable information from the original text. Currently, benchmarks for measuring the interpretability of such methods are limited. To address this, we designed a task and provided a benchmark dataset to assess the subjective interpretability of users' intent for the first time. The proposed method not only demonstrates outstanding interpretability in quantitative assessments but also enhances the classification performance of the pre-trained intent model. Our code and benchmark dataset are open-accessible in github.com/yuanxiaoheben/DENet.
Chaochen Wu, Guan Luo
ICMR2
2024 PI3D: Efficient Text-to-3D Generation with Pseudo-Image Diffusion
abstract
Diffusion models trained on large-scale text-image datasets have demonstrated a strong capability of con-trollable high-quality image generation from arbitrary text prompts. However, the generation quality and general-ization ability of 3D diffusion models is hindered by the scarcity of high-quality and large-scale 3D datasets. In this paper, we present PI3D, a framework that fully lever-ages the pre-trained text-to-image diffusion models' abil-ity to generate high-quality 3D shapes from text prompts in minutes. The core idea is to connect the 2D and 3D domains by representing a 3D shape as a set of Pseudo RGB Images. We fine-tune an existing text-to-image dif-fusion model to produce such pseudo-images using a small number of text-3D pairs. Surprisingly, we find that it can al-ready generate meaningful and consistent 3D shapes given complex text descriptions. We further take the generated shapes as the starting point for a lightweight iterative re-finement using score distillation sampling to achieve high-quality generation under a low budget. PI3D generates a single 3D shape from text in only 3 minutes and the quality is validated to outperform existing 3D generative models by a large margin.
Ying-Tian Liu, Guan Luo, Heyi Sun, Song-Hai Zhang
CVPR3
2024 3D Gaussian Editing with A Single Image
abstract
The modeling and manipulation of 3D scenes captured from the real world are pivotal in various applications, attracting growing research interest. While previous works on editing have achieved interesting results through manipulating 3D meshes, they often require accurately reconstructed meshes to perform editing, which limits their application in 3D content generation. To address this gap, we introduce a novel single-image-driven 3D scene editing approach based on 3D Gaussian Splatting, enabling intuitive manipulation via directly editing the content on a 2D image plane. Our method learns to optimize the 3D Gaussians to align with an edited version of the image rendered from a user-specified viewpoint of the original scene. To capture long-range object deformation, we introduce positional loss into the optimization process of 3D Gaussian Splatting and enable gradient propagation through reparameterization. To handle occluded 3D Gaussians when rendering from the specified viewpoint, we build an anchor-based structure and employ a coarse-to-fine optimization strategy capable of handling long-range deformation while maintaining structural stability. Furthermore, we design a novel masking strategy to adaptively identify non-rigid deformation regions for fine-scale modeling. Extensive experiments show the effectiveness of our method in handling geometric details, long-range, and non-rigid deformation, demonstrating superior editing flexibility and quality compared to previous approaches.
Guan Luo, Tian-Xing Xu, Ying-Tian Liu, Xiaoxiong Fan, Song-Hai Zhang
ACM Multimedia1
2024 VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
abstract
Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{generating} the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. \emph{The question is how to align the PV features with the generative models to facilitate the map estimation}. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on \url{https://github.com/Z1zyw/VQ-Map}.
Fudong Ge, Guan Luo, Bing Li 0001, Zhaoxiang Zhang 0001, Haibin Ling, Weiming Hu 0004
NeurIPS4
2024 Online biomedical named entities recognition by data and knowledge-driven model
Lulu Cao, Chaochen Wu, Guan Luo, Anni Zheng
Artif. Intell. Medicine3
2024 AdaPIP: Adaptive picture-in-picture guidance for 360° film watching
abstract
360° videos enable viewers to watch freely from different directions but inevitably prevent them from perceiving all the helpful information. To mitigate this problem, picture-in-picture (PIP) guidance was proposed using preview windows to show regions of interest (ROIs) outside the current view range. We identify several drawbacks of this representation and propose a new method for 360° film watching called AdaPIP. AdaPIP enhances traditional PIP by adaptively arranging preview windows with changeable view ranges and sizes. In addition, AdaPIP incorporates the advantage of arrow-based guidance by presenting circular windows with arrows attached to them to help users locate the corresponding ROIs more efficiently. We also adapted AdaPIP and Outside-In to HMD-based immersive virtual reality environments to demonstrate the usability of PIP-guided approaches beyond 2D screens. Comprehensive user experiments on 2D screens, as well as in VR environments, indicate that AdaPIP is superior to alternative methods in terms of visual experiences while maintaining a comparable degree of immersion.
Yi-Xiao Li, Guan Luo, Yi-Ke Xu, Yu He 0001, Song-Hai Zhang
Comput. Vis. Media2
2023 Health and Senior Care Video Moment Localization With Procedure Knowledge Distillation
abstract
With the aging population, health and senior care are becoming to be crucial issues for the whole world. Because the number of healthcare professionals is far from fulfilling increasing patients’ and seniors’ needs, seeking services from non-professional healthcare staff, such as home caregivers, is indispensable. Methods that support locating video moments with natural language queries can improve the normalization of operations for the non-professional healthcare staff and reduce their time expenditure on specific action moment retrieval. Addressing this problem, we propose a cross-modal neural network model for effective health and senior care video localization. Our model learns procedures in the video reference and uses procedure knowledge to improve the model’s localization performance. We conduct experiments on a dataset for health and senior care video localization and an open-accessible dataset about medical instruction. Experiment results show procedure knowledge can remarkably improve the model’s capacity for video moment localization. We hope our dataset and method could promote the development of cross-modal research and application for health and senior care.
Chaochen Wu, Yuna Jiang, Guan Luo
BIBM4
2022 Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation
abstract
Open- vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open- vocabulary two-stage detectors employ instance-level visual-to- visual knowledge distillation to align the visual space of the detector with the semantic space of the Pre-trained Visual-Language Model (PVLM). However, in the more efficient one-stage detector, the absence of class-agnostic object proposals hinders the knowledge distil-lation on unseen objects, leading to severe performance degradation. In this paper, we propose a hierarchical visual-language knowledge distillation method, i.e., Hi-erKD, for open-vocabulary one-stage detection. Specifi-cally, a global-level knowledge distillation is explored to transfer the knowledge of unseen categories from the PVLM to the detector. Moreover, we combine the proposed global-level knowledge distillation and the common instance-level knowledge distillation to learn the knowledge of seen and unseen categories simultaneously. Extensive experiments on MS-COCO show that our method significantly surpasses the previous best one-stage detector with 11.9% and 6.7% AP50 gains under the zero-shot detection and generalized zero-shot detection settings, and reduces the AP50performance gap from 14% to 7.3% compared to the best two-stage detector. Code will be released at this url11https://qithub.com/menqqiDyanqqe/HierKD.
Zongyang Ma, Guan Luo, Liang Li 0006, Shaoru Wang, Congxuan Zhang, Weiming Hu 0004
CVPR2
2022 Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive Clustering
abstract
This paper introduces an online self-supervised method that leverages inter- and intra-level variance for video representation learning. Most existing methods tend to focus on instance-level or inter-variance encoding but ignore the intra-variance existing in clips. The key observation to solving this problem is the underlying correlation between visual and audio, in which the distribution of flow patterns in feature space is diverse, but expresses complementary similar semantics. And in the semantic feature space, the horizontal dimension of the feature matrix could be regarded as cluster labels. These cluster labels should be consistent for different modalities of the same video clip. Based on this idea, we propose an end-to-end inter-intra cross-modality contrastive clustering scheme to simultaneously optimize the inter- and intra-level contrastive loss. Experiments show that our proposed approach is able to considerably outperform previous methods for self-supervised learning on HMDB51 and UCF101 when applied to video retrieval and action recognition tasks.
Jiutong Wei, Guan Luo, Bing Li 0001, Weiming Hu 0004
ICPR2
2020 An attention-based multi-task model for named entity recognition and intent analysis of Chinese online medical questions
Chaochen Wu, Guan Luo, Yin Ren, Anni Zheng
J. Biomed. Informatics2
2020 Tangent Fisher Vector on Matrix Manifolds for Action Recognition
abstract
In this paper, we address the problem of representing and recognizing human actions from videos on matrix manifolds. For this purpose, we propose a new vector representation method, named tangent Fisher vector, to describe video sequences in the Fisher kernel framework. We first extract dense curved spatio-temporal cuboids from each video sequence. Compared with the traditional 'straight cuboids', the dense curved spatio-temporal cuboids contain much more local motion information. Each cuboid is then described using a linear dynamical system (LDS) to simultaneously capture the local appearance and dynamics. Furthermore, a simple yet efficient algorithm is proposed to learn the LDS parameters and approximate the observability matrix at the same time. Each video sequence is thus represented by a set of LDSs. Considering that each LDS can be viewed as a point in a Grassmann manifold, we propose to learn an intrinsic GMM on the manifold to cluster the LDS points. Finally a tangent Fisher vector is computed by first accumulating all the tangent vectors in each Gaussian component, and then concatenating the normalized results across all the Gaussian components. A kernel is defined to measure the similarity between tangent Fisher vectors for classification and recognition of a video sequence. This approach is evaluated on the state-of-the-art human action benchmark datasets. The recognition performance is competitive when compared with current state-of-the-art results.
Guan Luo, Jiutong Wei, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Image Process.1
2019 Semi-interactive Attention Network for Answer Understanding in Reverse-QA
Qing Yin, Guan Luo, Qinghua Hu, Ou Wu 0001
PAKDD (2)2
2014 Learning Human Actions by Combining Global Dynamics and Local Appearance
abstract
In this paper, we address the problem of human action recognition through combining global temporal dynamics and local visual spatio-temporal appearance features. For this purpose, in the global temporal dimension, we propose to model the motion dynamics with robust linear dynamical systems (LDSs) and use the model parameters as motion descriptors. Since LDSs live in a non-Euclidean space and the descriptors are in non-vector form, we propose a shift invariant subspace angles based distance to measure the similarity between LDSs. In the local visual dimension, we construct curved spatio-temporal cuboids along the trajectories of densely sampled feature points and describe them using histograms of oriented gradients (HOG). The distance between motion sequences is computed with the Chi-Squared histogram distance in the bag-of-words framework. Finally we perform classification using the maximum margin distance learning method by combining the global dynamic distances and the local visual distances. We evaluate our approach for action recognition on five short clips data sets, namely Weizmann, KTH, UCF sports, Hollywood2 and UCF50, as well as three long continuous data sets, namely VIRAT, ADL and CRIM13. We show competitive results as compared with current state-of-the-art methods.
Guan Luo, Guodong Tian, Chunfeng Yuan, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Learning silhouette dynamics for human action recognition
abstract
In this paper, we address the problem of recognizing human actions with motion dynamics alone. For this purpose, we propose to use silhouette sequences to represent the human actions by discarding the appearance information, and then model the sequences with linear dynamical systems (LDSs). Recognition is achieved by directly comparing the distance between LDSs, rather than resorting to complex Bayesian learning and inference. In particular, we introduce an efficient optimization method to learn robust LDSs, and develop a shift invariant distance metric to measure the similarity on the LDSs space. We evaluate our approach on the human action data set and achieve comparable results.
Guan Luo, Weiming Hu 0004
ICIP1
2013 Action recognition using linear dynamic systems
Haoran Wang 0001, Chunfeng Yuan, Guan Luo, Weiming Hu 0004, Changyin Sun 0001
Pattern Recognit.3
2010 Occlusion Handling with ℓ1-Regularized Sparse Reconstruction
Wei Li 0034, Bing Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Hanzi Wang, Guan Luo
ACCV (4)6
2009 Human Action Recognition under Log-Euclidean Riemannian Metric
Chunfeng Yuan, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank, Guan Luo
ACCV (1)5
2009 Image spam filtering using Fourier-Mellin invariant features
abstract
Image spam is a new obfuscating method which spammers invented to more effectively bypass conventional text based spam filters. In this paper, a framework for filtering image spams by using the Fourier-Mellin invariant features is described. Fourier-Mellin features are robust for most kinds of image spam variations. A one-class classifier, the support vector data description (SVDD), is exploited to model the boundary of image spam class in the feature space without using information of legitimate emails. Experimental results demonstrate that our framework is effective for fighting image spam.
Haiqiang Zuo, Xi Li 0001, Ou Wu 0001, Weiming Hu 0004, Guan Luo
ICASSP5
2009 Detecting image spam using local invariant features and pyramid match kernel
abstract
Image spam is a new obfuscating method which spammers invented to more effectively bypass conventional text based spam filters. In this paper, we extract local invariant features of images and run a one-class SVM classifier which uses the pyramid match kernel as the kernel function to detect image spam. Experimental results demonstrate that our algorithm is effective for fighting image spam.
Haiqiang Zuo, Weiming Hu 0004, Ou Wu 0001, Yunfei Chen 0002, Guan Luo
WWW5
2008 Trajectory-Based Video Retrieval Using Dirichlet Process Mixture Models
abstract
In this paper, we present a trajectory-based video retrieval framework using Dirichlet process mixture models. The main contribution of this framework is four-fold. (1) We apply a Dirichlet process mixture model (DPMM) to unsupervised trajectory learning. DPMM is a countably infinite mixture model with its components growing by itself. (2) We employ a time-sensitive Dirichlet process mixture model (tDPMM) to learn trajectories ’ time-series characteristics. Furthermore, a novel likelihood estimation algorithm for tDPMM is proposed for the first time. (3) We develop a tDPMM-based probabilistic model matching scheme, which is empirically shown to be more error-tolerating and is able to deliver higher retrieval accuracy than the peer methods in the literature. (4) The framework has a nice scalability and adaptability in the sense that when new cluster data are presented, the framework automatically identifies the new cluster information without having to redo the training. Theoretic analysis and experimental evaluations against the state-of-the-art methods demonstrate the promise and effectiveness of the framework. 1
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
BMVC5
2007 Kernel-Bayesian Framework for Object Tracking
Xiaoqin Zhang 0002, Weiming Hu 0004, Guan Luo, Stephen J. Maybank
ACCV (1)3
2007 Robust Visual Tracking Based on Incremental Tensor Subspace Learning
abstract
Most existing subspace analysis-based tracking algorithms utilize a flattened vector to represent a target, resulting in a high dimensional data learning problem. Recently, subspace analysis is incorporated into the multilinear framework which offline constructs a representation of image ensembles using high-order tensors. This reduces spatio-temporal redundancies substantially, whereas the computational and memory cost is high. In this paper, we present an effective online tensor subspace learning algorithm which models the appearance changes of a target by incrementally learning a low-order tensor eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking then is led by the state inference within the framework in which a particle filter is used for propagating sample distributions over the time. A novel likelihood function, based on the tensor reconstruction error norm, is developed to measure the similarity between the test image and the learned tensor subspace model during the tracking. Theoretic analysis and experimental evaluations against a state-of-the-art method demonstrate the promise and effectiveness of this algorithm.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
ICCV5
2007 Dominant Sets-Based Action Recognition using Image Sequence Matching
abstract
Action recognition is one of the most active research fields in computer vision. In this paper, we propose a novel method for classifying human actions in a series of image sequences containing certain actions. Human action in image sequences can be recognized by a time-varying contour of human body. We first extract shape context of each contour to form the feature space. Then the dominant sets approach is used for feature clustering and classification to obtain the labeled sequences. Finally, we use a smoothing algorithm upon the labeled sequences to recognize human actions. The proposed dominant sets-based approach has been tested in comparison to three classical methods: K-means, mean shift, and fuzzy-C-mean. Experimental results demonstrate that the dominant sets-based approach achieves the best recognition performance. Moreover, our method is robust to non-rigid deformations, significant scale changes, high action irregularities, and low quality video.
Qingdi Wei, Weiming Hu 0004, Xiaoqin Zhang 0002, Guan Luo
ICIP (6)4