Guangxing Han

dblp:208/4894 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-8307-8716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 12 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting
Guangxing Han, Long Chen 0016, Jiawei Ma, Shiyuan Huang 0001, Rama Chellappa, Shih-Fu Chang
Int. J. Comput. Vis.1
2026 AddrProbe: An Internet-Wide Active IPv6 Address Probing System With Limited Seeds
abstract
With the large-scale deployment of IPv6, it is becoming more and more important to probe active IPv6 addresses on the global Internet. However, the vast address space and the random distribution of active addresses make the probing process full of challenges, especially for the probing of IPv6 prefixes without seed addresses. Furthermore, the widespread existence of IPv6 aliased prefixes also causes significant trouble for probing. In this paper, we presentAddrProbe, an active IPv6 address probing system, which dynamically probes all global routing prefixes based on learned fine-grained address patterns from limited seed addresses and quickly detects aliased prefixes during probing. The evaluation results show thatAddrProbeachieves a hit rate of 23%-45% with all routing prefixes announced by the BGP system, which is 6.6-13× that of current state-of-the-art approaches (no more than 4%). Moreover, we find 1.2×1033aliased addresses characterized by the detected aliased prefixes, covering 6,412 routing prefixes, which is a 107× and 5.9× improvement over existing methods, respectively. Finally, an IPv6 Hitlist is constructed based on the long-term probing results, which contains 562M addresses covering 190K routing prefixes and 29K ASes. These widely distributed addresses are meaningful for analyzing IPv6 address assignments and some other IPv6 measurement activities.
Daguo Cheng, Lin He 0004, Qilei Yin, Guangxing Han, Boran Jin, Ying Liu 0024, Guanglei Song, Jinlong E, Tiankai Yang 0001, Jiahai Yang 0001
IEEE Trans. Netw.4
2025 TIPS: Text-Image Pretraining with Spatial awareness
abstract
While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense vision applications (e.g. depth estimation, semantic segmentation), despite the lack of explicit supervisory signals. In this paper, we close this gap between image-text and self-supervised learning, by proposing a novel general-purpose image-text model, which can be effectively used off the shelf for dense and global vision tasks. Our method, which we refer to as Text-Image Pretraining with Spatial awareness (TIPS), leverages two simple and effective insights. First, on textual supervision: we reveal that replacing noisy web image captions by synthetically generated textual descriptions boosts dense understanding performance significantly, due to a much richer signal for learning spatially aware representations. We propose an adapted training method that combines noisy and synthetic captions, resulting in improvements across both dense and global understanding tasks. Second, on the learning technique: we propose to combine contrastive image-text learning with self-supervised masked image modeling, to encourage spatial coherence, unlocking substantial enhancements for downstream applications. Building on these two ideas, we scale our model using the transformer architecture, trained on a curated set of public images. Our experiments are conducted on $8$ tasks involving $16$ datasets in total, demonstrating strong off-the-shelf performance on both dense and global understanding, for several image-only and image-text tasks. Code and models are released at https://github.com/google-deepmind/tips .
Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh 0001, Arjun Karpur, Koert Chen, Bingyi Cao, Daniel Salz, Guangxing Han, Jan Dlabal, Dan Gnanapragasam, Mojtaba Seyedhosseini, Howard Zhou, André Araújo 0001
ICLR9
2024 Few-Shot Object Detection with Foundation Models
abstract
Few-shot object detection (FSOD) aims to detect objects with only a few training examples. Visual feature extraction and query-support similarity learning are the two critical components. Existing works are usually developed based on ImageNet pre-trained vision backbones and design sophis-ticated metric-learning networks for few-shot learning, but still have inferior accuracy. In this work, we study few-shot object detection using modern foundation models. First, vision-only contrastive pre-trained DINOv2 model is used for the vision backbone, which shows strong transferable performance without tuning the parameters. Second, Large Language Model (LLM) is employed for contextualized few-shot learning with the input of all classes and query image proposals. Language instructions are carefully designed to prompt the LLM to classify each proposal in context. The contextual information include proposal-proposal relations, proposal-class relations, and class-class relations, which can largely promote few-shot learning. We comprehensively evaluate the proposed model (FM-FSOD) in multiple FSOD benchmarks, achieving state-of-the-arts performance.
Guangxing Han, Ser-Nam Lim
CVPR1
2024 Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
abstract
The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tuning. However, due to the enormous diversity in input-output formats in the vision domain, existing general-purpose models fail to successfully integrate segmentation and multi-image inputs with coarse-level tasks into a single framework. In this work, we introduce VistaLLM, a powerful visual system that addresses coarse- and fine-grained VL tasks over single and multiple input images using a unified framework. VistaLLM utilizes an instruction-guided image tokenizer that filters global embeddings using task descriptions to extract compressed and refined features from numerous images. Moreover, VistaLLM employs a gradient-aware adaptive sampling technique to represent binary segmentation masks as sequences, significantly improving over previously used uniform sampling. To bolster the desired capability of VistaLLM, we curate CoinIt, a comprehensive coarse-to-fine instruction tuning dataset with 6.8M samples. We also address the lack of multi-image grounding datasets by introducing a novel task, AttCoSeg (Attribute-level Co-Segmentation), which boosts the model's reasoning and grounding capability over multiple input images. Extensive experiments on a wide range of V- and VL tasks demonstrate the effectiveness of VistaLLM by achieving consistent state-of-the-art performance over strong baselines across many downstream tasks. Our project page can be found at https://shramanpramanick.github.io/VistaLLM/.
Shraman Pramanick, Guangxing Han, Sayan Nag, Ser-Nam Lim, Nicolas Ballas, Qifan Wang 0001, Rama Chellappa, Amjad Almahairi
CVPR2
2024 One-Shot Unsupervised Cross-Domain Person Re-Identification
abstract
Cross-domain person re-identification is challenging due to the notorious domain shift problem. Most of the existing unsupervised cross-domain person ReID methods require a large number of unlabeled target-domain samples for adaptation. However, large scale of training data are not always available due to public privacy. Domain generalization methods have inferior adaptation ability without seeing any target domain data. Inspired by the few-shot learning capability of human vision system, we propose a novel setting, one-shot unsupervised cross-domain for person ReID and study the ability of adaptation using the minimum number of image in the target domain during training. Specifically, we first propose a novel Group Normalization (GN) based domain generalizable ReID model. We show that the GN based model could strike a better balance between model discrimination and generalization ability, compared with the Batch Normalization (BN) and Instance Normalization (IN) counterparts, and is more suitable for domain generalizable ReID baseline model. Then besides the supervised feature learning task in the source domain, we introduce two self-supervised learning tasks using the one-shot target domain data to further improve the generalization ability of the ReID model. We carefully design model architecture and perform model training to reduce overfitting to the one-shot target domain. Extensive experiments demonstrate the effectiveness of our approach for one-shot unsupervised cross-domain ReID. Our approach can be extended to few-shot setting and increasing the number of shot up to 1,000 images can steadily increase the performance, which provides practical values to the community.
Guangxing Han, Xuan Zhang 0006, Chongrong Li
IEEE Trans. Circuits Syst. Video Technol.1
2023 Supervised Masked Knowledge Distillation for Few-Shot Transformers
abstract
Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However, under few-shot learning (FSL) settings on small datasets with only a few labeled data, ViT tends to overfit and suffers from severe performance degradation due to its absence of CNN-alike inductive bias. Previous works in FSL avoid such problem either through the help of self-supervised auxiliary losses, or through the dextile uses of label information under supervised settings. But the gap between self-supervised and supervised few-shot Transformers is still unfilled. Inspired by recent advances in self-supervised knowledge distillation and masked image modeling (MIM), we propose a novel Supervised Masked Knowledge Distillation model (SMKD) for few-shot Transformers which incorporates label information into self-distillation frameworks. Compared with previous self-supervised methods, we allow intra-class knowledge distillation on both class and patch tokens, and introduce the challenging task of masked patch tokens reconstruction across intra-class images. Experimental results on four few-shot classification benchmark datasets show that our method with simple design outperforms previous methods by a large margin and achieves a new start-of-the-art. Detailed ablation studies confirm the effectiveness of each component of our model. Code for this paper is available here: https://github.com/HL-hanlin/SMKD.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Xudong Lin 0003, Shih-Fu Chang
CVPR2
2023 DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection
abstract
Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approaches enhance few-shot generalization with the sacrifice of base-class performance, or maintain high precision in base-class detection with limited improvement in novel-class adaptation. In this paper, we point out the reason is insufficient Discriminative feature learning for all of the classes. As such, we propose a new training framework, DiGeo, to learn Geometry-aware features of interclass separation and intra-class compactness. To guide the separation of feature clusters, we derive an offline simplex equiangular tight frame (ETF) classifier whose weights serve as class centers and are maximally and equally separated. To tighten the cluster for each class, we include adaptive class-specific margins into the classification loss and encourage the features close to the class centers. Experimental studies on two few-shot benchmark datasets (VOC, COCO) and one long-tail dataset (LVIS) demonstrate that, with a single model, our method can effectively improve generalization on novel classes without hurting the detection of base classes. Our code can be found here.
Jiawei Ma, Yulei Niu, Jincheng Xu, Shiyuan Huang 0001, Guangxing Han, Shih-Fu Chang
CVPR5
2023 TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Xudong Lin 0003, Guangxing Han, Shih-Fu Chang
ICLR6
2022 Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment
abstract
Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern object detectors. However, the quality of proposals generated for few-shot classes using existing methods is far worse than that of many-shot classes, e.g., missing boxes for few-shot classes due to misclassification or inaccurate spatial locations with respect to true objects. To address the noisy proposal problem, we propose a novel meta-learning based FSOD model by jointly optimizing the few-shot proposal generation and fine-grained few-shot proposal classification. To improve proposal generation for few-shot classes, we propose to learn a lightweight metric-learning based prototype matching network, instead of the conventional simple linear object/nonobject classifier, e.g., used in RPN. Our non-linear classifier with the feature fusion network could improve the discriminative prototype matching and the proposal recall for few-shot classes. To improve the fine-grained few-shot proposal classification, we propose a novel attentive feature alignment method to address the spatial misalignment between the noisy proposals and few-shot classes, thus improving the performance of few-shot object detection. Meanwhile we learn a separate Faster R-CNN detection head for many-shot base classes and show strong performance of maintaining base-classes knowledge. Our model achieves state-of-the-art performance on multiple FSOD benchmarks over most of the shots and metrics.
Guangxing Han, Shiyuan Huang 0001, Jiawei Ma, Yicheng He, Shih-Fu Chang
AAAI1
2022 Few-Shot Object Detection with Fully Cross-Transformer
abstract
Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and calculate the similarity between image regions and few-shot examples for detection. However, in previous works, the interaction between the two branches is only restricted in the detection head, while leaving the remaining hundreds of layers for separate feature extraction. Inspired by the recent work on vision transformers and vision-language transformers, we propose a novel Fully Cross-Transformer based model (FCT) for FSOD by incorporating cross-transformer into both the feature backbone and detection head. The asymmetric-batched cross-attention is proposed to aggregate the key information from the two branches with different batch sizes. Our model can improve the few-shot similarity learning between the two branches by introducing the multi-level interactions. Comprehensive experiments on both PASCAL VOC and MSCOCO FSOD benchmarks demonstrate the effectiveness of our model.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Shih-Fu Chang
CVPR1
2022 Task-Adaptive Negative Envision for Few-Shot Open-Set Recognition
abstract
We study the problem of few-shot open-set recognition (FSOR), which learns a recognition system capable of both fast adaptation to new classes with limited labeled exam-ples and rejection of unknown negative samples. Traditional large-scale open-set methods have been shown in-effective for FSOR problem due to data limitation. Current FSOR methods typically calibrate few-shot closed-set clas-sifiers to be sensitive to negative samples so that they can be rejected via thresholding. However, threshold tuning is a challenging process as different FSOR tasks may require different rejection powers. In this paper, we instead propose task-adaptive negative class envision for FSOR to integrate threshold tuning into the learning process. Specifically, we augment the few-shot closed-set classifier with additional negative prototypes generated from few-shot examples. By incorporating few-shot class correlations in the negative generation process, we are able to learn dynamic rejection boundaries for FSOR tasks. Besides, we extend our method to generalized few-shot open-set recognition (GF-SOR), which requires classification on both many-shot and few-shot classes as well as rejection of negative samples. Extensive experiments on public benchmarks validate our methods on both problems.11Code available at https://github.com/shiyuanh/TANE
Shiyuan Huang 0001, Jiawei Ma, Guangxing Han, Shih-Fu Chang
CVPR3
2022 Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding Across Heads
Jiawei Ma, Guangxing Han, Shiyuan Huang 0001, Yuncong Yang, Shih-Fu Chang
ECCV (26)2
2022 Explicit Image Caption Editing
Zhen Wang 0004, Long Chen 0016, Guangxing Han, Yulei Niu, Jian Shao 0001, Jun Xiao 0001
ECCV (36)4
2022 Weakly-Supervised Temporal Article Grounding
abstract
Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Long Chen 0016, Yulei Niu, Brian Chen 0001, Xudong Lin 0003, Guangxing Han, Christopher Thomas 0004, Hammad A. Ayyubi, Heng Ji 0001, Shih-Fu Chang
EMNLP5
2021 Query Adaptive Few-Shot Object Detection with Heterogeneous Graph Convolutional Networks
abstract
Few-shot object detection (FSOD) aims to detect never-seen objects using few examples. This field sees recent improvement owing to the meta-learning techniques by learning how to match between the query image and few-shot class examples, such that the learned model can generalize to few-shot novel classes. However, currently, most of the meta-learning-based methods perform parwise matching between query image regions (usually proposals) and novel classes separately, therefore failing to take into account multiple relationships among them. In this paper, we propose a novel FSOD model using heterogeneous graph convolutional networks. Through efficient message passing among all the proposal and class nodes with three different types of edges, we could obtain context-aware proposal features and query-adaptive, multiclass-enhanced prototype representations for each class, which could help promote the pairwise matching and improve final FSOD accuracy. Extensive experimental results show that our proposed model, denoted as QA-FewDet, outperforms the current state-of-the-art approaches on the PASCAL VOC and MSCOCO FSOD benchmarks under different shots and evaluation metrics.
Guangxing Han, Yicheng He, Shiyuan Huang 0001, Jiawei Ma, Shih-Fu Chang
ICCV1
2021 Partner-Assisted Learning for Few-Shot Image Classification
abstract
Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. In this paper, we focus on the design of training strategy to obtain an elemental representation such that the prototype of each novel class can be estimated from a few labeled samples. We propose a two-stage training scheme, Partner-Assisted Learning (PAL), which first trains a Partner Encoder to model pair-wise similarities and extract features serving as soft-anchors, and then trains a Main Encoder by aligning its outputs with soft-anchors while attempting to maximize classification performance. Two alignment constraints from logit-level and feature-level are designed individually. For each few-shot task, we perform prototype classification. Our method consistently outperforms the state-of-the-art methods on four benchmarks. Detailed ablation studies of PAL are provided to justify the selection of each component involved in training.
Jiawei Ma, Hanchen Xie, Guangxing Han, Shih-Fu Chang, Aram Galstyan, Wael Abd-Almageed
ICCV3
2020 Unsupervised Feature Propagation for Fast Video Object Detection Using Generative Adversarial Networks
Xuan Zhang 0006, Guangxing Han, Wenduo He
MMM (1)2
2020 A parallel optimization for energy and robustness of file distribution services
Dongchao Ma, Guangxing Han, Li Ma 0007, Chengan Zhao
Peer-to-Peer Netw. Appl.2
2018 Semi-Supervised DFF: Decoupling Detection and Feature Flow for Video Object Detectors
abstract
For efficient video object detection, our detector consists of a spatial module and a temporal module. The spatial module aims to detect objects in static frames using convolutional networks, and the temporal module propagates high-level CNN features to nearby frames via light-weight feature flow. Alternating the spatial and temporal module by a proper interval makes our detector fast and accurate. Then we propose a two-stage semi-supervised learning framework to train our detector, which fully exploits unlabeled videos by decoupling the spatial and temporal module. In the first stage, the spatial module is learned by traditional supervised learning. In the second stage, we employ both feature regression loss and feature semantic loss to learn our temporal module via unsupervised learning. Different to traditional methods, our method can largely exploit unlabeled videos and bridges the gap of object detectors in image and video domain. Experiments on the large-scale ImageNet VID dataset demonstrate the effectiveness of our method. Code will be made publicly available.
Guangxing Han, Xuan Zhang 0006, Chongrong Li
ACM Multimedia1
2017 Single shot object detection with top-down refinement
abstract
General object detection is one of the most challenging tasks in computer vision for it requires both high running speed and detection accuracy. In this paper, we propose a single shot object detector with top-down refinement, denoted as SSD-TDR. It not only runs at high speed and also detects multi-scale objects accurately. Concretely, original SSD directly adopts the built-in multi-scale hierarchy of convolutional neural networks for detection. However, object detection needs high semantic knowledge to recognise objects while low-level convolutional features do not have. We thus build a sequence of top-down refinement modules to transmit semantic knowledge backward such that all layers have rich semantics. Experiments on PASCAL VOC 2007 and 2012 demonstrate that our network achieves competitive results both in speed and accuracy compared to other VGG16 based networks.
Guangxing Han, Xuan Zhang 0006, Chongrong Li
ICIP1
2017 Revisiting Faster R-CNN: A Deeper Look at Region Proposal Network
Guangxing Han, Xuan Zhang 0006, Chongrong Li
ICONIP (3)1