Han Zhang 0048

dblp:26/4189-48 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
abstract
Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM_S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks.
Han Zhang 0048, Zhantao Yang, Fangyi Chen, Anudeepsekhar Bolimera, Marios Savvides
AAAI2
2024 A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute Editing
abstract
Facial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenges persist in achieving precise, 3D-aware attribute modifications, which are crucial for consistent and accurate representations of faces from different angles. Current methods struggle with semantic entanglement and lack effective guidance for incorporating attributes while maintaining image integrity. To address these issues, we introduce a novel framework that merges the strengths of latent-based and reference-based editing methods. Our approach employs a 3D GAN inversion technique to embed attributes from the reference image into a tri-plane space, ensuring 3D consistency and realistic viewing from multiple perspectives. We utilize blending techniques and predicted semantic masks to locate precise edit regions, merging them with the contextual guidance from the reference image. A coarse-to-fine inpainting strategy is then applied to preserve the integrity of untargeted areas, significantly enhancing realism. Our evaluations demonstrate superior performance across diverse editing tasks, validating our framework’s effectiveness in realistic and applicable facial attribute editing.
Yutong Zheng, Yen-Shuo Su, Anudeepsekhar Bolimera, Han Zhang 0048, Fangyi Chen, Marios Savvides
IJCB5
2023 Enhanced Training of Query-Based Object Detection via Selective Query Recollection
abstract
This paper investigates a phenomenon where query-based object detectors mispredict at the last decoding stage while predicting correctly at an intermediate stage. We review the training process and attribute the overlooked phenomenon to two limitations: lack of training emphasis and cascading errors from decoding sequence. We design and present Selective Query Recollection (SQR), a simple and effective training strategy for query-based object detectors. It cumulatively collects intermediate queries as decoding stages go deeper and selectively forwards the queries to the downstream stages aside from the sequential structure. Such-wise, SQR places training emphasis on later stages and allows later stages to work with intermediate queries from earlier stages directly. SQR can be easily plugged into various query-based object detectors and significantly enhances their performance while leaving the inference pipeline unchanged. As a result, we apply SQR on Adamixer, DAB-DETR, and Deformable-DETR across various settings (backbone, number of queries, schedule) and consistently brings 1.4 ~2.8 AP improvement. Code is available at https://github.com/Fangyi-Chen/SQR
Fangyi Chen, Han Zhang 0048, Kai Hu 0010, Chenchen Zhu, Marios Savvides
CVPR2
2022 Powering Finetuning in Few-Shot Learning: Domain-Agnostic Bias Reduction with Selected Sampling
abstract
In recent works, utilizing a deep network trained on meta-training set serves as a strong baseline in few-shot learning. In this paper, we move forward to refine novel-class features by finetuning a trained deep network. Finetuning is designed to focus on reducing biases in novel-class feature distributions, which we define as two aspects: class-agnostic and class-specific biases. Class-agnostic bias is defined as the distribution shifting introduced by domain difference, which we propose Distribution Calibration Module(DCM) to reduce. DCM owes good property of eliminating domain difference and fast feature adaptation during optimization. Class-specific bias is defined as the biased estimation using a few samples in novel classes, which we propose Selected Sampling(SS) to reduce. Without inferring the actual class distribution, SS is designed by running sampling using proposal distributions around support-set samples. By powering finetuning with DCM and SS, we achieve state-of-the-art results on Meta-Dataset with consistent performance boosts over ten datasets from different domains. We believe our simple yet effective method demonstrates its possibility to be applied on practical few-shot applications.
Ran Tao 0013, Han Zhang 0048, Yutong Zheng, Marios Savvides
AAAI2
2022 Unitail: Detecting, Reading, and Matching in Retail Scene
Fangyi Chen, Han Zhang 0048, Zaiwang Li, Jiachen Dou, Shentong Mo, Hao Chen 0102, Uzair Ahmed, Chenchen Zhu, Marios Savvides
ECCV (7)2
2020 Solving Missing-Annotation Object Detection with Background Recalibration Loss
abstract
This paper focuses on a novel and challenging detection scenario: A majority of true objects/instances is unlabeled in the datasets, so these missing-labeled areas will be regarded as the background during training. Previous art [1] on this problem has proposed to use soft sampling to re-weight the gradients of RoIs based on the overlaps with positive instances, while their method is mainly based on the two-stage detector (i.e. Faster RCNN) which is more robust and friendly for the missing label scenario. In this paper, we introduce a superior solution called Background Recalibration Loss (BRL) that can automatically re-calibrate the loss signals according to the pre-defined IoU threshold and input image. Our design is built on the one-stage detector which is faster and lighter. Inspired by the Focal Loss [2] formulation, we make several significant modifications to fit on the missing-annotation circumstance. We conduct extensive experiments on the curated PASCAL VOC [3] and MS COCO [4] datasets. The results demonstrate that our proposed method outperforms the baseline and other state-of-the-arts by a large margin.
Han Zhang 0048, Fangyi Chen, Qiqi Hao, Chenchen Zhu, Marios Savvides
ICASSP1
2020 NCMS: Towards accurate anchor free object detection through ℓ2 norm calibration and multi-feature selection
abstract
We present simple and flexible drop-in modules in feature pyramids for general object detection, which can be easily generalized to other anchor-free detectors without introducing extra parameters, and only involves negligible computational cost on training and testing. The proposed detector, called NCMS, inserts a simple norm calibration (NC) operation between the feature pyramids and detection head to alleviate and balance the norm bias caused by feature pyramid network (FPN). Furthermore, the NCMS leverages an enhanced multi-feature selective strategy (MS) during training to assign the ground-truth to particular feature pyramid levels as supervisions, in order to obtain more discriminative representation for objects. By generalizing to the state-of-the-art FSAF module (Zhu et al., 2019), our NCMS improves it by 1.6% on COCO val set without bells and whistles. The resulting best model achieves 44.0% mAP with single-model and single-scale testing, which is a fairly competitive result.
Fangyi Chen, Chenchen Zhu, Han Zhang 0048, Marios Savvides
Comput. Vis. Image Underst.4