VLDB 2026 Research / reviewers in the wild / expert
Zhaofeng He 0001
dblp:13/3992-1
· DBLP profile ↗
88ranked-venue papers
5as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 3 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 3 first-author · 30 since 2021Security and privacy · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkabstractThe impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model families. Moreover, overlooking the underlying toxic risks involved in the input prefix can lead to toxic accumulation during autoregressive generation. Existing methods rely on external strong attribute interventions to address this issue, which further exacerbates contextual semantic inconsistencies and makes it difficult to balance toxicity efficacy and generation quality. To address these concerns, we propose a novel Model-Agnostic Adaptive Detoxification (MAAD) framework. To address accumulating toxicity, we present prefix heuristics that serve as contextual signals, guiding the base LLM toward safer generation. Along this line, we construct an antidote dataset to support a lightweight model, Detoxifier, which steers the base LLM to make in-scope and reliable detoxifying distribution adjustments while preserving fluency and contextual understanding. Designed as an easy-to-deploy module, Detoxifier requires a small amount of data and can be seamlessly applied to various base LLMs with one-off training. Since over-purifying often reduces diversity, we also propose a dynamic truncation method called CW-cutoff sampling to trade off language model quality and diversity. Extensive experiments demonstrate that MAAD strikes a better balance between detoxification effectiveness and generation quality, while also maintaining model utility. Yuhu Shang, Xiang Cheng 0003, Yimeng Ren 0001, Huijia Wu, Xuexiong Luo, Kangkang Lu 0002, Jian Zhao 0018, Zhaofeng He 0001 |
AAAI | 8 |
| 2026 | Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace AdaptationabstractDianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He 0001 |
ACL (1) | 8 |
| 2026 | SCVQ: Sparse-Compensated Vector Quantization for Large Language ModelsabstractZixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He 0001 |
ACL (1) | 7 |
| 2026 | ForgeryMoE: Mixture of Experts for Image Forgery Detection under JPEG CompressionabstractAI-generated imagery threatens information integrity, making reliable forgery detection crucial. However, the pervasive use of JPEG compression throughout image sharing pipelines critically undermines the robustness and generalization capability of existing detectors. To address this, we propose ForgeryMoE, a new Mixture of Experts (MoE) framework designed for universal and robust image forgery detection. Our framework strategically integrates three complementary experts: a Frequency Domain Expert that analyzes wavelet-based artifacts, a Pre-trained model expert that leverages CLIP and DINOv2 for semantic inconsistencies, and a Visual Domain Expert that captures pixel-level features and reconstruction anomalies. A dual gating mechanism dynamically estimates expert reliability and adaptively fuses their decisions. Extensive experiments on the GenImage benchmark show our method achieves state-of-the-art performance in both pristine and JPEG compressed settings, demonstrating strong generalization across diverse generative models and compression qualities. Saihui Hou, Jian Zhao 0006, Zhaofeng He 0001 |
ICMR | 6 |
| 2026 | Knowledge-guided policy arbitration: A hierarchical cognitive framework for safety-critical decision-making under dynamic conflicting objectives
Fuqing Bie, Xingyang Chang, Leyan Wang, Dehua Ma, Songfu Xu, Shuodi Liu, Yingzhuo Liu, Liuyu Xiang, Zhaofeng He 0001 |
Expert Syst. Appl. | 9 |
| 2026 | Attacks in Adversarial Machine Learning: A Systematic Survey from the Lifecycle Perspective
Baoyuan Wu, Zihao Zhu 0001, Li Liu 0036, Qingshan Liu 0001, Zhaofeng He 0001, Siwei Lyu |
Int. J. Comput. Vis. | 5 |
| 2026 | RainbowArena: A multi-agent toolkit for reinforcement learning and large language models in tabletop games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
Knowl. Based Syst. | 8 |
| 2026 | Attention-assisted multilevel fusion framework for generalized iris presentation attack detection
Caiyong Wang, Fukang Guo, Zhaofeng He 0001, Zhenan Sun |
Pattern Recognit. | 5 |
| 2026 | Symmetric Image-Text Tuning With Entropy-Guided Fusion for Online Continual Learning in Non-Stationary Visual StreamsabstractOnline continual learning studies how models learn from continuous and non-stationary data streams. In this paper, we observe that CLIP models exhibit an asymmetric image-text interaction under online continual learning. Specifically, text features of previously seen classes may introduce unfavorable supervision when paired with visual features of newly observed data, leading to catastrophic forgetting. To alleviate this issue, we propose a simple yet effective symmetric image-text tuning (SIT) strategy that removes such asymmetric text supervision during online learning. We further introduce an entropy-guided fusion (EGF) mechanism that adaptively combines predictions from the pretrained and finetuned branches based on their relative uncertainty. This design allows the model to recover pretrained knowledge when the finetuned branch becomes unreliable, while still preserving plasticity on recently observed classes when confidence is high. In addition, we present MiD-Blurry, an online continual learning benchmark that combines multiple class distribution patterns to better reflect realistic data streams with blurred temporal boundaries. Extensive experiments on standard continual learning benchmarks and the MiD-Blurry setting evaluate inference-at-any-time performance and generalization to future data. The results show that the proposed approach maintains a practical balance between adapting to new data and preserving previously learned information in realistic online learning scenarios. Leyuan Wang, Liuyu Xiang, Yiwei Ru, Yunlong Wang 0003, Zhaofeng He 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Rethinking Class-Incremental Learning From a Dynamic Imbalanced Learning PerspectiveabstractDeep neural networks suffer from catastrophic forgetting when continually learning new concepts. In this paper, we analyze this problem from a data imbalance point of view. We argue that the imbalance between old task and new task data contributes to forgetting of the old tasks. Moreover, the increasing imbalance ratio during incremental learning further aggravates the problem. To address the dynamic imbalance issue, we propose Uniform Prototype Contrastive Learning (UPCL), where uniform and compact features are learned. Specifically, we generate a set of non-learnable uniform prototypes before each task starts. Then we assign these uniform prototypes to each class and guide the feature learning through prototype contrastive learning. We also dynamically adjust the relative margin between old and new classes so that the feature distribution will be maintained balanced and compact. Finally, we demonstrate through extensive experiments that the proposed method achieves state-of-the-art performance on several benchmark including CIFAR-100, ImageNet-100, TinyImageNet, Food-101, and CUB-200. Experimental results show that our approach not only effectively addresses the issue of imbalanced old data in memory but also tackles the problem of imbalanced new data distributions. Leyuan Wang, Liuyu Xiang, Yunlong Wang 0003, Huijia Wu, Huafeng Yang, Jingqian Liu, Zhaofeng He 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | UniOrch: A Unified Mixed Framework for High-Efficiency LLM Training on Heterogeneous AI ChipsabstractEfficient coordination of heterogeneous AI chips (GPU/NPU/DCU) in data centers is crucial for Large Language Model (LLM) training, but this process is hindered by architectural mismatches, protocol fragmentation, and network partitioning. Existing solutions fail to achieve unified resource management across different chips, resulting in severe resource fragmentation and reduced allreduce efficiency. To overcome these limitations, this paper proposes the unified coordination framework UniOrch, which not only integrates three core functionalities, including hardware abstraction, software standardization, and communication coordination, but also enables the training and inference of large models across heterogeneous AI chips. UniOrch's hardware-agnostic bare-metal cloud eliminates virtualization overhead through Border Gateway Protocol Ethernet Virtual Private Network (BGP EVPN) overlay networks and gateway-based chip integration; its PyTorch-based adaptation layer masks hardware differences and reduces migration costs; the Transformer Collective Communication Library (TCCL) unifies NCCL, HCCL, and OpenMPIprotocols to support seamless hybrid parallel training. Furthermore, the framework ' score scheduling mechanism is the Heterogeneous Hybrid Estimation Model (HHEM), which employs a two stage cost model combining static analysis with dynamic runtime feedback to dynamically allocate computing power based on Transformer task loads, ensuring cross-chip task synchronization, resource pooling, and dynamic allocation. Deployment verification in real-world production environments (e.g., China Construction Bank) shows that UniOrch achieves significant improvements: resource utilization of heterogeneous AI infrastructure is increased by 35%, cross-chip latency is reduced by 42%, accuracy loss in heterogeneous environments is <0.8%. Jia Wang 0038, Yang Zhai, Haojie Wang 0004, Wanxin Song, Wei Li 0008, Zhaofeng He 0001 |
IEEE Trans. Parallel Distributed Syst. | 11 |
| 2025 | Psyche-Wave: Fusing Vector-Quantized Morphology and LLM-Inferred Semantics from Millimeter-Wave SCG for Psychological State DecodingabstractThis paper introduces Psyche-Wave, a novel paradigm for non-contact psychological state assessment, addressing the challenge that existing methods struggle to reconcile signal representation robustness with deep physiological semantic understanding. The proposed framework is built upon high-fidelity Seismocardiogram (SCG) and respiratory signals, captured by a proprietary high-sampling-rate millimeter-wave (mmWave) radar system. Psyche-Wave features a parallel dual-branch architecture for complementary feature extraction. The first, a Data-Driven Morphological Branch, employs Vector Quantization (VQ) to encode the Mel spectrogram of the SCG signal into a codebook-based representation, yielding a noise-resilient morphological embedding. The second, a Knowledge-Driven Semantic Branch, leverages a Large Language Model (LLM) to infer deep contextual relationships from medically significant physiological parameters—including heart rate variability, cardiac time intervals, and cardiopulmonary coupling—outputting a rich semantic embedding. These complementary embeddings are then integrated through a dedicated fusion module and passed to a downstream classifier for precise emotion and personality trait evaluation. Comprehensive evaluations on a newly collected high-fidelity dataset, referred to as mmHeart-Pro, and the public ReMAP dataset demonstrate state-of-the-art performance. This work pioneers a new path that fuses data-driven morphological analysis with knowledge-driven semantic reasoning, significantly advancing the accuracy and interpretability of non-contact psychological sensing. Yiwei Ru, Zhenbo Xu, Yanlin Xu, Huijia Wu, Zhaofeng He 0001, Zhenan Sun |
BIBM | 6 |
| 2025 | Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable reasoning and planning capabilities, driving extensive research into task decomposition.Existing task decomposition methods focus primarily on memory, tool usage, and feedback mechanisms, achieving notable success in specific domains, but they often overlook the trade-off between performance and cost.In this study, we first conduct a comprehensive investigation on task decomposition, identifying six categorization schemes.Then, we perform an empirical analysis of three factors that influence the performance and cost of task decomposition: categories of approaches, characteristics of tasks, and configuration of decomposition and execution models, uncovering three critical insights and summarizing a set of practical principles.Building on this analysis, we propose the Select-Then-Decompose strategy, which establishes a closed-loop problemsolving process composed of three stages: selection, execution, and verification.This strategy dynamically selects the most suitable decomposition approach based on task characteristics and enhances the reliability of the results through a verification module.Comprehensive evaluations across multiple benchmarks show that the Select-Then-Decompose consistently lies on the Pareto frontier, demonstrating an optimal balance between performance and cost.Our code is publicly available at https://github.com/summervvind/ Select-Then-Decompose. Shuodi Liu, Yingzhuo Liu, Zi Wang 0014, Huijia Wu, Liuyu Xiang, Zhaofeng He 0001 |
EMNLP | 7 |
| 2025 | GaussianEnhancer: A General Rendering Enhancer for Gaussian SplattingabstractGaussian Splatting (GS) methods, including 3DGS and 2DGS, have demonstrated exceptional performance in real-time novel view synthesis (NVS), emerging as a transformative technology in the fields of explicit rendering and computer graphics. However, GS-based methods still face challenges in rendering high-quality image details. Even when using high-quality training frameworks, their outputs often exhibit severe rendering artifacts, such as noise and blurriness. A reasonable approach is to perform post-processing to restore clear details. In this paper, we propose GaussianEnhancer, a general network-agnostic post-processor that employs a degradation-driven view blending method to improve the rendering quality of GS models while preserving the original network’s performance. Specifically, we design a degradation modeling method tailored to the GS-style and construct a large-scale training dataset to effectively simulate the native rendering artifacts of GS, enabling efficient training. In addition, we introduce a spatial information fusion framework, consisting of view fusion and depth modulation modules, which can blend highly correlated high-quality training images and leverage the depth information of the target image to complete the rendering details. Through our GaussianEnhancer, we are able to effectively eliminate the rendering artifacts of GS models and generate highly realistic synthetic views. Chen Zou 0007, Qingsen Ma, Jia Wang 0038, Ming Lu 0002, Shanghang Zhang, Zhaofeng He 0001 |
ICASSP | 6 |
| 2025 | Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMsabstractFood recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the food recognition task still lags behind traditional vision models. To bridge this gap, this paper proposes two methods to improve the food recognition capabilities of LVLMs: Retrieval-Augmented Recognition (RAR) and Domain-Adaptive Recognition (DAR). On the one hand, the training-free RAR utilizes a vision model to retrieve relevant image-category pairs from an image-category memory pre-built from the training set, thus incorporating the categorical information into the input of LVLMs to enhance food recognition performance. On the other hand, DAR employs a two-stage training process by first pre-training LVLMs on diverse food analysis tasks and then fine-tuning LVLMs using food recognition data. Extensive evaluations on two large-scale food recognition datasets demonstrate that both RAR and DAR improve the food recognition performance of LVLMs and, compred to RAR, DAR achieves a higher precision that outperforms traditional vision models. Dehua Ma, Zhenbo Xu, Tianshun Xing, Huijia Wu, Zhaofeng He 0001 |
ICASSP | 8 |
| 2025 | Eye Movements as Images: A Multimodal Framework for Eye Movements RepresentationabstractEye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimulus text remains challenging. This paper proposes a text-guided eye movement representation framework that introduces a novel perspective by converting raw eye movement sequences into line graph images and encoding them with a powerful pre-trained vision transformer. To address the disparities between eye movements and text, we guide their temporal alignment using human reading order and combine Canonical Correlation Analysis with Optimal Transport to fuse the two modalities. This approach not only significantly simplifies the design of specialized models but also has the potential to become a universal representation for eye movements. Experimental results on six different domain tasks show that the proposed method achieves state-of-the-art performance. We release the source code at https://github.com/wulalahalala/VLEM. Dongsen Zhang, Peipei Li 0002, Zekun Li 0001, Yiwei Ru, Huijia Wu, Zhaofeng He 0001 |
ICASSP | 6 |
| 2025 | LAMAR: LLM-Guided Adaptive Perceptual Modeling for Micro-Action RecognitionabstractWe present LAMAR, a novel framework for Micro-Action Recognition that addresses the challenges of identifying subtle, ephemeral human movements lasting less than one-third of a second. Our approach leverages large language models to estimate semantic complexity of micro-actions and dynamically configure a hierarchical Vision Transformer architecture accordingly. LAMAR introduces: (1) a principled complexity estimation module that quantifies recognition difficulty by analyzing subtlety, noise susceptibility, and intra-class ambiguity; and (2) an adaptive perception pipeline that dynamically adjusts spatiotemporal resolution and attention mechanisms based on estimated complexity. Experiments on the MA-52 benchmark demonstrate that LAMAR outperforms state-of-the-art methods by 7.20% in accuracy while maintaining computational efficiency, establishing a new paradigm for context-aware visual analysis that intelligently allocates resources based on task difficulty. Yiwei Ru, Leyuan Wang, Ma He, Zhaofeng He 0001, Zhenan Sun |
IJCB | 4 |
| 2025 | Adaptive Articulated Object Manipulation on the Fly with Foundation Model Reasoning and Part GroundingabstractArticulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While existing works have attempted cross-category generalization in adaptive articulated object manipulation, two major challenges persist: (1) the geometric diversity of real-world articulated objects complicates visual perception and understanding, and (2) variations in object functions and mechanisms hinder the development of a unified adaptive manipulation strategy. To address these challenges, we propose AdaRPG, a novel framework that leverages foundation models to extract object parts, which exhibit greater local geometric similarity than entire objects, thereby enhancing visual affordance generalization for functional primitive skills. To support this, we construct a part-level affordance annotation dataset to train the affordance model. Additionally, AdaRPG utilizes the common knowledge embedded in foundation models to reason about complex mechanisms and generate high-level control codes that invoke primitive skill functions based on part affordance inference. Simulation and real-world experiments demonstrate AdaRPG's strong generalization ability across novel articulated object categories. Yuanfei Wang, Ruihai Wu, Kunqi Xu, Yu Li 0022, Liuyu Xiang, Hao Dong 0003, Zhaofeng He 0001 |
ICCV | 8 |
| 2025 | MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMsabstractCurrent multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods. Xuannan Liu, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He 0001 |
ICLR | 9 |
| 2025 | AdaManip: Adaptive Articulated Object Manipulation Environments and Policy LearningabstractArticulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios.
Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through complex relative motions. For example, a safe consists of a door, a handle, and a lock, where the door can only be opened when the latch is unlocked. The internal structure, such as the state of a lock or joint angle constraints, cannot be directly observed from visual observation. Consequently, successful manipulation of these objects requires adaptive adjustment based on trial and error rather than a one-time visual inference. However, previous datasets and simulation environments for articulated objects have primarily focused on simple manipulation mechanisms where the complete manipulation process can be inferred from the object's appearance. To enhance the diversity and complexity of adaptive manipulation mechanisms, we build a novel articulated object manipulation environment and equip it with 9 categories of objects. Based on the environment and objects, we further propose an adaptive demonstration collection and 3D visual diffusion-based imitation learning pipeline that learns the adaptive manipulation policy. The effectiveness of our designs and proposed method is validated through both simulation and real-world experiments. Yuanfei Wang, Ruihai Wu, Yu Li 0022, Yan Shen 0035, Mingdong Wu, Zhaofeng He 0001, Yizhou Wang 0001, Hao Dong 0003 |
ICLR | 7 |
| 2025 | Beyond Macro-Actions: A Bio-Inspired Framework for Fine-Grained Micro-Action RecognitionabstractHuman Action Recognition (HAR) is pivotal in advancing applications from surveillance to healthcare, but predominantly focuses on easily observable, macro-level actions such as running or jumping. Micro-Action Recognition (MAR), however, delves into the subtle, often involuntary motions like postural shifts, brief gestures, or faint facial twitches, which are critical for revealing underlying emotional states, intentions, or stress levels. MAR presents unique challenges due to the ephemeral nature of micro-actions, their fine-grained inter-class similarities, and significant class imbalance. To overcome these obstacles, our approach draws inspiration from the hierarchical and context-sensitive capabilities of the human visual system. We propose a biologically motivated, multi-pathway framework that cohesively integrates global context analysis, rapid temporal scanning, and meticulous fine-grained scrutiny. This framework combines skeletal dynamics, subtle motion amplitude cues, and RGB-based contextual features to enable a comprehensive and robust recognition of micro-actions, even in unconstrained environments. Our experimental results on the MA-52 dataset demonstrate leading performance, significantly advancing MAR research and broadening the spectrum of applications that require a nuanced understanding of human behavior. Yiwei Ru, Churan Yu, Dongsen Zhang, Mupei Li, Yongji Liu, Zhaofeng He 0001 |
ICME | 6 |
| 2025 | FoodWeight1.4M: A Large-scale Multi-modal Dataset for Weight EstimationabstractLarge vision language models (VLMs) excel in visual tasks but struggle with weight estimation, hindering 3D perception and embodied intelligence. To address the lack of large-scale weight datasets, we present FoodWeight1.4M, derived from real-world supermarket scenarios. It contains 1.4 million high-quality images across 1,550 food categories, with weights precisely measured and rigorously filtered, making it the first large-scale weight estimation dataset. The weight estimation performance of current VLMs were tested and found to be unsatisfactory, which can be significantly improved by instruction tuning using Food-Weight1.4M. Moreover, we propose two strategies, Category-Guided and Reference Calibration, to enhance weight estimation without fine-tuning. Experiments confirm their effectiveness in improving multi-modal weight perception. Furthermore, experimental results show that pre-training on FoodWeight1.4M can benefit other food analysis tasks. Our dataset will be publicly available soon. Zhenbo Xu, Dehua Ma, Liuyu Xiang, Huijia Wu, Zhaofeng He 0001 |
ICME | 7 |
| 2025 | RainbowArena: A Multi-Agent Toolkit for Reinforcement Learning and Large Language Models in Competitive Tabletop Games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
AAMAS | 8 |
| 2025 | GraphDiffusion: A Graph-conditioned Diffusion Model for Chip PlacementabstractPlacement is a crucial and time-intensive step in the modern Electronic Design Automation (EDA) chip design process, involving the allocation of millions of modules on a chip canvas. Previous studies have demonstrated the efficacy of machine learning methods, especially reinforcement learning (RL) in chip placement. However, existing RL-based methods still face challenges, including lengthy placement times due to placing only a single module at each step, and a lack of generalization capability. In response to these challenges, we propose a Graph Convolutional Network (GCN)-based diffusion model called GraphDiffusion. Our model presents a novel approach for chip placement by representing the positioning of movable nodes as a conditional denoising diffusion process. The GCN is utilized to extract feature information from the circuit, which then serves as an embedding to guide the diffusion model in generating the chip layout. We exploit the diffusion model’s capabilities to improve layout quality and generalizability. We conduct extensive experiments on the ISPD benchmark. The results demonstrate that our approach outperforms advanced placement methods while enhancing the transferability of the layout model. Upon training with expert data, GraphDiffusion can interpret the input circuit netlists and generate state-of-the-art chip layouts. Siyuan Fang, Liuyu Xiang, Wei Li 0032, Zhaofeng He 0001 |
ISCAS | 4 |
| 2025 | Learning Uniformly Distributed Embedding Clusters of Stylistic Skills for Physically Simulated Characters
Nian Liu 0003, Zi Wang 0014, Tengyu Liu, Hongzhao Xie, Xinyi Tong 0001, Libin Liu 0002, Yaodong Yang 0001, Zhaofeng He 0001 |
ACM Multimedia | 9 |
| 2025 | RecipeRAG: Advancing Recipe Generation with Reinforced Retrieval Augmented GenerationabstractGenerating accurate recipes from dish images is a challenging task that requires a deep understanding of food categories, ingredient combinations, cooking methods, and context. Current works mainly rely on the two-stage training method or supervised fine-tuning of vision-language models (VLMs). Two-stage models typically first predict ingredients from images and then generate recipes based on both ingredients and images. However, accumulated errors in ingredient prediction often lead to inaccurate recipes. Fine-tuning VLMs only fit the statistical patterns of the training data, lacking deep reasoning capabilities, which leads to severe hallucinations in the generated recipes. In this paper, we introduce a novel reinforced retrieval-augmented generation framework named RecipeRAG for recipe generation, and compare the supervised fine-tuning (SFT) paradigm and the reinforcement fine-tuning (RFT) paradigm. To effectively retrieve recipes relevant to the query image, we improve CLIP to obtain IR-CLIP as both our retriever and re-ranker by integrating metric learning and contrastive learning. The retrieved recipes are then used to enhance the generated results, improving accuracy and reducing hallucinations. However, the SFT VLM often fails to judge the quality of the retrieved recipe information and perform the complex recipe generation. Therefore, we furthermore investigate the two-phase RFT training framework. Firstly, the cold-start phase uses generated Chain-of-Thought (CoT) data for SFT to activate the reasoning capabilities of VLMs. Then, the reinforcement learning phase utilizes Group Relative Policy Optimization (GRPO) to generate multiple reasoning-answer pairs, further enhancing the generalization ability of VLMs in recipe generation tasks. Extensive evaluations on the large-scale Recipe1M dataset demonstrate that RecipeRAG outperforms all previous methods in recipe generation and exhibits strong generalization ability under the RL paradigm. Zhenbo Xu, Dehua Ma, Fei Liu 0008, Gong Huang, Zhaofeng He 0001 |
ACM Multimedia | 7 |
| 2025 | A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource SettingsabstractLarge Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thought, but this often leads to performance degradation. To address this problem, we introduce A*-Thought, an efficient tree search-based unified framework designed to identify and isolate the most essential thoughts from the extensive reasoning chains produced by these models. It formulates the reasoning process of LRMs as a search tree, where each node represents a reasoning span in the giant reasoning space. By combining the A* search algorithm with a cost function specific to the reasoning path, it can efficiently compress the chain of thought and determine a reasoning path with high information density and low cost. In addition, we also propose a bidirectional importance estimation mechanism, which further refines this search process and enhances its efficiency beyond uniform sampling. Extensive experiments on several advanced math tasks show that A*-Thought effectively balances performance and efficiency over a huge search space. Specifically, A*-Thought can improve the performance of QwQ-32B by 2.39$\times$ with low-budget and reduce the length of the output token by nearly 50\% with high-budget. The proposed method is also compatible with several other LRMs, demonstrating its generalization capability. The code can be accessed at: https://github.com/AI9Stars/AStar-Thought. Xiaoang Xu, Shuo Wang 0013, Zhenghao Liu 0001, Huijia Wu, Peipei Li 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Zhaofeng He 0001 |
NeurIPS | 9 |
| 2025 | CLANet: A Denoising-Driven Framework for Robust mmWave Radar Vital Sign Monitoring
Yiwei Ru, Yongji Liu, Mupei Li, Dongsen Zhang, Zhaofeng He 0001, Zhenan Sun |
PRCV (3) | 5 |
| 2025 | Cost-effective and real-time landslide monitoring method based on ultra-wideband using ultra-wideband transformer neural network
Yu Si, Zhaofeng He 0001, Haiqing Zheng |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Can large language models independently complete tasks? A dynamic evaluation framework for multi-turn task planning and completion
Junlin Cui, Huijia Wu, Liuyu Xiang, Xiangang Li, Yaodong Yang 0001, Zhaofeng He 0001 |
Neurocomputing | 9 |
| 2025 | Generalizable agent modeling for agent collaboration-competition adaptation with multi-retrieval and dynamic generation
Yonggang Jin, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018, Liuyu Xiang, Junge Zhang, Zhaofeng He 0001 |
Neurocomputing | 9 |
| 2025 | SecureXGB: A Secure and Efficient Multi-party Protocol for Vertical Federated XGBoostabstractExtreme Gradient Boosting (XGBoost) demonstrates excellent performance in practice and is widely used in both industry and academic research. This extensive application has led to a growing interest in employing multi-party data to develop more robust XGBoost models. In response to increasing concerns about privacy leakage, secure vertical federated XGBoost is proposed. It employs secure multi-party computation techniques, such as secret sharing (SS), to allow multiple parties holding vertically partitioned data, i.e., disjoint features on the same samples, to collaborate in constructing an XGBoost model. However, the running efficiency is the primary obstacle to the practical application of existing protocols, especially in multi-party settings. The reason is that these protocols not only require the execution of data-oblivious computations to protect intermediate results, leading to high computational complexity, but also involve a large number of SS-based non-linear operations with high overheads, e.g., division operations in gain score calculation and comparison operations in best split selection. To this end, we present a secure and efficient multi-party protocol for vertical federated XGBoost, called SecureXGB, which can perform the collaborative training of an XGBoost model in an SS-friendly manner. In SecureXGB, we first propose a parallelizable multi-party permutation method, which can secretly and efficiently permute all samples before model training to reduce the reliance on data-oblivious computations. Then, we design a linear gain score that can be evaluated without involving division operations and has equivalent utility to the original gain score. Finally, we develop a synchronous best split selection method to secretly identify the best split with the maximum gain score using a minimal number of comparison operations. Experimental results demonstrate that SecureXGB can achieve better training efficiency than state-of-the-art protocols without the loss of model accuracy. Zongda Han, Xiang Cheng 0003, Wenhong Zhao, Jiaxin Fu, Zhaofeng He 0001, Sen Su |
Proc. ACM Manag. Data | 5 |
| 2025 | JARVIS-1: Open-World Multi-Task Agents With Memory-Augmented Multimodal Language ModelsabstractAchieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could potentially be infinite and lack the capability to progressively enhance task completion as game time progresses. We introduceJARVIS-1, an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, all within the popular yet challenging open-world Minecraft universe. Specifically, we developJARVIS-1 on top of pre-trained multimodal language models, which map visual observations and textual instructions to plans. The plans will be ultimately dispatched to the goal-conditioned controllers. We outfitJARVIS-1 with a multimodal memory, which facilitates planning using both pre-trained knowledge and its actual game survival experiences.JARVIS-1 is the existing most general agent in Minecraft, capable of completing over 200 different tasks using control and observation space similar to humans. These tasks range from short-horizon tasks, e.g., “chopping trees” to long-horizon ones, e.g., “obtaining a diamond pickaxe”.JARVIS-1 performs exceptionally well in short-horizon tasks, achieving nearly perfect performance. In the classic long-term task ofObtainDiamondPickaxe,JARVIS-1 surpasses the reliability of current state-of-the-art agents by 5 times and can successfully complete longer-horizon and more challenging tasks. Furthermore, we show thatJARVIS-1 is able toself-improvefollowing a life-long learning paradigm thanks to multimodal memory, sparking a more general intelligence and improved autonomy. Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang 0007, Haowei Lin, Zhaofeng He 0001, Zilong Zheng, Yaodong Yang 0001, Xiaojian Ma 0001, Yitao Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | FreqFormer: Frequency-Enhanced Face Super-Resolution via Dual-Synergy LearningabstractIn this paper, we propose FreqFormer, a frequency-enhanced framework for face super-resolution (FSR) that synergizes spectral decomposition and dynamic feature modulation to address Transformers' inherent low-frequency bias. Unlike existing methods, FreqFormer leverages Empirical Mode Decomposition (EMD) to decompose high-resolution (HR) images into hierarchical Intrinsic Mode Functions (IMFs), providing explicit frequency anchors for progressive high-frequency recovery. Simultaneously, a lightweight prompt module dynamically injects degradation-aware textures into Transformer blocks through learnable feature interaction, compensating for real-world distortions. This work establishes a novel framework that bridges spectral fidelity and adaptive learning, significantly advancing FSR toward high-frequency-accurate restoration. Experiments on CelebA and Helen datasets demonstrate FreqFormer's superiority, which achieves 22.42 dB PSNR on$16\times$SR ($8\times 8 \rightarrow 128\times 128$) with higher high-frequency energy preservation than state-of-the-art methods. The framework's parameter efficiency and real-time performance enable practical deployment in security and biometric systems. Jia Wang 0038, Shuhan Xia, Chen Zou 0007, Guobin Wu 0001, Zhaofeng He 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | DrlGoFPGA: FPGA Global Placement Considering Input-Output Buffer Based on Deep Reinforcement Learning and Gradient OptimizationabstractThe placement of the input-output buffer (IOBUF) can impact the performance and power consumption of the FPGA. The existing global placement (GP) methods lack consideration for IOBUF, resulting in a decrease in placement and routing quality. To address this issue, we propose a GP framework, DrlGoFPGA, which combines IOBUF placement based on deep reinforcement learning (DRL) with other instances placement based on gradient optimization (GO). A policy network structure with multi-action sampling is designed to accelerate the running speed of DRL, and a parallelizable reward function is designed to optimize each IOBUF placement action and avoid sparse reward problems. Then, an IOBUF line-network relationship (ILNR) graph creation method is designed to improve the agent’s ability to explore optimal solutions, and the graph features of ILNR by capturing them through a graph neural network embedded in the convolutional neural network. Finally, an IOBUF placement legalization method is designed to ensure that the IOBUF position meets the FPGA architecture. The experimental results show that compared with the state-of-the-art placement tools based on GO, DrlGoFPGA can improve GP speed by 13.2%-7×, half-perimeter wirelength by 0.2%-2.6%, and wirelength by 0.2%-1.5% and the IOBUF placement model has good generalization. Jianwang Zhai, Liuyu Xiang, Zixi Huang, Zhaofeng He 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | Mini Honor of Kings: A Lightweight Environment for Multiagent Reinforcement LearningabstractGames are widely used as research environments for multiagent reinforcement learning (MARL), but they pose three significant challenges: limited customization, high computational demands, and oversimplification. To address these issues, we introduce the first publicly available map editor for the popular mobile gameHonor of Kingsand design a lightweight environment,Mini Honor of Kings(Mini HoK), for researchers to conduct experiments. Mini HoK is highly efficient, allowing experiments to be run on personal PCs or laptops while still presenting sufficient challenges for existing MARL algorithms. We have tested our environment on common MARL algorithms and demonstrated that these algorithms have yet to surpass the performance of rule based policies, indicating that current MARL methods are not able to solve this environment. This facilitates the dissemination and advancement of MARL methods within the research community. In addition, we hope that more researchers will leverage theHonor of Kingsmap editor to develop innovative and scientifically valuable new maps. Lin Liu 0016, Jian Zhao 0018, Zhengtao Cao, Youpeng Zhao 0001, Zhenbin Ye, Zhaofeng He 0001, Houqiang Li, Xia Lin, Lanxiao Huang |
IEEE Trans. Games | 9 |
| 2025 | Toward Real-World Remote Sensing Image Super-Resolution: A New Benchmark and an Efficient ModelabstractSuper-resolution (SR) is a fundamental and crucial task in remote sensing. It can improve low-resolution (LR) remote sensing images and has potential benefits for downstream tasks such as remote sensing object detection and recognition. Existing remote sensing image SR (RSISR) methods are trained on simulated paired datasets, in which LR images are obtained by a simple and uniform (i.e., bicubic) degradation from corresponding high-resolution (HR) images. However, since this simulated degradation usually deviates from the real degradation, the performance of the trained model is limited when applied to real scenarios. To address this issue, we construct a novel real-world RSISR (RRSISR) dataset to model the real-world degradation, which exploits the imaging characteristics of the spectral camera to capture paired LR-HR images of the same scene. To ensure the precise alignment of the paired images, algorithms such as image registration and geometric correction are utilized. In addition, considering the vast amount of data involved in the RSISR task and its requirement for higher efficiency, we divide the image into patches with different restoration difficulties and propose a reference table-based patch exiting (RPE) method to efficiently reduce the computation of SR. Specifically, this method incorporates a predictor to estimate the performance of the current layer and a lookup table to decide whether to exit. Extensive experiments show that models trained on the proposed RRSISR dataset produce more realistic images than models with simulated datasets and generalize well to other satellites. We also demonstrate the efficiency of our RPE. Jia Wang 0038, Liuyu Xiang, Jiaochong Xu, Peipei Li 0002, Qizhi Xu, Zhaofeng He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Exploring Near-Infrared Iris Image Sequences for High Throughput Iris RecognitionabstractHigh throughput is demanding in real-world iris recognition applications. The challenges mainly originate from the variability in image quality under high-throughput capture conditions. Most of the degraded images are typically filtered out by traditional iris systems through Image Quality Assessment (IQA) module, adversely affecting efficiency and leading to low throughput and poor user experience. Therefore, a better and practical solution is to make the utmost of degraded iris images. In order to investigate the key problems of high-throughput iris recognition, we collect a novel iris sequence dataset under Near-infrared (NIR) illumination. This dataset is specifically constructed for high-throughput evaluation, which faithfully simulates the process of iris sequence acquisition in real-world iris systems. Comprehensive evaluations were conducted to figure out the deficiencies of current iris recognition algorithms. To this end, a testing methodology along with specific evaluation metrics is proposed. It is capable of assessing the throughput performance, e.g., the newly proposed Frame Consumption per Match (FCM). Through performance analysis, several insights were gathered to guide potential directions for developing high-throughput iris recognition algorithms. Furthermore, we consider to leverage iris sequence features for better throughput performance. Continuity sequence criteria and cumulative sequence feature strategy are proposed to enhance the throughput performance of existing algorithms with minimal cost. In summary, this work provides valuable data and rational insights for high-throughput iris recognition studies. The datasets and evaluation toolkit are publicly available on our website1. Mupei Li, Yunlong Wang 0003, Kunbo Zhang, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | FDNet: A Frequency-Aware Decomposition Network for Robust Face Super-Resolution Against Adversarial AttacksabstractFace super-resolution (FSR) is a crucial step in the face analysis pipeline, achieving remarkable progress by applying deep neural networks (DNNs). However, DNN-based FSR models are not robust enough and may suffer significant performance degradation due to subtle adversarial perturbations. In addition, the high-frequency details of images restored by existing models are insufficient, especially at large upsampling factors. In this paper, we propose a frequency-aware decomposition network (FD-Net) for robust face super-resolution, which aims to defend against adversarial attacks and obtain face images with fidelity. Observing that the noise introduced by adversarial attacks is often intricately mixed with the high-frequency information of the input image, we decompose and process the features of different frequencies separately to eliminate harmful perturbations and enhance high-frequency information. Specifically, by leveraging the frequency-aware capability of empirical mode decomposition (EMD), we propose an EMD-based multi-branch structure. The framework implicitly compels different branches to adaptively extract features from distinct frequency bands, limiting the adversarial noise into decoupled components restricted to specific branches. It also improves the recovery of high-frequency information, which is conducive to producing more credible results. Furthermore, we introduce a high-frequency noise suppressor capable of randomly eliminating imperceptible noise in the high-frequency components. Quantitative and qualitative results demonstrate the superior robustness of our proposed method against adversarial attacks, showing better fidelity in image reconstruction compared to state-of-the-art FSR methods, especially for upscaling factors of 8 and 16. Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Rui Wang 0124, Zhaofeng He 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Distributed Policy Space Response Oracles in Two-Player Zero-Sum GamesabstractPolicy space response oracle (PSRO) is a population-based algorithm that can be used to solve two-player zero-sum games. In the PSRO solution framework, optimizing policy diversity is crucial for addressing nontransitive game problems, helping the agent population avoid exploitation by unfamiliar opponents. In addition, while deep reinforcement learning is highly effective in solving complex game environments, its integration with PSRO remains fragmented and lacking in effective coordination. In this study, we propose distributed PSRO to efficiently solve complex game scenarios. To enhance diversity while managing optimization costs, we introduce TOP-K truncation, which prioritizes high-quality opponents and limits the size of the policy pool during sampling. This approach not only reduces interference from less effective strategies but also ensures computational efficiency by seamlessly integrating with our distributed training framework. We also design the distributed training framework to incorporate diversity estimation directly into the sampling process, achieving diversity optimization without incurring additional computational overhead. Furthermore, we introduce the opponent first (OF) method, which enhances decision-making by leveraging opponent information during interaction sampling. We perform experimental validation using a nontransitive mixture model and AlphaStar888 to confirm the effectiveness of the TOP-K truncation approach. Finally, we demonstrate the feasibility and efficiency of the distributed training framework and the OF approach in a Google Research Football 11 versus 11 scenario. Hongsong Tang, Yingzhuo Liu, Letian Ni, Liuyu Xiang, Yaodong Yang 0001, Ke Bi, Zhaofeng He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent EnvironmentsabstractJunzhe Chen, Xuming Hu, Shuodi Liu, Shiyu Huang, Wei-Wei Tu, Zhaofeng He, Lijie Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junzhe Chen 0001, Xuming Hu, Shuodi Liu, Shiyu Huang 0001, Wei-Wei Tu, Zhaofeng He 0001, Lijie Wen 0001 |
ACL (1) | 6 |
| 2024 | HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsabstractThe Mixture of Experts (MoE) for language models has been proven effective in augmenting the capacity of models by dynamically routing each input token to a specific subset of experts for processing.Despite the success, most existing methods face a challenge for balance between sparsity and the availability of expert knowledge: enhancing performance through increased use of expert knowledge often results in diminishing sparsity during expert selection.To mitigate this contradiction, we propose HyperMoE, a novel MoE framework built upon Hypernetworks.This framework integrates the computational processes of MoE with the concept of knowledge transferring in multi-task learning.Specific modules generated based on the information of unselected experts serve as supplementary information, which allows the knowledge of experts not selected to be used while maintaining selection sparsity.Our comprehensive empirical evaluations across multiple datasets and backbones establish that HyperMoE significantly outperforms existing MoE methods under identical conditions concerning the number of experts. Zihan Qiu, Huijia Wu, Zhaofeng He 0001, Jie Fu 0001 |
ACL (1) | 5 |
| 2024 | MORE-3S: Multimodal-based Offline Reinforcement Learning with Shared Semantic SpacesabstractDrawing upon the intuition that aligning different modalities to the same semantic embedding space would allow models to understand states and actions more easily, we propose a new perspective to the offline reinforcement learning (RL) challenge. More concretely, we transform it into a supervised learning task by integrating multimodal and pre-trained language models. Our approach incorporates state information derived from images and action-related data obtained from text, thereby bolstering RL training performance and promoting long-term strategic thinking. We emphasize the contextual understanding of language and demonstrate how decision-making in RL can benefit from aligning states’ and actions’ representation with languages’ representation. Our method significantly outperforms current baselines as evidenced by evaluations conducted on Atari and OpenAI Gym environments. This contributes to advancing offline RL performance and efficiency while providing a novel perspective on offline RL. Tianyu Zheng, Ge Zhang 0009, Xingwei Qu, Ming Kuang, Wenhao Huang 0001, Zhaofeng He 0001 |
LREC/COLING | 6 |
| 2024 | INSTASTYLE: Inversion Noise of a Stylized Image is Secretly a Style Adviser
Xing Cui, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Xuannan Liu, Zhaofeng He 0001 |
ECCV (51) | 6 |
| 2024 | Enhancing Short-and Long-Term Sea Surface Temperature Forecasting with a Static and Dynamic Learnable Personalized Graph Convolution NetworkabstractSea surface temperature (SST) plays an important role in our Earth’s atmosphere, wielding significant influence over both local and global climates and profoundly impacting ecosystems. However, this task presents unique challenges due to the inherent complexity and uncertainty within ocean systems. Recently, deep learning techniques, such as Graph Neural Networks (GNNs), have been employed to address SST forecasting. These methods often grapple with substantial limitations in capturing dynamic spatiotemporal relationships between signals. To tackle this issue, this paper introduces the Static and Dynamic Learnable Personalized Graph Convolution Network (SD-LPGC). This innovative approach comprises two distinct graph learning layers designed to model both stable long-term and short-term evolutionary patterns within multivariate SST signals. Additionally, a learnable personalized convolution layer is integrated to fuse these insights. Our experiments, conducted on real SST datasets, highlight the state-of-the-art performance of the proposed SD-LPGC approach in the realm of SST forecasting, demonstrating its potential to revolutionize this critical task. Zhaofeng He 0001 |
ICASSP | 2 |
| 2024 | I3FDM: IRIS Inpainting Via Inverse Fusion of Diffusion ModelsabstractIris images captured in real-world scenarios are often occluded, which leads to a significant degradation for the iris recognition system. Therefore, it is necessary to propose an effective iris inpainting method. While generative adversarial network (GAN)-based image inpainting methods have shown promise, they often suffer from issues such as mode collapse and training instability. Recently, denoising diffusion probabilistic model (DDPM) has surpassed GAN in terms of image quality while maintaining stable training. Combining DDPM and the characteristics of iris image, I3FDM (Iris Inpainting via Inverse Fusion of Diffusion Models), a method that iteratively modifies intermediate variables in the generation process based on a given occluded image. Since these modifications introduce semantic differences, we introduce an inverse fusion module to enhance the performance of iris inpainting. I3FDM enables the processing of various types of occluded images using an unconditional DDPM without the need for additional learning. Extensive experimental results qualitatively and quantitatively demonstrate that our method produces iris images with richer texture information and improves the performance of iris recognition. Chenyang Li 0011, Peipei Li 0002, Zhaofeng He 0001 |
ICASSP | 4 |
| 2024 | Exploring 3D-aware Lifespan Face Aging via Disentangled Shape-Texture RepresentationsabstractExisting face aging methods often focus on modeling either texture aging or using an entangled shape-texture representation to achieve face aging. However, shape and texture are two distinct factors that mutually affect the human face aging process. In this paper, we propose 3D-STD, a novel 3D-aware Shape-Texture Disentangled face aging network that explicitly disentangles the facial image into shape and texture representations using 3D face reconstruction. Additionally, to facilitate high-fidelity texture synthesis, we propose a novel texture generation method based on Empirical Mode Decomposition (EMD). Extensive qualitative and quantitative experiments show that our method achieves state-of-the-art performance in terms of shape and texture transformation. Moreover, our method supports producing plausible 3D face aging results, which is rarely accomplished by current methods. Qianrui Teng, Rui Wang 0124, Xing Cui, Peipei Li 0002, Zhaofeng He 0001 |
ICME | 5 |
| 2024 | MQE: Unleashing the Power of Interaction with Multi-agent Quadruped EnvironmentabstractThe advent of deep reinforcement learning (DRL) has significantly advanced the field of robotics, particularly in the control and coordination of quadruped robots. However, the complexity of real-world tasks often necessitates the deployment of multi-robot systems capable of sophisticated interaction and collaboration. To address this need, we introduce the Multi-agent Quadruped Environment (MQE), a novel platform designed to facilitate the development and evaluation of multi-agent reinforcement learning (MARL) algorithms in realistic and dynamic scenarios. MQE emphasizes complex interactions between robots and objects, hierarchical policy structures, and challenging evaluation scenarios that reflect real-world applications. We present a series of collaborative and competitive tasks within MQE, ranging from simple coordination to complex adversarial interactions, and benchmark state-of-the-art MARL algorithms. Our findings indicate that hierarchical reinforcement learning can simplify task learning, but also highlight the need for advanced algorithms capable of handling the intricate dynamics of multi-agent interactions. MQE serves as a stepping stone towards bridging the gap between simulation and practical deployment, offering a rich environment for future research in multi-agent systems and robot learning. For open-sourced code and more details of MQE, please refer to https://ziyanx02.github.io/multiagent-quadruped-environment/. Ziyan Xiong, Shiyu Huang 0001, Wei-Wei Tu, Zhaofeng He 0001 |
IROS | 5 |
| 2024 | FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsabstractThe massive generation of multimodal fake news involving both text and images exhibits substantial distribution discrepancies, prompting the need for generalized detectors. However, the insulated nature of training restricts the capability of classical detectors to obtain open-world facts. While Large Vision-Language Models (LVLMs) have encoded rich world knowledge, they are not inherently tailored for combating fake news and struggle to comprehend local forgery details. In this paper, we propose FKA-Owl, a novel framework that leverages forgery-specific knowledge to augment LVLMs, enabling them to reason about manipulations effectively. The augmented forgery-specific knowledge includes semantic correlation between text and images, and artifact trace in image manipulation. To inject these two kinds of knowledge into the LVLM, we design two specialized modules to establish their representations, respectively. The encoded knowledge embeddings are then incorporated into LVLMs. Extensive experiments on the public benchmark demonstrate that FKA-Owl achieves superior cross-domain performance compared to previous methods. Code is publicly available at https://liuxuannan.github.io/FKA_Owl.github.io/. Xuannan Liu, Peipei Li 0002, Huaibo Huang, Zekun Li 0001, Xing Cui, Lixiong Qin, Weihong Deng, Zhaofeng He 0001 |
ACM Multimedia | 9 |
| 2024 | Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerabstractFlexible and accurate drag-based editing is a challenging task that has recently garnered significant attention. Current methods typically model this problem as automatically learning "how to drag" through point dragging and often produce one deterministic estimation, which presents two key limitations: 1) Overlooking the inherently ill-posed nature of drag-based editing, where multiple results may correspond to a given input, as illustrated in Fig.1; 2) Ignoring the constraint of image quality, which may lead to unexpected distortion.
To alleviate this, we propose LucidDrag, which shifts the focus from "how to drag" to "what-then-how" paradigm. LucidDrag comprises an intention reasoner and a collaborative guidance sampling mechanism. The former infers several optimal editing strategies, identifying what content and what semantic direction to be edited. Based on the former, the latter addresses "how to drag" by collaboratively integrating existing editing guidance with the newly proposed semantic guidance and quality guidance.
Specifically, semantic guidance is derived by establishing a semantic editing direction based on reasoned intentions, while quality guidance is achieved through classifier guidance using an image fidelity discriminator.
Both qualitative and quantitative comparisons demonstrate the superiority of LucidDrag over previous methods. Xing Cui, Peipei Li 0002, Zekun Li 0001, Xuannan Liu, Yueying Zou, Zhaofeng He 0001 |
NeurIPS | 6 |
| 2024 | SafeLLMs: A Benchmark for Secure Bilingual Evaluation of Large Language Models
Wenhan Liang, Huijia Wu, Yuhu Shang, Zhaofeng He 0001 |
NLPCC (2) | 5 |
| 2024 | Locally differentially private graph learning on decentralized social graph
Guanhong Zhang, Xiang Cheng 0003, Jiaan Pan, Zhaofeng He 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Leveraging Joint-Action Embedding in Multiagent Reinforcement Learning for Cooperative GamesabstractState-of-the-art multi-agent policy gradient (MAPG) methods have demonstrated convincing capability in many cooperative games. However, the exponentially growing joint-action space severely challenges the critic's value evaluation and hinders performance of MAPG methods. To address this issue, we augment Central-Q policy gradient with a joint-action embedding function and propose Mutual-information Maximization MAPG (M3APG). The joint-action embedding function makes joint-actions contain information of state transitions, which will improve the critic's generalization over the joint-action space by allowing it to infer joint-actions' outcomes. We theoretically prove that with a fixed joint-action embedding function, the convergence of M3APG is guaranteed. Experiment results on the StarCraft Multi-Agent Challenge (SMAC) demonstrate that M3APG gives evaluation results with better accuracy and outperform other MAPG basic models across various maps of multiple difficulty levels. We empirically show that our joint-action embedding model can be extended to value-based multi-agent reinforcement learning methods and state-of-the-art MAPG methods. Finally, we run ablation study to show that the usage of mutual information in our method is necessary and effective. Xingzhou Lou, Junge Zhang, Yali Du 0001, Chao Yu 0004, Zhaofeng He 0001, Kaiqi Huang |
IEEE Trans. Games | 5 |
| 2024 | Multi-Passage Machine Reading Comprehension Through Multi-Task Learning and Dual VerificationabstractMulti-passage machine reading comprehension (MRC) aims to answer a question by multiple passages. Existing multi-passage MRC approaches have shown that employing passages with and without golden answers (i.e., labeled and unlabeled passages) for model training can improve prediction accuracy. However, when using the unlabeled passages, they either incur the wrong labeling problem or treat the labeled and unlabeled passages equally. In addition, they ignore the original passage information to verify the correctness of the answer. In this paper, we present MLDV-MRC, a novel approach for multi-passage MRC viaMulti-taskLearning andDualVerification. MLDV-MRC adopts the extract-then-select framework, where an extractor is first used to predict answer candidates, then a selector is used to choose the final answer. For the extractor, we adopt multi-task learning with generative adversarial training to train it by using both labeled and unlabeled passages. To train the extractor by backpropagation, we propose a hybrid method which combines boundary-based and content-based extracting methods to produce the answer candidate set and its representation. For the selector, we propose to leverage both the information from answer candidates and original passages to verify the final answer. In particular, we propose a global-local memory-augmented neural network to build the representations of original passages, which fuses the passage-level information and word-level information. The experimental results on three open-domain QA datasets confirm the effectiveness of our approach. Xingyi Li 0006, Xiang Cheng 0003, Qiyu Ren, Zhaofeng He 0001, Sen Su |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueabstractSingle-turn Multi-turn Input Lipstick Pale skin Smiling Black Hair Figure 2: Comparison of previous repeated singleturn editing approaches and our proposed multiturn editing approach.The cascaded errors in the single-turn approach lead to unintended changes in gender and eye makeup. Xing Cui, Zekun Li 0001, Yibo Hu 0001, Hailin Shi, Chunshui Cao, Zhaofeng He 0001 |
EMNLP | 7 |
| 2023 | Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningabstractParameter-efficient fine-tuning (PEFT) has shown its effectiveness in adapting the pretrained language models to downstream tasks while only updating a small number of parameters.Despite the success, most existing methods independently adapt to each task without considering knowledge transfer between tasks and are limited to low-data regimes.To overcome this issue, we propose Prototype-based HyperAdapter (PHA), a novel framework built on the adapter-tuning and hypernetwork.It introduces an instance-dense retriever and a prototypical hypernetwork to generate the conditional modules in a sample-efficient manner.This leads to comparable performance improvements against existing PEFT methods on multi-task learning and few-shot transfer learning.More importantly, when the available data size gets smaller, our method outperforms other strong baselines by a large margin.Based on our extensive empirical experiments across various datasets, we demonstrate that PHA strikes a better trade-off between trainable parameters, accuracy on stream tasks, and sample efficiency.Our code is publicly available at https://github.com/Bumble666/PHA Jie Fu 0001, Zhaofeng He 0001 |
EMNLP | 3 |
| 2023 | Sclera-TransFuse: Fusing Swin Transformer and CNN for Accurate Sclera SegmentationabstractSclera segmentation is a crucial step in sclera recognition, which has been greatly advanced by Convolutional Neural Networks (CNNs). However, when dealing with non-ideal eye images, many existing CNN-based approaches are still prone to failure. One major reason is that due to the limited range of receptive fields, CNNs are difficult to effectively model global semantic relevance and thus robustly resist noise interference. To solve this problem, this paper proposes a novel two-stream hybrid model, named Sclera-TransFuse, to integrate classical ResNet-34 and recently emerging Swin Transformer encoders. Specially, the self-attentive Swin Transformer has shown a strong ability in capturing long-range spatial dependencies and has a hierarchical structure similar to CNNs. The dual encoders firstly extract coarse- and fine-grained feature representations at hierarchical stages, separately. Then a novel Cross-Domain Fusion (CDF) module based on information interaction and self-attention mechanism is introduced to efficiently fuse the multi-scale features extracted from dual encoders. Finally, the fused features are progressively upsampled and aggregated to predict the sclera masks in the decoder meanwhile deep supervision strategies are employed to learn intermediate feature representations better and faster. Experimental results show that Sclera-TransFuse achieves state-of-the-art performance on various sclera segmentation benchmarks. Additionally, a UBIRIS.v2 subset of 683 eye images with manually labeled sclera masks, and our codes are publicly available to the community through https://github.com/Ihqqq/Sclera-TransFuse. Caiyong Wang, Guangzhe Zhao, Zhaofeng He 0001, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2023 | Pluralistic Aging Diffusion AutoencoderabstractFace aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Pluralistic Aging Diffusion Autoencoder (PADA) to enhance the diversity of aging patterns. First, we employ diffusion models to generate diverse low-level aging details via a sequential denoising reverse process. Second, we present Probabilistic Aging Embedding (PAE) to capture diverse high-level aging patterns, which represents age information as probabilistic distributions in the common CLIP latent space. A text-guided KL-divergence loss is designed to guide this learning. Our method can achieve pluralistic face aging conditioned on open-world aging texts and arbitrary unseen face images. Qualitative and quantitative experiments demonstrate that our method can generate more diverse and high-quality plausible aging results. Peipei Li 0002, Rui Wang 0124, Huaibo Huang, Ran He 0001, Zhaofeng He 0001 |
ICCV | 5 |
| 2023 | Generative Iris Prior Embedded Transformer for Iris RestorationabstractIris restoration from complexly degraded iris images, aiming to improve iris recognition performance, is a challenging problem. Due to the complex degradation, directly training a convolutional neural network (CNN) without prior cannot yield satisfactory results. In this work, we propose a generative iris prior embedded Transformer model (Gformer), in which we build a hierarchical encoder-decoder network employing Transformer block and generative iris prior. First, we tame Transformer blocks to model long-range dependencies in target images. Second, we pretrain an iris generative adversarial network (GAN) to obtain the rich iris prior, and incorporate it into the iris restoration process with our iris feature modulator. Our experiments demonstrate that the proposed Gformer outperforms state-of-the-art methods. Besides, iris recognition performance has been significantly improved after applying Gformer. Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Peigang Li, Zhaofeng He 0001 |
ICME | 6 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 7 |
| 2023 | Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal ClassificationabstractWe present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance. Rui Wang 0124, Peipei Li 0002, Huaibo Huang, Chunshui Cao, Ran He 0001, Zhaofeng He 0001 |
NeurIPS | 6 |
| 2023 | MetaScleraSeg: an effective meta-learning framework for generalized sclera segmentation
Caiyong Wang, Wenhui Ma, Guangzhe Zhao, Zhaofeng He 0001 |
Neural Comput. Appl. | 5 |
| 2023 | Group Activity Representation Learning With Long-Short States Predictive TransformerabstractThe research goal of this paper is to learn the group activity representations in a self-supervised fashion instead of through the use of conventional methods that rely on manually annotated labels. It is essential for this task to better describe the complex group states and their future transitions. To this end, we propose a long-short state predictive Transformer (LSSPT), which mines the meaningful spatiotemporal features of group activities by predicting the future group states with long- and short-term historical state dynamics. LSSPT consists of an encoder that models diverse spatiotemporal state representations in the observation, together with a decoder that exploits rich dynamic patterns by attending to both the short-term spatial context and long-term history state evolutions to predict future group states. Furthermore, we consider the distinguishability and consistency of the predicted states and introduce a joint learning mechanism to optimize the models, enabling LSSPT to describe more reliable state transitions. Finally, extensive experiments are carried out to evaluate the learned representation on downstream tasks on the Volleyball, Collective Activity and VolleyTactic datasets, which showcases the method’s state-of-the-art performance over the existing self-supervised learning approaches. Longteng Kong, Duoxuan Pei, Zhaofeng He 0001, Di Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Progressive Task-Based Universal Network for Raw Infrared Remote Sensing Imagery Ship DetectionabstractInfrared remote sensing images are becoming increasingly popular due to their superior penetration and resistance to light interference. However, challenges still remain when applying them in real-world applications: 1) raw infrared images suffer from severe stripes interference, and the preprocessing techniques used to obtain standard image products for subsequent detection tasks tend to be time-consuming, which fails to meet the application requirements; 2) current destriping techniques may inevitably weaken the local contrast between some objects and the local background since they need to consider the gray consistency of the overall image; 3) in low-resolution images, dim and small infrared targets are challenging to discriminate, resulting in high false alarms. To address these challenges, we proposed a progressive task-based universal network for raw infrared image ship detection while simultaneously removing stripes. First, we built an integrated network consisting of two components: the stripe denoising component (SDC) and the object detection component (ODC). We also designed a feedback loss adjustment mechanism to enhance the focus of the SDC on the target area. Second, a directed two-branch network was constructed for efficient stripe noise removal, including anx-direction branch for feature enhancement and ay-direction branch for grayscale smoothing. Finally, a parallel network with two labels was designed to extract the inherent features of the target and the background, as well as their relationship features, to achieve refined ship detection. We conducted experiments on a self-assembled dataset from the GaoFen-1 satellite to validate our approach. The experimental results demonstrated that the proposed method outperformed other state-of-the-art methods in infrared image ship detection. Yuan Li 0037, Qizhi Xu, Zhaofeng He 0001, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Exploring Bias in Sclera Segmentation Models: A Group Evaluation ApproachabstractBias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias. Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 14 |
| 2022 | Achieving Consensus to Learn an Efficient and Robust Communication via Reinforcement Learning
Wei Qing, Zhaofeng He 0001, Junge Zhang, Luzhan Yuan, Wei Wang 0353 |
CogSci | 2 |
| 2022 | Generating Intra- and Inter-Class Iris Images by Identity ContrastabstractIris recognition is one of the most accurate and reliable biometric technologies. However, due to the high collection costs and privacy of the iris, it is difficult to build a large-scale iris image database for training iris recognition models. This paper proposes a novel iris image generation algorithm which can produce numerous intra- and inter-class iris images. By using contrastive learning, we disentangle identity-related features (e.g., iris texture, left or right eye) and condition-variant features (e.g., pupil size, iris exposure ratio) in the generated images. This disentanglement facilitates identity control over synthetic iris images. Since the iris has the multi-degree-of-freedom (MDOF) topology and high-entropy texture, we specially design the dual-channel input protocol to separate the topology and texture of the iris, so that the generator can infer multi-condition iris images while maintaining the unique texture details. Extensive experiments demonstrate that the proposed approach achieves impressive performance in both image quality and identity representation. Zhaofeng He 0001, Caiyong Wang |
IJCB | 2 |
| 2022 | Group Activity Representation Learning with Self-supervised Predictive Coding
Longteng Kong, Zhaofeng He 0001, Man Zhang 0005, Yunzhi Xue |
PRCV (3) | 2 |
| 2022 | Hierarchical Long-Short Transformer for Group Activity Recognition
Zhaofeng He 0001, Longteng Kong |
PRCV (3) | 2 |
| 2022 | Offline reinforcement learning with representations for actions
Xingzhou Lou, Qiyue Yin, Junge Zhang, Chao Yu 0004, Zhaofeng He 0001, Nengjie Cheng, Kaiqi Huang |
Inf. Sci. | 5 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 8 |
| 2021 | Pruning the Seg-Edge Bilateral Constraint Fully Convolutional Network for Iris Segmentation
Hui Zhang 0061, Junxing Hu, Jing Liu 0062, Zhaofeng He 0001, Lihu Xiao |
ICIG (2) | 4 |
| 2020 | Local Attention and Global Representation Collaborating for Fine-grained ClassificationabstractThe cosmetic contact lenses over an iris may change original iris textural pattern which is the foundation for iris recognition, making the cosmetic lenses a possible and easy-to-use iris presentation attack means. For practical application scenes, the cosmetic contact lenses detection still facing unsolved problems, due to the low image quality and difficulty in accurately iris localization. In this paper, we propose a novel framework called Weighted Region Network (WRN) to detect the cosmetic contact lenses. The WRN includes a local attention Weight Network and a global classification Region Network. With the inherent attention mechanism, the proposed network is able to find more discriminative regions, which reduces the requirement for target detection and improves the ability of classification. The Weight Network can be trained by using Rank loss and MSE loss without manual discriminative region annotations. Experiments are conducted on several public databases and a new collected low-quality iris image database. The proposed method outperforms state-of-the-art fake iris detection algorithms, and is also effective for the fine-grained image classification task. Yunming Bai, Hui Zhang 0061, Jing Liu 0062, Zhaofeng He 0001 |
ICPR | 6 |
| 2020 | Towards Complete and Accurate Iris Segmentation Using Deep Multi-Task Attention Network for Non-Cooperative Iris RecognitionabstractIris images captured in non-cooperative environments often suffer from adverse noise, which challenges many existing iris segmentation methods. To address this problem, this paper proposes a high-efficiency deep learning based iris segmentation approach, named IrisParseNet. Different from many previous CNN-based iris segmentation methods, which only focus on predicting accurate iris masks by following popular semantic segmentation frameworks, the proposed approach is a complete iris segmentation solution, i.e., iris mask and parameterized inner and outer iris boundaries are jointly achieved by actively modeling them into a unified multi-task network. Moreover, an elaborately designed attention module is incorporated into it to improve the segmentation performance. To train and evaluate the proposed approach, we manually label three representative and challenging iris databases, i.e., CASIA.v4-distance, UBIRIS.v2, and MICHE-I, which involve multiple illumination (NIR, VIS) and imaging sensors (long-range and mobile iris cameras), along with various types of noises. Additionally, several unified evaluation protocols are built for fair comparisons. Extensive experiments are conducted on these newly annotated databases, and results show that the proposed approach achieves state-of-the-art performance on various benchmarks. Further, as a general drop-in replacement, the proposed iris segmentation method can be used for any iris recognition methodology, and would significantly improve the performance of non-cooperative iris recognition. Caiyong Wang, Jawad Muhammad, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Efficient and Accurate Iris Detection and Segmentation Based on Multi-scale Optimized Mask R-CNN
Di Miao, Huanwei Liang, Hui Zhang 0061, Jing Liu 0062, Zhaofeng He 0001 |
ICIG (2) | 6 |
| 2019 | Toward practical remote iris recognition: A boosting based framework
Man Zhang 0005, Zhaofeng He 0001, Hui Zhang 0061, Tieniu Tan, Zhenan Sun |
Neurocomputing | 2 |
| 2018 | Generation Textured Contact Lenses Iris Images Based on 4DCycle-GANabstractWith the development of iris recognition, many identity authentication applications began to use this inherent biometric ID. Despite the breakthroughs in the identification with iris recognition technology, one primary problem remains unsolved: the presentation spoof attack. In this paper, we present a novel algorithm 4DCycle-GAN for expanding the spoof iris image database by synthesizing fake iris images wearing textured contact lenses. The proposed 4DCycle-GAN follows the Cycle-Consistent Adversarial Networks (Cycle-GAN) framework which translating between one kind images (genuine iris images) and one other kind images (textured contact lenses iris images). The 4DCycle-GAN introduces two more discriminators to improve the Cycle-GAN at the defect of lack of diversity. The two new discriminators `prefer' images generated by the generators, while the original discriminators in Cycle-GAN `prefer' real captured images. These new added confrontations make the 4DCycle-GAN avoid generating a certain kind of contact lenses texture which is larger percentage of the training iris database. The synthesized textured contact lenses iris images are used for spoofing iris detection training to improve the robustness of classification algorithm. Both the Cycle-GAN and the 4DCycle-GAN synthesizing images can improve the spoof classification results. Moreover, by using the 4DCycle-GAN, the spoof classification results are distinctly improved for unrelated non-homologous database experiments. Extensive experimental results show that the proposed method can improve the anti-spoof ability of iris recognition system. Hang Zou 0002, Hui Zhang 0061, Jing Liu 0062, Zhaofeng He 0001 |
ICPR | 5 |
| 2017 | Bin-based classifier fusion of iris and face biometrics
Di Miao, Man Zhang 0005, Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
Neurocomputing | 5 |
| 2013 | Robust spectral regression for face recognition
Yanqing Guo, Ran He 0001, Wei-Shi Zheng 0001, Xiangwei Kong 0001, Zhaofeng He 0001 |
Neurocomputing | 5 |
| 2011 | Bridge extraction based on on constrained Delaunay triangulation for panchromatic imageabstractNowadays, with the rapid increase of spatial resolution in remote sensing, accurate and efficient identification of man-made targets from large-scale high resolution panchromatic images plays a more and more important role in real applications. This paper presents an efficient bridge extraction method based on CDT. it consists of two steps: river segmentation and bridge extraction. First, river region is segmented from complicated background by texture analysis. Second, constrained Delaunay triangulation (CDT) is applied to segment river region into a triangular mesh. The medial axes of river is computed from the skeleton segments of triangles and therefore bridges are easily detected along the end points of river medial axes. Extensive experiments have been performed on panchromatic QuickBird images. The experiment results show that our method achieves detection accuracy and satisfies the speed request in large-scale high resolution panchromatic image. Feng Gao 0005, Lei Hu 0009, Zhaofeng He 0001 |
IGARSS | 3 |
| 2011 | Multiscale Contour Extraction Using a Level Set Method in Optical Satellite ImagesabstractThis letter presents a novel coarse-to-fine level set method for contour extraction in optical satellite images. To distinguish objects from a background, the undecimated wavelet transform is firstly adopted to extract image features, and a homogeneity metric is defined to measure the variation of the features inside and outside contours. In addition, the weight distribution ratio is proposed to adaptively tune the relative weight of the features. Based on the homogeneity metric and the weight distribution ratio, a novel energy functional is developed to model a contour extraction problem, and in order to reduce the computation burden, a coarse-to-fine scheme is applied to progressively extract contours in finer scale, during which a contour position constraint is introduced to limit contours evolving in a small space around the candidate contours extracted in coarser scale. Extensive experiments have been carried out on optical satellite images to validate the proposed method. Qizhi Xu, Bo Li 0006, Zhaofeng He 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2010 | Efficient and robust segmentation of noisy iris images for non-cooperative iris recognition
Tieniu Tan, Zhaofeng He 0001, Zhenan Sun |
Image Vis. Comput. | 2 |
| 2010 | Topology modeling for Adaboost-cascade based object detection
Zhaofeng He 0001, Tieniu Tan, Zhenan Sun |
Pattern Recognit. Lett. | 1 |
| 2009 | Toward Accurate and Fast Iris Segmentation for Iris BiometricsabstractIris segmentation is an essential module in iris recognition because it defines the effective image region used for subsequent processing such as feature extraction. Traditional iris segmentation methods often involve an exhaustive search of a large parameter space, which is time consuming and sensitive to noise. To address these problems, this paper presents a novel algorithm for accurate and fast iris segmentation. After efficient reflection removal, an Adaboost-cascade iris detector is first built to extract a rough position of the iris center. Edge points of iris boundaries are then detected, and an elastic model named pulling and pushing is established. Under this model, the center and radius of the circular iris boundaries are iteratively refined in a way driven by the restoring forces of Hooke's law. Furthermore, a smoothing spline-based edge fitting scheme is presented to deal with noncircular iris boundaries. After that, eyelids are localized via edge detection followed by curve fitting. The novelty here is the adoption of a rank filter for noise elimination and a histogram filter for tackling the shape irregularity of eyelids. Finally, eyelashes and shadows are detected via a learned prediction model. This model provides an adaptive threshold for eyelash and shadow detection by analyzing the intensity distributions of different iris regions. Experimental results on three challenging iris image databases demonstrate that the proposed algorithm outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Boosting ordinal features for accurate and fast iris recognitionabstractIn this paper, we present a novel iris recognition method based on learned ordinal features.Firstly, taking full advantages of the properties of iris textures, a new iris representation method based on regional ordinal measure encoding is presented, which provides an over-complete iris feature set for learning. Secondly, a novel Similarity Oriented Boosting (SOBoost) algorithm is proposed to train an efficient and stable classifier with a small set of features. Compared with Adaboost, SOBoost is advantageous in that it operates on similarity oriented training samples, and therefore provides a better way for boosting strong classifiers. Finally, the well-known cascade architecture is adopted to reorganize the learned SOBoost classifier into a dasiacascadepsila, by which the searching ability of iris recognition towards large-scale deployments is greatly enhanced. Extensive experiments on two challenging iris image databases demonstrate that the proposed method achieves state-of-the-art iris recognition accuracy and speed. In addition, SOBoost outperforms Adaboost (Gentle-Adaboost, JS-Adaboost, etc.) in terms of both accuracy and generalization capability across different iris databases. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
CVPR | 1 |
| 2008 | Robust 3D face recognition in uncontrolled environmentsabstractMost current 3D face recognition algorithms are designed based on the data collected in controlled situations, which leads to the un-guaranteed performance in practical systems. In this paper, we propose a Robust Local Log-Gabor Histograms (RLLGH) method to handle the uncontrolled problems encountered in 3D face recognition. In this challenging topic, large expressions and data noises are two main obstacles. To overcome the large expressions, we choose Log-Gabor features (LGF) to extract the distinctive and robust information embedded in 3D faces, which will be represented as 3D Log-Gabor faces. Data noises are summarized as distorted meshes, hair occlusions and misalignments. To overcome these problems, we introduce a robust local histogram (RLH) strategy, which takes advantage of the robustness of the accurate local statistical information. The combination of LGF and RLH leads to RLLGH. The novelties of this paper come from 1) Our work aims at studying 3D face recognition performance in uncontrolled environments; 2) We find that embedding LGF into the LVC framework leads to robustness in handling large expression variations; 3) The RLH strategy gives a promising way to solve the data noises problem. Our experiments are based on the large expression subset in FRGC2.0 3D face database and the expression subset in CASIA 3D face database. Experimental results show the efficiency, robustness and generalization of our proposed method. Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
CVPR | 4 |
| 2008 | Enhanced usability of iris recognition via efficient user interface and iris image restorationabstractIn this paper, we investigate the possibility of enhancing the usability of iris recognition via exploration of the specular spots in iris images. Firstly, the spatial configuration of the specular spots in iris images is utilized to estimate the distance between the user and the camera. Based on this a friendly user interface is established to assist users for their range adjustment. Furthermore, the estimated distance is used by an adaptive image restoration scheme to restore the blurred iris image, thereby increasing the depth of field of the iris camera. Experimental results show that the proposed method significantly enhances the usability of iris recognition without noticeable computation cost. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
ICIP | 1 |
| 2008 | Robust eyelid, eyelash and shadow localization for iris recognitionabstractEyelids, eyelashes and shadows are three major challenges for effective iris segmentation, which have not been adequately addressed in the current literature. In this paper, we present a novel method to localize each of them. First, a novel coarse-line to fine-parabola eyelid fitting scheme is developed for accurate and fast eyelid localization. Then, a smart prediction model is established to determine an appropriate threshold for eyelash and shadow detection. Experimental results on the challenging CASIA-IrisV3-Lamp iris image database demonstrate that the proposed method outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
ICIP | 1 |