Jing Huo

dblp:38/9090 · DBLP profile ↗
← Back
83ranked-venue papers
11as first author
57since 2021 · last 2026
0000-0002-8504-455XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 7 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Causality-Aware Efficient Exploration for Cooperative Multi-Agent Reinforcement Learning
abstract
Exploration is critical for cooperative multi agent reinforcement learning (MARL) to improve sample efficiency. However, existing intrinsic motivation based exploration strategies in MARL overlook the causal relationships among agents, global states, and rewards, suffering from interference by irrelevant factors and resulting in sample inefficiency. To address this issue, we propose Causality aware Efficient Exploration (CEE), a novel framework that enhances sample efficiency by inferring causal relationships between agents, global states with respect to rewards, thereby enabling causality guided exploration. Specifically, CEE operates through two components. First, CEE identifies causal relationships between global states and rewards, filtering out causally irrelevant state features that do not have a high impact on rewards to keep decision critical state information. Second, CEE discovers causal relationships between agents' behaviors and rewards to quantify each agent's contribution to collective performance. To achieve this, we introduce a causal entropy objective that promotes exploration aligned with decision critical aspects of the underlying causal structure. We provide comprehensive validation through experiments on 21 challenging tasks spanning SMAC, SMAC v2, and Google Research Football (GRF) environments. Our results demonstrate that CEE achieves superior performance in terms of sample efficiency and asymptotic performance compared to existing MARL methods.
Hongye Cao, Tianpei Yang, Hammadi Rafik Ouariachi, Yali Du 0001, Jing Huo, Yang Gao 0001
AAAI7
2026 ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation
abstract
One-shot imitation learning (OSIL) offers a promising way to teach robots new skills without large-scale data collection. However, current OSIL methods are primarily limited to short-horizon tasks, thus limiting their applicability to complex, long-horizon manipulations. To address this limitation, we propose ManiLong-Shot, a novel framework that enables effective OSIL for long-horizon prehensile manipulation tasks. ManiLong-Shot structures long-horizon tasks around physical interaction events, reframing the problem as sequencing interaction-aware primitives instead of directly imitating continuous trajectories. This primitive decomposition can be driven by high-level reasoning from a vision-language model (VLM) or by rule-based heuristics derived from robot state changes. For each primitive, ManiLong-Shot predicts invariant regions critical to the interaction, establishes correspondences between the demonstration and the current observation, and computes the target end-effector pose, enabling effective task execution. Extensive simulation experiments show that ManiLong-Shot, trained on only 10 short-horizon tasks, generalizes to 20 unseen long-horizon tasks across three difficulty levels via one-shot imitation, achieving a 22.8% relative improvement over the SOTA. Additionally, real-robot experiments validate ManiLong-Shot’s ability to robustly execute three long-horizon manipulation tasks via OSIL, confirming its practical applicability.
Chongkai Gao, Lin Shao 0002, Jieqi Shi, Jing Huo, Yang Gao 0001
AAAI5
2026 Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
abstract
Siyuan Gan, Jiaheng Liu, Boyan Wang, Tianpei Yang, Runqing Miao, Yuyao Zhang, Fanyu Meng, Junlan Feng, Linjian Meng, Jing Huo, Yang Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Siyuan Gan, Tianpei Yang, Runqing Miao, Junlan Feng, Linjian Meng, Jing Huo, Yang Gao 0001
ACL (1)10
2026 Mitigating Security Risks in Large Language Models: A Full Lifecycle Perspective
Yanming Wang, Zhixin Bai, Jing Huo, Hongye Cao, Yang Gao 0001
Mach. Learn.3
2026 A continual learning framework with long-term and multiple short-term memory networks
Shangge Liu, Lei Wang 0001, Rui Yan 0005, Jing Huo, Wenbin Li 0006, Yang Gao 0001
Neural Networks4
2026 AMPL: An adaptive meta-prompt learner for few-shot image classification
Zhiping Wu, Lian Huai, Zeyu Shangguan, Lei Wang 0001, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang
Neural Networks6
2026 Retrieval-augmented diffusion with acoustic priors for high-fidelity sonar image generation
Shaocong Yang, Zheng Gu 0001, Hongye Cao, Xiaolong Qi, Jing Huo, Yang Gao 0001
Pattern Recognit.6
2026 Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation
Kai Chen 0026, Hewei Wang 0001, Pinzhuo Tian, Jing Huo, Taihang Zhen, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Model-Based Offline Reinforcement Learning With Adversarial Data Augmentation
abstract
Model-based offline reinforcement learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble models, rolling out conservative estimation to mitigate extrapolation errors. However, the static data makes it challenging to develop a robust policy, and offline agents cannot access the environment to gather new data. To address these challenges, we introduce Model-based Offline Reinforcement learning with AdversariaL data augmentation (MORAL). In MORAL, we replace the fixed horizon rollout by employing adversarial data augmentation to execute alternating sampling with ensemble models to enrich training data. Specifically, this adversarial process dynamically selects ensemble models against policy for biased sampling, mitigating the optimistic estimation of fixed models, thus robustly expanding the training data for policy optimization. Moreover, a differential factor (DF) is integrated into the adversarial process for regularization, ensuring error minimization in extrapolations. This data-augmented optimization adapts to diverse offline tasks without rollout horizon tuning, showing remarkable applicability. Extensive experiments on the D4RL benchmark demonstrate that MORAL outperforms other model-based offline RL methods in terms of policy learning and sample efficiency.
Hongye Cao, Jing Huo, Shangdong Yang, Tianpei Yang, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2026 FISN: FInding Spatial Neighborhoods for Generalizable Novel View Synthesis
abstract
We present FISN, a generalizable novel view synthesis algorithm that enables feedforward inference of Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) from reference images. Unlike existing work that either separately model the 3D feature space on each view or process multiview reference features by 3D-point-based view aggregation, FISN integrates multi-reference 3D cost volumes into a unified high-dimensional entity. Specifically, we reconceptualize the generalizable novel view synthesis task as a feedforward process of FInding Spatial Neighborhoods across this unified 4D feature space, comprising both view and spatial dimensions, and introduce View-Spatial Convolutions for direct 4D feature aggregation. This enhances the correlation among multiview neighboring points in a window-to-window manner and incorporates 3D spatial awareness. However, this approach poses two intertwined challenges: high computational expense for high-dimensional features and degraded rendering performance with low-resolution features. To address these challenges, FISN constructs a new efficient convolution paradigm, Decomposable View-Spatial Convolution, which includes a Spatial Cross Decomposition strategy as well as a Feature Compression and Upscaling module. This paradigm maintains multiview geometric consistency better than existing decomposition methods and achieves a balance between efficiency and fine-grained spatial features. Furthermore, by integrating Depth Refinement modules based on this paradigm, FISN further improves global depth understanding. Comprehensive evaluations on mainstream datasets and benchmarks demonstrate that FISN achieves state-of-the-art performance for both NeRF and 3DGS, and remains robust in challenging scenarios where existing 3DGS-based methods struggle, such as those with noisy poses or dense references. The code will be released soon.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Vis. Comput. Graph.3
2025 Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibration
abstract
Few-shot Class-Incremental Learning (FSCIL) challenges models to adapt to new classes with limited samples, presenting greater difficulties than traditional class-incremental learning. While existing approaches rely heavily on visual models and require additional training during base or incremental phases, we propose a training-free framework that leverages pre-trained visual-language models like CLIP. At the core of our approach is a novel Bi-level Modality Calibration (BiMC) strategy. Our framework initially performs intra-modal calibration, combining LLM-generated fine-grained category descriptions with visual prototypes from the base session to achieve precise classifier estimation. This is further complemented by inter-modal calibration that fuses pre-trained linguistic knowledge with task-specific visual priors to mitigate modality-specific biases. To enhance prediction robustness, we introduce additional metrics and strategies that maximize the utilization of limited data. Extensive experimental results demonstrate that our approach significantly outperforms existing methods. Code is available at: https://github.com/yychen016/BiMC.
Tianyu Ding, Lei Wang 0001, Jing Huo, Yang Gao 0001, Wenbin Li 0006
CVPR4
2025 Enhancing Trust-Region Bayesian Optimization via Newton Methods
abstract
Bayesian Optimization (BO) has been widely applied to optimize expensive black-box functions while retaining sample efficiency. However, scaling BO to high-dimensional spaces remains challenging. Existing literature proposes performing standard BO in multiple local trust regions (TuRBO) for heterogeneous modeling of the objective function and avoiding over-exploration. Despite its advantages, using local Gaussian Processes (GPs) reduces sampling efficiency compared to a global GP. To enhance sampling efficiency while preserving heterogeneous modeling, we propose to construct multiple local quadratic models using gradients and Hessians from a global GP, and select new sample points by solving the bound-constrained quadratic program. Additionally, we address the issue of vanishing gradients of GPs in high-dimensional spaces. We provide a convergence analysis and demonstrate through experimental results that our method enhances the efficacy of TuRBO and outperforms a wide range of high-dimensional BO techniques on synthetic functions and real-world applications.
Quanlin Chen, Jing Huo, Tianyu Ding, Yang Gao 0001, Yuetong Chen
ECAI3
2025 A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models
abstract
We propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consultations based on the patient’s treatment background and symptoms. In the final summary and output stage of the consultation, it utilizes two experience databases—the Correct Consultation Experience Database and the Chain of Thought (CoT) Experience Database—which evolve with accumulated experience during consultations. This evolution drives the framework’s self-improvement, significantly enhancing the rationality and accuracy of the consultations. To ensure that the conclusions are safe, reliable, and aligned with human values, the final decisions undergo a safety review before being provided to the patient. This framework achieved accuracy rates of 89.2% and 83.1% on the MedQA and PubMedQA datasets, respectively.
Kai Chen 0026, Jing Huo, Pinzhuo Tian, Yang Gao 0001
ICASSP3
2025 Towards Empowerment Gain through Causal Structure Learning in Model-Based Reinforcement Learning
abstract
In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision. Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions. We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL. To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning. Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting. Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods. We evaluate ECL combined with $3$ causal discovery methods across $6$ environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance.
Hongye Cao, Shaokang Dong, Tianpei Yang, Jing Huo, Yang Gao 0001
ICLR6
2025 Causal Information Prioritization for Efficient Reinforcement Learning
abstract
Current Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to address this problem, they lack grounded modeling of reward-guided causal understanding of states and actions for goal-orientation, thus impairing learning efficiency. To tackle this issue, we propose a novel method named Causal Information Prioritization (CIP) that improves sample efficiency by leveraging factored MDPs to infer causal relationships between different dimensions of states and actions with respect to rewards, enabling the prioritization of causal information. Specifically, CIP identifies and leverages causal relationships between states and rewards to execute counterfactual data augmentation to prioritize high-impact state features under the causal understanding of the environments. Moreover, CIP integrates a causality-aware empowerment learning objective, which significantly enhances the agent's execution of reward-guided actions for more efficient exploration in complex environments. To fully assess the effectiveness of CIP, we conduct extensive experiments across $39$ tasks in $5$ diverse continuous control environments, encompassing both locomotion and manipulation skills learning with pixel-based and sparse reward settings. Experimental results demonstrate that CIP consistently outperforms existing RL methods across a wide range of scenarios.
Hongye Cao, Tianpei Yang, Jing Huo, Yang Gao 0001
ICLR4
2025 GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
abstract
Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to assist in understanding novel tasks, thereby mitigating this issue. However, these methods lack a task-specific learning process, which is essential for an accurate understanding of 3D environments, often leading to execution failures. In this paper, we introduce GravMAD, a sub-goal-driven, language-conditioned action diffusion framework that combines the strengths of imitation learning and foundation models. Our approach breaks tasks into sub-goals based on language instructions, allowing auxiliary guidance during both training and inference. During training, we introduce Sub-goal Keypose Discovery to identify key sub-goals from demonstrations. Inference differs from training, as there are no demonstrations available, so we use pre-trained foundation models to bridge the gap and identify sub-goals for the current task. In both phases, GravMaps are generated from sub-goals, providing GravMAD with more flexible 3D spatial guidance compared to fixed 3D positions. Empirical evaluations on RLBench show that GravMAD significantly outperforms state-of-the-art methods, with a 28.63\% improvement on novel tasks and a 13.36\% gain on tasks encountered during training. Evaluations on real-world robotic tasks further show that GravMAD can reason about real-world tasks, associate them with relevant visual information, and generalize to novel tasks. These results demonstrate GravMAD's strong multi-task learning and generalization in 3D manipulation. Video demonstrations are available at: https://gravmad.github.io.
Yangtao Chen, Junhui Yin, Jing Huo, Pinzhuo Tian, Jieqi Shi, Yang Gao 0001
ICLR4
2025 Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
abstract
Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF), self-play not only boosts Large Language Model (LLM) performance but also overcomes the limitations of traditional Bradley-Terry (BT) model assumptions by finding the Nash equilibrium (NE) of a preference-based, two-player constant-sum game. However, existing methods either guarantee only average-iterate convergence, incurring high storage and inference costs, or converge to the NE of a regularized game, failing to accurately reflect true human preferences. In this paper, we introduce Magnetic Preference Optimization (MPO), a novel approach capable of achieving last-iterate convergence to the NE of the original game, effectively overcoming the limitations of existing methods. Building upon Magnetic Mirror Descent (MMD), MPO attains a linear convergence rate, making it particularly suitable for fine-tuning LLMs. To ensure our algorithm is both theoretically sound and practically viable, we present a simple yet effective implementation that adapts the theoretical insights to the RLHF setting. Empirical results demonstrate that MPO can significantly enhance the performance of LLMs, highlighting the potential of self-play methods in alignment.
Chengdong Ma, Linjian Meng, Jiancong Xiao, Zhaowei Zhang 0001, Jing Huo, Weijie J. Su, Yaodong Yang 0001
ICLR8
2025 Advancing Safe Language Generation: Exploring Alternative Constrained RLHF
abstract
As Large Language Models (LLMs) have been increasingly deployed for various language generation tasks, ensuring the safety and appropriateness of generated content has become a critical challenge. While these models excel at tasks such as dialogue generation, text completion, and content creation, they may inadvertently generate harmful or inappropriate responses that compromise user safety and model reliability. Recent advancements have attempted to address this problem by incorporating safety constraints within the Reinforcement Learning with Human Feedback (RLHF) framework. However, these approaches struggle with the complexity of selecting appropriate compromise parameters, leading to suboptimal performance in safe language generation. To address this problem, we explore advanced alternative first-order constrained optimization strategies that dynamically adjust the balance of helpfulness and harmlessness during training. Specifically, we choose FOCOPS and P3O, which have demonstrated strong performance in various reinforcement learning tasks, and integrate them into the RLHF framework to balance the quality and safety of generated responses. Then we evaluate the performance of different constrained optimization methods for language generation. The results indicate that the P3O algorithm shows significant potential to advance safe language generation while maintaining the helpfulness of model responses.
Zhixin Bai, Yanming Wang, Jing Huo, Yang Gao 0001
ICME4
2025 HDBO-B: On Benchmarking High-Dimensional Bayesian Optimization
abstract
Bayesian optimization (BO) has been extensively studied and applied as a sample-efficient, black-box optimization method in neural network learning, particularly for hyperparameter optimization. However, high-dimensional Bayesian optimization (HDBO) remains a challenging research direction. While existing BO benchmarks focus on practical low-dimensional tasks and specific high-dimensional settings, current HDBO studies often rely on custom experimental settings, resulting in a lack of standardized benchmarks for comprehensive evaluation. To address this, we introduce the first standardized HDBO benchmark HDBO-B. HDBO-B encompasses diverse high-dimensional optimization tasks, synthetic functions tailored for common research settings, and a robust framework with standardized testing, method implementation, and extensive examples. By benchmarking nine representative optimization methods, we demonstrate HDBO-B’s validity and highlight new research directions in neural network methods. Codes are available at: https://github.com/Yiyuiii/HDBO-B.
Hongye Cao, Quanlin Chen, Jing Huo, Dong Li 0016, Yang Gao 0001
IJCNN4
2025 Multi-Agent Reinforcement Learning with Communication-Constrained Priors
abstract
Communication is one of the effective means to improve the learning of cooperative policy in multi-agent systems. However, in most real-world scenarios, lossy communication is a prevalent issue. Existing multi-agent reinforcement learning with communication, due to their limited scalability and robustness, struggles to apply to complex and dynamic real-world environments. To address these challenges, we propose a generalized communication-constrained model to uniformly characterize communication conditions across different scenarios. Based on this, we utilize it as a learning prior to distinguish between lossy and lossless messages for specific scenarios. Additionally, we decouple the impact of lossy and lossless messages on distributed decision-making, drawing on a dual mutual information estimatior, and introduce a communication-constrained multi-agent reinforcement learning framework, quantifying the impact of communication messages into the global reward. Finally, we validate the effectiveness of our approach across several communication-constrained benchmarks.
Guang Yang 0066, Tianpei Yang, Jingwen Qiao, Yanqing Wu, Jing Huo, Xingguo Chen, Yang Gao 0001
NeurIPS5
2025 REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.1
2025 ONNXPruner: ONNX-Based General Model Pruning Adapter
abstract
Recent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various models and platforms remains a significant challenge. To address this challenge, we propose ONNXPruner, a versatile pruning adapter designed for the ONNX format models. ONNXPruner streamlines the adaptation process across diverse deep learning frameworks and hardware platforms. A novel aspect of ONNXPruner is its use of node association trees, which automatically adapt to various model architectures. These trees clarify the structural relationships between nodes, guiding the pruning process, particularly highlighting the impact on interconnected nodes. Furthermore, we introduce a tree-level evaluation method. By leveraging node association trees, this method allows for a comprehensive analysis beyond traditional single-node evaluations, enhancing pruning performance without the need for extra operations. Experiments across multiple models and datasets confirm ONNXPruner's strong adaptability and increased efficacy. Our work aims to advance the practical application of model pruning.
Dongdong Ren, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Jing Huo, Hongbing Pan, Yang Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Automated C. elegans behavior analysis via deep learning-based detection and tracking
abstract
As a well-established and extensively utilized model organism, Caenorhabditis elegans (C. elegans) serves as a crucial platform for investigating behavioral regulation mechanisms and their biological significance. However, manually tracking the locomotor behavior of large numbers of C. elegans is both cumbersome and inefficient. To address the above challenges, we innovatively propose an automated approach for analyzing C. elegans behavior through deep learning-based detection and tracking. Building upon existing research, we developed an enhanced worm detection framework that integrates YOLOv8 with ByteTrack, enabling real-time, precise tracking of multiple worms. Based on the tracking results, we further established an automated high-throughput method for quantitative analysis of multiple movement parameters, including locomotion velocity, body bending angle, and roll frequency, thereby laying a robust foundation for high-precision, automated analysis of complex worm behaviors. including movement speed, body bending angle, and roll frequency, thereby laying a robust foundation for high-precision, automated analysis of complex worm behaviors. Comparative evaluations demonstrate that the proposed enhanced C. elegans detection framework outperforms existing methods, achieving a precision of 99.5%, recall of 98.7%, and mAP50 of 99.6%, with a processing speed of 153 frames per second (FPS). The established framework for worm detection, tracking, and automated behavioral analysis developed in this study delivers superior detection and tracking accuracy while enhancing tracking continuity and robustness. Unlike traditional labor-intensive measurement approaches, our framework supports simultaneous tracking of multiple worms while maintaining automated extraction of various behavioral parameters with high precision. Furthermore, our approach advances the standardization of C. elegans behavioral parameter analysis, which can analyze the behavioral data of multiple worms at the same time, significantly improving the experimental throughput and providing an efficient tool for drug screening, gene function research and other fields.
Xiaoke Liu, Wenjie Teng, Yuzhong Peng, Boao Li, Xiaoqing Han, Jing Huo
PLoS Comput. Biol.7
2025 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
abstract
3D Gaussian Splatting (3DGS) has emerged as a prominent technique with the potential to become a mainstream method for 3D representations. It can effectively transform multi-view images into explicit 3D Gaussian through efficient training, and achieve real-time rendering of novel views. This survey aims to analyze existing 3DGS-related works from multiple intersecting perspectives, including related tasks, technologies, challenges, and opportunities. The primary objective is to provide newcomers with a rapid understanding of the field and to assist researchers in methodically organizing existing technologies and challenges. Specifically, we delve into the optimization, application, and extension of 3DGS, categorizing them based on their focuses or motivations. Additionally, we summarize and classify nine types of technical modules and corresponding improvements identified in existing works. Based on these analyses, we further examine the common challenges and technologies across various tasks, proposing potential research opportunities.
Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Wenbin Li 0006, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Leveraging Frequency Analysis for Image Denoising Network Pruning
abstract
As a common model compression technique, network pruning is widely used to reduce storage and computational cost of deep models in the resource-constrained regime. However, most current pruning methods are designed for high-level vision tasks, with few developed for low-level vision tasks. We observed that the norm-based pruning criterion, originally designed for high-level vision tasks, is highly unsuitable for low-level image denoising networks. This difference arises because image denoising networks pursue distinct feature granularities and goals compared to typical high-level vision tasks. To address this issue, we propose a novel filter evaluation method, termed High-Frequency Components Pruning (HFCP), specifically tailored for image denoising network pruning. HFCP assesses filter importance based on high-frequency components. To the best of our knowledge, this is the first pruning method designed specifically for image denoising tasks, straightforward and applicable to various types of noise. Furthermore, HFCP enhances the pruned model's high-frequency information content with high reliability and interpretability. This facilitates the network's ability to distinguish high-frequency signals from noise. We comprehensively analyzed multiple image denoising networks and validated HFCP's effectiveness across four mainstream networks.
Dongdong Ren, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Hongbing Pan, Yang Gao 0001
IEEE Trans. Image Process.3
2025 Stitching, Fine-Tuning, and Re-Training: A SAM-Enabled Framework for Semi-Supervised 3D Medical Image Segmentation
abstract
Segment Anything Model (SAM) fine-tuning has shown remarkable performance in medical image segmentation in a fully supervised manner, but requires precise annotations. To reduce the annotation cost and maintain satisfactory performance, in this work, we leverage the capabilities of SAM for establishing semi-supervised medical image segmentation models. Rethinking the requirements of effectiveness, efficiency, and compatibility, we propose a three-stage framework, i.e., Stitching, Fine-tuning, and Re-training (SFR). The current fine-tuning approaches mostly involve 2D slice-wise fine-tuning that disregards the contextual information between adjacent slices. Our stitching strategy mitigates the mismatch between natural and 3D medical images. The stitched images are then used for fine-tuning SAM, providing robust initialization of pseudo-labels. Afterwards, we train a 3D semi-supervised segmentation model while maintaining the same parameter size as the conventional segmenter such as V-Net. Our SFR framework is plug-and-play, and easily compatible with various popular semi-supervised methods. We also develop an extended framework SFR+ with selective fine-tuning and re-training through confidence estimation. Extensive experiments validate that our SFR and SFR+ achieve significant improvements in both moderate annotation and scarce annotation across five datasets. In particular, SFR framework improves the Dice score of Mean Teacher from 29.68% to 74.40% with only one labeled data of LA dataset. The code is available at https://github.com/ShumengLI/SFR.
Shumeng Li, Lei Qi 0001, Qian Yu 0007, Jing Huo, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging4
2025 Dictionary Based Generative Adversarial Network for Multi-Collection Style Transfer
abstract
Most collection-based style transfer methods require training a separate model for each individual collection of styles, making the extension to multiple collections of styles less flexible. Besides, the existing collection-based methods are also less flexible in extending to new style collections in a continual manner. To address these issues, we propose a novelMultI-Dictionary Generative Adversarial Network framework (MID-GAN)for multi-collection style transfer. Specifically, we design a multi-dictionary architecture within a GAN, with each dictionary consisting of a set of local style codes for a specific style collection. Benefiting from the local style codes used in the dictionary, a stylization module with aligned skip connections is further proposed, which can better preserve both the local details and the overall image structure. The dictionary design allows a flexible extension to new style collections by readily adding new dictionaries and we propose a continual training strategy that can both preserve the style transfer ability of old styles and achieve good transfer results for newly added styles. Extensive experiments are performed to show that the proposed method is better than existing collection-based style transfer methods. We also demonstrate the proposed method can generate diverse meaningful style transfer results of the same style collection.
Jing Huo, Shiyin Jin, Jiashen Li, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
IEEE Trans. Multim.1
2025 State Abstraction via Deep Supervised Hash Learning
abstract
State abstraction is a widely used technique in reinforcement learning (RL) that compresses the state space to accelerate learning algorithms. However, designing an effective abstraction function in large-scale or high-dimensional state space problems remains a significant challenge. In this brief, we present a novel state abstraction method based on deep supervised hash learning (DSH) and provide a theoretical analysis of its near-optimal property. Furthermore, by leveraging the DSH-based representation as the optimization objective, we propose a direct and concise optimization method based on the target value. In addition, we construct an auxiliary learning task for state abstraction that can be combined with various RL algorithms. In particular, we apply the DSH-based state abstraction to both deep Q-learning (DQN) and soft actor-critic (SAC). Extensive experiments are conducted on Atari and several classic control benchmarks to evaluate the effectiveness of the DSH-based state abstraction method, showing that our method surpasses existing state abstraction algorithms in performance.
Guang Yang 0066, Jing Huo, Shangdong Yang, Tianyu Ding, Xingguo Chen, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Learning from Orthogonal Space with Multimodal Large Models for Generalized Few-shot Segmentation
abstract
Generalized Few-shot Segmentation (GFSS) aims to segment both base and novel classes in a query image, conditioning on richly annotated data of base classes and limited exemplars from novel classes. The learning of novel classes undoubtedly faces a disadvantage in this competition due to the highly unbalanced data, which skews the learned feature space toward the base classes. In this article, we present an innovative idea termed as “learning from orthogonal space” to avoid the conflict in the process of learning novel classes. Specifically, we first utilize textual modal information from labels to provide more distinguishable initial prototypes for different categories, ensuring that the prototypes for base and novel classes have distinct initial separations. Then, a simple but effective Feature Separating Module (FSM) is introduced to enhance the model’s ability to differentiate between base and novel classes through learning the novel features from orthogonal space. In addition, we propose a Trigger-Promoting Framework (TPF) during the testing stage to further boost performance. The prediction results from the FSM serve as a multimodal prompt to leverage information residing in large models, such as CLIP and SAM, to enhance performance. Comprehensive experiments on two benchmarks demonstrate that our method achieves superior performance on novel classes without sacrificing accuracy on base classes. Notably, our Feature Separating with Trigger-promoting Network (FS-TPNet) outperforms the current state-of-the-art method by \(12.8\%\) overall IoU on novel classes on PASCAL- \(5^{i}\) under the 1-shot scenario. Our codes will be available at https://github.com/returnZXJ/FS-TPNet .
Hang Yu 0006, Shengjie Yang, Jing Huo, Pinzhuo Tian
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Exploiting Inter-sample and Inter-feature Relations in Dataset Distillation
abstract
Dataset distillation has emerged as a promising approach in deep learning, enabling efficient training with small synthetic datasets derived from larger real ones. Particularly, distribution matching-based distillation methods attract attention thanks to its effectiveness and low computational cost. However, these methods face two primary limitations: the dispersed feature distribution within the same class in synthetic datasets, reducing class discrim-ination, and an exclusive focus on mean feature consistency, lacking precision and comprehensiveness. To address these challenges, we introduce two novel constraints: a class centralization constraint and a covariance matching constraint. The class centralization constraint aims to enhance class discrimination by more closely clustering samples within classes. The covariance matching constraint seeks to achieve more accurate feature distribution matching between real and synthetic datasets through local feature covariance matrices, particularly beneficial when sample sizes are much smaller than the number of features. Experiments demonstrate notable improvements with these constraints, yielding performance boosts of up to 6.6% on CIFAR10, 2.9% on SVHN, 2.5% on CIFAR100, and 2.5% on TinyImageNet, compared to the state-of-the-art relevant methods. In addition, our method maintains robust performance in cross-architecture settings, with a maximum performance drop of 1.7% on four architectures. Code is avail-able at https://github.com/VincenDen/IID.
Wenxiao Deng, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Kuihua Huang, Jing Huo, Yang Gao 0001
CVPR7
2024 Dynamic Replay Training for Class-Incremental Learning
abstract
Replay-based methods for Class-Incremental Learning (CIL) typically employ new classes and a limited subset of old classes stored in memory to facilitate the model training. However, these methods often lead to class imbalance and catastrophic forgetting, where the model forgets previously learned tasks. While some studies have attempted to address such a class imbalance issue, they do not fully consider the dynamic nature of forgetting in the model. In this paper, we propose a novel method called Dynamic Replay Training (DRT) to address the dynamic forgetting of previously learned tasks by the model. DRT replays memory data with dynamically changing frequencies, offering a novel perspective to tackle catastrophic forgetting and class imbalance. The proposed method is evaluated on CIFAR-100 and ImageNet-100 in various settings, showing significant improvements of 8.28% and 4.53% in terms of classification accuracy compared to the baseline method on the two datasets, respectively.
Dongdong Ren, Chenglei Peng, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICASSP4
2024 InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
abstract
Generalizing Neural Radiance Fields (NeRF) to new scenes is a significant challenge that existing approaches struggle to address without extensive modifications to vanilla NeRF framework. We introduce **InsertNeRF**, a method for **INS**tilling g**E**ne**R**alizabili**T**y into **NeRF**. By utilizing multiple plug-and-play HyperNet modules, InsertNeRF dynamically tailors NeRF's weights to specific reference scenes, transforming multi-scale sampling-aware features into scene-specific representations. This novel design allows for more accurate and efficient representations of complex appearances and geometries. Experiments show that this method not only achieves superior generalization performance but also provides a flexible pathway for integration with other NeRF-like systems, even in sparse input settings. Code will be available at: https://github.com/bbbbby-99/InsertNeRF.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICLR3
2024 Privacy-preserving scheme of cotton supply chain based on blockchain
abstract
With the popularization and development of artificial intelligence and intelligent agriculture, the supply chain has become an effective management method. However, the complexity and variety of data types and the large amount of data in the supply chain make its privacy and security a major challenge. Especially in the cotton supply chain, due to the large number of participants involved and the varying levels of data security protection of each participant, it is easy to leak the user's identity privacy and private data in the supply chain during data sharing. Therefore, this paper is dedicated to improving privacy security in the cotton supply chain. First, we define the data in the cotton supply chain and divide it into private data and shared data, which is conducive to the secure storage and sharing of data in the cotton supply chain. Second, we introduce blockchain and advanced cryptography techniques and propose a method of the pseudo-identity encryption based on SM9, which ensures the security of data in the supply chain while ensuring the privacy and security of the identities of the data owners and data receivers, and realizes the tracking of the identities of the malicious users who have uploaded false data. Finally, we adopt the method of dual-chain storage and introduce IPFS to reduce the storage burden of blockchain as much as possible and carry out security and experimental analyses of the proposed scheme to verify its feasibility.
Jing Huo, Xuehua Bi
ISPA3
2024 SCaR: Refining Skill Chaining for Long-Horizon Robotic Manipulation via Dual Regularization
abstract
Long-horizon robotic manipulation tasks typically involve a series of interrelated sub-tasks spanning multiple execution stages. Skill chaining offers a feasible solution for these tasks by pre-training the skills for each sub-task and linking them sequentially. However, imperfections in skill learning or disturbances during execution can lead to the accumulation of errors in skill chaining process, resulting in execution failures. In this paper, we investigate how to achieve stable and smooth skill chaining for long-horizon robotic manipulation tasks. Specifically, we propose a novel skill chaining framework called Skill Chaining via Dual Regularization (SCaR). This framework applies dual regularization to sub-task skill pre-training and fine-tuning, which not only enhances the intra-skill dependencies within each sub-task skill but also reinforces the inter-skill dependencies between sequential sub-task skills, thus ensuring smooth skill chaining and stable long-horizon execution. We evaluate the SCaR framework on two representative long-horizon robotic manipulation simulation benchmarks: IKEA furniture assembly and kitchen organization. Additionally, we conduct a simple real-world validation in tabletop robot pick-and-place tasks. The experimental results show that, with the support of SCaR, the robot achieves a higher success rate in long-horizon tasks compared to relevant baselines and demonstrates greater robustness to perturbations.
Ze Ji, Jing Huo, Yang Gao 0001
NeurIPS3
2024 Re-examining Supervised Dimension Reduction for High-Dimensional Bayesian Optimization
Quanlin Chen, Jing Huo, Tianyu Ding, Yang Gao 0001, Dong Li 0016
PPSN (2)2
2024 Task-Aware Few-Shot Image Generation via Dynamic Local Distribution Estimation and Sampling
Zheng Gu 0001, Wenbin Li 0006, Tianyu Ding, Jing Huo, Kuihua Huang, Yang Gao 0001
PRCV (2)5
2024 Making the Primary Task Primary: Boosting Few-Shot Classification by Gradient-Biased Multi-task Learning
Yunchen Wu, Boyao Shi, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Tinghao Yu
PRCV (1)3
2024 Towards efficient image and video style transfer via distillation and learnable feature transformation
Jing Huo, Meihao Kong, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001
Comput. Vis. Image Underst.1
2024 Attention-based investigation and solution to the trade-off issue of adversarial training
Chang-Bin Shao, Wenbin Li 0006, Jing Huo, Zhenhua Feng 0001, Yang Gao 0001
Neural Networks3
2024 Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
abstract
Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively. Our project webpage is available at https://analogist2d.github.io.
Zheng Gu 0001, Jing Liao 0001, Jing Huo, Yang Gao 0001
ACM Trans. Graph.4
2023 Modeling Inter-Class and Intra-Class Constraints in Novel Class Discovery
abstract
Novel class discovery (NCD) aims at learning a model that transfers the common knowledge from a class-disjoint labelled dataset to another unlabelled dataset and discovers new classes (clusters) within it. Many methods, as well as elaborate training pipelines and appropriate objectives, have been proposed and considerably boosted performance on NCD tasks. Despite all this, we find that the existing methods do not sufficiently take advantage of the essence of the NCD setting. To this end, in this paper, we propose to model both inter-class and intra-class constraints in NCD based on the symmetric Kullback-Leibler divergence (sKLD). Specifically, we propose an inter-class sKLD constraint to effectively exploit the disjoint relationship between labelled and unlabelled classes, enforcing the separability for different classes in the embedding space. In addition, we present an intra-class sKLD constraint to explicitly constrain the intra-relationship between a sample and its augmentations and ensure the stability of the training process at the same time. We conduct extensive experiments on the popular CIFAR10, CIFAR100 and ImageNet benchmarks and successfully demonstrate that our method can establish a new state of the art and can achieve significant performance improvements, e.g., 3.5%/3.7% clustering accuracy improvements on CIFAR100-50 dataset split under the task-aware/-agnostic evaluation protocol, over previous state-of-the-art methods. Code is available at https://github.com/FanZhichen/NCD-IIC.
Wenbin Li 0006, Zhichen Fan, Jing Huo, Yang Gao 0001
CVPR3
2023 Enhancing OOD Generalization in Offline Reinforcement Learning with Energy-Based Policy Optimization
abstract
Offline Reinforcement Learning (RL) is an important research domain for real-world applications because it can avert expensive and dangerous online exploration. Offline RL is prone to extrapolation errors caused by the distribution shift between offline datasets and states visited by behavior policy. Existing offline RL methods constrain the policy to offline behavior to prevent extrapolation errors. But these methods limit the generalization potential of agents in Out-Of-Distribution (OOD) regions and cannot effectively evaluate OOD generalization behavior. To improve the generalization of the policy in OOD regions while avoiding extrapolation errors, we propose an Energy-Based Policy Optimization (EBPO) method for OOD generalization. An energy function based on the distribution of offline data is proposed for the evaluation of OOD generalization behavior, instead of relying on model discrepancies to constrain the policy. The way of quantifying exploration behavior in terms of energy values can balance the return and risk. To improve the stability of generalization and solve the problem of sparse reward in complex environment, episodic memory is applied to store successful experiences that can improve sample efficiency. Extensive experiments on the D4RL datasets demonstrate that EBPO outperforms the state-of-the-art methods and achieves robust performance on challenging tasks that require OOD generalization.
Hongye Cao, Shangdong Yang, Jing Huo, Xingguo Chen, Yang Gao 0001
ECAI3
2023 Where and How: Mitigating Confusion in Neural Radiance Fields from Sparse Inputs
abstract
Neural Radiance Fields from Sparse inputs (NeRF-S) have shown great potential in synthesizing novel views with a limited number of observed viewpoints. However, due to the inherent limitations of sparse inputs and the gap between non-adjacent views, rendering results often suffer from over-fitting and foggy surfaces, a phenomenon we refer to as "CONFUSION" during volume rendering. In this paper, we analyze the root cause of this confusion and attribute it to two fundamental questions: "WHERE" and "HOW". To this end, we present a novel learning framework, WaH-NeRF, which effectively mitigates confusion by tackling the following challenges: (i) "WHERE" to Sample? in NeRF-S-we introduce a Deformable Sampling strategy and a Weight-based Mutual Information Loss to address sample-position confusion arising from the limited number of viewpoints; and (ii) "HOW" to Predict? in NeRF-S-we propose a Semi-Supervised NeRF learning Paradigm based on pose perturbation and a Pixel-Patch Correspondence Loss to alleviate prediction confusion caused by the disparity between training and testing viewpoints. By integrating our proposed modules and loss functions, WaH-NeRF outperforms previous methods under the NeRF-S setting. Code is available https://github.com/bbbbby-99/WaH-NeRF.
Yanqi Bao, Jing Huo, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
ACM Multimedia3
2023 LibFewShot: A Comprehensive Library for Few-Shot Learning
abstract
Few-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning.
Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Defensive Few-Shot Learning
abstract
This article investigates a new challenging problem called defensive few-shot learning in order to learn a robust few-shot model against adversarial attacks. Simply applying the existing adversarial defense methods to few-shot learning cannot effectively solve this problem. This is because the commonly assumed sample-level distribution consistency between the training and test sets can no longer be met in the few-shot setting. To address this situation, we develop a general defensive few-shot learning (DFSL) framework to answer the following two key questions: (1) how to transfer adversarial defense knowledge from one sample distribution to another? (2) how to narrow the distribution gap between clean and adversarial examples under the few-shot setting? To answer the first question, we propose an episode-based adversarial training mechanism by assuming a task-level distribution consistency to better transfer the adversarial defense knowledge. As for the second question, within each few-shot task, we design two kinds of distribution consistency criteria to narrow the distribution gap between clean and adversarial examples from the feature-wise and prediction-wise perspectives, respectively. Extensive experiments demonstrate that the proposed framework can effectively make the existing few-shot models robust against adversarial attacks. Code is available at https://github.com/WenbinLee/DefensiveFSL.git.
Wenbin Li 0006, Lei Wang 0001, Xingxing Zhang 0001, Lei Qi 0001, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Global- and local-aware feature augmentation with semantic orthogonality for few-shot image classification
Boyao Shi, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
Pattern Recognit.3
2023 A Multilayer Framework for Online Metric Learning
abstract
Online metric learning (OML) has been widely applied in classification and retrieval. It can automatically learn a suitable metric from data by restricting similar instances to be separated from dissimilar instances with a given margin. However, the existing OML algorithms have limited performance in real-world classifications, especially, when data distributions are complex. To this end, this article proposes a multilayer framework for OML to capture the nonlinear similarities among instances. Different from the traditional OML, which can only learn one metric space, the proposed multilayer OML (MLOML) takes an OML algorithm as a metric layer and learns multiple hierarchical metric spaces, where each metric layer follows a nonlinear layer for the complicated data distribution. Moreover, the forward propagation (FP) strategy and backward propagation (BP) strategy are employed to train the hierarchical metric layers. To build a metric layer of the proposed MLOML, a new Mahalanobis-based OML (MOML) algorithm is presented based on the passive-aggressive strategy and one-pass triplet construction strategy. Furthermore, in a progressively and nonlinearly learning way, MLOML has a stronger learning ability than traditional OML in the case of limited available training data. To make the learning process more explainable and theoretically guaranteed, theoretical analysis is provided. The proposed MLOML enjoys several nice properties, indeed learns a metric progressively, and performs better on the benchmark datasets. Extensive experiments with different settings have been conducted to verify these properties of the proposed MLOML.
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Adversarial Camera Alignment Network for Unsupervised Cross-Camera Person Re-Identification
abstract
In person re-identification (Re-ID), supervised methods usually need a large amount of expensive label information, while unsupervised ones are still unable to deliver satisfactory identification performance. In this paper, we introduce a novel person Re-ID task called unsupervised cross-camera person Re-ID, which only needs the within-camera (intra-camera) label information but not cross-camera (inter-camera) labels which are more expensive to obtain. In real-world applications, the intra-camera label information can be easily captured by tracking algorithms and few manual annotations. In this situation, the main challenge becomes the distribution discrepancy across different camera views, caused by the various body pose, occlusion, image resolution, illumination conditions, and background noises in different cameras. To address this situation, we propose a novel Adversarial Camera Alignment Network (ACAN) for unsupervised cross-camera person Re-ID. It consists of the camera-alignment task and the supervised within-camera learning task. To achieve the camera alignment, we develop a Multi-Camera Adversarial Learning (MCAL) to map images of different cameras into a shared subspace. Particularly, we investigate two different schemes, including the existing GRL (i.e., gradient reversal layer) scheme and the proposed scheme called “other camera equiprobability” (OCE), to conduct the multi-camera adversarial task. Based on this shared subspace, we then leverage the within-camera labels to train the network. Extensive experiments on five large-scale datasets demonstrate the superiority of ACAN over multiple state-of-the-art unsupervised methods that take advantage of labeled source domains and generated images by GAN-based models. In particular, we verify that the proposed multi-camera adversarial task does contribute to the significant improvement.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Xin Geng 0001, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 CAST: Learning Both Geometric and Texture Style Transfers for Effective Caricature Generation
abstract
Given a photo of a subject, ability to generate a caricature image that captures distinct characteristics of the subject but with certain exaggeration of their prominent features is of fundamental importance to image processing and facial recognition. There are two main challenges in this task: shape exaggeration and style transfer. The former morphs and exaggerates key facial features of the subject, while the latter generates caricature images in a certain artistic style. In this paper, we propose a CAricature Style Transfer (CAST) framework for caricature generation. There are two modules in the proposed framework. The first is a geometric warping module. Different from the existing style transfer methods, we incorporate the Whitening and Coloring Transformation (WCT) in the geometric style transfer. The WCT is learned on photo and caricature landmarks or the caricature landmark space of a specific artist and is capable of transforming input photo landmarks to caricature landmarks. The second module is a texture style rendering module. We propose a new style transfer method by considering a semantic region-aligned style transfer via affinity constraint. Given a reference caricature image as the style reference, this module is capable of transferring styles between the same or similar semantic regions in caricatures and photos. Furthermore, it can transfer visual attributes of the reference caricatures (such as mouth shape and expressions) to the output caricatures. Experiments have shown desirable effects of the proposed method in transferring both the geometric and artistic texture styles of caricatures. Both qualitative and quantitative results show that the CAST framework is more effective compared than the state-of-the-art caricature generation methods.
Jing Huo, Xiangde Liu, Wenbin Li 0006, Yang Gao 0001, Hujun Yin, Jiebo Luo 0001
IEEE Trans. Image Process.1
2022 CariMe: Unpaired Caricature Generation With Multiple Exaggerations
abstract
Caricature generation aims to translate real photos into caricatures with artistic styles and shape exaggerations while maintaining the identity of the subject. Different from generic image-to-image translation, drawing caricatures automatically is a more challenging task due to the existence of various spatial deformations. Previous caricature generation methods are obsessed with predicting definite image warping from a given photo while ignoring the intrinsic representation and distribution of geometric exaggerations in caricatures. This limits their ability on diverse exaggeration generation. In this paper, we generalize the caricature generation problem from instance-level warping prediction to distribution-level deformation modeling. Based on this assumption, we present the first exploration forunpaired CARIcature generation with Multiple Exaggerations (CariMe). Technically, we propose a Multi-exaggeration Warper network to learn the distribution-level mapping from photos to facial exaggerations. This makes it possible to generate diverse and reasonable exaggerations from randomly sampled warp codes given one input photo. To better represent the facial exaggeration and produce fine-grained warping, a deformation-field-based warping method is also proposed, which captures more detailed exaggerations than previous point-based warping methods. Experiments and two perceptual studies prove the superiority of our method comparing with other state-of-the-art methods, showing the improvement of our work on caricature generation. The source code is available athttps://github.com/edward3862/CariMe-pytorch.
Zheng Gu 0001, Chuanqi Dong, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Multim.3
2021 LoFGAN: Fusing Local Representations for Few-shot Image Generation
abstract
Given only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different images from a global perspective, making these works suffer from poor generation quality and diversity. To tackle this problem, we propose a novel Local-Fusion Generative Adversarial Network (LoFGAN) for fewshot image generation. Instead of using these available images as a whole, we first randomly divide them into a base image and several reference images. Next, LoFGAN matches local representations between the base and reference images based on semantic similarities, and replaces the local features with the closest related local features. In this way, LoFGAN can produce more realistic and diverse images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. Furthermore, a local reconstruction loss is also proposed, which can provide better training stability and generation quality. We conduct extensive experiments on three datasets, which successfully demonstrates the effectiveness of our proposed method for few-shot image generation and downstream visual applications with limited data. Code is available at https://github.com/edward3862/LoFGAN-pytorch.
Zheng Gu 0001, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001
ICCV3
2021 Manifold Alignment for Semantically Aligned Style Transfer
abstract
Most existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of local style patterns, where algorithms are designed to swap similar local features of content and style images. However, the limitation of these existing methods is that they neglect the semantic structure of the content image which may lead to corrupted content structure in the output. In this paper, we make a new assumption that image features from the same semantic region form a manifold and an image with multiple semantic regions follows a multi-manifold distribution. Based on this assumption, the style transfer problem is formulated as aligning two multi-manifold distributions and a Manifold Alignment based Style Transfer (MAST) framework is proposed. The proposed frame-work allows semantically similar regions between the output and the style image share similar style patterns. Moreover, the proposed manifold alignment method is flexible to allow user editing or using semantic segmentation maps as guidance for style transfer. To allow the method to be applicable to photorealistic style transfer, we propose a new adaptive weight skip connection network structure to preserve the content details. Extensive experiments verify the effectiveness of the proposed framework for both artistic and photorealistic style transfer. Code is available at https://github.com/NJUHuoJing/MAST.
Jing Huo, Shiyin Jin, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yinghuan Shi, Yang Gao 0001
ICCV1
2021 Learning 3D face reconstruction from a single sketch
Jing Wu 0004, Jing Huo, Yukun Lai, Yang Gao 0001
Graph. Model.3
2021 MetricUNet: Synergistic image- and voxel-level learning for precise prostate segmentation via online sampling
Kelei He, Chunfeng Lian, Ehsan Adeli-Mosabbeb, Jing Huo, Yang Gao 0001, Bing Zhang 0012, Dinggang Shen
Medical Image Anal.4
2021 Local descriptor-based multi-prototype network for few-shot Learning
Zhangkai Wu, Wenbin Li 0006, Jing Huo, Yang Gao 0001
Pattern Recognit.4
2021 MW-GAN: Multi-Warping GAN for Caricature Generation With Multi-Style Geometric Exaggeration
abstract
Given an input face photo, the goal of caricature generation is to produce stylized, exaggerated caricatures that share the same identity as the photo. It requires simultaneous style transfer and shape exaggeration with rich diversity, and meanwhile preserving the identity of the input. To address this challenging problem, we propose a novel framework called Multi-Warping GAN (MW-GAN), including a style network and a geometric network that are designed to conduct style transfer and geometric exaggeration respectively. We bridge the gap between the style/landmark space and their corresponding latent code spaces by a dual way design, so as to generate caricatures with arbitrary styles and geometric exaggeration, which can be specified either through random sampling of latent code or from a given caricature sample. Besides, we apply identity preserving loss to both image space and landmark space, leading to a great improvement in quality of generated caricatures. Experiments show that caricatures generated by MW-GAN have better quality than existing methods.
Haodi Hou, Jing Huo, Jing Wu 0004, Yukun Lai, Yang Gao 0001
IEEE Trans. Image Process.2
2021 GreyReID: A Novel Two-stream Deep Framework with RGB-grey Information for Person Re-identification
abstract
In this article, we observe that most false positive images (i.e., different identities with query images) in the top ranking list usually have the similar color information with the query image in person re-identification (Re-ID). Meanwhile, when we use the greyscale images generated from RGB images to conduct the person Re-ID task, some hard query images can obtain better performance compared with using RGB images. Therefore, RGB and greyscale images seem to be complementary to each other for person Re-ID. In this article, we aim to utilize both RGB and greyscale images to improve the person Re-ID performance. To this end, we propose a novel two-stream deep neural network with RGB-grey information, which can effectively fuse RGB and greyscale feature representations to enhance the generalization ability of Re-ID. First, we convert RGB images to greyscale images in each training batch. Based on these RGB and greyscale images, we train the RGB and greyscale branches, respectively. Second, to build up connections between RGB and greyscale branches, we merge the RGB and greyscale branches into a new joint branch. Finally, we concatenate the features of all three branches as the final feature representation for Re-ID. Moreover, in the training process, we adopt the joint learning scheme to simultaneously train each branch by the independent loss function, which can enhance the generalization ability of each branch. Besides, a global loss function is utilized to further fine-tune the final concatenated feature. The extensive experiments on multiple benchmark datasets fully show that the proposed method can outperform the state-of-the-art person Re-ID methods. Furthermore, using greyscale images can indeed improve the person Re-ID performance in the proposed deep framework.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression Ratio
abstract
Deep neural network compression is important and increasingly developed especially in resource-constrained environments, such as autonomous drones and wearable devices. Basically, we can easily and largely reduce the number of weights of a trained deep model by adopting a widely used model compression technique, e.g., pruning. In this way, two kinds of data are usually preserved for this compressed model, i.e., non-zero weights and meta-data, where meta-data is employed to help encode and decode these non-zero weights. Although we can obtain an ideally small number of non-zero weights through pruning, existing sparse matrix coding methods still need a much larger amount of meta-data (may several times larger than non-zero weights), which will be a severe bottleneck of the deploying of very deep models. To tackle this issue, we propose a layerwise sparse coding (LSC) method to maximize the compression ratio by extremely reducing the amount of meta-data. We first divide a sparse matrix into multiple small blocks and remove zero blocks, and then propose a novel signed relative index (SRI) algorithm to encode the remaining non-zero blocks (with much less meta-data). In addition, the proposed LSC performs parallel matrix multiplication without full decoding, while traditional methods cannot. Through extensive experiments, we demonstrate that LSC achieves substantial gains in pruned DNN compression (e.g., 51.03x compression ratio on ADMM-Lenet) and inference computation (i.e., time reduction and extremely less memory bandwidth), over state-of-the-art baselines.
Wenbin Li 0006, Jing Huo, Lili Yao, Yang Gao 0001
AAAI3
2020 Unsupervised Domain Attention Adaptation Network for Caricature Attribute Recognition
Kelei He, Jing Huo, Zheng Gu 0001, Yang Gao 0001
ECCV (8)3
2020 Learning Task-aware Local Representations for Few-shot Learning
abstract
Few-shot learning for visual recognition aims to adapt to novel unseen classes with only a few images. Recent work, especially the work based on low-level information, has achieved great progress. In these work, local representations (LRs) are typically employed, because LRs are more consistent among the seen and unseen classes. However, most of them are limited to an individual image-to-image or image-to-class measure manner, which cannot fully exploit the capabilities of LRs, especially in the context of a certain task. This paper proposes an Adaptive Task-aware Local Representations Network (ATL-Net) to address this limitation by introducing episodic attention, which can adaptively select the important local patches among the entire task, as the process of human recognition. We achieve much superior results on multiple benchmarks. On the miniImagenet, ATL-Net gains 0.93% and 0.88% improvements over the compared methods under the 5-way 1-shot and 5-shot settings. Moreover, ATL-Net can naturally tackle the problem that how to adaptively identify and weight the importance of different key local parts, which is the major concern of fine-grained recognition. Specifically, on the fine-grained dataset Stanford Dogs, ATL-Net outperforms the second best method with 5.39% and 9.69% gains under the 5-way 1-shot and 5-shot settings.
Chuanqi Dong, Wenbin Li 0006, Jing Huo, Zheng Gu 0001, Yang Gao 0001
IJCAI3
2020 Asymmetric Distribution Measure for Few-shot Learning
abstract
The core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https://github.com/WenbinLee/ADM.git.
Wenbin Li 0006, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001, Jiebo Luo 0001
IJCAI3
2020 Biased Feature Learning for Occlusion Invariant Face Recognition
abstract
To address the challenges posed by unknown occlusions, we propose a Biased Feature Learning (BFL) framework for occlusion-invariant face recognition. We first construct an extended dataset using a multi-scale data augmentation method. For model training, we modify the label loss to adjust the impact of normal and occluded samples. Further, we propose a biased guidance strategy to manipulate the optimization of a network so that the feature embedding space is dominated by non-occluded faces. BFL not only enhances the robustness of a network to unknown occlusions but also maintains or even improves its performance for normal faces. Experimental results demonstrate its superiority as well as the generalization capability with different network architectures and loss functions.
Chang-Bin Shao, Jing Huo, Lei Qi 0001, Zhenhua Feng 0001, Wenbin Li 0006, Chuanqi Dong, Yang Gao 0001
IJCAI2
2020 Cross-Domain Adversarial Autoencoder for Fine Grained Category Preserving Image Translation
abstract
Cross-domain image translation attempt to translate images from one domain to another domain, with the content of images preserved. Current approaches treat image's content as the underlying spatial structure, and translation only change image's style of color and texture. These methods can generate realistic results, but may not be able to preserve image's fine grained semantic category information and suffer from the lack of diversity in objects' shapes and viewing angles. In this paper, we propose the problem of fine grained category preserving image translation that aims at preserving image's fine grained category information in cross-domain translation. A novel framework called Cross-Domain Adversarial AutoEncoder (CDAAE) is proposed to solve the problem. CDAAE assumes that cross-domain images have shared content-latent-code space and separate style-latent-code spaces. The content latent code encodes image's basic category information, while the style latent code represents other domain-specific properties, including color, texture, shape, etc. Our experiments evaluate models from aspects of image's quality, diversity as well as category preserving ability, showing CDAAE's advantages over current methods. We also design an algorithm to apply CDAAE to domain adaptation. Experiments on benchmark datasets demonstrate that the proposed method achieves state-of-the-art results.
Haodi Hou, Jing Huo, Yang Gao 0001
IJCNN2
2020 CariGAN: Caricature generation through weakly paired adversarial learning
Wenbin Li 0006, Wei Xiong 0008, Haofu Liao, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
Neural Networks4
2020 Progressive Cross-Camera Soft-Label Learning for Semi-Supervised Person Re-Identification
abstract
In this paper, we focus on the semi-supervised person re-identification (Re-ID) case, which only has the intra-camera (within-camera) labels but not inter-camera (cross-camera) labels. In real-world applications, these intra-camera labels can be readily captured by tracking algorithms or few manual annotations, when compared with cross-camera labels. In this case, it is very difficult to explore the relationships between cross-camera persons in the training stage due to the lack of cross-camera label information. To deal with this issue, we propose a novel Progressive Cross-camera Soft-label Learning (PCSL) framework for the semi-supervised person Re-ID task, which can generate cross-camera soft-labels and utilize them to optimize the network. Concretely, we calculate an affinity matrix based on person-level features and adapt them to produce the similarities between cross-camera persons (i.e., cross-camera soft-labels). To exploit these soft-labels to train the network, we investigate the weighted cross-entropy loss and the weighted triplet loss from the classification and discrimination perspectives, respectively. Particularly, the proposed framework alternately generates progressive cross-camera soft-labels and gradually improves feature representations in the whole learning course. Extensive experiments on five large-scale benchmark datasets show that PCSL significantly outperforms the state-of-the-art unsupervised methods that employ labeled source domains or the images generated by the GANs-based models. Furthermore, the proposed method even has a competitive performance with respect to deep supervised Re-ID methods.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Distribution Consistency Based Covariance Metric Networks for Few-Shot Learning
abstract
Few-shot learning aims to recognize new concepts from very few examples. However, most of the existing few-shot learning methods mainly concentrate on the first-order statistic of concept representation or a fixed metric on the relation between a sample and a concept. In this work, we propose a novel end-to-end deep architecture, named Covariance Metric Networks (CovaMNet). The CovaMNet is designed to exploit both the covariance representation and covariance metric based on the distribution consistency for the few-shot classification tasks. Specifically, we construct an embedded local covariance representation to extract the second-order statistic information of each concept and describe the underlying distribution of this concept. Upon the covariance representation, we further define a new deep covariance metric to measure the consistency of distributions between query samples and new concepts. Furthermore, we employ the episodic training mechanism to train the entire network in an end-to-end manner from scratch. Extensive experiments in two tasks, generic few-shot image classification and fine-grained fewshot image classification, demonstrate the superiority of the proposed CovaMNet. The source code can be available from https://github.com/WenbinLee/CovaMNet.git.
Wenbin Li 0006, Jinglin Xu, Jing Huo, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001
AAAI3
2019 Revisiting Local Descriptor Based Image-To-Class Measure for Few-Shot Learning
abstract
Few-shot learning in image classification aims to learn a classifier to classify images when only few training examples are available for each class. Recent work has achieved promising classification performance, where an image-level feature based measure is usually used. In this paper, we argue that a measure at such a level may not be effective enough in light of the scarcity of examples in few-shot learning. Instead, we think a local descriptor based image-to-class measure should be taken, inspired by its surprising success in the heydays of local invariant features. Specifically, building upon the recent episodic training mechanism, we propose a Deep Nearest Neighbor Neural Network (DN4 in short) and train it in an end-to-end manner. Its key difference from the literature is the replacement of the image-level feature based measure in the final layer by a local descriptor based image-to-class measure. This measure is conducted online via a k-nearest neighbor search over the deep local descriptors of convolutional feature maps. The proposed DN4 not only learns the optimal deep local descriptors for the image-to-class measure, but also utilizes the higher efficiency of such a measure in the case of example scarcity, thanks to the exchangeability of visual patterns across the images in the same class. Our work leads to a simple, effective, and computationally efficient framework for few-shot learning. Experimental study on benchmark datasets consistently shows its superiority over the related state-of-the-art, with the largest absolute improvement of 17% over the next best. The source code can be available from https://github.com/WenbinLee/DN4.git.
Wenbin Li 0006, Lei Wang 0001, Jinglin Xu, Jing Huo, Yang Gao 0001, Jiebo Luo 0001
CVPR4
2019 A Novel Unsupervised Camera-Aware Domain Adaptation Framework for Person Re-Identification
abstract
Unsupervised cross-domain person re-identification (Re-ID) faces two key issues. One is the data distribution discrepancy between source and target domains, and the other is the lack of discriminative information in target domain. From the perspective of representation learning, this paper proposes a novel end-to-end deep domain adaptation framework to address them. For the first issue, we highlight the presence of camera-level sub-domains as a unique characteristic in person Re-ID, and develop a “camera-aware” domain adaptation method via adversarial learning. With this method, the learned representation reduces distribution discrepancy not only between source and target domains but also across all cameras. For the second issue, we exploit the temporal continuity in each camera of target domain to create discriminative information. This is implemented by dynamically generating online triplets within each batch, in order to maximally take advantage of the steadily improved representation in training process. Together, the above two methods give rise to a new unsupervised domain adaptation framework for person Re-ID. Extensive experiments and ablation studies conducted on benchmark datasets demonstrate its superiority and interesting properties.
Lei Qi 0001, Lei Wang 0001, Jing Huo, Luping Zhou, Yinghuan Shi, Yang Gao 0001
ICCV3
2019 A Mask Based Deep Ranking Neural Network for Person Retrieval
abstract
Person retrieval faces many challenges including cluttered background, appearance variations (e.g., illumination, pose, occlusion) among different camera views and the similarity among different person's images. To address these issues, we put forward a novel mask based deep ranking neural network with a skipped fusing layer. Firstly, to alleviate the problem of cluttered background, masked images with only the foreground regions are incorporated as input in the proposed neural network. Secondly, to reduce the impact of the appearance variations, the multi-layer fusion scheme is developed to obtain more discriminative fine-grained information. Lastly, considering person retrieval is a special image retrieval task, we propose a novel ranking loss to optimize the whole network. The proposed ranking loss can further mitigate the interference problem of similar negative samples when producing ranking results. The extensive experiments validate the superiority of the proposed method compared with the state-of-the-art methods on many benchmark datasets.
Lei Qi 0001, Jing Huo, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001
ICME2
2019 A Novel Deep Multi-Modal Feature Fusion Method for Celebrity Video Identification
abstract
In this paper, we develop a novel multi-modal feature fusion method for the 2019 iQIYI Celebrity Video Identification Challenge, which is held in conjunction with ACM MM 2019. The purpose of this challenge is to retrieve all the video clips of a given identity in the testing set. In this challenge, the multi-modal features of a celebrity are encouraged to be combined for a promising performance, such as face features, head features, body features, and audio features. As we know, the features from different modalities usually have their own influences on the results. To achieve better results, a novel weighted multi-modal feature fusion method is designed to obtain the final feature representation. After many experimental verification, we found that different feature fusion weights for training and testing make the method robust to multi-modal person identification. Experiments on the iQIYI-VID-2019 dataset show that our multi-modal feature fusion strategy effectively improves the accuracy of person identification. Specifically, for competition, we use a single model to get the result of 0.8952 in mAP, which ranks TOP-5 among all the competitive results.
Jianrong Chen, Jing Huo, Yinghuan Shi, Yang Gao 0001
ACM Multimedia4
2019 DeepMEF: A Deep Model Ensemble Framework for Video Based Multi-modal Person Identification
abstract
The goal of video based multi-modal person identification is to identify a person of interest using multi-modal video features, such as person's face, body, audio or head features. This task is challenging due to many factors, for example, variant body or face poses, poor face image quality, low frame resolution, etc. To address these problems, we propose a deep model ensemble framework, namely DeepMEF. Specifically, the proposed framework includes three novel modules, i.e., the video feature fusion module, the multi-modal feature fusion module and the model ensemble module. The first and second module form the basic deep model for ensemble, with the video feature fusion module fuses facial features from different frames as one. Then the multi-modal feature fusion module further fuses the face feature and features of other modalities for identification. In this work, we adopt the scene feature extracted by ourselves as the additional input of the multi-modal module. At last, the model ensemble module promotes the overall performance by combining the predictions of multiple multi-modal learners. The proposed method achieves a competitive result of 89.86% in mAP on the iQIYI-VID-2019 dataset, which helps us win the third place in the 2019 iQIYI Celebrity Video Identification Challenge.
Chuanqi Dong, Zheng Gu 0001, Zhonghao Huang, Jing Huo, Yang Gao 0001
ACM Multimedia5
2019 NeoLOD: A Novel Generalized Coupled Local Outlier Detection Model Embedded Non-IID Similarity Metric
Yang Gao 0001, Jing Huo, Xiaolong Qi
PAKDD (1)3
2019 MIDCN: A Multiple Instance Deep Convolutional Network for Image Classification
Kelei He, Jing Huo, Yinghuan Shi, Yang Gao 0001, Dinggang Shen
PRICAI (1)2
2018 A Joint Local and Global Deep Metric Learning Method for Caricature Recognition
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001
ACCV (4)2
2018 WebCaricature: a benchmark for caricature recognition
Jing Huo, Wenbin Li 0006, Yinghuan Shi, Yang Gao 0001, Hujun Yin
BMVC1
2018 Joint Multi-field Siamese Recurrent Neural Network for Entity Resolution
Lei Qi 0001, Jing Huo, Hao Wang 0013, Yang Gao 0001
PRICAI3
2018 OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi
Pattern Recognit.5
2018 Heterogeneous Face Recognition by Margin-Based Cross-Modality Metric Learning
abstract
Heterogeneous face recognition deals with matching face images from different modalities or sources. The main challenge lies in cross-modal differences and variations and the goal is to make cross-modality separation among subjects. A margin-based cross-modality metric learning (MCM2L) method is proposed to address the problem. A cross-modality metric is defined in a common subspace where samples of two different modalities are mapped and measured. The objective is to learn such metrics that satisfy the following two constraints. The first minimizes pairwise, intrapersonal cross-modality distances. The second forces a margin between subject specific intrapersonal and interpersonal cross-modality distances. This is achieved by defining a hinge loss on triplet-based distance constraints for efficient optimization. It allows the proposed method to focus more on optimizing distances of those subjects whose intrapersonal and interpersonal distances are hard to separate. The proposed method is further extended to a kernelized MCM2L (KMCM2L). Both methods have been evaluated on an ID card face dataset and two other cross-modality benchmark datasets. Various feature extraction methods have also been incorporated in the study, including recent deep learned features. In extensive experiments and comparisons with the state-of-the-art methods, the MCM2L and KMCM2L methods achieved marked improvements in most cases.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
IEEE Trans. Cybern.1
2018 Cross-Modal Metric Learning for AUC Optimization
abstract
Cross-modal metric learning (CML) deals with learning distance functions for cross-modal data matching. The existing methods mostly focus on minimizing a loss defined on sample pairs. However, the numbers of intraclass and interclass sample pairs can be highly imbalanced in many applications, and this can lead to deteriorating or unsatisfactory performances. The area under the receiver operating characteristic curve (AUC) is a more meaningful performance measure for the imbalanced distribution problem. To tackle the problem as well as to make samples from different modalities directly comparable, a CML method is presented by directly maximizing AUC. The method can be further extended to focus on optimizing partial AUC (pAUC), which is the AUC between two specific false positive rates (FPRs). This is particularly useful in certain applications where only the performances assessed within predefined false positive ranges are critical. The proposed method is formulated as a log-determinant regularized semidefinite optimization problem. For efficient optimization, a minibatch proximal point algorithm is developed. The algorithm is experimentally verified stable with the size of sampled pairs that form a minibatch at each iteration. Several data sets have been used in evaluation, including three cross-modal data sets on face recognition under various scenarios and a single modal data set, the Labeled Faces in the Wild. Results demonstrate the effectiveness of the proposed methods and marked improvements over the existing methods. Specifically, pAUC-optimized CML proves to be more competitive for performance measures such as Rank-1 and verification rate at FPR = 0.1%.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Hujun Yin
IEEE Trans. Neural Networks Learn. Syst.1
2016 Ensemble of Sparse Cross-Modal Metrics for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition aims to identify or verify person identity by matching facial images of different modalities. In practice, it is known that its performance is highly influenced by modality inconsistency, appearance occlusions, illumination variations and expressions. In this paper, a new method named as ensemble of sparse cross-modal metrics is proposed for tackling these challenging issues. In particular, a weak sparse cross-modal metric learning method is firstly developed to measure distances between samples of two modalities. It learns to adjust rank-one cross-modal metrics to satisfy two sets of triplet based cross-modal distance constraints in a compact form. Meanwhile, a group based feature selection is performed to enforce that features in the same position of two modalities are selected simultaneously. By neglecting features that attribute to "noise" in the face regions (eye glasses, expressions and so on), the performance of learned weak metrics can be markedly improved. Finally, an ensemble framework is incorporated to combine the results of differently learned sparse metrics into a strong one. Extensive experiments on various face datasets demonstrate the benefit of such feature selection especially when heavy occlusions exist. The proposed ensemble metric learning has been shown superiority over several state-of-the-art methods in heterogeneous face recognition.
Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin
ACM Multimedia1
2014 A Novel Ego-Centered Academic Community Detection Approach via Factor Graph Model
Yusheng Jia, Yang Gao 0001, Wanqi Yang, Jing Huo, Yinghuan Shi
IDEAL4
2014 Multi-Instance Dictionary Learning for Detecting Abnormal Events in Surveillance Videos
abstract
In this paper, a novel method termed Multi-Instance Dictionary Learning (MIDL) is presented for detecting abnormal events in crowded video scenes. With respect to multi-instance learning, each event (video clip) in videos is modeled as a bag containing several sub-events (local observations); while each sub-event is regarded as an instance. The MIDL jointly learns a dictionary for sparse representations of sub-events (instances) and multi-instance classifiers for classifying events into normal or abnormal. We further adopt three different multi-instance models, yielding the Max-Pooling-based MIDL (MP-MIDL), Instance-based MIDL (Inst-MIDL) and Bag-based MIDL (Bag-MIDL), for detecting both global and local abnormalities. The MP-MIDL classifies observed events by using bag features extracted via max-pooling over sparse representations. The Inst-MIDL and Bag-MIDL classify observed events by the predicted values of corresponding instances. The proposed MIDL is evaluated and compared with the state-of-the-art methods for abnormal event detection on the UMN (for global abnormalities) and the UCSD (for local abnormalities) datasets and results show that the proposed MP-MIDL and Bag-MIDL achieve either comparable or improved detection performances. The proposed MIDL method is also compared with other multi-instance learning methods on the task and superior results are obtained by the MP-MIDL scheme.
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
Int. J. Neural Syst.1
2012 Abnormal Event Detection via Multi-Instance Dictionary Learning
Jing Huo, Yang Gao 0001, Wanqi Yang, Hujun Yin
IDEAL1