Yue Deng 0001

dblp:35/8109-1 · DBLP profile ↗
← Back
56ranked-venue papers
15as first author
29since 2021 · last 2026
0000-0003-2871-8922ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Constrained Particle Seeking: Solving Diffusion Inverse Problems with Just Forward Passes
abstract
Diffusion models have gained prominence as powerful generative tools for solving inverse problems due to their ability to model complex data distributions. However, existing methods typically rely on complete knowledge of the forward observation process to compute gradients for guided sampling, limiting their applicability in scenarios where such information is unavailable. In this work, we introduce *Constrained Particle Seeking (CPS)*, a novel gradient-free approach that leverages all candidate particle information to actively search for the optimal particle while incorporating constraints aligned with high-density regions of the unconditional prior. Unlike previous methods that passively select promising candidates, CPS reformulates the inverse problem as a constrained optimization task, enabling more flexible and efficient particle seeking. We demonstrate that CPS can effectively solve both image and scientific inverse problems, achieving results comparable to gradient-based methods while significantly outperforming gradient-free alternatives.
Hongkun Dou, Zike Chen, Hongjue Li, Yue Deng 0001
AAAI6
2026 Pseudo-Spiking Neurons: A Noise-Based Training Framework for Heterogeneous-Latency Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) promise significant energy efficiency by processing information via sparse, event-driven spikes. However, realizing this potential is hindered by the conventional use of a rigid, uniform timestep, T. This constraint imposes a challenging trade-off between accuracy and latency, while also incurring the prohibitive training costs of Backpropagation Through Time (BPTT). To overcome this limitation, we introduce the Pseudo-Spiking Neuron (PseudoSN), a novel training proxy that conceptualizes latency as an intrinsic, learnable parameter for each neuron. Building on the efficiency of rate-based methods, the PseudoSN models temporal dynamics in a single, BPTT-free pass. It employs a learnable probabilistic noise scheme to emulate the discretization effects of spike generation (e.g., clipping and quantization), making the neuron-specific timestep—and thus latency—directly optimizable via backpropagation. Integrated into a hardware-aware objective, our framework trains heterogeneous-latency SNNs that autonomously learn to optimize the trade-offs among accuracy, latency and energy, establishing a new state-of-the-art on major benchmarks.
Hongjue Li, Yue Deng 0001, Wen Yao 0001
AAAI4
2026 Bi-Spectrum Distillation: Addressing Spectral Mismatch in ANN-SNN Knowledge Transfer
abstract
Knowledge distillation from Artificial Neural Networks (ANNs) to Spiking Neural Networks (SNNs) is a prominent training paradigm. However, its efficacy is fundamentally limited by a spectral mismatch: SNNs, with their intrinsic low-pass filtering characteristics, struggle to learn high-frequency details from their ANN teachers, creating a bottleneck in knowledge transfer at both the feature and logit levels. To address this, we propose Bi-Spectrum Distillation (BSD), a novel framework that mitigates the mismatch from two complementary perspectives. First, at the feature level, our Spectral Residual Distillation (SRD) enhances the student SNN's features with a parameter-efficient, learnable filter that adaptively compensates for high-frequency information loss, which transforms the student's output to better match the teacher's rich spectral target. Second, at the logits level, our Spectral Semantic Distillation (SSD) enhances fine-grained classification by distilling high-frequency components from teacher-ordered logits. Extensive experiments on CIFAR-10/100, ImageNet, and CIFAR10-DVS demonstrate that BSD achieves new state-of-the-art performance across both CNN and Transformer-based SNNs, validating its effectiveness and broad applicability.
Wen Yao 0001, Yue Deng 0001, Hongjue Li
AAAI4
2026 You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling With Gradient Shortcuts
abstract
Diffusion models (DMs) have recently demonstrated remarkable success in modeling large-scale data distributions. However, many downstream tasks require guiding the generated content based on specific differentiable metrics, typically necessitating backpropagation during the generation process. This approach is computationally expensive, as generating with DMs often demands tens to hundreds of recursive network calls, resulting in high memory usage and significant time consumption. In this paper, we propose a more efficient alternative that approaches the problem from the perspective of parallel denoising. We show that full backpropagation throughout the entire generation process is unnecessary. The downstream metrics can be optimized by retaining the computational graph of only one step during generation, thus providing a shortcut for gradient propagation. The resulting method, which we call Shortcut Diffusion Optimization (SDO), is generic, high-performance, and computationally lightweight, capable of optimizing all parameter types in diffusion sampling. We demonstrate the effectiveness of SDO on several real-world tasks, including controlling generation by optimizing latent and aligning the DMs by fine-tuning network parameters. Compared to full backpropagation, our approach reduces computational costs by $\sim\! 90\%$∼90% while maintaining superior performance. Code is available at https://github.com/deng-ai-lab/SDO.
Hongkun Dou, Xingyu Jiang 0003, Hongjue Li, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Global Modeling Matters: A Fast, Lightweight, and Effective Baseline for Efficient Image Restoration
abstract
Natural image quality is often degraded by adverse weather conditions, significantly impairing the performance of downstream tasks. Image restoration has emerged as a core solution to this challenge and has been widely discussed in the literature. Although recent transformer-based approaches have made remarkable progress in image restoration, their increasing system complexity poses significant challenges for real-time processing, particularly in real-world deployment scenarios. To this end, most existing methods attempt to simplify the self-attention mechanism, such as by channel self-attention or state space model. However, these methods primarily focus on network architecture while neglecting the inherent characteristics of image restoration itself. In this context, we explore a pyramid Wavelet-Fourier iterative pipeline to demonstrate the potential of Wavelet-Fourier processing for image restoration. Inspired by the above findings, we propose a novel and efficient restoration baseline, named Pyramid Wavelet-Fourier Network (PW-FNet). Specifically, PW-FNet features two key design principles: 1) at the inter-block level, integrates a pyramid wavelet-based multi-input multi-output structure to achieve multi-scale and multi-frequency bands decomposition; and 2) at the intra-block level, incorporates Fourier transforms as an efficient alternative to self-attention mechanisms, effectively reducing computational complexity while preserving global modeling capability. Extensive experiments on tasks such as image deraining, raindrop removal, image super-resolution, motion deblurring, image dehazing, image desnowing and underwater/low-light enhancement demonstrate that PW-FNet not only surpasses state-of-the-art methods in restoration quality but also achieves superior efficiency, with significantly reduced parameter size, computational cost and inference time. The code is available at: https://github.com/deng-ai-lab/PW-FNet.
Xingyu Jiang 0003, Ning Gao 0004, Hongkun Dou, Xiuhui Zhang, Xiaoqing Zhong, Yue Deng 0001, Hongjue Li
IEEE Trans. Image Process.6
2026 MTRAG: Multi-Target Referring and Grounding via Hybrid Semantic-Spatial Integration
abstract
Fine-grained visual referring and grounding are critical for enhancing scene understanding and enabling various real-world vision-language applications. Although recent studies have extended multimodal large language models (MLLMs) to these tasks, they still face significant challenges in fine-grained multi-target scenarios. To address this, we propose MTRAG, a pixel-level multi-target referring and grounding framework that leverages semantic-spatial collaboration. Specifically, we introduce a Channel Extension Mechanism (CEM) that enables a global image encoder to extract global semantics and multi-region representations while retaining background context, without extra region feature extractors. Moreover, we introduce a grounding branch for pixel-level grounding and design a Hybrid Adapter (HA) to fuse semantic features from the MLLM branch with spatial information from the grounding branch, thereby enhancing the semantic-spatial alignment. For training, we meticulously curate MTRAG-D, a dataset comprising single- and multi-target referring and grounding samples derived from existing datasets and newly synthesized free-form multi-target referring instruction-following data. We also present MTR-Bench, a benchmark for systematic evaluation of multi-target referring. Extensive experiments across five core tasks, including single- and multi-target referring and grounding as well as image-level captioning, show that MTRAG consistently outperforms strong baselines on both multi- and single-target tasks, while maintaining competitive image-level understanding. The code is available at https://github.com/deng-ai-lab/MTRAG.
Yili Ren, Jinyang Du, Qianxiao Su, Yue Deng 0001, Hongjue Li
IEEE Trans. Image Process.5
2025 DPoser-X: Diffusion Model as Robust 3D Whole-Body Human Pose Prior
abstract
We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limitations, we introduce a Diffusion model as body Pose prior (DPoser) and extend it to DPoser-X for expressive whole-body human pose modeling. Our approach unifies various pose-centric tasks as inverse problems, solving them through variational diffusion sampling. To enhance performance on downstream applications, we introduce a novel truncated timestep scheduling method specifically designed for pose data characteristics. We also propose a masked training mechanism that effectively combines whole-body and part-specific datasets, enabling our model to capture interdependencies between body parts while avoiding overfitting to specific actions. Extensive experiments demonstrate DPoser-X's robustness and versatility across multiple benchmarks for body, hand, face, and full-body pose modeling. Our model consistently outperforms state-of-the-art alternatives, establishing a new benchmark for whole-body human pose prior modeling.
Junzhe Lu 0001, Hongkun Dou, Ailing Zeng, Yue Deng 0001, Zhongang Cai, Lei Yang 0059, Yulun Zhang 0001, Haoqian Wang, Ziwei Liu 0002
ICCV5
2025 Hybrid Regularization Improves Diffusion-based Inverse Problem Solving
abstract
Diffusion models, recognized for their effectiveness as generative priors, have become essential tools for addressing a wide range of visual challenges. Recently, there has been a surge of interest in leveraging Denoising processes for Regularization (DR) to solve inverse problems. However, existing methods often face issues such as mode collapse, which results in excessive smoothing and diminished diversity. In this study, we perform a comprehensive analysis to pinpoint the root causes of gradient inaccuracies inherent in DR. Drawing on insights from diffusion model distillation, we propose a novel approach called Consistency Regularization (CR), which provides stabilized gradients without the need for ODE simulations. Building on this, we introduce Hybrid Regularization (HR), a unified framework that combines the strengths of both DR and CR, harnessing their synergistic potential. Our approach proves to be effective across a broad spectrum of inverse problems, encompassing both linear and nonlinear scenarios, as well as various measurement noise statistics. Experimental evaluations on benchmark datasets, including FFHQ and ImageNet, demonstrate that our proposed framework not only achieves highly competitive results compared to state-of-the-art methods but also offers significant reductions in wall-clock time and memory consumption.
Hongkun Dou, Jinyang Du, Wen Yao 0001, Yue Deng 0001
ICLR6
2025 Value-aligned Behavior Cloning for Offline Reinforcement Learning via Bi-level Optimization
abstract
Offline reinforcement learning (RL) aims to optimize policies under pre-collected data, without requiring any further interactions with the environment. Derived from imitation learning, Behavior cloning (BC) is extensively utilized in offline RL for its simplicity and effectiveness. Although BC inherently avoids out-of-distribution deviations, it lacks the ability to discern between high and low-quality data, potentially leading to sub-optimal performance when facing with poor-quality data. Current offline RL algorithms attempt to enhance BC by incorporating value estimation, yet often struggle to effectively balance these two critical components, specifically the alignment between the behavior policy and the pre-trained value estimations under in-sample offline data. To address this challenge, we propose the Value-aligned Behavior Cloning via Bi-level Optimization (VACO), a novel bi-level framework that seamlessly integrates an inner loop for weighted supervised behavior cloning (BC) with an outer loop dedicated to value alignment. In this framework, the inner loop employs a meta-scoring network to evaluate and appropriately weight each training sample, while the outer loop maximizes value estimation for alignment with controlled noise to facilitate limited exploration. This bi-level structure allows VACO to identify the optimal weighted BC policy, ultimately maximizing the expected estimated return conditioned on the learned value function. We conduct a comprehensive evaluation of VACO across a variety of continuous control benchmarks in offline RL, where it consistently achieves superior performance compared to existing state-of-the-art methods.
Xingyu Jiang 0003, Ning Gao 0004, Xiuhui Zhang, Hongkun Dou, Yue Deng 0001
ICLR5
2025 Physics-aligned field reconstruction with diffusion bridge
abstract
The reconstruction of physical fields from sparse measurements is pivotal in both scientific research and engineering applications. Traditional methods are increasingly supplemented by deep learning models due to their efficacy in extracting features from data. However, except for the low accuracy on complex physical systems, these models often fail to comply with essential physical constraints, such as governing equations and boundary conditions. To overcome this limitation, we introduce a novel data-driven field reconstruction framework, termed the Physics-aligned Schr\"{o}dinger Bridge (PalSB). This framework leverages a diffusion bridge mechanism that is specifically tailored to align with physical constraints. The PalSB approach incorporates a dual-stage training process designed to address both local reconstruction mapping and global physical principles. Additionally, a boundary-aware sampling technique is implemented to ensure adherence to physical boundary conditions. We demonstrate the effectiveness of PalSB through its application to three complex nonlinear systems: cylinder flow from Particle Image Velocimetry experiments, two-dimensional turbulence, and a reaction-diffusion system. The results reveal that PalSB not only achieves higher accuracy but also exhibits enhanced compliance with physical constraints compared to existing methods. This highlights PalSB's capability to generate high-quality representations of intricate physical interactions, showcasing its potential for advancing field reconstruction techniques. The source code can be found at https://github.com/lzy12301/PalSB.
Hongkun Dou, Shen Fang, Wang Han, Yue Deng 0001
ICLR5
2025 Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction
abstract
Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatiotemporal degradations in underwater imaging, including caustics, flickering, attenuation, and backscattering, frequently result in inaccurate geometry and appearance in existing 3D reconstruction methods. While a few recent works have explored underwater degradation-aware reconstruction, they often address either spatial or temporal degradation alone, falling short in more real-world underwater scenarios where both types of degradation occur. We propose MarineSTD-GS, a novel 3D Gaussian Splatting-based framework that explicitly models both temporal and spatial degradations for realistic underwater scene reconstruction. Specifically, we introduce two paired Gaussian primitives: Intrinsic Gaussians represent the true scene, while Degraded Gaussians render the degraded observations. The color of each Degraded Gaussian is physically derived from its paired Intrinsic Gaussian via a Spatiotemporal Degradation Modeling (SDM) module, enabling self-supervised disentanglement of realistic appearance from degraded images. To ensure stable training and accurate geometry, we further propose a Depth-Guided Geometry Loss and a Multi-Stage Optimization strategy. We also construct a simulated benchmark with diverse spatial and temporal degradations and ground-truth appearances for comprehensive evaluation. Experiments on both simulated and real-world datasets show that MarineSTD-GS robustly handles spatiotemporal degradations and outperforms existing methods in novel view synthesis with realistic, water-free scene appearances.
Shaohua Liu 0003, Ning Gao 0004, Zuoya Gu, Hongkun Dou, Yue Deng 0001, Hongjue Li
ACM Multimedia5
2025 RF-Agent: Automated Reward Function Design via Language Agent Tree Search
abstract
Designing efficient reward functions for low-level control tasks is a challenging problem. Recent research aims to reduce reliance on expert experience by using Large Language Models (LLMs) with task information to generate dense reward functions. These methods typically rely on training results as feedback, iteratively generating new reward functions with greedy or evolutionary algorithms. However, they suffer from poor utilization of historical feedback and inefficient search, resulting in limited improvements in complex control tasks. To address this challenge, we propose RF-Agent, a framework that treats LLMs as language agents and frames reward function design as a sequential decision-making process, enhancing optimization through better contextual reasoning. RF-Agent integrates Monte Carlo Tree Search (MCTS) to manage the reward design and optimization process, leveraging the multi-stage contextual reasoning ability of LLM. This approach better utilizes historical information and improves search efficiency to identify promising reward functions. Outstanding experimental results in 17 diverse low-level control tasks demonstrate the effectiveness of our method.
Ning Gao 0004, Xiuhui Zhang, Xingyu Jiang 0003, Mukang You, Mohan Zhang, Yue Deng 0001
NeurIPS6
2025 Progress Reward Model for Reinforcement Learning via Large Language Models
abstract
Traditional reinforcement learning (RL) algorithms face significant limitations in handling long-term tasks with sparse rewards. Recent advancements have leveraged large language models (LLMs) to enhance RL by utilizing their world knowledge for task planning and reward generation. However, planning-based approaches often depend on pre-defined skill libraries and fail to optimize low-level control policies, while reward-based methods require extensive human feedback or exhaustive searching due to the complexity of tasks. In this paper, we propose the Progress Reward Model for RL (PRM4RL), a novel framework that integrates task planning and dense reward to enhance RL. For high-level planning, a complex task is decomposed into a series of simple manageable subtasks, with a subtask-oriented, fine-grained progress function designed to monitor task execution progress. For low-level reward generation, inspired by potential-based reward shaping, we use the progress function to construct a Progress Reward Model (PRM), providing theoretically grounded optimality and convergence guarantees, thereby enabling effective policy optimization. Experimental results on robotics control tasks demonstrate that our approach outperforms both LLM-based planning and reward methods, achieving state-of-the-art performance.
Xiuhui Zhang, Ning Gao 0004, Xingyu Jiang 0003, Yuheng Pan, Mohan Zhang, Yue Deng 0001
NeurIPS7
2025 Image-to-Image Bayesian Flow Networks With Structurally Informative Priors
abstract
Generative models represented by diffusion models have recently shown great potential in image generation. They usually use a reverse iteration process to map noise into the data. However, for many real-world applications such as image restoration and translation, the model input comes from a distribution that is not random noise, making it difficult for these models to adapt directly to these tasks. In this paper, we introduce Image-to-Image Bayesian Flow Networks (I2I-BFNs), a novel framework for general-purpose image-to-image translation (I2I) that operates within the parameter space of distributions. This method upholds Gaussian distributions over pixel intensities, refining distribution parameters through closed-form Bayesian inference, steered by the network's predictions for the target image. An essential aspect of our approach is the utilization of the conditional image as a robust prior parameter, initializing the translation process from a deterministic, clean image to reduce variance and produce interpretable generation. Additionally, we introduce a skip sampling technique that enhances the efficiency of I2I-BFNs, facilitating rapid translation in diverse image restoration and general I2I tasks. Our experimental evaluations showcase the model's competitive edge in various settings, underscoring its efficacy and adaptability. This work contributes new insights and opportunities for the large-scale development of efficient conditional generation systems.
Hongkun Dou, Jinyang Du, Xingyu Jiang 0003, Hongjue Li, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Image Process.6
2025 SEGSID: A Semantic-Guided Framework for Sonar Image Despeckling
abstract
Sonar imagery is substantially degraded by speckle noise, making the task of despeckling crucial for improving image quality. Self-supervised despeckling methods, represented by blind-spot networks (BSNs), have shown promise in this regard. However, these methods consistently face significant challenges due to the spatial correlation of speckle noise and the inherent information loss within BSNs. In this paper, we introduce SEGSID, a BSN-based, semantic-guided sonar despeckling framework designed to address these challenges. Specifically, the SEGSID framework primarily comprises a Receptive Field Augmentation (RFA) module and a Global Semantic Enhancement (GSE) module. To address the noise spatial correlation, the RFA module is crafted to strategically extract valuable local information while avoiding the exploitation of noise-correlated pixels. Concurrently, the GSE module extracts the global semantic information from entire images and injects it into the extracted local features. This enhances BSNs' ability to harness more comprehensive image information and compensates for their inherent information loss. Furthermore, to bolster efficiency, we employ knowledge distillation techniques to transfer the expertise from the trained SEGSID into a more streamlined network suitable for broader practical applications. Extensive experiments on three distinct sonar datasets demonstrate that SEGSID outperforms both traditional despeckling methods and state-of-the-art self-supervised despeckling techniques. The implementation is publicly accessible at https://github.com/deng-ai-lab/SEGSID.
Shaohua Liu 0003, Junzhe Lu 0001, Hongkun Dou, Yue Deng 0001
IEEE Trans. Image Process.5
2025 Score-Based Neural Processes
abstract
Neural processes (NPs) have recently emerged as a powerful meta-learning framework capable of making predictions based on an arbitrary number of context points. However, the learning of NPs and their variants is hindered by the need for explicit reliance on the log-likelihood of predictive distributions, which complicates the training process. To tackle this problem, we introduce score-based NP (SNP) models, drawing inspiration from recently developed score-based generative models (SGMs) that restore data from noise by reversing a perturbation process. With denoising score matching (DSM) techniques, the SNPs bypass the intractable log-likelihood calculations, learning parameterized score functions instead. We also demonstrate that score functions possess excellent attributes that enable us to represent a wide family of conditional distributions naturally. Moreover, as data points are inherently unordered, it is crucial to incorporate appropriate inductive biases into SNPs. To this end, we propose building blocks for parameterizing permutation equivariant score functions, which induce the SNPs with the desired properties. Through extensive experimentation on both synthetic and real-world datasets, our SNPs exhibit remarkable performance and outperform existing state-of-the-art NP approaches.
Hongkun Dou, Junzhe Lu 0001, Xiaoqing Zhong, Wen Yao 0001, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 A Quantum Spatial Graph Convolutional Neural Network Model on Quantum Circuits
abstract
This article proposes a quantum spatial graph convolutional neural network (QSGCN) model that is implementable on quantum circuits, providing a novel avenue to processing non-Euclidean type data based on the state-of-the-art parameterized quantum circuit (PQC) computing platforms. Four basic blocks are constructed to formulate the whole QSGCN model, including the quantum encoding, the quantum graph convolutional layer, the quantum graph pooling layer, and the network optimization. In particular, the trainability of the QSGCN model is analyzed through discussions on the barren plateau phenomenon. Simulation results from various types of graph data are presented to demonstrate the learning, generalization, and robustness capabilities of the proposed quantum neural network (QNN) model.
Qing Gao 0001, Maciej Ogorzalek, Jinhu Lü 0001, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Efficient Frequency-Domain Image Deraining with Contrastive Regularization
Ning Gao 0004, Xingyu Jiang 0003, Xiuhui Zhang, Yue Deng 0001
ECCV (41)4
2024 When Fast Fourier Transform Meets Transformer for Image Restoration
Xingyu Jiang 0003, Xiuhui Zhang, Ning Gao 0004, Yue Deng 0001
ECCV (45)4
2024 SPADIX: A Highly Efficient Accelerator for Solving 3-D Partial Differential Equations
abstract
Solving partial differential equations (PDEs) holds immense significance in numerous scientific and engineering fields. While analytical solutions to PDEs are often restricted to simple cases, numerical methods offer powerful techniques to approximate solutions for complex PDEs. Previous works have proposed customized accelerators to address the compute-and memory-intensive aspects of numerical PDE solvers. However, these approaches primarily focus on 2D PDEs and encounter challenges in scaling to support 3D PDEs due to increased complexity and computational demands. In this paper, we introduce Spadix, a highly efficient hardware accelerator designed for numerical 3D PDE solvers. Spadix leverages a customized Processing Element array architecture specifically tailored to the compute and data access patterns in 3D PDEs. The PE incorporates techniques such as temporal and spatial data reuse to minimize data accesses, enhancing overall performance and energy efficiency. Additionally, Spadix supports the checkerboard method for numerical PDE solvers, which exhibits a faster convergence rate compared to the Jacobi method without compromising parallelism. Our evaluation demonstrates that Spadix achieves an average 8.4× speedup with 9.2× energy reduction over NVIDIA RTX3090 GPU and a 9.7× speedup with 3.2× energy reduction over Alrescha, the state-of-the-art PDE-solving accelerator.
Yue Deng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 A Subgraph-Based Hierarchical Q-Learning Approach to Optimal Resource Scheduling for Complex Industrial Networks
abstract
This paper proposes a subgraph-based hierarchical Q-learning network (SgHQN) approach to solve the optimal resource scheduling problem for complex industrial networks. In the industrial network, each connection between two individual stations has limited communication bandwidth, while each station has limited computing and storage capability and is only accessible to its local information. The resource packages flowing within the industrial network are treated as agents that have different sizes and different levels of decision-making priority. This makes the industrial resource scheduling problem on the industrial network a multi-level decision-making problem with information asymmetry. Specifically, the resource packages with lower decision-making priority have knowledge of the decisions made by those with higher priority, but not vice versa. To solve this resource scheduling problem with information asymmetry, an SgHQN model is developed by exploiting partial observations. It is found that the proposed SgHQN can be used to solve resource scheduling problems for general industrial networks. Numerical experiments simulating industrial scheduling scenarios demonstrate the effectiveness and advantages of our method.
Kexin Zhang 0005, Qing Gao 0001, Jinhu Lü 0001, Maciej Ogorzalek, Yue Deng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Task-aware world model learning with meta weighting via bi-level optimization
abstract
Aligning the world model with the environment for the agent’s specific task is crucial in model-based reinforcement learning. While value-equivalent models may achieve better task awareness than maximum-likelihood models, they sacrifice a large amount of semantic information and face implementation issues. To combine the benefits of both types of models, we propose Task-aware Environment Modeling Pipeline with bi-level Optimization (TEMPO), a bi-level model learning framework that introduces an additional level of optimization on top of a maximum-likelihood model by incorporating a meta weighter network that weights each training sample. The meta weighter in the upper level learns to generate novel sample weights by minimizing a proposed task-aware model loss. The model in the lower level focuses on important samples while maintaining rich semantic information in state representations. We evaluate TEMPO on a variety of continuous and discrete control tasks from the DeepMind Control Suite and Atari video games. Our results demonstrate that TEMPO achieves state-of-the-art performance regarding asymptotic performance, training stability, and convergence speed.
Huining Yuan 0002, Hongkun Dou, Xingyu Jiang 0003, Yue Deng 0001
NeurIPS4
2023 Hyperspectral Image Instance Segmentation Using Spectral-Spatial Feature Pyramid Network
abstract
In recent years, hyperspectral image (HSI) classification and detection techniques based on deep learning have been widely applied to various aspects, such as environmental monitoring, urban planning, and energy surveys. As an important image content analysis method, instance segmentation can provide important support for the extraction of ground object information and monomeric application of HSI. This article introduces instance segmentation into HSI interpretation for the first time. In this article, we create the hyperspectral instance segmentation dataset (HS-ISD), which contains a total of 56 images, each with a size of$298\times301$and a number of channels of 48. More than 1000 architectural examples are annotated to apply to the research of HSI instance segmentation. In addition, considering that HSI contains rich spectral and spatial information, and the traditional instance segmentation network model cannot well utilize both types of information effectively, we propose the spectral–spatial feature pyramid network (Spectral–Spatial FPN). The Spectral–Spatial FPN can integrate multiscale spectral information and multiscale spatial information in the feature extraction stage through attention mechanism and bidirectional feature pyramid structure, so as to better improve the performance of the network model by spectral information and spatial information and realize the end-to-end instance segmentation of HSI. The experimental results conducted on the HS-ISD show that the proposed Spectral–Spatial FPN can achieve state-of-the-art results.
Leyuan Fang, Yinglong Yan, Jun Yue 0004, Yue Deng 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Toward the Vectorization of Hyperspectral Imagery
abstract
Hyperspectral images (HSIs) can provide rich spectral-spatial information that has been widely utilized in many fields, such as national defense, mineralogy and agriculture. Most of the recent HSI interpretation methods are conducted in the raster pattern, which results in high memory costs, amplification distortion, and difficulties in topological editing. To address this issue, a novel end-to-end vectorization framework is proposed, called as the HSI Vectorization Network (HSI-VecNet), which learns a vector representation from spectral-spatial information through cross-level interactions. Specifically, this framework integrates low-level geometry information and high-level semantic instance information, which consists of two branches: the HSI Semantic Instance Segmentation (HSIS) and the Spectral-Spatial Junction Prediction (SSJP). The HSIS conducts the raster-based classification and extracts the semantic information of each object in the HSI. In addition, the SSJP exploits spectral-spatial information to predict the positions of junctions in the HSI. The instance information of each object and the relations of junctions are then fused to vectorize the HSI. To verify the effectiveness of the proposed method, four hyperspectral datasets are vectorially labeled. Experimental results on these datasets demonstrate that the proposed end-to-end HSI-VecNet outperforms existing post-process vectorization methods. Our model and datasets will be made publicly available at https://github.com/yyyyll0ss/HSI-VecNet.
Leyuan Fang, Yinglong Yan, Jun Yue 0004, Yue Deng 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Efficient Layer Compression Without Pruning
abstract
Network pruning is one of the chief means for improving the computational efficiency of Deep Neural Networks (DNNs). Pruning-based methods generally discard network kernels, channels, or layers, which however inevitably will disrupt original well-learned network correlation and thus lead to performance degeneration. In this work, we propose an Efficient Layer Compression (ELC) approach to efficiently compress serial layers by decoupling and merging rather than pruning. Specifically, we first propose a novel decoupling module to decouple the layers, enabling us readily merge serial layers that include both nonlinear and convolutional layers. Then, the decoupled network is losslessly merged based on the equivalent conversion of the parameters. In this way, our ELC can effectively reduce the depth of the network without destroying the correlation of the convolutional layers. To our best knowledge, we are the first to exploit the mergeability of serial convolutional layers for lossless network layer compression. Experimental results conducted on two datasets demonstrate that our method retains superior performance with a FLOPs reduction of 74.1% for VGG-16 and 54.6% for ResNet-56, respectively. In addition, our ELC improves the inference speed by 2× on Jetson AGX Xavier edge device.
Jie Wu 0035, Dingshun Zhu, Leyuan Fang, Yue Deng 0001, Zhun Zhong
IEEE Trans. Image Process.4
2023 Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion Models
abstract
Color plays an important role in human visual perception, reflecting the spectrum of objects. However, the existing infrared and visible image fusion methods rarely explore how to handle multi-spectral/channel data directly and achieve high color fidelity. This paper addresses the above issue by proposing a novel method with diffusion models, termed as Dif-Fusion, to generate the distribution of the multi-channel input data, which increases the ability of multi-source information aggregation and the fidelity of colors. In specific, instead of converting multi-channel images into single-channel data in existing fusion methods, we create the multi-channel data distribution with a denoising network in a latent space with forward and reverse diffusion process. Then, we use the the denoising network to extract the multi-channel diffusion features with both visible and infrared information. Finally, we feed the multi-channel diffusion features to the multi-channel fusion module to directly generate the three-channel fused image. To retain the texture and intensity information, we propose multi-channel gradient loss and intensity loss. Along with the current evaluation metrics for measuring texture and intensity fidelity, we introduce Delta E as a new evaluation metric to quantify color fidelity. Extensive experiments indicate that our method is more effective than other state-of-the-art image fusion methods, especially in color fidelity. The source code is available at https://github.com/GeoVectorMatrix/Dif-Fusion.
Jun Yue 0004, Leyuan Fang, Shaobo Xia, Yue Deng 0001, Jiayi Ma 0001
IEEE Trans. Image Process.4
2022 Boosting Supervised Dehazing Methods via Bi-level Patch Reweighting
Xingyu Jiang 0003, Hongkun Dou, Chengwei Fu, Bingquan Dai, Tianrun Xu, Yue Deng 0001
ECCV (18)6
2021 Gated Value Network for Multilabel Classification
abstract
We introduce a gated value network (GVN) for general multilabel classification (MLC) tasks. GVN was motivated by deep value network (DVN) that directly exploits the "compatibility" metric as the learning pursuit for MLC. Meanwhile, it further improves traditional DVN on twofold. First, GVN relaxes the complex variable optimization steps in DVN inference by incorporating a feedforward predictor for straightforward multilabel prediction. Second, GVN also introduces the gating mechanism to block confounding factors from the input data that allows more precise compatibility evaluations for data and their potential multilabels. The whole GVN framework is trained in an end-to-end manner with policy gradient approaches. We show the effectiveness and generalization of GVN on diverse learning tasks, including document classification, audio tagging, and image attribute prediction.
Yimin Hou 0001, Sen Wan, Feng Bao 0002, Zhiquan Ren, Yunfeng Dong, Qionghai Dai, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2021 Human-in-the-Loop Low-Shot Learning
abstract
We consider a human-in-the-loop scenario in the context of low-shot learning. Our approach was inspired by the fact that the viability of samples in novel categories cannot be sufficiently reflected by those limited observations. Some heterogeneous samples that are quite different from existing labeled novel data can inevitably emerge in the testing phase. To this end, we consider augmenting an uncertainty assessment module into low-shot learning system to account into the disturbance of those out-of-distribution (OOD) samples. Once detected, these OOD samples are passed to human beings for active labeling. Due to the discrete nature of this uncertainty assessment process, the whole Human-In-the-Loop Low-shot (HILL) learning framework is not end-to-end trainable. We hence revisited the learning system from the aspect of reinforcement learning and introduced the REINFORCE algorithm to optimize model parameters via policy gradient. The whole system gains noticeable improvements over existing low-shot learning approaches.
Sen Wan, Yimin Hou 0001, Feng Bao 0002, Zhiquan Ren, Yunfeng Dong, Qionghai Dai, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2020 Cooperative Deep Reinforcement Learning for Large-Scale Traffic Grid Signal Control
abstract
Exploiting reinforcement learning (RL) for traffic congestion reduction is a frontier topic in intelligent transportation research. The difficulty in this problem stems from the inability of the RL agent simultaneously monitoring multiple signal lights when taking into account complicated traffic dynamics in different regions of a traffic system. Such challenge is even more outstanding when forming control decisions on a large-scale traffic grid, where the RL action space grows exponentially with the number of intersections within the traffic grid. In this paper, we tackle such a problem by proposing a cooperative deep reinforcement learning (Coder) framework. The intuition behind Coder is to decompose the original difficult RL task as a number of subproblems with relatively easy RL goals. Accordingly, we implement Coder with multiple regional agents and a centralized global agent. Each regional agent learns its own RL policy and value functions over a small region with limited actions. Then, the centralized global agent hierarchically aggregates RL achievements from different regional agents and forms the final Q -function over the entire large-scale traffic grid. The experimental investigations demonstrate that the proposed Coder could reduce on average 30% congestions in terms of the number of waiting vehicles during high density traffic flows in simulations.
Tian Tan 0003, Feng Bao 0002, Yue Deng 0001, Alex Jin, Qionghai Dai, Jie Wang 0006
IEEE Trans. Cybern.3
2020 Learning Deep Landmarks for Imbalanced Classification
abstract
We introduce a deep imbalanced learning framework called learning DEep Landmarks in laTent spAce (DELTA). Our work is inspired by the shallow imbalanced learning approaches to rebalance imbalanced samples before feeding them to train a discriminative classifier. Our DELTA advances existing works by introducing the new concept of rebalancing samples in a deeply transformed latent space, where latent points exhibit several desired properties including compactness and separability. In general, DELTA simultaneously conducts feature learning, sample rebalancing, and discriminative learning in a joint, end-to-end framework. The framework is readily integrated with other sophisticated learning concepts including latent points oversampling and ensemble learning. More importantly, DELTA offers the possibility to conduct imbalanced learning with the assistancy of structured feature extractor. We verify the effectiveness of DELTA not only on several benchmark data sets but also on more challenging real-world tasks including click-through-rate (CTR) prediction, multi-class cell type classification, and sentiment analysis with sequential inputs.
Feng Bao 0002, Yue Deng 0001, Youyong Kong, Zhiquan Ren, Jin-Li Suo, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.2
2020 A New Concept of Multiple Neural Networks Structure Using Convex Combination
abstract
In this article, a new concept of convex-combined multiple neural networks (NNs) structure is proposed. This new approach uses the collective information from multiple NNs to train the model. Based on both theoretical and experimental analyses, the new approach is shown to achieve faster training convergence with a similar or even better test accuracy than a conventional NN structure. Two experiments are conducted to demonstrate the performance of our new structure: the first one is a semantic frame parsing task for spoken language understanding (SLU) on the Airline Travel Information System (ATIS) data set and the other is a handwritten digit recognition task on the Mixed National Institute of Standards and Technology (MNIST) data set. We test this new structure using both the recurrent NN and convolutional NNs through these two tasks. The results of both experiments demonstrate a 4× - 8× faster training speed with better or similar performance by using this new concept.
Yu Wang 0091, Yue Deng 0001, Yilin Shen, Hongxia Jin
IEEE Trans. Neural Networks Learn. Syst.2
2019 Adversarial Multi-label Prediction for Spoken and Visual Signal Tagging
abstract
We introduce an adversarial multi-label classification (ADMLC) framework to improve the robustness and performance of existing algorithms on multi-domain signals. The core contribution of our ADMLC is the innovation of an `adversarial module' that serves as a critic to provide augmenting information to improve supervised learning in multi label classification (MLC) tasks. Our approach is not intended to be regarded as an emerging competitor for many well-established algorithms in the field. In fact, many existing deep and shallow architectures can all be adopted as building blocks integrated in the ADMLC framework. We show the performance and generalization ability of ADMLC on diverse tasks including audio and image tagging.
Yue Deng 0001, KaWai Chen, Yilin Shen, Hongxia Jin
ICASSP1
2019 Learning Assistance from an Adversarial Critic for Multi-Outputs Prediction
abstract
We introduce an adversarial-critic-and-assistant (ACA) learning framework to improve the performance of existing supervised learning with multiple outputs. The core contribution of our ACA is the innovation of two novel modules, i.e. an `adversarial critic' and a `collaborative assistant', that are jointly designed to provide augmenting information for facilitating general learning tasks. Our approach is not intended to be regarded as an emerging competitor for tons of well-established algorithms in the field. In fact, most existing approaches, while implemented with different learning objectives, can all be adopted as building blocks seamlessly integrated in the ACA framework to accomplish various real-world tasks. We show the performance and generalization ability of ACA on diverse learning tasks including multi-label classification, attributes prediction and sequence-to-sequence generation.
Yue Deng 0001, Yilin Shen, Hongxia Jin
IJCAI1
2018 Adversarial Active Learning for Sequences Labeling and Generation
abstract
We introduce an active learning framework for general sequence learning tasks including sequence labeling and generation. Most existing active learning algorithms mainly rely on an uncertainty measure derived from the probabilistic classifier for query sample selection. However, such approaches suffer from two shortcomings in the context of sequence learning including 1) cold start problem and 2) label sampling dilemma. To overcome these shortcomings, we propose a deep-learning-based active learning framework to directly identify query samples from the perspective of adversarial learning. Our approach intends to offer labeling priorities for sequences whose information content are least covered by existing labeled data. We verify our sequence-based active learning approach on two tasks including sequence labeling and sequence generation.
Yue Deng 0001, KaWai Chen, Yilin Shen, Hongxia Jin
IJCAI1
2018 Training Recurrent Neural Network through Moment Matching for NLP Applications
Yue Deng 0001, Yilin Shen, KaWai Chen, Hongxia Jin
INTERSPEECH1
2018 Interactive recommendation via deep neural memory augmented contextual bandits
abstract
Personalized recommendation with user interactions has become increasingly popular nowadays in many applications with dynamic change of contents (news, media, etc.). Existing approaches model user interactive recommendation as a contextual bandit problem to balance the trade-off between exploration and exploitation. However, these solutions require a large number of interactions with each user to provide high quality personalized recommendations. To mitigate this limitation, we design a novel deep neural memory augmented mechanism to model and track the history state for each user based on his previous interactions. As such, the user's preferences on new items can be quickly learned within a small number of interactions. Moreover, we develop new algorithms to leverage large amount of all users' history data for offline model training and online model fine tuning for each user with the focus of policy evaluation. Extensive experiments on different synthetic and real-world datasets validate that our proposed approach consistently outperforms a variety of state-of-the-art approaches.
Yilin Shen, Yue Deng 0001, Avik Ray, Hongxia Jin
RecSys2
2018 Probabilistic natural mapping of gene-level tests for genome-wide association studies
abstract
Genome-wide association studies (GWASs) generally focus on a single marker, which limits the elucidation of the genetic architecture of complex traits. Herein, we present a new computational framework, termed probabilistic natural mapping (PALM), for performing gene-level association tests. PALM robustly reveals the inherent genomic structures of genes and generates feature representations that can be seamlessly incorporated into conventional statistic tests. Our approach substantially improves the effectiveness of uncovering associations derived from a subgroup of variants with weak effects, which represents a known challenge associated with existing methods. We applied PALM in a gastric cancer GWAS and identified two additional gastric cancer-associated susceptibility genes, NOC3L and RUNDC2A. The robust susceptibility discoveries of PALM are widely supported by existing studies from other biological perspectives. PALM will be useful for further GWAS analytical strategies that use gene-level analyses.
Feng Bao 0002, Yue Deng 0001, Mulong Du, Zhiquan Ren, Yanyu Zhao, Jin-Li Suo, Meilin Wang, Qionghai Dai
Briefings Bioinform.2
2018 ACID: Association Correction for Imbalanced Data in GWAS
abstract
Genome-wide association study (GWAS) has been widely witnessed as a powerful tool for revealing suspicious loci from various diseases. However, real world GWAS tasks always suffer from the data imbalance problem of sufficient control samples and limited case samples. This imbalance issue can cause serious biases to the result and thus leads to losses of significance for true causal markers. To tackle this problem, we proposed a computational framework to perform association correction for imbalanced data (ACID) that could potentially improve the performance of GWAS under the imbalance condition. ACID is inspired by the imbalance learning theory but is particularly modified to address the task of association discovery from sequential genomic data. Simulation studies demonstrate ACID can dramatically improve the power of traditional GWAS method on the dataset with severe imbalances. We further applied ACID to two imbalanced datasets (gastric cancer and bladder cancer) to conduct genome wide association analysis. Experimental results indicate that our method has better abilities in identifying suspicious loci than the regression approach and shows consistencies with existing discoveries.
Feng Bao 0002, Yue Deng 0001, Qionghai Dai
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 Disguise Adversarial Networks for Click-through Rate Prediction
abstract
We introduced an adversarial learning framework for improving CTR prediction in Ads recommendation. Our approach was motivated by observing the extremely low click-through rate and imbalanced label distribution in the historical Ads impressions. We hence proposed a Disguise-Adversarial-Networks (DAN) to improve the accuracy of supervised learning with limited positive-class information. In the context of CTR prediction, the rationality behind DAN could be intuitively understood as ``non-clicked Ads makeup''. DAN disguises the disliked Ads impressions (non-clicks) to be interesting ones and encourages a discriminator to classify these disguised Ads as positive recommendations. In an adversarial aspect, the discriminator should be sober-minded which is optimized to allocate these disguised Ads to their inherent classes according to an unsupervised information theoretic assignment strategy. We applied DAN to two Ads datasets including both mobile and display Ads for CTR prediction. The results showed that our DAN approach significantly outperformed other supervised learning and generative adversarial networks (GAN) in CTR prediction.
Yue Deng 0001, Yilin Shen, Hongxia Jin
IJCAI1
2017 Discriminant Kernel Assignment for Image Coding
abstract
This paper proposes discriminant kernel assignment (DKA) in the bag-of-features framework for image representation. DKA slightly modifies existing kernel assignment to learn width-variant Gaussian kernel functions to perform discriminant local feature assignment. When directly applying gradient-descent method to solve DKA, the optimization may contain multiple time-consuming reassignment implementations in iterations. Accordingly, we introduce a more practical way to locally linearize the DKA objective and the difficult task is cast as a sequence of easier ones. Since DKA only focuses on the feature assignment part, it seamlessly collaborates with other discriminative learning approaches, e.g., discriminant dictionary learning or multiple kernel learning, for even better performances. Experimental evaluations on multiple benchmark datasets verify that DKA outperforms other image assignment approaches and exhibits significant efficiency in feature coding.
Yue Deng 0001, Yanyu Zhao, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Cybern.1
2017 A Hierarchical Fused Fuzzy Deep Neural Network for Data Classification
abstract
Deep learning (DL) is an emerging and powerful paradigm that allows large-scale task-driven feature learning from big data. However, typical DL is a fully deterministic model that sheds no light on data uncertainty reductions. In this paper, we show how to introduce the concepts of fuzzy learning into DL to overcome the shortcomings of fixed representation. The bulk of the proposed fuzzy system is a hierarchical deep neural network that derives information from both fuzzy and neural representations. Then, the knowledge learnt from these two respective views are fused altogether forming the final data representation to be classified. The effectiveness of the model is verified on three practical tasks of image categorization, high-frequency financial data prediction and brain MRI segmentation that all contain high level of uncertainties in the raw data. The fuzzy dDL paradigm greatly outperforms other nonfuzzy and shallow learning approaches on these tasks.
Yue Deng 0001, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Fuzzy Syst.1
2017 Deep Direct Reinforcement Learning for Financial Signal Representation and Trading
abstract
Can we train the computer to beat experienced traders for financial assert trading? In this paper, we try to address this challenge by introducing a recurrent deep neural network (NN) for real-time financial signal representation and trading. Our model is inspired by two biological-related learning concepts of deep learning (DL) and reinforcement learning (RL). In the framework, the DL part automatically senses the dynamic market condition for informative feature learning. Then, the RL module interacts with deep representations and makes trading decisions to accumulate the ultimate rewards in an unknown environment. The learning system is implemented in a complex NN that exhibits both the deep and recurrent structures. Hence, we propose a task-aware backpropagation through time method to cope with the gradient vanishing issue in deep training. The robustness of the neural system is verified on both the stock and the commodity future markets under broad testing conditions.
Yue Deng 0001, Feng Bao 0002, Youyong Kong, Zhiquan Ren, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.1
2016 Local visual feature fusion via maximum margin multimodal deep neural network
Zhiquan Ren, Yue Deng 0001, Qionghai Dai
Neurocomputing2
2016 Directed Adaptive Graphical Lasso for causality inference
Zhiquan Ren, Feng Bao 0002, Yue Deng 0001, Qionghai Dai
Neurocomputing4
2016 Deep and Structured Robust Information Theoretic Learning for Image Analysis
abstract
This paper presents a robust information theoretic (RIT) model to reduce the uncertainties, i.e., missing and noisy labels, in general discriminative data representation tasks. The fundamental pursuit of our model is to simultaneously learn a transformation function and a discriminative classifier that maximize the mutual information of data and their labels in the latent space. In this general paradigm, we, respectively, discuss three types of the RIT implementations with linear subspace embedding, deep transformation, and structured sparse learning. In practice, the RIT and deep RIT are exploited to solve the image categorization task whose performances will be verified on various benchmark data sets. The structured sparse RIT is further applied to a medical image analysis task for brain magnetic resonance image segmentation that allows group-level feature selections on the brain tissues.
Yue Deng 0001, Feng Bao 0002, XueSong Deng, Ruiping Wang 0001, Youyong Kong, Qionghai Dai
IEEE Trans. Image Process.1
2015 Discriminative Clustering and Feature Selection for Brain MRI Segmentation
abstract
Automatic segmentation of brain tissues from MRI is of great importance for clinical application and scientific research. Recent advancements in supervoxel-level analysis enable robust segmentation of brain tissues by exploring the inherent information among multiple features extracted on the supervoxels. Within this prevalent framework, the difficulties still remain in clustering uncertainties imposed by the heterogeneity of tissues and the redundancy of the MRI features. To cope with the aforementioned two challenges, we propose a robust discriminative segmentation method from the view of information theoretic learning. The prominent goal of the method is to simultaneously select the informative feature and to reduce the uncertainties of supervoxel assignment for discriminative brain tissue segmentation. Experiments on two brain MRI datasets verified the effectiveness and efficiency of the proposed approach.
Youyong Kong, Yue Deng 0001, Qionghai Dai
IEEE Signal Process. Lett.2
2015 Sparse Coding-Inspired Optimal Trading System for HFT Industry
abstract
The financial industry has witnessed an exceptionally fast progress of incorporating information processing techniques in designing knowledge-based automated systems for high-frequency trading (HFT). This paper proposes a sparse coding-inspired optimal trading (SCOT) system for real-time high-frequency financial signal representation and trading. Mathematically, SCOT simultaneously learns the dictionary, sparse features, and the trading strategy in a joint optimization, yielding optimal feature representations for the specific trading objective. The learning process is modeled as a bilevel optimization and solved by the online gradient descend method with fast convergence. In this dynamic context, the system is tested on the real financial market to trade the index futures in the Shanghai exchange center.
Yue Deng 0001, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Ind. Informatics1
2014 Visual Words Assignment Via Information-Theoretic Manifold Embedding
abstract
Codebook-based learning provides a flexible way to extract the contents of an image in a data-driven manner for visual recognition. One central task in such frameworks is codeword assignment, which allocates local image descriptors to the most similar codewords in the dictionary to generate histogram for categorization. Nevertheless, existing assignment approaches, e.g., nearest neighbors strategy (hard assignment) and Gaussian similarity (soft assignment), suffer from two problems: 1) too strong Euclidean assumption and 2) neglecting the label information of the local descriptors. To address the aforementioned two challenges, we propose a graph assignment method with maximal mutual information (GAMI) regularization. GAMI takes the power of manifold structure to better reveal the relationship of massive number of local features by nonlinear graph metric. Meanwhile, the mutual information of descriptor-label pairs is ultimately optimized in the embedding space for the sake of enhancing the discriminant property of the selected codewords. According to such objective, two optimization models, i.e., inexact-GAMI and exact-GAMI, are respectively proposed in this paper. The inexact model can be efficiently solved with a closed-from solution. The stricter exact-GAMI nonparametrically estimates the entropy of descriptor-label pairs in the embedding space and thus leads to a relatively complicated but still trackable optimization. The effectiveness of GAMI models are verified on both the public and our own datasets.
Yue Deng 0001, Yanjun Qian, Xiangyang Ji, Qionghai Dai
IEEE Trans. Cybern.1
2014 Joint Non-Gaussian Denoising and Superresolving of Raw High Frame Rate Videos
abstract
High frame rate cameras capture sharp videos of highly dynamic scenes by trading off signal-noise-ratio and image resolution, so combinational super-resolving and denoising is crucial for enhancing high speed videos and extending their applications. The solution is nontrivial due to the fact that two deteriorations co-occur during capturing and noise is nonlinearly dependent on signal strength. To handle this problem, we propose conducting noise separation and super resolution under a unified optimization framework, which models both spatiotemporal priors of high quality videos and signal-dependent noise. Mathematically, we align the frames along temporal axis and pursue the solution under the following three criterion: 1) the sharp noise-free image stack is low rank with some missing pixels denoting occlusions; 2) the noise follows a given nonlinear noise model; and 3) the recovered sharp image can be reconstructed well with sparse coefficients and an over complete dictionary learned from high quality natural images. In computation aspects, we propose to obtain the final result by solving a convex optimization using the modern local linearization techniques. In the experiments, we validate the proposed approach in both synthetic and real captured data.
Jin-Li Suo, Yue Deng 0001, Liheng Bian, Qionghai Dai
IEEE Trans. Image Process.2
2013 Free-Viewpoint Video of Human Actors Using Multiple Handheld Kinects
abstract
We present an algorithm for creating free-viewpoint video of interacting humans using three handheld Kinect cameras. Our method reconstructs deforming surface geometry and temporal varying texture of humans through estimation of human poses and camera poses for every time step of the RGBZ video. Skeletal configurations and camera poses are found by solving a joint energy minimization problem, which optimizes the alignment of RGBZ data from all cameras, as well as the alignment of human shape templates to the Kinect data. The energy function is based on a combination of geometric correspondence finding, implicit scene segmentation, and correspondence finding using image features. Finally, texture recovery is achieved through jointly optimization on spatio-temporal RGB data using matrix completion. As opposed to previous methods, our algorithm succeeds on free-viewpoint video of human actors under general uncontrolled indoor scenes with potentially dynamic background, and it succeeds even if the cameras are moving.
Genzhi Ye, Yebin Liu, Yue Deng 0001, Nils Hasler, Xiangyang Ji, Qionghai Dai, Christian Theobalt
IEEE Trans. Cybern.3
2013 Low-Rank Structure Learning via Nonconvex Heuristic Recovery
abstract
In this paper, we propose a nonconvex framework to learn the essential low-rank structure from corrupted data. Different from traditional approaches, which directly utilizes convex norms to measure the sparseness, our method introduces more reasonable nonconvex measurements to enhance the sparsity in both the intrinsic low-rank structure and the sparse corruptions. We will, respectively, introduce how to combine the widely used ℓp norm (0 < p < 1) and log-sum term into the framework of low-rank structure learning. Although the proposed optimization is no longer convex, it still can be effectively solved by a majorization-minimization (MM)-type algorithm, with which the nonconvex objective function is iteratively replaced by its convex surrogate and the nonconvex problem finally falls into the general framework of reweighed approaches. We prove that the MM-type algorithm can converge to a stationary point after successive iterations. The proposed model is applied to solve two typical problems: robust principal component analysis and low-rank representation. Experimental results on low-rank structure learning demonstrate that our nonconvex heuristic methods, especially the log-sum heuristic recovery algorithm, generally perform much better than the convex-norm-based method (0 < p < 1) for both data with higher rank and with denser corruptions.
Yue Deng 0001, Qionghai Dai, Risheng Liu, Zengke Zhang, Sanqing Hu
IEEE Trans. Neural Networks Learn. Syst.1
2012 Visual words assignment on a graph via minimal mutual information loss
abstract
Visual codewords assignment plays an important role in many Bag of Features (BoF) models for image understanding and visual recognition. It allocates image descriptors to the most similar codewords in the pre-configured visual dictionary to generate descriptive histogram for the consequent categorization. Nevertheless, existing assignment approaches, e.g. nearest neighbors strategy and Gaussian similarity, suffer from two problems:1) too strong Euclidean assumption and 2) neglecting the label information of the local features. Accordingly, in this paper, we propose an assignment method to simultaneously consider the above two issues in a unified model via graph learning and information theoretic criterions. For learning, the proposed model can be efficiently solved in a closed-form with the reasonable graph topology invariant approximation. Moreover, the learned projections enable us to extend the assignment ability to the out-of-sample visual features beyond the initial training graph. Experiments on our own manifold dataset and two benchmarks verify the effectiveness of the proposed graph assignment method.
Yanjun Qian, Yue Deng 0001, Qionghai Dai, Guihua Er
BMVC2
2012 Commute time guided transformation for feature extraction
Yue Deng 0001, Qionghai Dai, Ruiping Wang 0001, Zengke Zhang
Comput. Vis. Image Underst.1
2011 Graph Laplace for Occluded Face Completion and Recognition
abstract
This paper proposes a spectral-graph-based algorithm for face image repairing, which can improve the recognition performance on occluded faces. The face completion algorithm proposed in this paper includes three main procedures: 1) sparse representation for partially occluded face classification; 2) image-based data mining; and 3) graph Laplace (GL) for face image completion. The novel part of the proposed framework is GL, as named from graphical models and the Laplace equation, and can achieve a high-quality repairing of damaged or occluded faces. The relationship between the GL and the traditional Poisson equation is proven. We apply our face repairing algorithm to produce completed faces, and use face recognition to evaluate the performance of the algorithm. Experimental results verify the effectiveness of the GL method for occluded face completion.
Yue Deng 0001, Qionghai Dai, Zengke Zhang
IEEE Trans. Image Process.1
2009 Partially occluded face completion and recognition
abstract
This paper proposes a spectral graph based algorithm for face image repairing, which can improve the recognition performance on occluded faces. Our algorithm is called `guided label-learning', so named from graphical models, and can achieve a high-quality repairing of damaged or occluded faces. We apply our face repairing algorithm in order to produce completed faces, and then use face recognition to evaluate the performance of our algorithm. Experiment results show that, at most, a nearly 30%-increase in the recognition rate can be achieved for occluded faces with the use of our algorithm.
Yue Deng 0001, Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Qionghai Dai
ICIP1