Feng Chen 0007

dblp:21/3047-7 · DBLP profile ↗
← Back
60ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0003-4813-2494ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 2 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4Theory of computation · 3Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Chain-of-Detection: Enhancing Cross-Granularity Robotic Perception for Object Manipulation
abstract
In robotic perception, cross-granularity object detection is essential for identifying and localizing targets at varying levels of detail. Traditional detection methods often struggle to bridge the gap between coarse object detection and fine-grained component localization, limiting their ability to associate parts, such as a cup and its handle. Vision-language models (VLMs), while effective in spatial reasoning, face challenges in fine-grained detection due to the scarcity of annotated datasets. To address these issues, we first propose the chain-of-detection (CoD) framework, which focuses on guiding detection in a step-by-step manner from coarse recognition to fine-grained localization. During this process, we observe that existing detectors still lack sufficient capability in recognizing fine-grained components. To overcome this limitation, we further combine the CoD framework with Monte Carlo tree search (MCTS) to automatically generate fine-grained datasets, eliminating the need for manual labeling and significantly improving detector performance. Experiments show that our approach achieves an average improvement of 17.31% in robotic manipulation success rates for common objects, 51.39% for larger object operations, and about 50% in simulated environments. These results demonstrate the effectiveness of CoD in advancing cross-granularity detection and enhancing precise robotic manipulation. The implementation is publicly available at https://github.com/tinnel123666888/CoD and the CoD dataset is released at https://huggingface.co/datasets/tinnel123/CoD_dataset.
Tianrun Xu, Haichuan Gao, Changlin Chen, Shiyuan Xu, Shangqi Guo, Feng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.7
2025 OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene Understanding
Tianrun Xu, Yuxin Xi, Zeyu Mu, Haichuan Gao, Feng Chen 0007
ICCV9
2025 When do neural networks learn world models?
abstract
Humans develop world models that capture the underlying generation process of data. Whether neural networks can learn similar world models remains an open problem. In this work, we present the first theoretical results for this problem, showing that in a multi-task setting, models with a low-degree bias provably recover latent data-generating variables under mild assumptions–even if proxy tasks involve complex, non-linear functions of the latents. However, such recovery is sensitive to model architecture. Our analysis leverages Boolean models of task solutions via the Fourier-Walsh transform and introduces new techniques for analyzing invertible Boolean transforms, which may be of independent interest. We illustrate the algorithmic implications of our results and connect them to related research areas, including self-supervised learning, out-of-distribution generalization, and the linear representation hypothesis in large language models.
Feng Chen 0007
ICML3
2025 Adaptive Fission: Post-training Encoding for Low-latency Spike Neural Networks
abstract
Spiking Neural Networks (SNNs) often rely on rate coding, where high-precision inference depends on long time-steps, leading to significant latency and energy cost—especially for ANN-to-SNN conversions. To address this, we propose Adaptive Fission, a post-training encoding technique that selectively splits high-sensitivity neurons into groups with varying scales and weights. This enables neuron-specific, on-demand precision and threshold allocation while introducing minimal spatial overhead. As a generalized form of population coding, it seamlessly applies to a wide range of pretrained SNN architectures without requiring additional training or fine-tuning. Experiments on neuromorphic hardware demonstrate up to 80\% reductions in latency and power consumption without degrading accuracy.
Yizhou Jiang, Feng Chen 0007, Yuqian Liu, Haichuan Gao
NeurIPS2
2025 Causal dreamer for partially observable model-based reinforcement learning
Haichuan Gao, Tianrun Xu, Chujie Zhao, Jinsheng Ren, Yizhou Jiang, Shangqi Guo, Feng Chen 0007
Neurocomputing9
2025 Iterative compression towards in-distribution features in domain generalization
Yizhou Jiang, Feng Chen 0007
Neurocomputing5
2025 Multi-core token mixer: a novel approach for underwater image enhancement
Tianrun Xu, Shiyuan Xu, Feng Chen 0007, Hongjue Li
Mach. Vis. Appl.4
2024 Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers
abstract
Spiking neural networks (SNNs) are energy-efficient and hold great potential for large-scale inference. Since training SNNs from scratch is costly and has limited performance, converting pretrained artificial neural networks (ANNs) to SNNs is an attractive approach that retains robust performance without additional training data and resources. However, while existing conversion methods work well on convolution networks, emerging Transformer models introduce unique mechanisms like self-attention and test-time normalization, leading to non-causal non-linear interactions unachievable by current SNNs. To address this, we approximate these operations in both temporal and spatial dimensions, thereby providing the first SNN conversion pipeline for Transformers. We propose \textit{Universal Group Operators} to approximate non-linear operations spatially and a \textit{Temporal-Corrective Self-Attention Layer} that approximates spike multiplications at inference through an estimation-correction approach. Our algorithm is implemented on a pretrained ViT-B/32 from CLIP, inheriting its zero-shot classification capabilities, while improving control over conversion losses. To our knowledge, this is the first direct training-free conversion of a pretrained Transformer to a purely event-driven SNN, promising for neuromorphic hardware deployment.
Yizhou Jiang, Kunlin Hu, Haichuan Gao, Yuqian Liu, Feng Chen 0007
ICLR7
2024 Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize
abstract
Learning representations that generalize under distribution shifts is critical for building robust machine learning models. However, despite significant efforts in recent years, algorithmic advances in this direction have been limited. In this work, we seek to understand the fundamental difficulty of out-of-distribution generalization with deep neural networks. We first empirically show that perhaps surprisingly, even allowing a neural network to explicitly fit the representations obtained from a teacher network that can generalize out-of-distribution is insufficient for the generalization of the student network. Then, by a theoretical study of two-layer ReLU networks optimized by stochastic gradient descent (SGD) under a structured feature model, we identify a fundamental yet unexplored feature learning proclivity of neural networks, feature contamination: neural networks can learn uncorrelated features together with predictive features, resulting in generalization failure under distribution shifts. Notably, this mechanism essentially differs from the prevailing narrative in the literature that attributes the generalization failure to spurious correlations. Overall, our results offer new insights into the non-linear feature learning dynamics of neural networks and highlight the necessity of considering inductive biases in out-of-distribution generalization.
Chujie Zhao, Yizhou Jiang, Feng Chen 0007
ICML5
2023 Fast Counterfactual Inference for History-Based Reinforcement Learning
abstract
Incorporating sequence-to-sequence models into history-based Reinforcement Learning (RL) provides a general way to extend RL to partially-observable tasks. This method compresses history spaces according to the correlations between historical observations and the rewards. However, they do not adjust for the confounding correlations caused by data sampling and assign high beliefs to uninformative historical observations, leading to limited compression of history spaces. Counterfactual Inference (CI), which estimates causal effects by single-variable intervention, is a promising way to adjust for confounding. However, it is computationally infeasible to directly apply the single-variable intervention to a huge number of historical observations. This paper proposes to perform CI on observation sub-spaces instead of single observations and develop a coarse-to-fine CI algorithm, called Tree-based History Counterfactual Inference (T-HCI), to reduce the number of interventions exponentially. We show that T-HCI is computationally feasible in practice and brings significant sample efficiency gains in various challenging partially-observable tasks, including Maze, BabyAI, and robot manipulation tasks.
Haichuan Gao, Zhile Yang, Jinsheng Ren, Shangqi Guo, Feng Chen 0007
AAAI7
2023 MER: Modular Element Randomization for robust generalizable policy in deep reinforcement learning
Jinsheng Ren, Feng Chen 0007
Knowl. Based Syst.5
2023 Adjacency Constraint for Efficient Hierarchical Reinforcement Learning
abstract
Goal-conditioned Hierarchical Reinforcement Learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is large. Searching in a large goal space poses difficulty for both high-level subgoal generation and low-level policy learning. In this article, we show that this problem can be effectively alleviated by restricting the high-level action space from the whole goal space to a k-step adjacent region of the current state using an adjacency constraint. We theoretically prove that in a deterministic Markov Decision Process (MDP), the proposed adjacency constraint preserves the optimal hierarchical policy, while in a stochastic MDP the adjacency constraint induces a bounded state-value suboptimality determined by the MDP's transition structure. We further show that this constraint can be practically implemented by training an adjacency network that can discriminate between adjacent and non-adjacent subgoals. Experimental results on discrete and continuous control tasks including challenging simulated robot locomotion and manipulation tasks show that incorporating the adjacency constraint significantly boosts the performance of state-of-the-art goal-conditioned HRL approaches.
Shangqi Guo, Tian Tan 0003, Xiaolin Hu 0001, Feng Chen 0007
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Partial Consistency for Stabilizing Undiscounted Reinforcement Learning
abstract
Undiscounted return is an important setup in reinforcement learning (RL) and characterizes many real-world problems. However, optimizing an undiscounted return often causes training instability. The causes of this instability problem have not been analyzed in-depth by existing studies. In this article, this problem is analyzed from the perspective of value estimation. The analysis result indicates that the instability originates from transient traps that are caused by inconsistently selected actions. However, selecting one consistent action in the same state limits exploration. For balancing exploration effectiveness and training stability, a novel sampling method called last-visit sampling (LVS) is proposed to ensure that a part of actions is selected consistently in the same state. The LVS method decomposes the state-action value into two parts, i.e., the last-visit (LV) value and the revisit value. The decomposition ensures that the LV value is determined by consistently selected actions. We prove that the LVS method can eliminate transient traps while preserving optimality. Also, we empirically show that the method can stabilize the training processes of five typical tasks, including vision-based navigation and manipulation tasks.
Haichuan Gao, Zhile Yang, Tian Tan 0003, Jinsheng Ren, Shangqi Guo, Feng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.8
2022 Interacting Attention Graph for Single Image Two-Hand Reconstruction
abstract
Graph convolutional network (GCN) has achieved great success in single hand reconstruction task, while interacting two-hand reconstruction by GCN remains unexplored. In this paper, we present Interacting Attention Graph Hand (IntagHand), the first graph convolution based network that reconstructs two interacting hands from a single RGB image. To solve occlusion and interaction challenges of two-hand reconstruction, we introduce two novel attention based modules in each upsampling step of the original GCN. The first module is the pyramid image feature attention (PIFA) module, which utilizes multiresolution features to implicitly obtain vertex-to-image alignment. The second module is the cross hand attention (CHA) module that encodes the coherence of interacting hands by building dense cross-attention between two hand vertices. As a result, our model outperforms all existing two-hand re-construction methods by a large margin on InterHand2.6M benchmark. Moreover, ablation studies verify the effectiveness of both PIFA and CHA modules for improving the reconstruction accuracy. Results on in-the-wild images and live video streams further demonstrate the generalization ability of our network. Our code is available at https://github.com/Dw1010/IntagHand.
Mengcheng Li, Liang An 0001, Hongwen Zhang 0001, Lianpeng Wu, Feng Chen 0007, Tao Yu 0007, Yebin Liu
CVPR5
2022 Primitive-contrastive network: data-efficient self-supervised learning from robot demonstration videos
Zhile Yang, Shangqi Guo, Feng Chen 0007
Appl. Intell.5
2022 State-Temporal Compression in Reinforcement Learning With the Reward-Restricted Geodesic Metric
abstract
It is difficult to solve complex tasks that involve large state spaces and long-term decision processes by reinforcement learning (RL) algorithms. A common and promising method to address this challenge is to compress a large RL problem into a small one. Towards this goal, the compression should be state-temporal and optimality-preserving (i.e., the optimal policy of the compressed problem should correspond to that of the uncompressed problem). In this paper, we propose a reward-restricted geodesic (RRG) metric, which can be learned by a neural network, to perform state-temporal compression in RL. We prove that compression based on the RRG metric is approximately optimality-preserving for the raw RL problem endowed with temporally abstract actions. With this compression, we design an RRG metric-based reinforcement learning (RRG-RL) algorithm to solve complex tasks. Experiments in both discrete (2D Minecraft) and continuous (Doom) environments demonstrated the superiority of our method over existing RL approaches.
Shangqi Guo, Qi Yan 0005, Xiaolin Hu 0001, Feng Chen 0007
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Revealing Fine Structures of the Retinal Receptive Field by Deep-Learning Networks
abstract
Deep convolutional neural networks (CNNs) have demonstrated impressive performance on many visual tasks. Recently, they became useful models for the visual system in neuroscience. However, it is still not clear what is learned by CNNs in terms of neuronal circuits. When a deep CNN with many layers is used for the visual system, it is not easy to compare the structure components of CNNs with possible neuroscience underpinnings due to highly complex circuits from the retina to the higher visual cortex. Here, we address this issue by focusing on single retinal ganglion cells with biophysical models and recording data from animals. By training CNNs with white noise images to predict neuronal responses, we found that fine structures of the retinal receptive field can be revealed. Specifically, convolutional filters learned are resembling biological components of the retinal circuit. This suggests that a CNN learning from one single retinal cell reveals a minimal neural network carried out in this cell. Furthermore, when CNNs learned from different cells are transferred between cells, there is a diversity of transfer learning performance, which indicates that CNNs are cell specific. Moreover, when CNNs are transferred between different types of input images, here white noise versus natural images, transfer learning shows a good performance, which implies that CNNs indeed capture the full computational ability of a single retinal cell for different inputs. Taken together, these results suggest that CNNs could be used to reveal structure components of neuronal circuits, and provide a powerful model for neural system identification.
Qi Yan 0005, Yajing Zheng, Shanshan Jia 0001, Yichen Zhang 0002, Zhaofei Yu, Feng Chen 0007, Yonghong Tian 0001, Tiejun Huang 0001, Jian K. Liu
IEEE Trans. Cybern.6
2022 Orientation-Preserving Rewards' Balancing in Reinforcement Learning
abstract
Auxiliary rewards are widely used in complex reinforcement learning tasks. However, previous work can hardly avoid the interference of auxiliary rewards on pursuing the main rewards, which leads to the destruction of the optimal policy. Thus, it is challenging but essential to balance the main and auxiliary rewards. In this article, we explicitly formulate the problem of rewards' balancing as searching for a Pareto optimal solution, with the overall objective of preserving the policy's optimization orientation for the main rewards (i.e., the policy driven by the balanced rewards is consistent with the policy driven by the main rewards). To this end, we propose a variant Pareto and show that it can effectively guide the policy search toward more main rewards. Furthermore, we establish an iterative learning framework for rewards' balancing and theoretically analyze its convergence and time complexity. Experiments in both discrete (grid word) and continuous (Doom) environments demonstrated that our algorithm can effectively balance rewards, and achieve remarkable performance compared with those RLs with heuristically designed rewards. In the ViZDoom platform, our algorithm can learn expert-level policies.
Jinsheng Ren, Shangqi Guo, Feng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.3
2021 CRIL: Continual Robot Imitation Learning via Generative and Prediction Model
abstract
Imitation learning (IL) algorithms have shown promising results for robots to learn skills from expert demonstrations. However, they need multi-task demonstrations to be provided at once for acquiring diverse skills, which is difficult in real world. In this work we study how to realize continual imitation learning ability that empowers robots to continually learn new tasks one by one, thus reducing the burden of multitask IL and accelerating the process of new task learning at the same time. We propose a novel trajectory generation model that employs both a generative adversarial network and a dynamics-aware prediction model to generate pseudo trajectories from all learned tasks in the new task learning process. Our experiments on both simulation and real-world manipulation tasks demonstrate the effectiveness of our method.
Chongkai Gao, Haichuan Gao, Shangqi Guo, Feng Chen 0007
IROS5
2020 Task Understanding from Confusing Multi-task Data
abstract
Beyond machine learning’s success in the specific tasks, research for learning multiple tasks simultaneously is referred to as multi-task learning. However, existing multi-task learning needs manual definition of tasks and manual task annotation. A crucial problem for advanced intelligence is how to understand the human task concept using basic input-output pairs. Without task definition, samples from multiple tasks are mixed together and result in a confusing mapping challenge. We propose Confusing Supervised Learning (CSL) that takes these confusing samples and extracts task concepts by differentiating between these samples. We theoretically proved the feasibility of the CSL framework and designed an iterative algorithm to distinguish between tasks. The experiments demonstrate that our CSL methods could achieve a human-like task understanding without task labeling in multi-function regression problems and multi-task recognition problems.
Yizhou Jiang, Shangqi Guo, Feng Chen 0007
ICML4
2020 Adaptability Preserving Domain Decomposition for Stabilizing Sim2Real Reinforcement Learning
abstract
In sim-to-real transfer of Reinforcement Learning (RL) policies for robot tasks, Domain Randomization (DR) is a widely used technique for improving adaptability. However, in DR there is a conflict between adaptability and training stability, and heavy DR tends to result in instability or even failure in training. To relieve this conflict, we propose a new algorithm named Domain Decomposition (DD) that decomposes the randomized domain according to environments and trains a separate RL policy for each part. This decomposition stabilizes the training of each RL policy, and as we prove theoretically, the adaptability of the overall policy can be preserved. Our simulation results verify that DD really improves stability in training while preserving ideal adaptability. Further, we complete a complex real-world vision-based patrolling task using DD, which demonstrates DD’s practicality. A video is attached as supplementary material.
Haichuan Gao, Zhile Yang, Tian Tan 0003, Feng Chen 0007
IROS5
2020 Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning
abstract
Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is often large. Searching in a large goal space poses difficulties for both high-level subgoal generation and low-level policy learning. In this paper, we show that this problem can be effectively alleviated by restricting the high-level action space from the whole goal space to a k-step adjacent region of the current state using an adjacency constraint. We theoretically prove that the proposed adjacency constraint preserves the optimal hierarchical policy in deterministic MDPs, and show that this constraint can be practically implemented by training an adjacency network that can discriminate between adjacent and non-adjacent subgoals. Experimental results on discrete and continuous control tasks show that incorporating the adjacency constraint improves the performance of state-of-the-art HRL approaches in both deterministic and stochastic environments.
Shangqi Guo, Tian Tan 0003, Xiaolin Hu 0001, Feng Chen 0007
NeurIPS5
2020 Cycle representation-disentangling network: learning to completely disentangle spatial-temporal features in video
Shangqi Guo, Feng Chen 0007
Appl. Intell.4
2020 Emergent Inference of Hidden Markov Models in Spiking Neural Networks Through Winner-Take-All
abstract
Hidden Markov models (HMMs) underpin the solution to many problems in computational neuroscience. However, it is still unclear how to implement inference of HMMs with a network of neurons in the brain. The existing methods suffer from the problem of being nonspiking and inaccurate. Here, we build a precise equivalence between the inference equation of HMMs with time-invariant hidden variables and the dynamics of spiking winner-take-all (WTA) neural networks. We show that the membrane potential of each spiking neuron in the WTA circuit encodes the logarithm of the posterior probability of the hidden variable in each state, and the firing rate of each neuron is proportional to the posterior probability of the HMMs. We prove that the time course of the neural firing rate can implement posterior inference of HMMs. Theoretical analysis and experimental results show that the proposed WTA circuit can get accurate inference results of HMMs.
Zhaofei Yu, Shangqi Guo, Fei Deng 0001, Qi Yan 0005, Keke Huang, Jian K. Liu, Feng Chen 0007
IEEE Trans. Cybern.7
2020 Generative Memory for Lifelong Learning
abstract
Lifelong learning is a crucial issue in advanced artificial intelligence. It requires the learning system to learn and accumulate knowledge from sequential tasks. The learning system needs to deal with increasingly more domains and tasks. We consider that the key to an effective and efficient lifelong learning system is the ability to memorize and recall the learned knowledge using neural networks. Following this idea, we propose Generative Memory (GM) as a novel memory module, and the resulting lifelong learning system is referred to as the GM Net (GMNet). To make the GMNet feasible, we propose a novel learning mechanism, referred to as P -invariant learning method. It replaces the memory of the real data by a memory of the data distribution, which makes it possible for the learning system to accurately and continuously accumulate the learned experiences. We demonstrate that GMNet achieves the state-of-the-art performance on lifelong learning tasks.
Shangqi Guo, Tian Tan 0003, Feng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.4
2020 Neural Hand Reconstruction Using A Single RGB Image
abstract
We present a neural hand reconstruction method for monocular 3D hand pose and shape estimation in this paper. Instead of directly representing hand with 3D data, a novel UV position map is introduced to represent hand pose and shape with 2D data, which maps 3D hand surface points to 2D image space. Furthermore, an encoder-decoder neural network is proposed to infer such UV position map from only single image. To train such network with the lack of ground truth training pairs, we propose a novel MANOReg module which employs MANO model as shape prior to constrain high-dimensional space of UV position map. Both quantitative and qualitative experiments demonstrate the effectiveness of our UV position map representation and MANOReg module.
Mengcheng Li, Liang An 0001, Tao Yu 0007, Yangang Wang 0001, Feng Chen 0007, Yebin Liu
Virtual Real. Intell. Hardw.5
2019 Noise helps optimization escape from saddle points in the neural dynamics
Zhaofei Yu, Feng Chen 0007
ESANN3
2019 A unified neural circuit of causal inference and multisensory integration
Zhaofei Yu, Jian K. Liu, Feng Chen 0007
Neurocomputing4
2019 Hierarchical Bayesian Inference and Learning in Spiking Neural Networks
abstract
Numerous experimental data from neuroscience and psychological science suggest that human brain utilizes Bayesian principles to deal the complex environment. Furthermore, hierarchical Bayesian inference has been proposed as an appropriate theoretical framework for modeling cortical processing. However, it remains unknown how such a computation is organized in the network of biologically plausible spiking neurons. In this paper, we propose a hierarchical network of winner-take-all circuits which can carry out hierarchical Bayesian inference and learning through a spike-based variational expectation maximization (EM) algorithm. Particularly, we show how the firing activities of spiking neurons in response to the input stimuli and the spike-timing-dependent plasticity rule can be understood, respectively, as variational E-step and M-step of variational EM. Finally, we demonstrate the utility of this spiking neural network on the MNIST benchmark for unsupervised classification of handwritten digits.
Shangqi Guo, Zhaofei Yu, Fei Deng 0001, Xiaolin Hu 0001, Feng Chen 0007
IEEE Trans. Cybern.5
2018 Unification of MAP Estimation and Marginal Inference in Recurrent Neural Networks
abstract
Numerous experimental data show that human brain can represent probability distributions and perform Bayesian inference. However, it remains unclear how the brain implements probabilistic inference in the form of neural circuits. Several models have been proposed that aim at explaining how the network of neurons carry out maximum a posterior inference (MAP) estimation and marginal inference, but they are all task specific in that they treat MAP estimation and marginal inference separately. In this brief, we propose that human brain could implement MAP estimation and marginal inference in the same network of neurons. We illustrate our result in hidden Markov models and prove that a recurrent neural network (RNN) implementation of belief propagation can be tuned to perform approximate Bayesian inference (to provide posterior or conditional distribution over the latent causes of observations) or identify the MAP or peak of the joint distribution. The key tuning parameter is a temperature parameter that controls the precision of probability distributions that are optimized. Theoretical analyses and experimental results demonstrate that RNNs can carry out near-optimal MAP estimation and marginal inference.
Zhaofei Yu, Feng Chen 0007, Fei Deng 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 A Robust Static Sign Language Recognition System Based on Hand Key Points Estimation
Feng Chen 0007, Guijin Wang, Jinsheng Ren, Jianwu Dong
ISDA2
2016 Sampling-based causal inference in cue combination and its neural implementation
Zhaofei Yu, Feng Chen 0007, Jianwu Dong, Qionghai Dai
Neurocomputing2
2016 Signal-dependent noise removal for color videos using temporal and cross-channel priors
Jin-Li Suo, Liheng Bian, Feng Chen 0007, Qionghai Dai
J. Vis. Commun. Image Represent.3
2015 Efficient approximate linear programming for factored MDPs
Feng Chen 0007, Qiang Shawn Cheng, Jianwu Dong, Zhaofei Yu, Wenli Xu
Int. J. Approx. Reason.1
2015 Object Tracking With Joint Optimization of Representation and Classification
abstract
We present a novel algorithm that exploits joint optimization of representation and classification for robust tracking in which the goal is to minimize the least-squares reconstruction errors and discriminative penalties with regularized constraints. In this formulation, an object is represented by the sparse coefficients of local patches based on an overcomplete dictionary, and a classifier is learned to discriminate the target object from the background. To locate the target object in each frame, we propose a deterministic approach to solve the optimization problem. We show that the proposed algorithm can be considered as a generalization of several tracking methods with effectiveness. To account for appearance change of the target and the background, the classifier is adaptively updated with new tracking results. Compared with the most recent tracking algorithms based on sparse representation, the proposed formulation has more discriminative power due to the use of background information and is much faster due to the use of deterministic optimization. Qualitative and quantitative experiments on a variety of challenging sequences show favorable performance of the proposed algorithm against several state-of-the-art methods.
Qing Wang 0017, Feng Chen 0007, Wenli Xu, Ming-Hsuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2015 Simultaneous Phase Unwrapping and Removal of Chemical Shift (SPURS) Using Graph Cuts: Application in Quantitative Susceptibility Mapping
abstract
Quantitative susceptibility mapping (QSM) is a magnetic resonance imaging technique that reveals tissue magnetic susceptibility. It relies on having a high quality field map, typically acquired with a relatively long echo spacing and long final TE. Applications of QSM outside the brain require the removal of fat contributions to the total signal phase. However, current water/fat separation methods applied on typical data acquired for QSM suffer from three issues: inadequacy when using large echo spacing, over-smoothing of the field maps and high computational cost. In this paper, the general phase wrap and chemical shift problem is formulated using a single species fitting and is solved using graph cuts with conditional jump moves. This method is referred as simultaneous phase unwrapping and removal of chemical shift (SPURS). The result from SPURS is then used as the initial guess for a voxel-wise iterative decomposition of water and fat with echo asymmetric and least-squares estimation (IDEAL). The estimated 3-D field maps are used to compute QSM in body regions outside of the brain, such as the liver. Experimental results show substantial improvements in field map estimation, water/fat separation and reconstructed QSM compared to two existing water/fat separation methods on 1.5T and 3T magnetic resonance human data with long echo spacing and rapid field map variation.
Jianwu Dong, Feng Chen 0007, Alexey Dimov, Ashish Raj, Qiang Shawn Cheng, Pascal Spincemaille, Yi Wang 0028
IEEE Trans. Medical Imaging3
2014 Supervised feature subset selection with ordinal optimization
Dingcheng Feng, Feng Chen 0007, Wenli Xu
Knowl. Based Syst.2
2013 Variational Planning for Graph-based MDPs
abstract
Markov Decision Processes (MDPs) are extremely useful for modeling and solving sequential decision making problems. Graph-based MDPs provide a compact representation for MDPs with large numbers of random variables. However, the complexity of exactly solving a graph-based MDP usually grows exponentially in the number of variables, which limits their application. We present a new variational framework to describe and solve the planning problem of MDPs, and derive both exact and approximate planning algorithms. In particular, by exploiting the graph structure of graph-based MDPs, we propose a factored variational value iteration algorithm in which the value function is first approximated by the multiplication of local-scope value functions, then solved by minimizing a Kullback-Leibler (KL) divergence. The KL divergence is optimized using the belief propagation algorithm, with complexity exponential in only the cluster size of the graph. Experimental comparison on different models shows that our algorithm outperforms existing approximation algorithms at finding good policies.
Qiang Shawn Cheng, Qiang Liu 0001, Feng Chen 0007, Alexander Ihler
NIPS3
2013 Decomposition and Approximation of Loopy Bayesian Networks
abstract
This paper proposes a new method, conditional probability table (CPT) decomposition, to analyze the independent and deterministic components of CPT. This method can be used to approximate and analyze Baysian networks. The decomposition of Bayesian networks is accomplished by representing CPTs as a linear combination of extreme CPTs, which forms a new framework to conduct inference. Based on this new framework, inference in Bayesian networks can be done by decomposing them into less connected and weighted subnetworks. We can achieve exact inference if the original network is decomposed into singly-connected subnetworks. Besides, approximate inference can be done by discarding the subnetworks with small weights or by a partial decomposition and application of belief propagation (BP) on the still multiply-connected subnetworks. Experiments show that the decomposition-based approximation outperforms BP in most cases.
Jianwu Dong, Feng Chen 0007, Yanyan Huo
Fundam. Informaticae2
2013 Energy distribution view for monotonic dual decomposition
Qiang Shawn Cheng, Feng Chen 0007, Jianwu Dong, Wenli Xu
Int. J. Approx. Reason.2
2013 A new criterion for choosing planar subproblems in MAP-MRF inference
Jianwu Dong, Feng Chen 0007, Qiang Shawn Cheng, Song Wang 0002
Neurocomputing2
2013 Learning linear non-Gaussian networks: A new view from matrix identification
abstract
Identifying causal structures from observations is fundamental in many applications. Probabilistic graphical models provide a unifying framework for capturing complex causal dependencies among random variables. Recently, a linear non-Gaussian acyclic model (LiNGAM) and some smart algorithms (ICA-LiNGAM, DirectLiNGAM) have been proposed, which outperform previous graphical models and learning methods in identifying variable orders. We propose new solutions (TMFLiNGAM, SchurLiNGAM and RCLiNGAM) to the LiNGAM learning task from the perspective of matrix identification. TMFLiNGAM and SchurLiNGAM recover orders more directly, and RCLiNGAM can improve the accuracy of previous algorithms on uniform and sparse structures. The perspective also facilitates the learning of sparse models where the performance of all independent component analysis-based algorithms can be improved by reconstructing variable orders from the inverse of separation matrices. Experimental results under various settings provide average evaluations over the learning methods, and verify the effectiveness of our perspective and algorithms.
Dingcheng Feng, Feng Chen 0007, Wenli Xu
J. Exp. Theor. Artif. Intell.2
2013 Video content categorization using the double decomposition
Youtian Du, Feng Chen 0007, Wenli Xu, Xueming Qian
Multim. Tools Appl.2
2012 Approximating the Sum Operation for Marginal-MAP Inference
abstract
We study the marginal-MAP problem on graphical models, and present a novel approximation method based on direct approximation of the sum operation. A primary difficulty of marginal-MAP problems lies in the non-commutativity of the sum and max operations, so that even in highly structured models, marginalization may produce a densely connected graph over the variables to be maximized, resulting in an intractable potential function with exponential size. We propose a chain decomposition approach for summing over the marginalized variables, in which we produce a structured approximation to the MAP component of the problem consisting of only pairwise potentials. We show that this approach is equivalent to the maximization of a specific variational free energy, and it provides an upper bound of the optimal probability. Finally, experimental results demonstrate that our method performs favorably compared to previous methods.
Qiang Shawn Cheng, Feng Chen 0007, Jianwu Dong, Wenli Xu, Alexander Ihler
AAAI2
2012 Online discriminative object tracking with local sparse representation
abstract
We propose an online algorithm based on local sparse representation for robust object tracking. Local image patches of a target object are represented by their sparse codes with an over-complete dictionary constructed online, and a classifier is learned to discriminate the target from the background. To alleviate the visual drift problem often encountered in object tracking, a two-stage algorithm is proposed to exploit both the ground truth information of the first frame and observations obtained online. Different from recent discriminative tracking methods that use a pool of features or a set of boosted classifiers, the proposed algorithm learns sparse codes and a linear classifier directly from raw image patches. In contrast to recent sparse representation based tracking methods which encode holistic object appearance within a generative framework, the proposed algorithm employs a discrimination formulation which facilitates the tracking task in complex environments. Experiments on challenging sequences with evaluation of the state-of-the-art methods show effectiveness of the proposed algorithm.
Qing Wang 0017, Feng Chen 0007, Wenli Xu, Ming-Hsuan Yang 0001
WACV2
2012 Recursive sum-product algorithm for generalized outer-planar graphs
Qiang Shawn Cheng, Feng Chen 0007, Wenli Xu, Song Wang 0002
Inf. Process. Lett.2
2012 Learning robust principal components from L1-normmaximization
abstract
Principal component analysis (PCA) is fundamental in many pattern recognition applications. Much research has been performed to minimize the reconstruction error in L1-norm based reconstruction error minimization (L1-PCA-REM) since conventional L2-norm based PCA (L2-PCA) is sensitive to outliers. Recently, the variance maximization formulation of PCA with L1-norm (L1-PCA-VM) has been proposed, where new greedy and nongreedy solutions are developed. Armed with the gradient ascent perspective for optimization, we show that the L1-PCA-VM formulation is problematic in learning principal components and that only a greedy solution can achieve robustness motivation, which are verified by experiments on synthetic and real-world datasets.
Dingcheng Feng, Feng Chen 0007, Wenli Xu
J. Zhejiang Univ. Sci. C2
2012 Object Tracking via Partial Least Squares Analysis
abstract
We propose an object tracking algorithm that learns a set of appearance models for adaptive discriminative object representation. In this paper, object tracking is posed as a binary classification problem in which the correlation of object appearance and class labels from foreground and background is modeled by partial least squares (PLS) analysis, for generating a low-dimensional discriminative feature subspace. As object appearance is temporally correlated and likely to repeat over time, we learn and adapt multiple appearance models with PLS analysis for robust tracking. The proposed algorithm exploits both the ground truth appearance information of the target labeled in the first frame and the image observations obtained online, thereby alleviating the tracking drift problem caused by model update. Experiments on numerous challenging sequences and comparisons to state-of-the-art methods demonstrate favorable performance of the proposed tracking algorithm.
Qing Wang 0017, Feng Chen 0007, Wenli Xu, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.2
2012 Transferring Visual Prior for Online Object Tracking
abstract
Visual prior from generic real-world images can be learned and transferred for representing objects in a scene. Motivated by this, we propose an algorithm that transfers visual prior learned offline for online object tracking. From a collection of real-world images, we learn an overcomplete dictionary to represent visual prior. The prior knowledge of objects is generic, and the training image set does not necessarily contain any observation of the target object. During the tracking process, the learned visual prior is transferred to construct an object representation by sparse coding and multiscale max pooling. With this representation, a linear classifier is learned online to distinguish the target from the background and to account for the target and background appearance variations over time. Tracking is then carried out within a Bayesian inference framework, in which the learned classifier is used to construct the observation model and a particle filter is used to estimate the tracking result sequentially. Experiments on a variety of challenging sequences with comparisons to several state-of-the-art methods demonstrate that more robust object tracking can be achieved by transferring visual prior.
Qing Wang 0017, Feng Chen 0007, Jimei Yang, Wenli Xu, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.2
2011 Analysis of Markov Boundary Induction in Bayesian Networks: A New View From Matroid Theory
abstract
Learning Markov boundaries from data without having to learn a Bayesian network first can be viewed as a feature subset selection problem and has received much attention due to its significance in the wide applications of AI techniques. Popular constraint based methods suffer from high computational complexity and are usually unstable in spaces of high dimensionality. We propose a new perspective from matroid theory towards the discovery of Markov boundaries of random variable in the domain, and develop a learning algorithm which guarantees to recover the true Markov boundaries by a greedy learning algorithm. Then we use the precision matrix of the original distribution as a measure of independence to make our algorithm feasible in large scale problems, which is essentially an approximation of the probabilistic relations with Gaussians and can find possible variables in Markov boundaries with low computational complexity. Experimental results on standard Bayesian networks show that our analysis and approximation can efficiently and accurately identify Markov boundaries in complex networks from data.
Dingcheng Feng, Feng Chen 0007, Wenli Xu
Fundam. Informaticae2
2011 Adaptive multi-cue tracking by online appearance learning
Qing Wang 0017, Feng Chen 0007, Wenli Xu
Neurocomputing2
2011 Object tracking via appearance modeling and sparse representation
Feng Chen 0007, Qing Wang 0017, Song Wang 0002, Wenli Xu
Image Vis. Comput.1
2011 Tracking by Third-Order Tensor Representation
abstract
This paper proposes a robust tracking algorithm by third-order tensor representation and adaptive appearance modeling. In this method, the target in each video frame is represented by a third-order tensor. This representation preserves the spatial correlation inside the target region and can integrate multiple appearance cues for target description. Based on this representation, a multilinear subspace is learned online to model the target appearance variations during tracking. Compared to other methods, our approach can detect local spatial structure in the target tensor space and fuse information from different feature spaces. Therefore, the learned appearance model is more discriminative when there are significant appearance variations of the target or when the background gets cluttered. Applying the multilinear algebra, our appearance model can efficiently be learned and updated online, without causing high-dimensional data-learning problems. Then, tracking is implemented in the Bayesian inference framework, where a likelihood model is defined to measure the similarity between a test sample and the learned appearance model, and a particle filter is used to recursively estimate the target state over time. Theoretic analysis and experiments compared with other state-of-the-art methods demonstrate the effectiveness of the proposed approach.
Qing Wang 0017, Feng Chen 0007, Wenli Xu
IEEE Trans. Syst. Man Cybern. Part B2
2010 Saliency selection for robust visual tracking
abstract
This paper proposes a robust visual tracking approach based on saliency selection. In this method, salient patches and their spatial context inside the object region are exploited for object representation and appearance modeling. Tracking is then implemented by a hybrid stochastic and deterministic mechanism, which needs a small number of samples for particle filtering and escapes local minimum in conventional deterministic tracking. As time progresses, the selected salient patches and their spatial context are updated online to adapt the appearance model to both object and environmental changes. We carry out experiments on several challenging sequences and compare our method with the state-of-the-art algorithm to show its improvement in terms of tracking performance.
Qing Wang 0017, Feng Chen 0007, Wenli Xu
ICIP2
2010 Learning interactions among multi-channel sequences with dynamical influence models
Feng Chen 0007, Wenli Xu
Sci. China Inf. Sci.2
2009 Hierarchical Control Models for Multimodal Process Modeling
abstract
The multimodal and hierarchical structure characteristics of a system make process modeling quite difficult. In this paper, we present a hierarchical control model (HCM) for hierarchically multimodal processing. From multiple streams, a control layer extracts the inherent group process that denotes the evolution of the system and controls the evolution of every modality. HCMs model the influences of the group on modalities and represent the hierarchical structure of the system by a multilayer network. To estimate the state order of the model, we also present a new information criterion that corrects the preference of traditional criteria for more complex models and proves the rationality of HCMs. Comparisons with other models on multiagent activity recognition show that HCMs are reliable and efficient.
Feng Chen 0007, Wenli Xu
IEEE Trans. Syst. Man Cybern. Part B2
2008 Activity recognition through multi-scale motion detail analysis
Youtian Du, Feng Chen 0007, Wenli Xu
Neurocomputing2
2008 Hierarchical group process representation in multi-agent activity recognition
Feng Chen 0007, Wenli Xu, Youtian Du
Signal Process. Image Commun.2
2007 Human Interaction Representation and Recognition Through Motion Decomposition
abstract
Human action recognition is one of the most important problems in video content analysis and computer vision. In this letter, we propose a novel framework of human interaction recognition through motion decomposition. Interactions contain not only motions corresponding to each person but also motion details on different scales. Hence, we decompose an interaction into multiple interacting stochastic processes in the above two aspects. Under the framework, we present a Coupled Hierarchical Durational-State Dynamic Bayesian Network (CHDS-DBN) to model interactions by modeling the multiple stochastic processes. The effectiveness of the approach is demonstrated by experiments of two-person interaction recognition.
Youtian Du, Feng Chen 0007, Wenli Xu
IEEE Signal Process. Lett.2
2006 Real-Time Video Intelligent Surveillance System
abstract
With the rapid development of hardware equipments, it is now economically and technically feasible to build a video surveillance system. This paper presents the system architecture of VISS, a video intelligent surveillance system deployed in parking lots. In VISS we adopt robust moving object detecting and tracking algorithm, and we present a novel activity recognition framework based on layer hidden semi-Markov model (LHSMM) which is used for modeling activities. The experimental results on real-time video shows the system is effective and robust in complex activity recognition
Feng Chen 0007, Wenli Xu, Enwei Zhang
ICME2