Hao Shen 0002

dblp:26/2210-2 · DBLP profile ↗
← Back
34ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-0091-4155ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 54% Optimization for machine learning · 14% Multi-agent systems · 12%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
embodied foundation models
0.912025
Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review · IJCAI 2025
Knowledge, reasoning and agents › Multi-agent systems
embodied multi-agent system
0.912025
Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review · IJCAI 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.722020
Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Towards a Unified Framework of Contrastive Learning for Disentangled Representations · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.712023
Towards a Unified Framework of Contrastive Learning for Disentangled Representations · NeurIPS 2023
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability
0.712023
Towards a Unified Framework of Contrastive Learning for Disentangled Representations · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.522017
Dynamical Textures Modeling via Joint Video Dictionary Learning · IEEE Trans. Image Process. 2017
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
low-dimensional representation learning
0.412020
Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Machine learning › Optimization for machine learning › optimization landscape
critical point analysis
0.312018
Towards a Mathematical Understanding of the Difficulty in Learning With Feedforward Neural Networks · CVPR 2018
Machine learning › Deep learning architectures and training › feedforward neural network
feedforward neural network training
0.312018
Towards a Mathematical Understanding of the Difficulty in Learning With Feedforward Neural Networks · CVPR 2018
Machine learning › Optimization for machine learning
non-convex optimization
0.312018
Towards a Mathematical Understanding of the Difficulty in Learning With Feedforward Neural Networks · CVPR 2018
Machine learning › Optimization for machine learning
optimization landscape
0.312018
Towards a Mathematical Understanding of the Difficulty in Learning With Feedforward Neural Networks · CVPR 2018
Computer vision › Video understanding and tracking › motion analysis
dynamic texture modeling
0.312017
Dynamical Textures Modeling via Joint Video Dictionary Learning · IEEE Trans. Image Process. 2017
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212016
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016

Methods — techniques the papers use, named apart from their topics

generative agents · 0.9foundation model · 0.9sparse coding · 0.7noise contrastive estimation · 0.7InfoNCE · 0.7trace quotient criterion · 0.4conjugate gradient · 0.4generalized gauss-newton algorithm · 0.3approximate newton's method · 0.3markov random process · 0.3
YearPublicationVenuePosition
2026 Enhance Language Model-based Repair for Memory-related Vulnerabilities via Knowledge-and Semantic-guided Analysis
abstract
Memory-related vulnerabilities often result in system crashes and performance drops, imposing significant risks for embedded systems. Despite the potential of Language Models (LMs) in program repair, existing LM-based approaches struggle with these vulnerabilities due to two primary limitations: i) LMs do not possess adequate domain knowledge concerning program analysis and the characteristics of memory-related vulnerabilities, and ii) LMs face constraints in managing contexts as the size of programs increases. To address this issue, we introduce MVRepair, a novel lightweight Language Model (ℓLM)-driven framework built upon a domain-specific knowledge library that is developed through the examination of 7,935 real-world memory-related vulnerabilities. By using our proposed knowledge-based analysis strategy and semantic-guided segmentation mechanism, MVRepair can substantially enhance the LM’s ability to repair programs with memory-related vulnerabilities. Comprehensive experimental results on 8,118 real-world memory-related vulnerabilities demonstrate that, compared with state-of-the-art LM-based approaches, MVRepair yields improvements of a minimum of 23.8% in EM, 31.9% in BLEU-4, and 16.7% in CodeBLEU.
Hao Shen 0002, Ming Hu 0003, Yanxin Yang, Xiaofei Xie, Mingsong Chen 0001
DATE1
2025 Multi-Agent Credit Assignment with Pretrained Language Models
abstract
The difficulty of appropriately assigning credit is particularly heightened in cooperative MARL with sparse reward, due to the concurrent time and structural scales involved. Automatic subgoal generation (ASG) has recently emerged as a viable MARL approach inspired by utilizing subgoals in intrinsically motivated reinforcement learning. However, end-to-end learning of complex task planning from sparse rewards without prior knowledge, undoubtedly requires massive training samples. Moreover, the diversity-promoting nature of existing ASG methods can lead to the "over-representation" of subgoals, generating numerous spurious subgoals of limited relevance to the actual task reward and thus decreasing the sample efficiency of the algorithm. To address this problem and inspired by the disentangled representation learning, we propose a novel "disentangled" decision-making method, Semantically Aligned task decomposition in MARL (SAMA), that prompts pretrained language models with chain-of-thought that can suggest potential goals, provide suitable goal decomposition and subgoal allocation as well as self-reflection-based replanning. Additionally, SAMA incorporates language-grounded MARL to train each agent’s subgoal-conditioned policy. SAMA demonstrates considerable advantages in sample efficiency compared to state-of-the-art ASG methods, as evidenced by its performance on two challenging sparse-reward tasks, Overcooked and MiniRTS. The code is available at \url{https://anonymous.4open.science/r/SAMA/.}
Wenhao Li 0001, Baoxiang Wang 0001, Xiangfeng Wang 0001, Hao Shen 0002, Bo Jin 0003, Hongyuan Zha
AISTATS6
2025 Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review
abstract
Embodied multi-agent systems (EMAS) have attracted growing attention for their potential to address complex, real-world challenges in areas such as logistics and robotics. Recent advances in foundation models pave the way for generative agents capable of richer communication and adaptive problem-solving. This survey provides a systematic examination of how EMAS can benefit from these generative capabilities. We propose a taxonomy that categorizes EMAS by system architectures and embodiment modalities, emphasizing how collaboration spans both physical and virtual contexts. Central building blocks, perception, planning, communication, and feedback, are then analyzed to illustrate how generative techniques bolster system robustness and flexibility. Through concrete examples, we demonstrate the transformative effects of integrating foundation models into embodied, multi-agent frameworks. Finally, we discuss challenges and future directions, underlining the significant promise of EMAS to reshape the landscape of AI-driven collaboration.
Xian Wei, Guang Chen 0001, Hao Shen 0002, Bo Jin 0003
IJCAI4
2025 Cross-Level Fusion: Integrating Object Lists with Raw Sensor Data for 3D Object Tracking
abstract
Smart sensors and Vehicle-To-Everything (V2X) modules are commonly utilized in automotive perception systems, which primarily provide processed object lists rather than raw data. However, high-level fusion approaches suffer from significant information loss and representational misalignment due to the inherently abstract and sparse nature of these high-level outputs. We propose a novel cross-level fusion paradigm that enables bidirectional information flow between object lists and raw vision features within an end-to-end Transformer framework for 3D object detection and tracking. Our approach extracts inherent positional and dimensional cues from object lists to generate two outputs: structured query features that are fused with the initial learnable queries in the Transformer decoder, and soft Gaussian attention masks that guide feature extraction. This integrated mechanism not only improves tracking accuracy by synergistically combining object priors with fine-grained vision data but also promotes hardware economy and AI model sustainability by adapting legacy sensors to evolving sensor setups. To overcome the lack of dedicated datasets, we develop a pseudo object list generation pipeline that simulates realistic sensor tracking behavior. Experiments on the nuScenes dataset demonstrate significant performance gains over vision-only baselines and robust generalization across diverse noise levels, validating the efficacy of our cross-level fusion strategy. The code is available at: https://github.com/CesarLiu/DNF.git.
Xiangzhong Liu, Xihao Wang, Hao Shen 0002
IROS3
2025 Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
abstract
In automotive sensor fusion systems, smart sensors and Vehicle-to-Everything (V2X) modules are commonly utilized. Sensor data from these systems are typically available only as processed object lists rather than raw sensor data from traditional sensors. Instead of processing other raw data sep-arately and then fusing them at object level, we propose an end-to-end cross-level fusion concept with Transformer, which integrates highly abstract object list information with raw camera images for 3D object detection. Object lists are fed into a Transformer as denoising queries and propagated together with learnable queries through the latter feature aggregation process. Additionally, a deformable Gaussian mask, derived from the positional and size dimensional priors from the object lists, is explicitly integrated into the Transformer decoder. This directs attention toward the target area of interest and accelerates model training convergence. Furthermore, as there is no public dataset containing object lists as a standalone modality, we propose an approach to generate pseudo object lists from ground-truth bounding boxes by simulating state noise and false positives and negatives. As the first work to conduct cross-level fusion, our approach shows substantial performance improvements over the vision-based baseline on the nuScenes dataset. It demonstrates its generalization capability over diverse noise levels of simulated object lists and real detectors.
Xiangzhong Liu, Hao Shen 0002
IV3
2023 Potential-based Credit Assignment for Cooperative RL-based Testing of Autonomous Vehicles
abstract
While autonomous vehicles (AVs) may perform remarkably well in generic real-life cases, their irrational action in some unforeseen cases leads to critical safety concerns. This paper introduces the concept of collaborative reinforcement learning (RL) to generate challenging test cases for AV planning and decision-making module. One of the critical challenges for collaborative RL is the credit assignment problem, where a proper assignment of rewards to multiple agents interacting in the traffic scenario, considering all parameters and timing, turns out to be non-trivial. In order to address this challenge, we propose a novel potential-based reward-shaping approach inspired by counterfactual analysis for solving the credit-assignment problem. The evaluation in a simulated environment demonstrates the superiority of our proposed approach against other methods using local and global rewards.
Utku Ayvaz, Chih-Hong Cheng, Hao Shen 0002
IJCNN3
2023 Scene Understanding for Autonomous Driving Using Visual Question Answering
abstract
This paper investigates the feasibility of dot-products present in self-attention mechanisms as an explainability technique for autonomous driving. A Visual Question Answering (VQA) framework is implemented with three types of questions pertaining to the presence or absence of road signs and traffic lights. The models are evaluated for the encoding of uni- and multimodal encodings: a standard version and a modified version of the Learning Cross-Modality Encoder Representations from Transformers (LXMERT) framework. We present numerical results for the two model architectures on the question answering task, with overall accuracies of 79.7% and 78.5% respectively, and overall F1-scores of 0.749 and 0.693 respectively. Moreover, we show that these questions, despite containing and asking no information on the objects' positions, indirectly tune the model such that the self-attention dot-products provide scenic understanding to the questions in the form of a visual map. The choice of pooling for the model's output and the plotting parameters for the visual maps influence the reliability and accuracy of visualization. Finally, an argument is made for the benefit of this approach to autonomous driving.
Adrien Wantiez, Tianming Qiu, Stefan Matthes, Hao Shen 0002
IJCNN4
2023 Towards a Unified Framework of Contrastive Learning for Disentangled Representations
abstract
Contrastive learning has recently emerged as a promising approach for learning data representations that discover and disentangle the explanatory factors of the data. Previous analyses of such approaches have largely focused on individual contrastive losses, such as noise-contrastive estimation (NCE) and InfoNCE, and rely on specific assumptions about the data generating process. This paper extends the theoretical guarantees for disentanglement to a broader family of contrastive methods, while also relaxing the assumptions about the data distribution. Specifically, we prove identifiability of the true latents for four contrastive losses studied in this paper, without imposing common independence assumptions. The theoretical findings are validated on several benchmark datasets. Finally, practical limitations of these methods are also investigated.
Stefan Matthes, Zhiwei Han, Hao Shen 0002
NeurIPS3
2023 Couplformer: Rethinking Vision Transformer with Coupling Attention
abstract
With the development of the self-attention mechanism, the Transformer model has demonstrated its outstanding performance in the computer vision domain. However, the massive computation brought from the full attention mechanism became a heavy burden for memory consumption. Sequentially, the limitation of memory consumption hinders the deployment of the Transformer model on the embedded system where the computing resources are limited. To remedy this problem, we propose a novel memory economy attention mechanism named Couplformer, which decouples the attention map into two sub-matrices and generates the alignment scores from spatial information. Our method enables the Transformer model to improve time and memory efficiency while maintaining expressive power. A series of different scale image classification tasks are applied to evaluate the effectiveness of our model. The result of experiments shows that on the ImageNet-1K classification task, the Couplformer can significantly decrease 42% memory consumption compared with the regular Transformer. Meanwhile, it accesses sufficient accuracy requirements, which outperforms 0.56% on Top-1 accuracy and occupies the same memory footprint. Besides, the Couplformer achieves state-of-art performance in MS COCO 2017 object detection and instance segmentation tasks. As a result, the Couplformer can serve as an efficient backbone in visual tasks and provide a novel perspective on deploying attention mechanisms for researchers.
Xihao Wang, Hao Shen 0002, Peidong Liang, Xian Wei
WACV3
2023 Identifying influential users in unknown social networks for adaptive incentive allocation under budget restriction
Shiqing Wu 0001, Weihua Li 0007, Hao Shen 0002, Quan Bai 0001
Inf. Sci.3
2022 Adaptive Fusion CNN Features for RGBT Object Tracking
abstract
Thermal sensors play an important role in intelligent transportation system. This paper studies the problem of RGB and thermal (RGBT) tracking in challenging situations by leveraging multimodal data. A RGBT object tracking method is proposed in correlation filter tracking framework based on short term historical information. Given the initial object bounding box, hierarchical convolutional neural network (CNN) is employed to extract features. The target is tracked for RGB and thermal modalities separately. Then the backward tracking is implemented in the two modalities. The difference between each pair is computed, which is an indicator of the tracking quality in each modality. Considering the temporal continuity of sequence frames, we also incorporate the history data into the weights computation to achieve a robust fusion of different source data. Experiments on three RGBT datasets show the proposed method achieves comparable results to state-of-the-art methods.
Yong Wang 0032, Xian Wei, Hao Shen 0002, Huanlong Zhang
IEEE Trans. Intell. Transp. Syst.4
2021 Dynamic Texture Recognition via Nuclear Distances on Kernelized Scattering Histogram Spaces
abstract
Distance-based dynamic texture recognition is an important research field in multimedia processing with applications ranging from retrieval to segmentation of video data. Based on the conjecture that the most distinctive characteristic of a dynamic texture is the appearance of its individual frames, this work proposes to describe dynamic textures as kernelized spaces of frame-wise feature vectors computed using the Scattering transform. By combining these spaces with a basis-invariant metric, we get a framework that produces competitive results for nearest neighbor classification and state-of-the-art results for nearest class center classification.
Alexander Sagel, Julian Wörmann, Hao Shen 0002
ICASSP3
2021 Performance evaluation of low resolution visual tracking for unmanned aerial vehicles
Yong Wang 0032, Xian Wei, Hao Shen 0002, Jilin Hu, Lingkun Luo
Neural Comput. Appl.3
2020 Dynamic Variational Autoencoders for Visual Process Modeling
abstract
This work studies the problem of modeling visual processes by leveraging deep generative architectures for learning linear, Gaussian representations from observed sequences. We propose a joint learning framework, combining a vector autoregressive model and a Variational Autoencoder. This results in an architecture that allows Variational Autoencoders to simultaneously learn a non-linear observation as well as a linear state model from sequences of frames. We validate our approach on synthesis of artificial sequences and dynamic textures. To this end, we use our architecture to learn a statistical model of each visual process, and generate a new sequence from each learned visual process model.
Alexander Sagel, Hao Shen 0002
ICASSP2
2020 CNN tracking based on data augmentation
Yong Wang 0032, Xian Wei, Hao Shen 0002
Knowl. Based Syst.4
2020 Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations
abstract
This work studies the problem of learning appropriate low dimensional image representations. We propose a generic algorithmic framework, which leverages two classic representation learning paradigms, i.e., sparse representation and the trace quotient criterion, to disentangle underlying factors of variation in high dimensional images. Specifically, we aim to learn simple representations of low dimensional, discriminant factors by applying the trace quotient criterion to well-engineered sparse representations. We construct a unified cost function, coined as the SPARse LOW dimensional representation (SparLow) function, for jointly learning both a sparsifying dictionary and a dimensionality reduction transformation. The SparLow function is widely applicable for developing various algorithms in three classic machine learning scenarios, namely, unsupervised, supervised, and semi-supervised learning. In order to develop efficient joint learning algorithms for maximizing the SparLow function, we deploy a framework of sparse coding with appropriate convex priors to ensure the sparse representations to be locally differentiable. Moreover, we develop an efficient geometric conjugate gradient algorithm to maximize the SparLow function on its underlying Riemannian manifold. Performance of the proposed SparLow algorithmic framework is investigated on several image processing tasks, such as 3D data visualization, face/digit recognition, and object/scene categorization.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Adaptive model updating for robust object tracking
Yong Wang 0032, Xian Wei, Hao Shen 0002
Signal Process. Image Commun.3
2019 Reconstructible Nonlinear Dimensionality Reduction via Joint Dictionary Learning
abstract
This paper presents a parametric low-dimensional (LD) representation learning method that allows to reconstruct high-dimensional (HD) input vectors in an unsupervised manner. Under the assumption that the HD data and its LD representation share the same or similar local sparse structure, the proposed method achieves reconstructible dimensionality reduction via jointly learning dictionaries in both the original HD data space and its LD representation space. By regarding the sparse representation as a smooth function with respect to a specific dictionary, we construct an encoding-decoding block for learning LD representations from sparse coefficients of HD data. It is expected that this learning process preserves the desirable structure of HD data in the LD representation space, and simultaneously allows a reliable reconstruction from the LD space back to the original HD space. In addition, the proposed single layer encoding-decoding block can be easily extended to deep learning structures. Numerical experiments on both synthetic data sets and real images show that the proposed method achieves strongly competitive and robust performance in data DR, reconstruction, and synthesis, even on heavily corrupted data. The proposed method can be used as an alternative approach to compressive sensing (CS); however, it can outperform the traditional CS methods in: 1) task-driven learning problems, such as 2-D/3-D data visualization, and 2) data reconstruction at a lower dimensional space.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey
IEEE Trans. Neural Networks Learn. Syst.2
2018 Towards a Mathematical Understanding of the Difficulty in Learning With Feedforward Neural Networks
abstract
Training deep neural networks for solving machine learning problems is one great challenge in the field, mainly due to its associated optimisation problem being highly non-convex. Recent developments have suggested that many training algorithms do not suffer from undesired local minima under certain scenario, and consequently led to great efforts in pursuing mathematical explanations for such observations. This work provides an alternative mathematical understanding of the challenge from a smooth optimisation perspective. By assuming exact learning of finite samples, sufficient conditions are identified via a critical point analysis to ensure any local minimum to be globally minimal as well. Furthermore, a state of the art algorithm, known as the Generalised Gauss-Newton (GGN) algorithm, is rigorously revisited as an approximate Newton's algorithm, which shares the property of being locally quadratically convergent to a global minimum under the condition of exact learning.
Hao Shen 0002
CVPR1
2017 Formation control using GQ(λ) reinforcement learning
abstract
Formation control is an important subtask for autonomous robots. From flying drones to swarm robotics, many applications need their agents to control their group behavior. Especially when moving autonomously in humanrobot teams, motion and formation control of a group of agents is a critical and challenging task. In this work, we propose a method of applying the GQ(λ) reinforcement learning algorithm to a leader-follower formation control scenario on the e-puck robot platform. In order to allow control via classical reinforcement learning, we present how we modeled a formation control problem as a Markov decision making process. This allows us to use the Greedy-GQ(λ) algorithm for learning a leader-follower control law. The applicability and performance of this control approach is investigated in simulation as well as on real robots. In both experiments, the followers are able to move behind the leader. Additionally, the algorithm improves the smoothness of the follower's path online, which is beneficial in the context of human-robot interaction.
Martin Knopp, Can Aykin, Johannes Feldmaier, Hao Shen 0002
RO-MAN4
2017 Joint learning sparsifying linear transformation for low-resolution image synthesis and recognition
Xian Wei, Hao Shen 0002, Weidong Xiang, Yi Lu Murphey
Pattern Recognit.3
2017 Dynamical Textures Modeling via Joint Video Dictionary Learning
abstract
Video representation is an important and challenging task in the computer vision community. In this paper, we consider the problem of modeling and classifying video sequences of dynamic scenes which could be modeled in a dynamic textures (DTs) framework. At first, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named joint video dictionary learning (JVDL), to model a video adaptively. By treating the sparse coefficients of image frames over a learned dictionary as the underlying "states", we learn an efficient and robust linear transition matrix between two adjacent frames of sparse events in time series. Hence, a dynamic scene sequence is represented by an appropriate transition matrix associated with a dictionary. In order to ensure the stability of JVDL, we impose several constraints on such transition matrix and dictionary. The developed framework is able to capture the dynamics of a moving scene by exploring both the sparse properties and the temporal correlations of consecutive video frames. Moreover, such learned JVDL parameters can be used for various DT applications, such as DT synthesis and recognition. Experimental results demonstrate the strong competitiveness of the proposed JVDL approach in comparison with the state-of-the-art video representation methods. Especially, it performs significantly better in dealing with DT synthesis and recognition on heavily corrupted data.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Zhongfeng Wang 0001
IEEE Trans. Image Process.3
2016 Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations
abstract
This paper presents an algorithm that allows to learn low dimensional representations of images in an unsupervised manner. The core idea is to combine two criteria that play important roles in unsupervised representation learning, namely sparsity and trace quotient. The former is known to be a convenient tool to identify underlying factors, and the latter is known as a disentanglement of underlying discriminative factors. In this work, we develop a generic cost function for learning jointly a sparsifying dictionary and a dimensionality reduction transformation. It leads to several counterparts of classic low dimensional representation methods, such as Principal Component Analysis, Local Linear Embedding, and Laplacian Eigenmap. Our proposed optimisation algorithm leverages the efficiency of geometric optimisation on Riemannian manifolds and a closed form solution to the elastic net problem.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
CVPR2
2016 Joint learning dictionary and discriminative features for high dimensional data
abstract
Recently, sparse representation (SR) over a redundant dictionary has become a popular way of representing the data. It has been verified as an efficient and useful tool to promote the discrimination between signals. This work develops a joint learning approach to find the low dimensional discriminative features for high dimensional data. To avoid the high computational cost of direct sparse coding on large scale input data, we first learn SR in an orthogonal projected space over a task-driven sparsifying dictionary. We then exploit the discriminative projection on SR. The whole learning process is treated as an optimization problem of trace quotient maximization, which involves an orthogonal projection on original data space, a dictionary and a discriminative projection on sparse codes. The related cost function is well defined on a product manifold of the Stiefel manifold, the Oblique manifold and the Grassmann manifold. Finally, we employ a stochastic gradient descent algorithm on the smooth product manifold to maximize the cost function. Our numerical experiments on visual recognition demonstrate the effectiveness of the proposed algorithm, in comparison with the state of the arts.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey
ICPR3
2014 Accelerated gradient temporal difference learning algorithms
abstract
In this paper we study Temporal Difference (TD) Learning with linear value function approximation. The classic TD algorithm is known to be unstable with linear function approximation and off-policy learning. Recently developed Gradient TD (GTD) algorithms have addressed this problem successfully. Despite their prominent properties of good scalability and convergence to correct solutions, they inherit the potential weakness of slow convergence as they are a stochastic gradient descent algorithm. Accelerated stochastic gradient descent algorithms have been developed to speed up convergence, while still keeping computational complexity low. In this work, we develop an accelerated stochastic gradient descent method for minimizing the Mean Squared Projected Bellman Error (MSPBE), and derive a bound for the Lipschitz constant of the gradient of the MSPBE, which plays a critical role in our proposed accelerated GTD algorithms. Our comprehensive numerical experiments demonstrate promising performance in solving the policy evaluation problem, in comparison to the GTD]algorithm family. In particular, accelerated TDC surpasses state-of-the-art algorithms.
Dominik Meyer, Rémy Degenne, Ahmed Omrane, Hao Shen 0002
ADPRL4
2014 An adaptive dictionary learning approach for modeling dynamical textures
abstract
Video representation is an important and challenging task in the computer vision community. In this paper, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named adaptive video dictionary learning (AVDL), to model a video adaptively. The developed framework is able to capture the dynamics of a moving scene by exploring both sparse properties and the temporal correlations of consecutive video frames. The proposed method is compared with state of the art video processing methods on several benchmark data sequences, which exhibit appearance changes and heavy occlusions.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
ICASSP2
2013 Averaging complex subspaces via a Karcher mean approach
Knut Hüper, Martin Kleinsteuber, Hao Shen 0002
Signal Process.3
2012 HRTF-based localization and separation of multiple sound sources
abstract
The human auditory system excels at pinpointing and distinguishing multiple sound sources in noisy and reverberant environments. Mobile robotic platforms implement such capabilities with varying success, classically solving localization and separation independently. This paper presents an algorithm utilizing Head-Related Transfer Function (HRTF) based localization to aid the task of separation. HRTFs for robotic binaural hearing represent the digital emulation of a human's innate direction-dependent filtering for solving the localization problem in a compact and robust manner. The overall result of the presented algorithm for robotic binaural hearing is an HRTF-based localization and separation system, capable of dynamically and intelligently processing simultaneously active sound sources.
Martin Rothbucher, Marko Durkovic, Tim Habigt, Hao Shen 0002, Klaus Diepold
RO-MAN4
2012 Blind Source Separation With Compressively Sensed Linear Mixtures
abstract
This work studies the problem of simultaneously separating and reconstructing signals from compressively sensed linear mixtures. We assume that all source signals share a common sparse representation basis. The approach combines classical Compressive Sensing (CS) theory with a linear mixing model. It allows the mixtures to be sampled independently of each other. If samples are acquired in the time domain, this means that the sensors need not be synchronized. Since Blind Source Separation (BSS) from a linear mixture is only possible up to permutation and scaling, factoring out these ambiguities leads to a minimization problem on the so-called oblique manifold. We develop a geometric conjugate subgradient method that scales to large systems for solving the problem. Numerical results demonstrate the promising performance of the proposed algorithm compared to several state of the art methods.
Martin Kleinsteuber, Hao Shen 0002
IEEE Signal Process. Lett.2
2009 Block Jacobi-type methods for non-orthogonal joint diagonalisation
abstract
In this paper, we study the problem of non-orthogonal joint diagonalisation of a set of real symmetric matrices via simultaneous conjugation. A family of block Jacobi-type methods are proposed to optimise two popular cost functions for the non-orthogonal joint diagonalisation, namely, the off-norm function and the log-likelihood function. By exploiting the appropriate underlying manifold, namely the so-called oblique manifold, rigorous analysis shows that, under the exact non-orthogonal joint diagonalisation setting, the proposed methods converge locally quadratically fast to a joint diagonaliser. Finally, performance of our methods is investigated by numerical experiments for both exact and approximate non-orthogonal joint diagonalisation.
Hao Shen 0002, Knut Hüper
ICASSP1
2008 Local Convergence Analysis of FastICA and Related Algorithms
abstract
The FastICA algorithm is one of the most prominent methods to solve the problem of linear independent component analysis (ICA). Although there have been several attempts to prove local convergence properties of FastICA, rigorous analysis is still missing in the community. The major difficulty of analysis is because of the well-known sign-flipping phenomenon of FastICA, which causes the discontinuity of the corresponding FastICA map on the unit sphere. In this paper, by using the concept of principal fiber bundles, FastICA is proven to be locally quadratically convergent to a correct separation. Higher order local convergence properties of FastICA are also investigated in the framework of a scalar shift strategy. Moreover, as a parallelized version of FastICA, the so-called QR FastICA algorithm, which employs the QR decomposition (Gram-Schmidt orthonormalization process) instead of the polar decomposition, is shown to share similar local convergence properties with the original FastICA.
Hao Shen 0002, Martin Kleinsteuber, Knut Hüper
IEEE Trans. Neural Networks1
2007 Generalised Fastica for Independent Subspace Analysis
abstract
Independent subspace analysis (ISA) was developed as an extension of independent component analysis (ICA) when statistical independences are assumed to exist between groups of components rather than between individual components. Due to the superiority of FastICA against other linear ICA algorithms, an intuitive analogy, the so-called FastISA algorithm, has been proposed to solve the problem of ISA. Experimental evidences so far have shown the capability of FastISA, regardless of any independence criterion. Since standard FastICA can be viewed as a special case of an approximate Newton ICA method and moreover can be generalised as a scalar shifted fixed point algorithm, in this work, we propose two new classes of ISA algorithms, an approximate Newton-like ISA method and a matrix shifted fixed point ISA algorithm on the Grabmann manifold. As an aside, FastISA is a special case in the class of matrix shifted fixed point ISA algorithms. Performances of the proposed algorithms are investigated by numerical experiments.
Hao Shen 0002, Knut Hüper
ICASSP (4)1
2006 Local Convergence Properties of Fastica and Some Generalisations
abstract
In recent years, algorithms to perform Independent Component Analysis in blind identification, localisation of sources or more general in data analysis have been developed. Prominent example certainly is the socalled FastICA algorithms from the Finnish school. In this paper we will generalise the FastICA algorithm considered as a discrete dynamical system on the unit sphere to the case where all units converge simultaneously, i.e., we consider some kind of parallel FastICA algorithm living on the orthogonal group. In addition we present a local convergence analysis for the algorithms proposed in this paper building on earlier work. It turns out that one can treat these type of algorithms in a similar manner as the Rayleigh quotient iteration, well known in numerical linear algebra, i.e. considering the algorithm as a discrete dynamical system on a suitable manifold. The algorithms presented here are compared by several numerical experiments and simulations.
Knut Hüper, Hao Shen 0002, Abd-Krim Seghouane
ICASSP (5)2
2006 Newton-Like Methods for Nonparametric Independent Component Analysis
Hao Shen 0002, Knut Hüper, Alexander J. Smola
ICONIP (1)1