EDBT 2026 Demo / reviewers in the wild / expert
Xuguang Lan
dblp:86/6892
· DBLP profile ↗
97ranked-venue papers
9as first author
45since 2021 · last 2026
0000-0002-3422-944XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 2 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 4 first-author · 9 since 2021Systems, architecture and hardware · 19 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust adaptive dynamic programming for morphing air-breathing hypersonic vehicles under unmatched uncertainty
Wenxin Guo, Jieyu Liu, Weiwei Qin, Xuguang Lan, Hongyang Bai |
Sci. China Inf. Sci. | 4 |
| 2026 | High-performance multi-agent path finding in high-obstacle-density and large-size maps
Shiguang Sun, Chang Tang, Shi-tao Chen, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan |
Neurocomputing | 7 |
| 2026 | Relative depth knowledge distillation for generalizable monocular depth estimation
Mankun Li, Meng Yang 0002, Xuguang Lan, Ce Zhu |
Neurocomputing | 4 |
| 2026 | MInCo: Mitigating conflicting objectives in distracted visual model-based reinforcement learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan |
Knowl. Based Syst. | 7 |
| 2026 | Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement LearningabstractThe increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase. Zeyang Liu 0001, Lipeng Wan 0003, Shiguang Sun, Xue Sui, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | DualSkill: Unifying Discrete Stability and Continuous Flexibility for Embodied Control
Ziru Wang, Haowen Sun 0003, Zeyang Liu 0001, Xuguang Lan |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | D2TriPO-DETR: Dual-Decoder Triple-Parallel-Output Detection TransformerabstractVision-based grasping, though widely employed for industrial and household applications, still struggles with object stacking scenarios. Current methods face three major challenges: limited inter-object relationship understanding; poor grasping adaptation across different viewpoints; and error propagation. To address the above challenges, we propose D2TriPO-DETR, a dual-decoder transformer with three outputs, of which are object detection, manipulation relationship, and grasp detection. Specifically, a distributed attention perception module and a rotation attention invariance module are designed to address limited interobject relationship understanding and poor grasping adaptation across different viewpoints. These two modules are respectively integrated into the two parallel decoders to output the triple results simultaneously, partly eliminating task-level error propagation. Experimental results on the visual manipulation relationship dataset indicate that D2TriPO-DETR outperforms existing state-of-the-art methods across all metrics, e.g., +6.1% object detection recall, +6.7% manipulation relationship image accuracy, and +1.5% grasp detection accuracy. Extensive real-world experiments and quantitative results validate D2TriPO-DETR’s effectiveness. Menghao Pu, Chaoqun Han, Zhiping Chai, Pu Wen, Jihong Zhu 0002, Chao Wang 0096, Han Ding 0002, Xuguang Lan |
IEEE Trans. Ind. Informatics | 9 |
| 2026 | ETLight: An Evolution Transformer for Efficient Traffic Signal ControlabstractTraffic signal control (TSC) is still one of the most challenging and promising research issues in the field of transportation. Since traditional methods have difficulty in handling dynamically changing traffic flows, reinforcement learning (RL) methods have been introduced into TSC. However, the cost of practical application is critically high due to multiple sampling trials and long learning process. The Transformer architecture has recently attained remarkable results in natural language processing (NLP), but when applied to the field of RL, the standard Transformer architecture is difficult to optimize and faces the problem of hyperparameter sensitivity. In the paper, we transform TSC into a sequence modeling issue and propose a new evolution Transformer architecture to adjust the autoregressive model through reward, past states and actions in the traffic environment to directly generate the best predicted action. In addition, we use the feature evolution module (FEM) instead of residual connections to make the learning process more stable and efficient. Through experiments on public datasets, we demonstrate that our ETLight model achieves a state-of-the-art (SOTA): 1) It achieves the overall best performance on average travel time (ATT) metric, with improvements of up to 6.85%, 3.73% and 3.10% over the best conventional, RL and Transformer methods, respectively; 2) It has a more stable learning process, faster learning speed and better convergence compared to published TSC methods so far; and; 3) it has good robustness and is less sensitive to hyperparameter selection. Meiqin Liu 0001, Senlin Zhang, Ronghao Zheng, Shanling Dong, Xuguang Lan |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Bootstrapped Model Predictive ControlabstractModel Predictive Control (MPC) has been demonstrated to be effective in continuous control tasks. When a world model and a value function are available, planning a sequence of actions ahead of time leads to a better policy. Existing methods typically obtain the value function and the corresponding policy in a model-free manner. However, we find that such an approach struggles with complex tasks, resulting in poor policy learning and inaccurate value estimation. To address this problem, we leverage the strengths of MPC itself. In this work, we introduce Bootstrapped Model Predictive Control (BMPC), a novel algorithm that performs policy learning in a bootstrapped manner. BMPC learns a network policy by imitating an MPC expert, and in turn, uses this policy to guide the MPC process. Combined with model-based TD-learning, our policy learning yields better value estimation and further boosts the efficiency of MPC. We also introduce a lazy reanalyze mechanism, which enables computationally efficient imitation learning. Our method achieves superior performance over prior works on diverse continuous control tasks. In particular, on challenging high-dimensional locomotion tasks, BMPC significantly improves data efficiency while also enhancing asymptotic performance and training stability, with comparable training time and smaller network sizes. Code is available at https://github.com/wertyuilife2/bmpc. Hanwei Guo, Xuguang Lan |
ICLR | 5 |
| 2025 | Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought LeapsabstractHumor is previously regarded as a gift exclusive to humans for the following reasons. Humor is a culturally nuanced aspect of human language, presenting challenges for its understanding and generation.
Humor generation necessitates a multi-hop reasoning process, with each hop founded on proper rationales. Although many studies, such as those related to GPT-o1, focus on logical reasoning with reflection and correction, they still fall short in humor generation. Due to the sparsity of the knowledge graph in creative thinking, it is arduous to achieve multi-hop reasoning.
Consequently, in this paper, we propose a more robust framework for addressing the humor reasoning task, named LoL. LoL aims to inject external information to mitigate the sparsity of the knowledge graph, thereby enabling multi-hop reasoning. In the first stage of LoL, we put forward an automatic instruction-evolution method to incorporate the deeper and broader thinking processes underlying humor.
Judgment-oriented instructions are devised to enhance the model's judgment capability, dynamically supplementing and updating the sparse knowledge graph. Subsequently, through reinforcement learning, the reasoning logic for each online-generated response is extracted using GPT-4o. In this process, external knowledge is re-introduced to aid the model in logical reasoning and the learning of human preferences.
Finally, experimental results indicate that the combination of these two processes can enhance both the model's judgment ability and its generative capacity.
These findings deepen our comprehension of the creative capabilities of large language models (LLMs) and offer approaches to boost LLMs' creative abilities for cross-domain innovative applications. Sinbadliu, Xuguang Lan |
ICLR | 6 |
| 2025 | Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan |
AAMAS | 8 |
| 2025 | State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect SimulatorabstractIn reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL to jointly learn transferable policies from limited offline data and imperfect simulators. However, due to the unrestricted exploration in the imperfect simulator, the hybrid offline-and-online RL methods inevitably suffer from low sample efficiency and insufficient state-action space coverage during training. To solve this problem, we propose a State Revisit and Re-exploration (SR2) hybrid offline-and-online RL framework. In particular, the proposed algorithm employs a meta-policy and a sub-policy, where the meta-policy aims to find high-quality states in the offline trajectories for online exploration, and the sub-policy learns the robot skill using mixed offline and online data. By introducing the state revisit and explore mechanism, our approach efficiently improves performance on a set of sim-to-real robotic tasks. Through extensive simulation and real-world tasks, we demonstrate the superior performance of our approach against other state-of-the-art methods. Xingyu Chen 0001, Jiayi Xie, Ruixun Liu, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan |
IJCAI | 8 |
| 2025 | Consistent Feature Alignment for Cross-Modal Knowledge Distillation in Monocular 3D Object DetectionabstractCross-modal knowledge distillation (CMKD) in monocular 3D object detection transfers LiDAR’s accurate depth information to compensate for the limitations of camera model. However, current methods directly align the intermediate features of the teacher and student networks, in which the modality gap between LiDAR and camera hinders their effectiveness. To mitigate this issue, we design two modules, namely, Consistent Alignment Module (CAM) and Deformable Adapter Module (DAM) to reduce the modality gap of CMKD. The CAM transforms intermediate features of LiDAR and camera into some consistent features through a lightweight Target Head. It is based on the observation that some high-level features such as heatmaps and depths are highly correlated in CMKD, though modality gap appears between LiDAR and camera. Therefore, these features can be effectively transferred from teacher to student in CMKD. The DAM introduces a deformable adapter for the intermediate features of the student network to reduce background noise in CMKD. This helps to dynamically align its intermediate features with the teacher network. We then propose a Consistent Feature Alignment network (MonoCFA) for CMKD to boost monocular 3D object detection. Our network integrates the two designed modules at different levels of the teacher and student networks, in order to align the intermediate features of LiDAR and camera more accurately and reliably. Our model can be widely applied to existing monocular 3D object detection models. For validation, we choose the representative MonoDLE, GUPNet, and DID-M3D as base models. Experiments on the KITTI benchmark show that our method significantly outperforms the three base models by 39%, 15.5%, and 15%, respectively, and achieves state-of-the-art when compared to other CMKD models. Meng Yang 0002, Xuguang Lan |
IROS | 4 |
| 2025 | Towards Extrinsic Dexterity Grasping in Unrestricted EnvironmentsabstractGrasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls or table edges to facilitate object grasping. However, they are confined to task-specific policies while neglecting semantic perception and planning to identify optimal pre-grasp configurations. This limits their operational versatility, impeding effective adaptation to varied extrinsic dexterity constraints. In this work, we present ExDiff, a robot manipulation approach for extrinsic dexterity grasping in unrestricted environments. It utilizes Vision-Language Models (VLMs) to perceive the environmental state and generate instructions, followed by a Goal-Conditioned Action Diffusion (GCAD) model to predict the sequence of low-level actions. This diffusion model learns the low-level policy, conditioned on high-level instructions and cumulative rewards, which improves the generation of robot actions. Simulation experiments and real-world deployment results demonstrate that ExDiff effectively performs ungraspable tasks and generalizes to previously unseen target objects and scenes. Videos at - https://exdiff.github.io/index.html Chengzhong Ma, Houxue Yang, Hanbo Zhang, Zeyang Liu 0001, Xuguang Lan, Nanning Zheng 0001 |
IROS | 7 |
| 2025 | Relationship detection for manipulation in object stacking scene with fully connected CRF
Mengyuan Ding, Chenjie Yang, Xuguang Lan, Nanning Zheng 0001 |
Neurocomputing | 4 |
| 2025 | Dual Graph Attention Networks for Multi-View Visual Manipulation Relationship Detection and Robotic GraspingabstractVisual manipulation relationship detection facilitates robots to achieve safe, orderly, and efficient grasping tasks. However, most existing algorithms only model object-level or relational-level dependency individually, lacking sufficient global information, which is difficult to handle different types of reasoning errors, especially in complex environments with multi-object stacking and occlusion. To solve the above problems, we propose Dual Graph Attention Networks (Dual-GAT) for visual manipulation relationship detection, with an object-level graph network for capturing object-level dependencies and a relational-level graph network for capturing relational triplets-level interactions. The attention mechanism assigns different weights to different dependencies, obtains more accurate global context information for reasoning, and gets a manipulation relationship graph. In addition, we use multi-view feature fusion to improve the occluded object features, then enhance the relationship detection performance in multi-object scenes. Finally, our method is deployed on the robot to construct a multi-object grasping system, which can be well applied to stacking environments. Experimental results on the datasets VMRD and REGRAD show that our method significantly outperforms others. Note to Practitioners—The motivation of this research paper is to design efficient and accurate visual manipulation relationship reasoning methods to accomplish relevant grasping tasks in complex stacked scenes with multiple objects. The problem requires judging the positional relationship between objects in a stacked scene and determining the appropriate grasping order. In order to accurately grasp the target objects, this paper proposes a Dual Graph Attention Network for visual manipulation relationship detection, which utilizes object-level and relationship-level dependencies to obtain accurate global information for better performance. Meanwhile, multi-view feature fusion effectively improves the object occlusion problem. Our method can be combined with grasping detection to realize the robot grasping task in a real environment with multiple objects stacked in cluttered scenes. Practitioners can apply our method to real-time robot operating systems. Mengyuan Ding, Yaorui Shi, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Flight Mastery in Turbulent Skies: Shared Control and Curriculum Reinforcement Learning for Crosswind LandingabstractLanding in crosswind conditions poses significant challenges for aircraft, as traditional control methods often fail to ensure stability in rapidly changing wind environments. While reinforcement learning offers a promising alternative, it typically suffers from low sample efficiency and limited generalization under stochastic wind fields. To address these challenges, we propose a Shared Control and Curriculum Reinforcement Learning framework. We model the crosswind landing task as a Markov Decision Process (MDP), explicitly defining the state space, action space, and wind field representation. To initialize learning, we decompose the multi-objective landing task into four sub-tasks—altitude, attitude, heading, and speed control—and train expert policies for each. These are then distilled into a shared control model via behavior cloning, providing a pre-trained policy with basic flight control capabilities. We further fine-tune this model using curriculum reinforcement learning, progressively increasing the complexity of wind conditions to enhance robustness and generalization. Experimental results across multiple aircraft and wind scenarios show that our method improves landing success rates and trajectory smoothness, while generalizing more effectively to unseen wind conditions, outperforming PID controllers, imitation learning, and mainstream RL baselines. Zechen Shi, Xingyu Chen 0001, Zeyang Liu 0001, Chi Zhang 0020, Yimeng Yu, Junbin You, Xuguang Lan |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | Improving Offline Reinforcement Learning With in-Sample Advantage Regularization for Robot ManipulationabstractOffline reinforcement learning (RL) aims to learn the possible policy from a fixed dataset without real-time interactions with the environment. By avoiding the risky exploration of the robot, this approach is expected to significantly improve the robot's learning efficiency and safety. However, due to errors in value estimation from out-of-distribution actions, most offline RL algorithms constrain or regularize the policy to the actions contained within the dataset. The cost of such methods is the introduction of new hyperparameters and additional complexity. In this article, we aim to adapt offline RL to robotic manipulation with minimal changes and to avoid evaluating out-of-distribution actions as much as possible. Therefore, we improve offline RL with in-sample advantage regularization (ISAR). To mitigate the impact of unseen actions, the ISAR learns the state-value function only with the dataset sample to regress the optimal action-value function. Our method calculates the advantage function of action-state pairs based on in-sample value estimation and adds a behavior cloning (BC) regularization term in the policy update. This improves sample efficiency with minimal changes, resulting in a simple and easy-to-implement method. The experiments of the D4RL robot benchmark and multigoal sparse rewards robotic tasks show that the ISAR achieves excellent performance comparable to current state-of-the-art algorithms without the need for complex parameter tuning and too much training time. In addition, we demonstrate the effectiveness of our method on a real-world robot platform. Chengzhong Ma, Deyu Yang, Zeyang Liu 0001, Houxue Yang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement LearningabstractPrioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings. Ziru Wang, Wanli Jiang, Ru Peng, Qian Kou, Lipeng Wan 0003, Xuguang Lan |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement LearningabstractEffective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models. Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan |
AAAI | 6 |
| 2024 | Relation DETR: Exploring Explicit Position Relation Prior for Object Detection
Xiuquan Hou, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen, Xuguang Lan |
ECCV (50) | 6 |
| 2024 | Grasp Manipulation Relationship Detection based on Graph Sample and AggregationabstractIn multi-object stacking scenarios, exploring the relationships among objects and determining the correct sequence of operations are crucial for robotic manipulation. However, previous algorithms inefficiently combine global and local information, often focusing solely on the local features of objects or the interactions of object features at a global level. This approach leads to imbalanced distribution of features and the generation of redundant or missing relationships in complex scenes, such as multi-object stacking and partial occlusion. To address this issue, we have developed a grasp manipulation relationship detection algorithm called Graph Sampling Aggregation Network for Visual Manipulation Relationship Detection (GSAGED). This algorithm assists robots in detecting targets in complex scenes and determining the appropriate grasping order. Firstly, the Positional Encoding Module in GSAGED enhances object feature information by considering global contexts. Secondly, the Graph Sampling Aggregation method effectively integrates global and local information, relieving imbalanced distribution of features. Finally, we applied the developed algorithm to a physical robot for grasping. Experimental results on the Visual Manipulation Relationship Dataset (VMRD) and the large-scale relational grasp dataset named REGRAD demonstrate that our method significantly improves the accuracy of relationship detection in complex scenes and exhibits robust generalization capabilities in real-world applications. Jiayuan Luo, Mengyuan Ding, Xuguang Lan |
ICRA | 5 |
| 2024 | Towards Unified Interactive Visual Grounding in The WildabstractInteractive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user’s input by active information gathering. Previous approaches often rely on predefined templates to ask disambiguation questions, resulting in performance reduction in realistic interactive scenarios. In this paper, we propose TiO, an end-to-end system for interactive visual grounding in human-robot interaction. Benefiting from a unified formulation of visual dialog and grounding, our method can be trained on a joint of extensive public data, and show superior generality to diversified and challenging open-world scenarios. In the experiments, we validate TiO on GuessWhat?! and InViG benchmarks, setting new state-of-the-art performance by a clear margin. Moreover, we conduct HRI experiments on the carefully selected 150 challenging scenes as well as real-robot platforms. Results show that our method demonstrates superior generality to diversified visual and language inputs with a high success rate. Codes and demos are available on https://jxu124.github.io/TiO/. Hanbo Zhang, Qingyi Si, Xuguang Lan, Tao Kong |
ICRA | 5 |
| 2024 | Experience Consistency Distillation Continual Reinforcement Learning for Robotic Manipulation TasksabstractContinual reinforcement learning, which aims to help robots acquire skills without catastrophic forgetting, obviating the need to re-learn all tasks from scratch. In order to enable lifelong acquisition of skills in robots, replay-based continual reinforcement learning has emerged as a promising research direction. These techniques replay data from previous tasks to mitigate forgetting when learning new skills. However, existing replay-based methods store poor representative experience, and the experience utilization of old tasks is inefficient. To address these issues, we propose an experience consistency distillation method for robot continual reinforcement learning to improve the data efficiency of the experience. Specifically, the experience of old tasks are distilled to obtain Markov Decision Process (MDP) data with high compression ratio and information content. To ensure consistent data distributions before and after distillation, we further utilize a Fréchet Inception Distance (FID) loss as a regularization constraint. In order to improve experience utilization efficiency, the policy is then trained using both the distilled data and current task data, with policy distillation performed based on uncertainty metrics. Our method is validated in the continual reinforcement learning simulation platform and real scene with a UR5e robot arm. Experimental results indicate that our method achieves higher success and lower buffer size requirement compared to other methods. Ru Peng, Xingyu Chen 0001, Xuguang Lan |
ICRA | 6 |
| 2024 | Grounded Answers for Multi-agent Decision-making Problem through Generative World ModelabstractRecent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error experience and reasoning as humans. To address this limitation, we explore a paradigm that integrates a language-guided simulator into the multi-agent reinforcement learning pipeline to enhance the generated answer. The simulator is a world model that separately learns dynamics and reward, where the dynamics model comprises an image tokenizer as well as a causal transformer to generate interaction transitions autoregressively, and the reward model is a bidirectional transformer learned by maximizing the likelihood of trajectories in the expert demonstrations under language guidance. Given an image of the current state and the task description, we use the world model to train the joint policy and produce the image sequence as the answer by running the converged policy on the dynamics model. The empirical results demonstrate that this framework can improve the answers for multi-agent decision-making problems by showing superior performance on the training and unseen tasks of the StarCraft Multi-Agent Challenge benchmark. In particular, it can generate consistent interaction sequences and explainable reward functions at interaction states, opening the path for training generative models of the future. Zeyang Liu 0001, Shiguang Sun, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan |
NeurIPS | 7 |
| 2024 | Robust control for affine nonlinear systems under the reinforcement learning framework
Wenxin Guo, Weiwei Qin, Xuguang Lan, Jieyu Liu, Zhaoxiang Zhang 0007 |
Neurocomputing | 3 |
| 2024 | A causality guided loss for imbalanced learning in scene graph generation
Ru Peng, Xingyu Chen 0001, Ziru Wang, Xuguang Lan |
Neurocomputing | 7 |
| 2024 | Optimal bipartite graph matching-based goal selection for policy-based hindsight learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan |
Neurocomputing | 5 |
| 2024 | Multi-agent evaluation for energy management by practically scaling α-rankabstractCurrently, decarbonization has become an emerging trend in the power system arena. However, the increasing number of photovoltaic units distributed into a distribution network may result in voltage issues, providing challenges for voltage regulation across a large-scale power grid network. Reinforcement learning based intelligent control of smart inverters and other smart building energy management (EM) systems can be leveraged to alleviate these issues. To achieve the best EM strategy for building microgrids in a power system, this paper presents two large-scale multi-agent strategy evaluation methods to preserve building occupants’ comfort while pursuing system-level objectives. The EM problem is formulated as a general-sum game to optimize the benefits at both the system and building levels. The α -rank algorithm can solve the general-sum game and guarantee the ranking theoretically, but it is limited by the interaction complexity and hardly applies to the practical power system. A new evaluation algorithm (TcEval) is proposed by practically scaling the α -rank algorithm through a tensor complement to reduce the interaction complexity. Then, considering the noise prevalent in practice, a noise processing model with domain knowledge is built to calculate the strategy payoffs, and thus the TcEval-AS algorithm is proposed when noise exists. Both evaluation algorithms developed in this paper greatly reduce the interaction complexity compared with existing approaches, including ResponseGraphUCB (RG-UCB) and α InformationGain ( α -IG). Finally, the effectiveness of the proposed algorithms is verified in the EM case with realistic data. Yiyun Sun, Senlin Zhang, Meiqin Liu 0001, Ronghao Zheng, Shanling Dong, Xuguang Lan |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2024 | Knowledge Graph Enhancement for Fine-Grained Zero-Shot Learning on ImageNet21KabstractFine-grained Zero-shot Learning on the large-scale dataset ImageNet21K is an important task that has promising perspectives in many real-world scenarios. One typical solution is to explicitly model the knowledge passing using a Knowledge Graph (KG) to transfer knowledge from seen to unseen instances. By analyzing the hierarchical structure and the word descriptions on ImageNet21K, we find that the noisy semantic information, the sparseness of seen classes, and the lack of supervision of unseen classes make the knowledge passing insufficient, which limits the KG-based fine-grained ZSL. To resolve this problem, in this paper, we enhance the knowledge passing from three aspects. First, we use more powerful models such as the Large Language Model and Vision-Language Model to get more reliable semantic embeddings. Then we propose a strategy that globally enhances the knowledge graph based on the convex combination relationship of the semantic embeddings. It effectively connects the edges between the non-kinship seen and unseen classes that have strong correlations while assigning an importance score to each edge. Based on the enhanced knowledge graph, we further present a novel regularizer that locally enhances the knowledge passing during training. We extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches. Xingyu Chen 0001, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked ScenesabstractManipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from a predefined view, which has limitation in visual dislocation in unstructured environments. Multi-view data provide more comprehensive information in space, while a challenge of multi-view MRD is domain shift. In this paper, we propose a novel multi-view fusion framework, namely multi-view MRD network (MMRDN), which is trained by 2D and 3D multi-view data. We project the 2D data from different views into a common hidden space and fit the embeddings with a set of Von-Mises-Fisher distributions to learn the consistent representations. Besides, taking advantage of position information within the 3D data, we select a set of$K$Maximum Vertical Neighbors (KMVN) points from the point cloud of each object pair, which encodes the relative position of these two objects. Finally, the features of multi-view 2D and 3D data are concatenated to predict the pairwise relationship of objects. Experimental results on the challenging REGRAD dataset show that MMRDN outperforms the state-of-the-art methods in multi-view MRD tasks. The results also demonstrate that our model trained by synthetic data is capable to transfer to real-world scenarios. Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
ICRA | 5 |
| 2023 | Deep Hierarchical Communication Graph in Multi-Agent Reinforcement LearningabstractSharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message-passing update in the undirected graph with cycles cannot guarantee convergence. To solve this problem, we propose Deep Hierarchical Communication Graph (DHCG) to learn the dependency relationships between agents based on their messages. The relationships are formulated as directed acyclic graphs (DAGs), where the selection of the proper topology is viewed as an action and trained in an end-to-end fashion. To eliminate the cycles in the graph, we apply an acyclicity constraint as intrinsic rewards and then project the graph in the admissible solution set of DAGs. As a result, DHCG removes redundant communication edges for cost improvement and guarantees convergence. To show the effectiveness of the learned graphs, we propose policy-based and value-based DHCG. Policy-based DHCG factorizes the joint policy in an auto-regressive manner, and value-based DHCG factorizes the joint value function to individual value functions and pairwise payoff functions. Empirical results show that our method improves performance across various cooperative multi-agent tasks, including Predator-Prey, Multi-Agent Coordination Challenge, and StarCraft Multi-Agent Challenge. Zeyang Liu 0001, Lipeng Wan 0003, Xue Sui, Zhuoran Chen, Kewu Sun, Xuguang Lan |
IJCAI | 6 |
| 2023 | Prioritized Planning for Target-Oriented Manipulation via Hierarchical Stacking Relationship PredictionabstractIn scenarios involving grasping multiple targets, the learning of stacking relationships between objects is fundamental for robots to execute safely and efficiently. However, current methods lack subdivision for the hierarchy of stacking relationship types. In scenes where objects are mostly stacked in an orderly manner, they are incapable of performing human-like and high-efficient grasping decisions. This paper proposes a perception-planning method to distinguish different stacking forms between objects and generate prioritized manipulation sequences based on given target designations. We utilize a Hierarchical Stacking Relationship Network (HSRN) to discriminate the hierarchy of stacking and generate a refined Stacking Relationship Tree (SRT) for relationship description. Considering objects with high stacking stability can be processed together if necessary, we introduce an elaborate decision-making planner based on Partially Observable Markov Decision Process (POMDP), which leverages observations and generates the least grasp-consuming decision chain with robustness and is suitable for simultaneously specifying multiple targets. To verify our work, we set the scene to the dining table and augment REGRAD dataset for network training. Experiments show that our method effectively generates grasping decisions that conform to human requirements, and improves the implementation efficiency compared with existing methods on the basis of guaranteeing success rate. Zewen Wu, Xingyu Chen 0001, Chengzhong Ma, Xuguang Lan, Nanning Zheng 0001 |
IROS | 5 |
| 2022 | Constrained Contrastive Reinforcement Learning
Xuguang Lan |
ACML | 4 |
| 2022 | Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement LearningabstractDue to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the best team performance). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure the optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and eliminates the non-optimal STNs via superior experience replay. Theoretical proofs and empirical results demonstrate that given the true Q values, GVR ensures the optimal consistency under sufficient exploration. Besides, in tasks where the true Q values are unavailable, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks. Lipeng Wan 0003, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
ICML | 4 |
| 2022 | A Continuous Learning Approach for Probabilistic Human Motion PredictionabstractHuman Motion Prediction (HMP) plays a crucial role in safe Human-Robot-Interaction (HRI). Currently, the majority of HMP algorithms are trained by massive pre-collected data. As the training data only contains a few pre-defined motion patterns, these methods cannot handle the unfamiliar motion patterns. Moreover, the pre-collected data are usually non-interactive, which does not consider the real-time responses of collaborators. As a result, these methods usually perform unsatisfactorily in real HRI scenarios. To solve this problem, in this paper, we propose a novel Continual Learning (CL) approach for probabilistic HMP which makes the robot continually learns during its interaction with collaborators. The proposed approach consists of two steps. First, we leverage a Bayesian Neural Network to model diverse uncertainties of observed human motions for collecting online interactive data safely. Then we take Experience Replay and Knowledge Distillation to elevate the model with new experiences while maintaining the knowledge learned before. We first evaluate our approach on a large-scale benchmark dataset Human3.6m. The experimental results show that our approach achieves a lower prediction error compared with the baselines methods. Moreover, our approach could continually learn new motion patterns without forgetting the learned knowledge. We further conduct real-scene experiments using Kinect DK. The results show that our approach can learn the human kinematic model from scratch, which effectively secures the interaction. Shihong Wang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
ICRA | 5 |
| 2022 | Visual Manipulation Relationship Detection based on Gated Graph Neural Network for Robotic GraspingabstractExploring the relationship among objects and giving the correct operation sequence is vital for robotic manipulation. However, most previous algorithms only model the relationship between pairs of objects independently, ignoring the interaction effect between them, which may generate redundant or missing relations in complex scenes, such as multi-object stacking and partial occlusion. To solve this problem, a Gated Graph Neural Network (GGNN) is designed for visual manipulation relationship detection, which can help robots detect targets in complex scenes and obtain the appropriate grasping order. Firstly, the robot extracts feature from the input image and estimate object categories. Then GGNN is used to effectively capture the dependencies between objects in the whole scene, update the relevant features, and output the grasping sequence. In addition, by embedding positional encoding into pair object features, accurate context information is obtained to reduce the adverse effects of complex scenes. Finally, the constructed algorithm is applied to the physical robot for grasping. Experiment results on the Visual Ma-nipulation Relationship Dataset (VMRD) and the large-scale relational grasp dataset named REGRAD show that our method significantly improves the accuracy of relationship detection in complex scenes, and can be well generalized in the real world. Mengyuan Ding, Chenjie Yang, Xuguang Lan |
IROS | 4 |
| 2022 | Deep semantic space guided multi-scale neural style transfer
Jiachen Yu, Youzi Xiao, Xuguang Lan |
Multim. Tools Appl. | 6 |
| 2022 | Depth Map Recovery Based on a Unified Depth Boundary Distortion ModelabstractDepth maps acquired by either physical sensors or learning methods are often seriously distorted due to boundary distortion problems, including missing, fake, and misaligned boundaries (compared with RGB images). An RGB-guided depth map recovery method is proposed in this paper to recover true boundaries in seriously distorted depth maps. Therefore, a unified model is first developed to observe all these kinds of distorted boundaries in depth maps. Observing distorted boundaries is equivalent to identifying erroneous regions in distorted depth maps, because depth boundaries are essentially formed by contiguous regions with different intensities. Then, erroneous regions are identified by separately extracting local structures of RGB image and depth map with Gaussian kernels and comparing their similarity on the basis of the SSIM index. A depth map recovery method is then proposed on the basis of the unified model. This method recovers true depth boundaries by iteratively identifying and correcting erroneous regions in recovered depth map based on the unified model and a weighted median filter. Because RGB image generally includes additional textural contents compared with depth maps, texture-copy artifacts problem is further addressed in the proposed method by restricting the model works around depth boundaries in each iteration. Extensive experiments are conducted on five RGB-depth datasets including depth map recovery, depth super-resolution, depth estimation enhancement, and depth completion enhancement. The results demonstrate that the proposed method considerably improves both the quantitative and visual qualities of recovered depth maps in comparison with fifteen competitive methods. Most object boundaries in recovered depth maps are corrected accurately, and kept sharply and well aligned with the ones in RGB images. Haotian Wang 0009, Meng Yang 0002, Xuguang Lan, Ce Zhu, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Generalized Zero-Shot Learning Via Multi-Modal Aggregated Posterior Aligning Neural NetworkabstractThe visual-semantic gap between the visual space (visual features) and semantic space (semantic attributes) is one of the main problems in the Generalized Zero-Shot Learning (GZSL) task. The essence of this problem is that the structure of manifolds in these two spaces is inconsistent, which makes it difficult to learn embeddings that unify visual features and semantic attributes for similarity measurement. In this work, we tackle this problem by proposing a multi-modal aggregated posterior aligning neural network based on Wasserstein Auto-encoders (WAE) which learns a shared latent space for visual features and semantic attributes. The key to our approach is that the aggregated posterior distribution of the latent representations encoded from visual features of each class is encouraged to be aligned with a Gaussian distribution predicted by the corresponding semantic attribute in the latent space. On one hand, requiring the latent manifolds of visual features and semantic attributes to be consistent preserves the inter-class association between seen and unseen classes. On the other hand, the aggregated posterior of each class is directly defined as a Gaussian in the latent space, which provides a reliable way to synthesize latent features for training classification models. Using the AWA1, AWA2, CUB, aPY, FLO, and SUN benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches. Xingyu Chen 0001, Jin Li 0011, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Neighborhood Geometric Structure-Preserving Variational Autoencoder for Smooth and Bounded Data SourcesabstractMany data sources, such as human poses, lie on low-dimensional manifolds that are smooth and bounded. Learning low-dimensional representations for such data is an important problem. One typical solution is to utilize encoder-decoder networks. However, due to the lack of effective regularization in latent space, the learned representations usually do not preserve the essential data relations. For example, adjacent video frames in a sequence may be encoded into very different zones across the latent space with holes in between. This is problematic for many tasks such as denoising because slightly perturbed data have the risk of being encoded into very different latent variables, leaving output unpredictable. To resolve this problem, we first propose a neighborhood geometric structure-preserving variational autoencoder (SP-VAE), which not only maximizes the evidence lower bound but also encourages latent variables to preserve their structures as in ambient space. Then, we learn a set of small surfaces to approximately bound the learned manifold to deal with holes in latent space. We extensively validate the properties of our approach by reconstruction, denoising, and random image generation experiments on a number of data sources, including synthetic Swiss roll, human pose sequences, and facial expression images. The experimental results show that our approach learns more smooth manifolds than the baselines. We also apply our approach to the tasks of human pose refinement and facial expression image interpolation where it gets better results than the baselines. Xingyu Chen 0001, Chunyu Wang 0001, Xuguang Lan, Nanning Zheng 0001, Wenjun Zeng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Probabilistic Human Motion Prediction via A Bayesian Neural NetworkabstractHuman motion prediction is an important and challenging topic that has promising prospects in efficient and safe human-robot-interaction systems. Currently, the majority of the human motion prediction algorithms are based on deterministic models, which may lead to risky decisions for robots. To solve this problem, we propose a probabilistic model for human motion prediction in this paper. The key idea of our approach is to extend the conventional deterministic motion prediction neural network to a Bayesian one. On one hand, our model could generate several future motions when given an observed motion sequence. On the other hand, by calculating the Epistemic Uncertainty and the Heteroscedastic Aleatoric Uncertainty, our model could tell the robot if the observation has been seen before and also give the optimal result among all possible predictions. We extensively validate our approach on a large scale benchmark dataset Human3.6m. The experiments show that our approach performs better than deterministic methods. We further evaluate our approach in a Human-Robot-Interaction (HRI) scenario. The experimental results show that our approach makes the interaction more efficient and safer. Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001 |
ICRA | 3 |
| 2021 | REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point CloudsabstractReliable robotic grasping in unstructured environments is a crucial but challenging task. The main problem is to generate the optimal grasp of novel objects from partial noisy observations. This paper presents an end-to-end grasp detection network taking one single-view point cloud as input to tackle the problem. Our network includes three stages: Score Network (SN), Grasp Region Network (GRN), and Refine Network (RN). Specifically, SN regresses point grasp confidence and selects positive points with high confidence. Then GRN conducts grasp proposal prediction on the selected positive points. RN generates more accurate grasps by refining proposals predicted by GRN. To further improve the performance, we propose a grasp anchor mechanism, in which grasp anchors with assigned gripper orientations are introduced to generate grasp proposals. Experiments demonstrate that REGNet achieves a success rate of 79.34% and a completion rate of 96% in real-world clutter, which significantly outperforms several state-of-the-art point-cloud based methods, including GPD, PointNetGPD, and S4G. The code is available at https://github.com/zhaobinglei/REGNet for 3D Grasping. Binglei Zhao 0002, Hanbo Zhang, Xuguang Lan, Nanning Zheng 0001 |
ICRA | 3 |
| 2021 | Hindsight Trust Region Policy OptimizationabstractReinforcement Learning (RL) with sparse rewards is a major challenge. We pro- pose Hindsight Trust Region Policy Optimization (HTRPO), a new RL algorithm that extends the highly successful TRPO algorithm with hindsight to tackle the challenge of sparse rewards. Hindsight refers to the algorithm’s ability to learn from information across goals, including past goals not intended for the current task. We derive the hindsight form of TRPO, together with QKL, a quadratic approximation to the KL divergence constraint on the trust region. QKL reduces variance in KL divergence estimation and improves stability in policy updates. We show that HTRPO has similar convergence property as TRPO. We also present Hindsight Goal Filtering (HGF), which further improves the learning performance for suitable tasks. HTRPO has been evaluated on various sparse-reward tasks, including Atari games and simulated robot control. Experimental results show that HTRPO consistently outperforms TRPO, as well as HPG, a state-of-the-art policy 14 gradient algorithm for RL with sparse rewards. Hanbo Zhang, Cedar Site Bai, Xuguang Lan, David Hsu, Nanning Zheng 0001 |
IJCAI | 3 |
| 2021 | A Real-Time Robotic Grasping Approach With Oriented Anchor BoxabstractGrasping is an essential skill for robots to interact with humans and the environment. In this paper, we build a vision-based, robust, and real-time robotic grasping approach with fully convolutional neural network. The main component of our approach is a grasp detection network with oriented anchor boxes as detection priors. Because the orientation of detected grasps is significant, which determines the rotation angle configuration of the gripper, we propose the orientation anchor box mechanism to regress grasp angle based on predefined assumption instead of classification or regression without any priors. With oriented anchor boxes, the grasps can be predicted more accurately and efficiently. Besides, to accelerate the network training and further improve the performance of angle regression, angle matching is proposed during training instead of Jaccard index matching. Fivefold cross-validation results demonstrate that our proposed algorithm achieves an accuracy of 98.8% and 97.8% in image-wise split and object-wise split, respectively, and the speed of our detection algorithm is 67 frames per second (FPS) with GTX 1080Ti, outperforming all the current state-of-the-art grasp detection algorithms on Cornell Dataset both in speed and accuracy. Robotic experiments demonstrate the robustness and generalization ability in unseen objects and real-world environment, with the average success rate of 90.0% and 84.2% of familiar things and unseen things, respectively, on Baxter robot platform. Hanbo Zhang, Xinwen Zhou, Xuguang Lan, Jin Li 0011, Nanning Zheng 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | A Boundary Based Out-of-Distribution Classifier for Generalized Zero-Shot Learning
Xingyu Chen 0001, Xuguang Lan, Fuchun Sun 0001, Nanning Zheng 0001 |
ECCV (24) | 2 |
| 2020 | Multiresolution Mixture Generative Adversarial Network For Image Super-ResolutionabstractWith regard to the problem of image super-resolution (SR), generative adversarial network (GAN) can make generated images have more details and better effect on perceptual quality than other methods. However, GAN-based methods may lose the contour of object in some texture-intensive areas. In order to recover contour better and further enhance perceptual quality, we propose a Multiresolution Mixture Generative Adversarial Network for Image Super-Resolution (MRMGAN), which employs a multiresolution mixture network (MRMNet) for image super-resolution. The MRMNet is able to have multiple resolution feature maps at the same time when training. Meanwhile, we propose a residual fluctuation loss, which aims to reduce the overall fluctuation of residual between SR image and high-resolution (HR) image. We evaluated the proposed method on benchmark datasets. Experimental results show that the proposed MRMGAN can get satisfactory performance. Yudiao Wang, Xuguang Lan, Yinshu Zhang, Ruixue Miao |
ICME | 2 |
| 2020 | Autonomous Tool Construction with Gated Graph Neural NetworkabstractAutonomous tool construction is a significant but challenging task in robotics. This task can be interpreted as when given a reference tool, selecting some available candidate parts to reconstruct it. Most of the existing works perform tool construction in the form of action part and grasp part, which is only a specific construction pattern and limits its application to some extent. In general scenarios, a tool can be constructed in various patterns with different part pairs. Therefore, whether a part pair is most suitable for constructing the tool depends not only on itself, but on other parts in the same scene. To solve this problem, we construct a Gated Graph Neural Network (GGNN) to model the relations between all part pairs, so that we can select the candidate parts in consideration of the global information. Afterwards, we embed the constructed GGNN into a RCNN-like structure to finally accomplish tool construction. The whole model will be named Tool Construction Graph RCNN (TC-GRCNN). In addition, we develop a mechanism that can generate large-scale training and testing data in simulation environments, by which we can save the time of data collection and annotation. Finally, the proposed model is deployed on the physical robot. The experiment results show that TC-GRCNN can perform well in the general scenarios of tool construction. Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001 |
ICRA | 2 |
| 2020 | A review of object detection based on deep learning
Youzi Xiao, Jiachen Yu, Yinshu Zhang, Shaoyi Du, Xuguang Lan |
Multim. Tools Appl. | 7 |
| 2020 | Visual manipulation relationship recognition in object-stacking scenes
Hanbo Zhang, Xuguang Lan, Xinwen Zhou, Nanning Zheng 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | A Joint Label Space for Generalized Zero-Shot ClassificationabstractThe fundamental problem of Zero-Shot Learning (ZSL) is that the one-hot label space is discrete, which leads to a complete loss of the relationships between seen and unseen classes. Conventional approaches rely on using semantic auxiliary information, e.g. attributes, to re-encode each class so as to preserve the inter-class associations. However, existing learning algorithms only focus on unifying visual and semantic spaces without jointly considering the label space. More importantly, because the final classification is conducted in the label space through a compatibility function, the gap between attribute and label spaces leads to significant performance degradation. Therefore, this paper proposes a novel pathway that uses the label space to jointly reconcile visual and semantic spaces directly, which is named Attributing Label Space (ALS). In the training phase, one-hot labels of seen classes are directly used as prototypes in a common space, where both images and attributes are mapped. Since mappings can be optimized independently, the computational complexity is extremely low. In addition, the correlation between semantic attributes has less influence on visual embedding training because features are mapped into labels instead of attributes. In the testing phase, the discrete condition of label space is removed, and priori one-hot labels are used to denote seen classes and further compose labels of unseen classes. Therefore, the label space is very discriminative for the Generalized ZSL (GZSL), which is more reasonable and challenging for real-world applications. Extensive experiments on five benchmarks manifest improved performance over all of compared state-of-the-art methods. Jin Li 0011, Xuguang Lan, Yang Long 0001, Yang Liu 0069, Xingyu Chen 0001, Ling Shao 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Learning to Refine 3D Human Pose SequencesabstractWe present a basis approach to refine noisy 3D human pose sequences by jointly projecting them onto a non-linear pose manifold, which is represented by a number of basis dictionaries with each covering a small manifold region. We learn the dictionaries by jointly minimizing the distance between the original poses and their projections on the dictionaries, along with the temporal jittering of the projected poses. During testing, given a sequence of noisy poses which are probably off the manifold, we project them to the manifold using the same strategy as in training for refinement. We apply our approach to the monocular 3D pose estimation and the long term motion prediction tasks. The experimental results on the benchmark dataset shows the estimated 3D poses are notably improved in both tasks. In particular, the smoothness constraint helps generate more robust refinement results even when some poses in the original sequence have large errors. Jieru Mei, Xingyu Chen 0001, Chunyu Wang 0001, Alan L. Yuille, Xuguang Lan, Wenjun Zeng 0001 |
3DV | 5 |
| 2019 | Compressing Unknown Images With Product Quantizer for Efficient Zero-Shot ClassificationabstractFor Zero-Shot Learning (ZSL), the Nearest Neighbor (NN) search is generally conducted for classification, which may cause unacceptable computational complexity for large-scale datasets. To compress zero-shot classes by the trained quantizer for efficient search, it tends to induce large quantization error because distributions between seen and unseen classes are different. However, as semantic attributes of classes are available in ZSL, both seen and unseen classes have the same distribution for one specific property, e.g., animals have or not have spots. Based on this intuition, a Product Quantization Zero-Shot Learning (PQZSL) method is proposed to learn embeddings as well as quantizers to compress visual features into compact codes for Approximate NN (ANN) search. Particularly, visual features are projected into an orthogonal semantic space, and then the Product Quantization (PQ) is utilized to quantize individual properties. Experimental results on five benchmark datasets demonstrate that unseen classes are represented by the Cartesian product of quantized properties with little quantization error. As classes in orthogonal common space are more discriminative, the classification based on PQZSL achieves state-of-the-art performance in Generalized Zero-Shot Learning (GZSL) task, meanwhile, the speed of ANN search is 10-100 times higher than traditional NN search. Jin Li 0011, Xuguang Lan, Yang Liu 0069, Le Wang 0003, Nanning Zheng 0001 |
CVPR | 2 |
| 2019 | Preference Relationship-Based CrossCMN Scheme for Answer Ranking in Community QAabstractCommunity question answering (CQA) systems aim to provide users with high-quality answers. Nevertheless, unreliable answers are often returned to users in CQA systems, and the phenomenon causes that users have to browse multiple answers to find the best one. To improve such problem, we design a novel scheme, named PW-CrossCMN. The scheme ranks the candidate answers by pair-wise approach based on numerous historical documents. In the scheme, we apply the preference relationship into deep learning framework. Specifically, the scheme consists of two phases. In phase 1, the scheme extracts the features via automated feature engineering to construct the preference vectors and then divides the vectors into balanced positive and negative training samples based on the preference relationship. In phase 2, we build the CrossCMN model, which implements the multi-network parallel convolution and the cross forward propagation of full-connected layers, to achieve training and prediction tasks. Moreover, the multi-layer perception (MLP) is introduced to extract combination features in the prediction module. We perform extensive experiments on two typical datasets, and the results show that our scheme has more excellent performance in answer ranking task compared with several state-of-the-art baselines. In addition, we have released the relevant codes. Jianji Wang 0001, Xuguang Lan, Nanning Zheng 0001 |
ICDM | 3 |
| 2019 | Task-oriented Grasping in Object Stacking Scenes with CRF-based Semantic ModelabstractIn task-oriented grasping, the robot is supposed to manipulate the objects in a task-compatible manner, which is more important but more challenging than just stably grasping. However, most of existing works perform task-oriented grasping only in single object scenes. This greatly limits their practical application in real world scenes, in which there are usually multiple stacked objects with serious overlaps and occlusions. To perform task-oriented grasping in object stacking scenes, in this paper, we firstly build a synthetic dataset named Object Stacking Grasping Dataset (OSGD) for task-oriented grasping in object stacking scenes. Secondly, a Conditional Random Field (CRF) is constructed to model the semantic contents in object regions. The modelled semantic contents can be illustrated as incompatibility of task labels and continuity of task regions. This proposed approach can greatly reduce the interference of overlaps and occlusions in object stacking scenes. To embed the CRF-based semantic model into our grasp detection network, we implement the inference process of CRFs as a RNN so that the whole model, Task-oriented Grasping CRFs (TOG-CRFs) can be trained end to end. Finally, in object stacking scenes, the constructed model can help robot achieve 69.4% success rate for task-oriented grasping. Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001 |
IROS | 2 |
| 2019 | A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking ScenesabstractAutonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to grasp; (2) it is hard to deal with the situations in which the target is hidden or covered by other objects. In this paper, we propose a multi-task convolutional neural network for autonomous robotic grasping, which can help the robot find the target, make the plan for grasping and finally grasp the target step by step in object stacking scenes. We integrate vision-based robotic grasping detection and visual manipulation relationship reasoning in one single deep network and build the autonomous robotic grasping system. Experimental results demonstrate that with our model, Baxter robot can autonomously grasp the target with a success rate of 90.6%, 71.9% and 59.4% in object cluttered scenes, familiar stacking scenes and complex stacking scenes respectively. Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Lipeng Wan 0003, Chenjie Yang, Nanning Zheng 0001 |
IROS | 2 |
| 2019 | ROI-based Robotic Grasp Detection for Object Overlapping ScenesabstractGrasp detection considering the affiliations between grasps and their owner in object overlapping scenes is a necessary and challenging task for the practical use of the robotic grasping approach. In this paper, a robotic grasp detection algorithm named ROI-GD is proposed to provide a feasible solution to this problem based on Region of Interest (ROI), which is the region proposal for objects. ROI-GD uses features from ROIs to detect grasps instead of the whole scene. It has two stages: the first stage is to provide ROIs in the input image and the second-stage is the grasp detector based on ROI features. We also contribute a multi-object grasp dataset, (a) which is much larger than Cornell Grasp Dataset, by labeling Visual Manipulation Relationship Dataset. Experimental results demonstrate that ROI-GD performs much better in object overlapping scenes and at the meantime, remains comparable with state-of-the-art grasp detection algorithms on Cornell Grasp Dataset and Jacquard Dataset. Robotic experiments demonstrate that ROI-GD can help robots grasp the target in single-object and multi-object scenes with the overall success rates of 92.5% and 83.8% respectively. Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Xinwen Zhou, Nanning Zheng 0001 |
IROS | 2 |
| 2019 | Cross-View Person Identification Based on Confidence-Weighted Human Pose MatchingabstractCross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interest recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D pose estimation on videos/images collected in the wild. To address this problem, we first introduce a new metric of confidence to the estimated location of each human-body joint in 3D human pose estimation. Then, a mapping function, which can be hand-crafted or learned directly from the datasets, is proposed to combine the inaccurately estimated human pose and the inferred confidence metric to accomplish CVPI. Specifically, the joints with higher confidence are weighted more in the pose matching for CVPI. Finally, the estimated pose information is integrated into the appearance and motion features to boost the CVPI performance. In the experiments, we evaluate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. The experimental results show the effectiveness of the proposed confidence metric, and the integration of pose, appearance, and motion produces a new state-of-the-art CVPI performance. Guoqiang Liang 0001, Xuguang Lan, Xingyu Chen 0001, Song Wang 0002, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Efficient Estimation of View Synthesis Distortion for Depth Coding OptimizationabstractDepth coding in depth-based three-dimensional (3-D) video is unique in that its quality is measured by view synthesis distortion (VSD) rather than the depth distortion itself, which further complicates the coding optimization as the VSD is related to quality of both the associated depth and texture videos. In this paper, an efficient VSD estimation scheme is developed to measure the effect of depth errors on the VSD for a block given its depth distortion in mean-squared error. Unlike other relevant VSD models which involve computationally intensive parameter training or Fourier transform, the proposed scheme is free of parameter training, while taking the advantage of integer 4 × 4 discrete Cosine transform to replace Fourier transform, thus well-saving computational cost and diminishing sensitivity to training dataset of video. The proposed scheme is then incorporated on the coding unit basis into the rate-distortion optimization for depth coding optimization, coupled with adapting quantization parameter accordingly to accommodate local effect of the depth errors on the VSD. Experimental results show that our solution obtains better results in depth coding than three testing solutions, on the platform of H.264/AVC reference software JM16.0. Benefiting from the efficiency of the VSD estimation, low coding complexity is obtained as well. The proposed solution is further evaluated on the reference software HTM13.0 of the latest 3-D high-efficiency video coding standard, exhibiting better and comparable results compared against the HTM codec with the view synthesis optimization disabled and enabled, respectively. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Cross-View Person Identification by Matching Human Poses Estimated With Confidence on Each Body JointabstractCross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interests recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D human pose estimation on videos/images collected in the wild. In this paper, we introduce a new metric of confidence to the 3D human pose estimation and show that the combination of the inaccurately estimated human pose and the inferred confidence metric can be used to boost the CVPI performance---the estimated pose information can be integrated to the appearance and motion features to achieve the new state-of-the-art CVPI performance. More specifically, the estimated confidence metric is measured at each human-body joint and the joints with higher confidence are weighted more in the pose matching for CVPI. In the experiments, we validate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. Guoqiang Liang 0001, Xuguang Lan, Song Wang 0002, Nanning Zheng 0001 |
AAAI | 2 |
| 2018 | Fully Convolutional Grasp Detection Network with Oriented Anchor BoxabstractIn this paper, we present a real-time approach to predict multiple grasping poses for a parallel-plate robotic gripper using RGB images. A model with oriented anchor box mechanism is proposed and a new matching strategy is used during the training process. An end-to-end fully convolutional neural network is employed in our work. The network consists of two parts: the feature extractor and multi-grasp predictor. The feature extractor is a deep convolutional neural network. The multi-grasp predictor regresses grasp rectangles from predefined oriented rectangles, called oriented anchor boxes, and classifies the rectangles into graspable and ungraspable. On the standard Cornell Grasp Dataset, our model achieves an accuracy of 97.74% and 96.61% on image-wise split and object-wise split respectively, and outperforms the latest state-of-the-art approach by 1.74% on image-wise split and 0.51% on object-wise split. Xinwen Zhou, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001 |
IROS | 2 |
| 2018 | Leveraging Spatio-Temporal Evidence and Independent Vision Channel to Improve Multi-Sensor Fusion for Vehicle Environmental PerceptionabstractFor intelligent vehicles, multi-sensor fusion is of great importance to perceive traffic environment with high accuracy and robustness. In this paper, we propose two effective methods, i.e. spatio-temporal evidence generating and independent vision channel, to improve multi-sensor track-level fusion for vehicle environmental perception. The spatio-temporal evidence includes instantaneous evidence, tracking evidence and tracks matching evidence to improve existence fusion. Independent vision channel leverages the specific advantage of vision processing on object recognition to improve classification fusion. The proposed methods are evaluated by using the multi-sensor dataset collected from real traffic environment. Experimental results demonstrate that the proposed methods can significantly improve the multi-sensor track-level fusion in terms of both detection accuracy and classification accuracy. Juwang Shi, Wenxiu Wang, Xiao Wang 0002, Hongbin Sun 0001, Xuguang Lan, Jingmin Xin, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 5 |
| 2018 | A Limb-Based Graphical Model for Human Pose EstimationabstractModeling the relationship among human joints is one of the most important components in human pose estimation. Most of previous methods define this relationship as a geometric constraint on the relative locations of two neighboring joints. In this constraint, the local appearance of the region connecting two neighboring joints is ignored. However, discarding this image appearance leads to some severe problems, such as double-counting and localization failure when the human pose is rare in the training dataset. Moreover, this image appearance, called human limb, plays an important role in human pose estimation in human visual system. Due to these reasons, we propose to solve a new task: human limb detection, which aims at detecting and representing this local image appearance. We combine this task with human joint localization as a unified framework. After getting the initial detections, we design a two-steps graphical model to capture the spatial relationship among human joints and limbs in a coarse to fine way. We evaluate the proposed method on two widely used datasets for human pose estimation: 1) frame labeled in cinema and 2) leeds sports pose datasets. The experiments results show the effectiveness of our method. Guoqiang Liang 0001, Xuguang Lan, Jiang Wang 0001, Jianji Wang 0001, Nanning Zheng 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2017 | Pose-and-illumination-invariant face representation via a triplet-loss trained deep reconstruction model
Xingyu Chen 0001, Xuguang Lan, Guoqiang Liang 0001, Nanning Zheng 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Fast additive quantization for vector compression in nearest neighbor search
Jin Li 0011, Xuguang Lan, Jiang Wang 0001, Meng Yang 0002, Nanning Zheng 0001 |
Multim. Tools Appl. | 2 |
| 2017 | A new compressive sensing video coding framework based on Gaussian mixture model
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
Signal Process. Image Commun. | 2 |
| 2017 | Online Variable Coding Length Product Quantization for Fast Nearest Neighbor Search in Mobile RetrievalabstractQuantization methods are crucial for efficient nearest neighbor search in many applications such as image, music, or product search. As mobile devices are becoming increasingly more popular, the quantization methods on mobile devices are more important, because a large portion of the search queries are becoming performed on mobile devices. One important characteristic of the communication on mobile devices is the inherent unreliability of their communication channels. In order to adapt the quality changes of the communication channels, we need to change the coding length of the quantization accordingly. The existing quantization methods use fixed-length codebooks, and it is expensive to retrain another codebook with different coding length. In this paper, we propose a novel variable length product quantization framework that consists of a set of fast universal scalar quantizers. The framework is capable of producing variable length quantization without retraining the codebook. Each data vector is transformed into a new space to reduce the correlation across dimensions. A proper number of bits is allocated to represent the scalar component in each dimension according to the given coding length. For each component, we estimate its probability density function (PDF) and design an efficient universal scalar quantizer based on the PDF and the allocated bits. To reduce distortion, we learn a Gaussian mixture model for the data. The experimental results show that, compared to state-of-the-art product quantization methods, our approach can construct the codebooks online for variable coding lengths and achieve the comparable performance. Jin Li 0011, Xuguang Lan, Xiangwei Li, Jiang Wang 0001, Nanning Zheng 0001, Ying Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Human pose estimation based on human limbsabstractModeling the relationship among human joints is one of the most important components in human pose estimation. Previous methods usually define this relationship as geometric constraints on the relative location of two neighboring joints. In this definition, the local image appearance of the region connecting two neighboring joints is ignored. In fact, this image appearance, called human limb, plays an important role in human joint localization in human visual system. To make full use of this local image appearance, we propose to solve a new task: human limb detection. We combine it with human joint localization in one deep convolutional neural network. After getting coarse results, we employ a graphical model to remove false positive detections. Besides, shallow and deep features are combined in this model. We evaluate our method on the FLIC and LSP datasets. The experiments results show the effectiveness of our method. Guoqiang Liang 0001, Xuguang Lan, Jiang Wang 0001, Nanning Zheng 0001 |
ICPR | 2 |
| 2016 | Depth Map Coding by Modeling the Locality and Local Correlation of View Synthesis Distortion in 3-D Video
Qiong Xue, Xuguang Lan, Meng Yang 0002 |
MMM (1) | 2 |
| 2016 | Efficient detail-enhanced exposure correction based on auto-fusion for LDR imageabstractWe consider the problem of how to simultaneously and well correct the over- and under-exposure regions in a single low dynamic range (LDR) image. Recent methods typically focus on global visual quality but cannot well-correct much potential details in extremely wrong exposure areas, and some are also time consuming. In this paper, we propose a fast and detail-enhanced correction method based on automatic fusion which combines a pair of complementarily corrected images, i.e. backlight & highlight correction images (BCI &HCI). A BCI with higher visual quality in details is quickly produced based on a proposed faster multi-scale retinex algorithm; meanwhile, a HCI is generated through contrast enhancement method. Then, an automatic fusion algorithm is proposed to create a color-protected exposure mask for fusing BCI and HCI when avoiding potential artifacts on the boundary. The experiment results show that the proposed method can fast correct over/under-exposed regions with higher detail quality than existing methods. Xuguang Lan, Meng Yang 0002 |
MMSP | 2 |
| 2016 | Efficient compressive sensing video compression method based on Gaussian mixture modelsabstractIn this paper, we propose an efficient lossy compression method for the compressive sensing video that utilizes Gaussian mixture models (GMM). The GMM is used to model the compressive sensing video (CSV) frames. Then we design an efficient lossy compression method based on the GMM. Each CSV frame can be efficiently compressed by the proposed method. The proposed method is for better compromise of compression efficiency and computational complexity. And it achieves a significant Bjontegaard-Delta (BD)-PSNR improvement about 8.84~11.81dB in average compared with existing low complexity compression solutions for compressing the CSV sequence. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 2 |
| 2016 | An efficient depth image-based rendering with depth reliability maps for view synthesis
Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Parameter-free view synthesis distortion model with application to depth video codingabstractDepth coding in 3D video is unique in that its quality is measured by the view synthesis distortion (VSD) rather than its own depth distortion, which further complicates the coding optimization as VSD is related to both the depth and texture quality. We propose a parameter-free VSD model to directly estimate the impact of depth errors on VSD on the small block basis, given its depth distortion. The proposed model is incorporated into rate-distortion optimization of depth coding, by adapting the Lagrange multiplier and optimizing the selection of quantization parameters. Simulation results show that our proposed scheme improves the BDPSNR and BDBR by 0.6dB PSNR and 20% bits saving on average compared with H.264 coding standard in depth coding, while keeping the implementation easy. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
ISCAS | 3 |
| 2015 | Optimized truncation model for adaptive compressive sensing acquisition of imagesabstractThe sparsity of the input signal is important for compressive sensing (CS) reconstruction in CS system. In this paper, we establish an optimized truncation model to determine the number of the sparsified coefficients to be truncated in CS acquisition according to the sampling rate. The proposed truncation model suits for signals of any dimension. With the truncation model, the sparsity of the signal can be optimized by properly truncating the small elements of the sparsified coefficients. Furthermore we propose an adaptive CS acquisition solution based on the truncation model to reduce the noise folding effect. The proposed solution is verified for CS acquisition of natural images. Simulation results show that the proposed solution achieves significant improvement of the reconstructed image quality by 0.7~1.4 dB on average compared with existing solutions. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 2 |
| 2014 | Video object segmentation with shape cue based on spatiotemporal superpixel neighbourhoodabstractIn this study, the authors present a method to extract moving objects in image sequences. The proposed approach is based on a graph cuts algorithm defined on a spatiotemporal superpixel neighbourhood. Presegmented superpixels are partitioned into foreground and background while preserving temporal and spatial coherence. It achieves this goal by three steps. First, instead of operating at pixel level, the superpixels are advocated as basic units of the authors segmentation scheme. Second, within the graph cuts framework, two superpixel‐based data terms and two superpixel‐based smoothness terms are proposed to solve segmentation problem. Finally, the proposed method yields the segmentation of all the superpixels within video volume by the graph cuts algorithm. To illustrate the advantages of this approach, the quantitative and qualitative results are compared with other state‐of‐the‐art methods. The experimental results show that the proposed method gives better performance of segmentation with respect to these methods. Nanning Zheng 0001, Jianru Xue, Xuguang Lan, Ce Li 0001 |
IET Comput. Vis. | 4 |
| 2014 | Image parsing by loopy dynamic programming
Shizhou Zhang, Jinjun Wang, Yihong Gong, Xinzi Zhang, Xuguang Lan |
Neurocomputing | 6 |
| 2014 | Object segmentation and key-pose based summarization for motion video
Jianru Xue, Xuguang Lan, Ce Li 0001, Nanning Zheng 0001 |
Multim. Tools Appl. | 3 |
| 2013 | Universal and low-complexity quantizer design for compressive sensing image codingabstractCompressive sensing imaging (CSI) is a new framework for image coding, which enables acquiring and compressing a scene simultaneously. The CS encoder shifts the bulk of the system complexity to the decoder efficiently. Ideally, implementation of CSI provides lossless compression in image coding. In this paper, we consider the lossy compression of the CS measurements in CSI system. We design a universal quantizer for the CS measurements of any input image. The proposed method firstly establishes a universal probability model for the CS measurements in advance, without knowing any information of the input image. Then a fast quantizer is designed based on this established model. Simulation result demonstrates that the proposed method has nearly optimal rate-distortion (R~D) performance, meanwhile, maintains a very low computational complexity at the CS encoder. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 2 |
| 2013 | Adaptively post-encoding multiple description video coding
Xuguang Lan, Meng Yang 0002, Yuan Yuan 0001, Songlin Zhao, Nanning Zheng 0001 |
Neurocomputing | 1 |
| 2013 | Adaptive multiple description coding for hybrid networks with dynamic PLR and BER
Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
Signal Process. Image Commun. | 2 |
| 2012 | Improved View Synthesis with Depth Reliability MapsabstractAn improved view interpolation scheme with depth reliability maps is presented to enhance the quality of synthesized images in this paper. First of all, the left depth map and the right one can be estimated by a graph cut approach from the input reference images. Then, the predict depth maps are derived by forward mapping, and then a median filter is applied to eliminate small blank points caused by forward mapping. Next, the left predict image and the right one are synthesized by reverse mapping using the reference image with the associated predict depth map. After that, in order to obtain better visual quality, the post-processing operations are developed. On the one hand, the segmentation mask can be derived from the depth reliability maps, which are produced from the depth maps. On the other hand, based on the segmentation mask, a virtual view image is generated by a weighted interpolation scheme from the left and right predict images. Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
DCC | 2 |
| 2012 | Disocclusion using depth reliability map for view synthesisabstractIn this paper, a novel disocclusion scheme is proposed based on the depth reliability maps for virtual view synthesis. One important aspect of the developed method is that the occlusion map is estimated with the depth reliability maps. Therefore, the holes in a virtual view image can be filled in by blending the two predict images based on the occlusion map. The experimental results demonstrate that the addressed method has better performance for view synthesis than the other three conventional methods through the assessment of image quality and algorithm cost. Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
ICASSP | 2 |
| 2011 | Design of error-resilient M-description codec over wireless broadcasting networksabstractA practical error-resilient M-description codec scheme is designed to combat the bit errors of the wireless broadcasting networks and raise the quality of the reconstructed signal. The signal is coded into large number of mutually refinable descriptions by robust staggered M-description scalar quantizer (RSMDSQ). Then an index assignment method is used to enhance the error-resilient capacity of any subset of all generated descriptions by enlarging the hamming distance of the codewords. Accordingly, at the terminal, an enhanced decoding scheme is used to recover the signal utilizing the robust correlation among the descriptions. Simple simulations have been done on MATLAB. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
CCNC | 2 |
| 2011 | Explicit Network-Adaptive Robust Multiple Description CodingabstractThe data delivery performance of multiple description coding (MDC) over unreliable network with capacity constraints is related with three factors: redundancy rate, packet loss rate (PLR), and bit error rate (BER). We simplified this network-adaptive delivery problem to only relate with redundancy rate. The proposed scheme is an extension of scalar quantization (SQ) based MDC. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
DCC | 2 |
| 2011 | 3D spatio-temporal graph cuts for video objects segmentationabstractIn this paper, we present a method to extract moving objects in monocular image sequences. The proposed method is based on graph cuts defined on a spatio-temporal region adjacency graph (RAG). First, we initially over-segment each frame in the video, and take the over-segmented regions as the vertices in the 3D spatio-temporal graph. Second, multiple cues are fused together to extract objects accurately. Finally, accurate foreground/background segmentation are efficiently achieved by binary graph cut. The experimental results showed that the proposed method improved the performance of segmentation with respect to the popular methods. Jianru Xue, Nanning Zheng 0001, Xuguang Lan, Ce Li 0001 |
ICIP | 4 |
| 2011 | Auto-generated strokes for motion segmentationabstractWe propose a new approach to motion segmentation that is based on auto-generated strokes. The novelty of the approach is twofold. First, inspired by recent work of other researchers we formulate the problem as that of interactive segmentation. Instead of inputting the strokes by the user, the strokes in our approach are auto-generated. The second novelty of the paper is formulation in which, unlike in many other motion segmentation algorithms, we do not use complex algorithm which fuses the output of multiple cues to segment foreground objects, a simple and effective maximum hybrid similarity method is presented. The maximum hybrid similarity does not need to set the threshold in advance. Experimental results have shown the superiority of the proposed method in extracting moving objects. Jianru Xue, Ce Li 0001, Xuguang Lan, Nanning Zheng 0001 |
ISCAS | 4 |
| 2011 | Key object-based static video summarizationabstractIn this paper, we present a system for object-based video summarization facilitated by an efficient video object segmentation system. We eliminate the redundancy not only from spatial and temporal domain, but also from content domain. First, we detect shot boundaries and extract video objects by a 3D graph-based algorithm. Once the objects are obtained, the shape of the objects need to be represented. The key objects are extracted in a global manner by K-means clustering of shapes. Experimental results on the proposed object-based scheme combined with efficient video object segmentation show desirable summarization. Jianru Xue, Xuguang Lan, Ce Li 0001, Nanning Zheng 0001 |
ACM Multimedia | 3 |
| 2011 | Consecutive redundancy control for robust multiple description coding over unreliable networksabstractDifferent system may have different request on coding efficiency and security, even the characteristics of the unreliable networks are fixed, such as bit error rate, packet loss rate and bandwidth. In this paper, a novel multiple description coding (MDC) scheme is presented for the unreliable networks, considering both the network characteristics and system request. An error-resilient problem based on MDSQ method is proposed and analyzed, and then an iterative redundancy-control method is designed based on ERMDC by index pair ordering. Accordingly, a fast IA method is proposed only considering the high redundancy case. Simulation results show that the redundancy can be easily and precisely adjusted to meet all the requests. Meanwhile, the error resilience and R-D bound of MDC is self-adaptively guaranteed well enough for all cases. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
WCNC | 2 |
| 2011 | Arbitrary ROI-based wavelet video coding
Xuguang Lan, Nanning Zheng 0001, Yuan Yuan 0001 |
Neurocomputing | 1 |
| 2010 | Arbitrary ROI-Based Wavelet Video CodingabstractAn arbitrary shape region of interest (ROI) coding is presented for scalable wavelet video codec in this paper. The padding of macroblock and polygon matching are employed to estimate the motion of the ROI. The motion vectors derived are set as the motion trajectory of the samples to generate one-dimensional temporal signal, filtered to reduce the temporal redundancy by using motion compensated temporal filtering for arbitrary shape ROI. The Reconstructed quality of the ROI coding can be significantly improved at low bit rate, compared to non-ROI coding. The efficiency of the MCTF based on arbitrary ROI is compared with that of the video object coding in MPEG-4. The ability of the MCTF to reduce the temporal redundancy is better than or comparable to that MPEG-4 to some extent. Xuguang Lan, Nanning Zheng 0001, Miao Hui, Jianru Xue |
DCC | 1 |
| 2009 | A Peer-to-Peer Architecture for Live Streaming with DRMabstractDRM is becoming more and more important for P2P live streaming. In this paper, a manageable overlay network architecture with DRM, is proposed for live streaming. The system consists of register server, index servers, supernodes and peers. The register server authorizes the peers, assign the key and index server list; The index server acts as the centralized index server to store the peer list, program list, and buffer information of peers. Supernodes are special peers which store the bigger buffer of live streaming. Each peer periodically exchanges data availability information with the assigned index server. The peer can retrieve correspondingly unavailable data from partners given by index server using the proposed scheduling algorithm. It also supports DRM in register server. The proposed system has been demonstrated based on CERNET of China. Good streaming quality can be achieved due to its global optimization and the digital right of video contents can be protected to some extent. Xuguang Lan, Jianru Xue, Lihua Tian, Nanning Zheng 0001 |
CCNC | 1 |
| 2009 | Joint Network-Source Video Coding Based on Lagrangian Rate AllocationabstractJoint network-source video coding (JNSC) is targeted to achieve the optimum delivery of a video source to a number of destinations over network with capacity constraints. In this paper, a practical scalable multiple description coding is proposed for JNSC, based on Lagrangian rate allocation and scalable video coding. After the spatiotemporal wavelet transformation of input video sequence and the bit plane coding and context-based adaptive binary arithmetic coding, jointing network-source coding is performed on the coding passes of the code blocks (CB) using Lagrangian rate allocation. Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Ce Li 0001, Songlin Zhao |
DCC | 1 |
| 2008 | Scalable Multiple Description Coding Based on Bitplane and Lagrangian Rate AllocationabstractScalable video coding is applied to meet the heterogeneity of networks, fluctuation of bandwidth, and diversity of end users. But if the base layer is damaged or missed, the scalable video coding will no longer be effective. To address this problem, a scalable multiple description of video is proposed based on the bit plane and rate allocation of scalable wavelet video coding in an error-prone transmission environment. After spatiotemporal transformation, three-dimensional bit planes of the wavelet frame are sampled in odd and even order except the most significant bit to generate multiple descriptions, which are encoded into scalable video streaming by using entropy coding and Lagragian rate allocation. Combined with multiple path transport, the performance of proposed scalable multiple description coding is demonstrated in Ns2. Xuguang Lan, Nanning Zheng 0001, Songlin Zhao, Weike Chen, Jianru Xue |
CCNC | 1 |
| 2008 | A Peer-to-Peer Architecture Based on Scalable Video CodingabstractCombining the advantages of the centralized P2P structure and the data-driven structure with scalable video coding, we propose an adaptive P2P architecture for live video based on scalable wavelet video coding over Internet. The core operations are simple: every peer periodically exchanges data availability and bandwidth information with the central server, acting as the centralized index, which selects and sends a set of partners that have expected data to the demanding node. And the central server classifies one peer to a certain level according to the peer's downloading bandwidth which coordinates with the layer level of the video data encoded using scalable wavelet coding. The peer retrieves correspondingly unavailable data from partners according to availability information. There are three principal advantages of this architecture: 1) easy to manage, as the server authorizes, classifies, and clusters each new peer according to its bandwidth as it joins the overlay network, and thus the servers maintain a global structure; 2) efficiency in dynamically heterogeneous networks, as users with different processing ability under heterogeneous networks can retrieve the adaptive data from their partner peers according to their downloading bandwidths, and 3) robustness and resilience, as the partner peers can adapt to quick switching among multi-suppliers, and the scalable video transmission is adaptive. A scheduling algorithm is proposed to enable efficient and continuous scalable streaming of low to high bandwidth content with different service levels over heterogeneous networks. Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Weike Chen, Songlin Zhao |
DCC | 1 |
| 2008 | Linear regression models for DCT domain approximate filtering and deblurringabstractThis paper presents two linear regression models by exploring the relationships between DCT coefficients of the original image and filtered image. The first model is to scale the DCT coefficients of the original image in order to approximate the operation of 2-D spatial domain filtering. The second model is to predict the original image from the filtered image in a similar manner. We show that the first model is used for DCT domain filtering, while the second model can be used for fast DCT domain image deblurring. Both of them are easy to implement on compressed formats of DCT-based compression methods (JPEG, MPEG, H.26X) by using decoding quantization tables that are different from the encoding quantization tables. Nanning Zheng 0001, Jianru Xue, Xuguang Lan |
ICME | 4 |
| 2008 | Manageable peer-to-peer architecture for video-on-demandabstractAn efficiently manageable overlay network architecture, called AIRVoD, is proposed for video-on- demand, based on a distributed-centralized P2P network. The system consists of distributed servers, peers, SuperNodes and sub-SuperNodes. The distributed central servers act as the centralized index server to store the peer list, program list, and buffer information of peers. Each newly joined peer in AIRVoD periodically exchanges data availability information with the central server. Some powerful peers are selected to be sub-SuperNodes which store a larger part of the demanded program. The demanding peers can retrieve correspondingly unavailable data from partners selected from the central server, sub- SuperNodes and SuperNodes that have the original programs to supply the available data, by using the proposed parallel scheduling algorithm. There are four characteristics of this architecture: 1) easy to globally balance and manage: central server can identify, and cluster each newly joined peer, and allocate the load in the whole peer network; 2) efficient for dynamic networks: data transmission is dynamically determined according to data availability which can be derived from central server; 3) resilient, as the partnerships can adapt to quick switching among multi-suppliers under global balance; and 4) highly cost-effective: the powerful peers are taken full advantage to be larger suppliers. AIRVoD has been demonstrated based on CERNET OF CHINA. Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Weike Chen |
IPDPS | 1 |
| 2007 | A peer-to-peer architecture for efficient live scalable media streaming on internetabstractThis paper presents a manageable overlay network architecture SVCP2P for live scalable media streaming. Every peer in SVCP2P periodically exchanges data availability information with one of distributed central servers which act as the centralized index for storing peer list, program list and buffer information of peers. An efficient scheduling algorithm is proposed, which achieves real-time and continuous transmission of the scalable streaming. There are three characteristics of this architecture: 1) easy management; 2) efficient to heterogeneous network because of the scalable media streaming adapting to the heterogeneous demand; and 3) robust and resilient. We have examined the SVCP2P which has been implemented based on the IP Internet over LAN, and the results demonstrate the efficiency of SVCP2P. Xuguang Lan, Nanning Zheng 0001, Jianru Xue, Xiaoguang Wu |
ACM Multimedia | 1 |