EDBT 2026 Demo / reviewers in the wild / expert
Yuehu Liu
dblp:50/6184
· DBLP profile ↗
109ranked-venue papers
1as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 1 first-author · 20 since 2021Artificial intelligence and machine learning · 48 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hindsight-based state space exploration via counterfactual intrinsic reward assignment
Fukai Zhang, Cong Wang 0007, Yuehu Liu |
Neural Networks | 5 |
| 2026 | Perception-Failure-Induced Test Scenario Searching via Online Causal Reinforcement LearningabstractDespite compliance with safety standards such as ISO 26262, the perception systems of autonomous vehicles still face numerous challenges during real-world operation. In this work, we propose a framework to identify safety-critical scenarios specifically induced by perception failures, aiming to support more targeted and effective scenario-based testing. Importantly, the key challenge lies in disentangling whether perception failures directly lead to critical outcomes. To address this, we propose an Online Causal Reinforcement Searching (OCRS) framework that simultaneously performs scenario search and causal reasoning. Focusing on visual unawareness and perception degradation, OCRS employs an LSTM-RNN controller to identify accident-prone scenarios linked to perception failures, which are then verified in a counterfactual world to determine causal responsibility. To improve efficiency, an online scenario classifier with passive-aggressive updates is introduced to dynamically filter out non-critical cases. The experimental results demonstrate both the effectiveness and efficiency of the proposed approach, achieving a 33% reduction in execution time for searchingperception-failure-induced corner cases. Furthermore, we evaluate the method under adverse weather conditions such as snow and fog, confirming that OCRS remains effective in identifying various failure modes. Chi Zhang 0020, Tingting Long, Linhai Xu, Mingwen Bi, Xingyu Chen 0001, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Key Instance-Based Spatio-Temporal Network for Group Activity RecognitionabstractGroup activity recognition involves detecting the collective actions performed by a group of individuals, where identifying the key actors and key frames is crucial for understanding the group’s behavior. To tackle this challenge, we propose a spatio-temporal reasoning framework that leverages key instances. Our key instance identification module effectively detects key roles and frames from video sequences, while a graph-based reasoning model dynamically aggregates the features of related actors. We extract joint features and RGB features from video sequences, and these are fused using our multi-modal fusion TCT module, which improves the representation power of the original features. To better understand group activity through spatio-temporal correlations, we further utilize an enhanced cross-transformer module for spatio-temporal synchronous reasoning, considering both time and space dimensions. Our method has been tested on two public datasets, showing that it achieves high accuracy and surpasses many state-of-the-art approaches. Yaochen Li, Haoting He, Yutong Wang 0008, Gaojie Li, Yuehu Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2026 | SCALE-Pose: Skeletal Correction and Language Knowledge-assisted for 3D Human Pose EstimationabstractTransformer-based 3D human pose estimation methods typically use 2D joint sequences as inputs, leveraging spatial and temporal transformer encoders to model the 3D human pose. However, these methods often fail to incorporate skeletal constraints to limit joint motion. The integration of prior category knowledge to enhance joint representations is also neglected. To address these challenges, a novel approach named SCALE-Pose is proposed in this article. Our method first constructs a feature extraction network based on spatiotemporal skeleton refinement, where the skeletal correction modules are designed within both the spatial and temporal skeleton encoders to enhance the backbone network’s understanding of skeletal features. Meanwhile, a weighted average joint position error loss function is applied to improve the network’s ability to represent the joints with varying difficulty levels. A new radian-based loss function for skeletal joint angles is also designed, further enhancing the model’s capability to capture subtle skeletal movements. Furthermore, a training strategy based on large language model (LLM) priors is proposed to generate category-specific prior semantic knowledge from category keywords, which is then incorporated as auxiliary information to extract motion features. The experimental results based on the Human3.6M and MPI-INF-3DHP datasets well demonstrate the effectiveness of the proposed method. Yaochen Li, Xinnan Ma, Limeng Zhao, ChenXu Zhou, Yuehu Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style TransferabstractRecent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especially when applied to complex traffic scenes with diverse objects, layouts, and stylistic variations, current diffusion models tend to exhibit Style Neglection, i.e., failing to generate the required style in the prompt. To address this issue, we propose Style Nursing, which directs the model to focus on style subject tokens in the text prompt and excites their corresponding visual activations. Moreover, we introduce Spatial and Semantic Guidance to guide the preservation of content after editing, which utilizes spatial features from the DDIM sampling process together with attention maps from the semantic reconstruction. To evaluate the performance of zero-shot style transfer methods in traffic scenes, we present STREET-6K, a new benchmark dataset comprising 6000 images showcasing diverse traffic scenes and style transfer variations, accompanied by comprehensive annotations and evaluation metrics. Our approach beats state-of-the-art image translation methods in comprehensive quantitative metrics and human evaluations on traffic scene image synthesis while seamlessly generalizing to various other types of images without training or fine-tuning. Further experiments on detection and segmentation tasks show that fine-tuning perception models on our synthesized images improves Recall and mean Intersection over Union (mIoU) by over 10% and 3% respectively in rarely-seen traffic scenes. Zihang Lin, Yuehu Liu, Chi Zhang 0020 |
AAAI | 4 |
| 2025 | Mitigating Shortcut Learning in Online Action Detection and Anticipation via Cross-Modal Semantic AlignmentabstractOnline Action Detection (OAD) and Online Action Anticipation (OAA) are conventionally framed as multi-class classification tasks reliant on action representation learning. However, existing methods easily overfit to appearance features corre-lated with specific categories, neglecting semantically meaningful action features, i.e., shortcut learning. This results in a misinter-pretation of visually similar actions with distinct semantics. We argue that shortcut learning stems from the one-hot label supervision, which simplifies task objectives from semantic recognition to category differentiation. Inspired by advances in text-supervised visual representation learning (e.g., CLIP), we propose a CLIP for Online Action (CLIP40A), a unified model that formulates OAD and OAA as Video-Text Retrieval tasks. This approach mitigates shortcut learning by supervising the alignment between actions and their corresponding labels. Specifically, CLIP40A extracts action contexts from distant videos across multiple time scales, and then learns current and future action representations based on these contexts and recent videos. For labels, CLIP40A introduces a learnable prompt mechanism to compensate for the lack of label contexts. Subsequently, CLIP40A leverages the pretrained CLIP text encoder to convert labels into representations, preserving label semantics and semantic associations with visual data. By similarity comparisons, the labels most similar to current and future actions are the results of OAD and OAA. CLIP40A achieves superior performance on THUMOS'14 and TVSeries. Yuehu Liu, Chi Zhang 0020 |
CBMI | 2 |
| 2025 | Semantic Graph Embedded Energy Minimization Learning for Scene Graph GenerationabstractThe performance of current scene graph generation models is affected by training with cross-entropy loss, exacerbating the problem of prediction bias stemming from biased training data. Energy-based model adopts a learning method for joint image and scene graph to alleviate this challenge. However, this method only focuses on the visual features of images, neglecting the rich relation information contained in the semantic space. To address this issue, we innovatively employ powerful pre-trained large models to realize a simple yet effective semantic graph embedded energy minimization framework for the SGG task. Specifically, we use large models to generate image descriptions and extract relation triplets, which are then transformed into semantic graphs with entities as nodes and relations as edges. Moreover, by mapping these graphs into the same space using GNN to learn the minimal energy value, our approach enables SGG model to learn structural information in both semantic and visual spaces. We validate the effectiveness and efficiency of our method on the SGG benchmark Visual Genome dataset. Compared with prevailing models and EBM, we achieve a significant performance improvement of up to 2.99% and 2.24%, respectively. Jinghang Chen, Chi Zhang 0020, Yuehu Liu, Le Wang 0003 |
ICASSP | 3 |
| 2025 | Worst Perception Scenario Prediction for Testing Autonomous Driving PerceptionabstractRecent studies have suggested that potential short-comings of certain perception modules can be discovered by analyzing the performance of worst scenarios. However, finding the worst perception scenario (WPS) requires datasets with rich semantic annotations of the scenes in visual perception tasks and it is time-consuming to label all the scenario data. To address this, we proposed a method of prediction for WPS, which utilized prior information to predict the model performance under the absence of annotations. Specifically, this paper introduced a scenario matcher based on hybrid re-ranking, which combined the labeled and unlabeled data to generate the pseudo-sample set. In addition, we designed a sample reorganization module to update this sample set through the nearest neighbor retrieval. We also discussed the distribution relationship between labeled and unlabeled data, categorizing it into three cases, and validated the effectiveness of the proposed method on the KITTI and ApolloScape datasets. Liheng Xu, Chi Zhang 0020, Yuehu Liu, Li Li 0013 |
IV | 4 |
| 2025 | 3D Shape Adaptation Across Datasets for Weakly Supervised Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is a key yet challenging task that usually involves extensive and expensive manual annotation of 3D boxes. To eliminate the dependence on 3D box labels, weakly supervised M3D (WM3D) has been explored using only 2D annotations, which necessitate the use of extra resources, like LiDAR data, stereo images, and video sequences. However, the strict correspondence and complex calibration between the target image and additional resources limit their applicability. In this work, we propose a simple yet effective framework, 3D Shape Adaptation across datasets for Weakly supervised Monocular 3D Detection (SAWM3D). We observed that directly applying a source-dataset detector to the target dataset results in a significant domain gap, with the primary contribution coming from the 3D location, while orientation and dimensions have a smaller impact. This enables us to view WM3D as 3D shape adaptation optimization on the target dataset. Directly scaling the predicted shape results in a significant reduction of the adaptation gap; fine-tuning on the target dataset using only 2D supervision also yields impressive results. Experiments on the KITTI benchmark demonstrate the effectiveness of our strategies. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 5 |
| 2025 | Rethinking SSIM-Based Optimization in Neural Field TrainingabstractThe Structural Similarity (SSIM) index is a widely used metric for evaluating image quality, with broad applications in areas such as image restoration, 3D reconstruction, and novel view synthesis. A number of previous works have introduced SSIM-based optimization into neural field training to enhance the model's performance. Despite its widespread use, there has been limited research on how to effectively incorporate SSIM loss into the training process. In this work, we explore this gap and provide insights into the role of SSIM loss in neural field training. Our key finding is that SSIM loss is particularly beneficial during the early phase of training, before the model fully learns the luminance information. We show that SSIM loss acts as an effective “guidance” mechanism in the initial training phase, and removing it after the model has learned the luminance does not harm the final performance-in fact, it may improve it. Our experiments demonstrate the effectiveness of our strategy, offering new insights into how SSIM loss can be more efficiently used in neural field training. We believe these findings will not only enhance SSIM's application in neural field training but also inspire further research into more adaptive loss functions for deep learning models. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 5 |
| 2025 | 3D Shape Transfer Learning for Enhanced Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is challenging due to the lack of depth information in the RGB image. Existing works resort to various additional resources to enhance detection performance, including LiDAR data, depth information, CAD models, stereo images, and others, where strict correspondence or synchronization may limit their applicability and scalability. In this work, we propose a simple yet effective framework, 3D Shape Transfer Learning for Enhanced Monocular 3D Object Detection (STLM3D). It views M3D as 3D shape reconstruction and leverages transfer learning across datasets to enhance shape reconstruction capability, thereby enhancing M3D performance. In addition, we design a plug-and-play 3D detection branch for 3D attributes prediction. Experimental results on the KITTI benchmark demonstrate that the proposed method achieves state-of-the-art performance compared to existing approaches. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 6 |
| 2025 | BiOMamba: Mamba-based Forward-Then-Backward Temporal Modeling for Online Action Detection and AnticipationabstractGiven that action evolution follows temporal progression, recent studies for Online Action Detection (OAD) and Online Action Anticipation (OAA) generally adopt forward temporal modeling to capture dependencies in observable video sequences. However, the strictly sequential nature of forward temporal modeling prevents subsequent frames from being used to enhance the earlier modeling process. In particular, the current frame, the last observable frame in the online video stream, serves as the direct visual cue for ongoing action recognition and the informative context for future action anticipation. As modeling errors accumulate over time, the resulting representations may progressively deviate from the actual semantics. Findings in cognitive neuroscience show that the hippocampus performs backward replay after observation to reinforce and correct the interpretation of previous observations. Inspired by this, we propose to incorporate backward temporal modeling following forward temporal modeling, enabling the model to leverage backward temporal modeling to enhance forward temporal modeling. Based on this idea, we propose a unified model for OAD and OAA, named Bidirectional Online Mamba (BiOMamba). Specifically, to address the excessive length and relevance imbalance in observable sequences, BiOMamba compresses distant long-term memory and preserves recent short-term memory. Then, BiOMamba sequentially model both forward and backward temporal dependencies in the whole memory. Finally, according to the temporal modeling result, BiOMamba generates representations for current and future actions. BiOMamba achieves state-of-the-art performance on THUMOS'14 (OAD: 73.3% mAP, OAA: 59.7% mAP) and TVSeries (OAD: 89.9% mcAP, OAA: 83.7% mcAP). Yuehu Liu, Chi Zhang 0020 |
ACM Multimedia | 2 |
| 2025 | Adding Multi-Scale Priors for 3D Human Pose EstimationabstractRecently solutions based on human prior information have been introduced to estimate 3D human pose from 2D keypoint sequence. These methods aim to learn human skeleton knowledge by analyzing joint connections within the body structure. However, we observe that previous methods cannot capture motion integrity and effectively model action correlation, resulting in the lack of smoothness and coherence in the estimated pose sequence. Therefore, we propose to use multi-scale priors for 3D pose estimation, namely Body-Scale Prior (BSP) and Action-Scale Prior (ASP). These modules take advantage of the body structure and action semantics, to enhance the network’s capability of understanding human pose. BSP models human body knowledge by constructing the skeleton joint structure and relationship between frames, while ASP learns high-dimensional action semantics by constructing topological encoding of the action sequence. By fusing prior features from these two scales, our method enables simultaneous modeling of skeleton joints and integral actions. Extensive experiments are conducted on the most popular benchmark: Human3.6M. The results show that our model achieves better performance in comparison to state-of-the-art methods. Jia Dang, Yuehu Liu |
SMC | 3 |
| 2025 | Prototype-based multi-domain self-distillation for unbiased scene graph generation
Yaochen Li, Yujie Zang, Jingze Liu, Yuehu Liu |
Neurocomputing | 5 |
| 2025 | UniCuboid: Cuboid-based dense shape supervision for monocular 3D object detection
Yuanqi Su, Haoyue Shi 0002, Haoang Lu, Yuehu Liu, Le Wang 0003 |
Neurocomputing | 5 |
| 2025 | Long and Short-Term Collaborative Decision-Making Transformer for Online Action Detection and Anticipation
Chi Zhang 0020, Le Wang 0003, Yuehu Liu |
Pattern Recognit. | 4 |
| 2025 | Autonomous and Adaptive Role Selection for Multi-Robot Collaborative Area Search Based on Deep Reinforcement LearningabstractIn the tasks of multi-robot collaborative area search, we propose the unified approach for simultaneous mapping for sensing more targets (exploration) while searching and locating the targets (coverage). Specifically, we implement a hierarchical multi-agent reinforcement learning algorithm to decouple task planning from task execution. The role concept is integrated into the upper-level task planning for role selection, which enables robots to learn the role based on the state status from the upper-view. Besides, an intelligent role switching mechanism enables the role selection module to function between two timesteps, promoting both exploration and coverage interchangeably. Then, the primitive policy learns how to plan based on their assigned roles and local observation for sub-task execution. The well-designed experiments show the scalability and generalization of our method compared with state-of-the-art approaches in scenes with varying complexity and numbers of robots. Our code is released at https://github.com/linaug/Role_selection. Jiyu Cheng, Hao Zhang 0113, Zhichao Cui, Wei Zhang 0021, Yuehu Liu |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | ESE-GAN: Zero-Shot Food Image Classification Based on Low Dimensional Embedding of Visual FeaturesabstractExisting zero-shot learning based image classification methods transform the zero-shot learning problem into supervised learning by applying generative adversarial network (GAN) to synthesize visual features of unseen classes. However, the visual features generated by the generator tend to be biased towards seen classes, and the discriminator is too weak to generate high-quality image features. To solve these problems, we propose a novel zero-shot food image classification method based on low dimensional embedding of visual features. Our method applies reinforced semantic guidance to increase the discriminative ability of the model by enhancing the strong distribution of input features. Moreover, the visual space is utilized as the embedding space to reduce the bias towards seen classes by reducing the distance between semantic information and visual features in the embedding space. Finally, the feature distribution of unseen classes is further specified by improving the prototype similarity function. Extensive experiments on three food datasets and four general benchmark datasets demonstrate the effectiveness of the proposed method. Gaojie Li, Yaochen Li, Jingle Liu, Wenneng Tang, Yuehu Liu |
IEEE Trans. Multim. | 6 |
| 2024 | SEIT: Structural Enhancement for Unsupervised Image Translation in Frequency DomainabstractFor the task of unsupervised image translation, transforming the image style while preserving its original structure remains challenging. In this paper, we propose an unsupervised image translation method with structural enhancement in frequency domain named SEIT. Specifically, a frequency dynamic adaptive (FDA) module is designed for image style transformation that can well transfer the image style while maintaining its overall structure by decoupling the image content and style in frequency domain. Moreover, a wavelet-based structure enhancement (WSE) module is proposed to improve the intermediate translation results by matching the high-frequency information, thus enriching the structural details. Furthermore, a multi-scale network architecture is designed to extract the domain-specific information using image-independent encoders for both the source and target domains. The extensive experimental results well demonstrate the effectiveness of the proposed method. Zhifeng Zhu, Yaochen Li, Jinhuo Yang, Peijun Chen, Yuehu Liu |
AAAI | 6 |
| 2024 | An Interactive Navigation Method with Effect-oriented AffordanceabstractVisual navigation is to let the agent reach the target according to the continuous visual input. In most previous works, visual navigation is usually assumed to be done in a static and ideal environment: the target is always reachable with no need to alter the environment. However, the “messy” environments are more general and practical in our daily lives, where the agent may get blocked by obstacles. Thus Interactive Navigation (InterNav) is introduced to navigate to the objects in more realistic “messy” environments according to the object interaction. Prior work on InterNav learns shortterm interaction through extensive trials with reinforcement learning. However, interaction does not guarantee efficient navigation, that is, plan-ning obstacle interactions that make shorter paths and con-sume less effort is also crucial. In this paper, we introduce an effect-oriented affordance map to enable longterm interactive navigation, extending the existing map-based nav-igation framework to the domain of dynamic environment. We train a set of affordance functions predicting available interactions and the time cost of removing obstacles, which informatively support an interactive modular system to ad-dress interaction and longterm planning. Experiments on the ProcTHOR simulator demonstrate the capability of our affordance-driven system in longterm navigation in complex dynamic environments. Yuehu Liu, Xinhang Song, Yuyi Liu, Sixian Zhang, Shuqiang Jiang |
CVPR | 2 |
| 2024 | Image Caption Method from Coarse to Fine Based On Dual Encoder-Decoder FrameworkabstractEncoders are widely used in the field of image caption, but the statements generated by the current image caption method may miss the target and the generated description statements are not appropriate enough for the image content. In order to solve the above problems, we propose a coarse-fine image caption method based on dual encoder-decoder framework, which provides a mechanism for discovering and correcting omissions and enables the model to generate a complete image description. Firstly, an image feature extractor based on global and local information is designed, which can extract global information and local information of image and obtain more abundant image representation. Secondly, a dual encoder-decoder framework is designed, which consists of a coarse-grained encoder-decoder and a fine-grained encoder-decoder. Coarse-grained encoder-decoder requires only the original image features as input, which is processed by transformer to produce a coarse text description. In addition, an image feature auto-enhancement module is proposed to detect missing objects in coarse text and enhance their feature expression. Finally, the fine-grained encoder-decoder uses both the image feature and the coarse text caption as input, and generates the final fine-grained caption after multi-modal information fusion. Experimental results on MSCOCO datasets show that our proposed method outperforms previous image caption methods and achieves a performance of 39.7 BLEU-4 score and 121.6 CIDEr-A score. Zefeng Li, Yuehu Liu, Yonghong Song |
IJCNN | 2 |
| 2024 | Computationally Efficient Imitation Learning via K-Timestep Adaptive Action ApproximationabstractA key challenge for training control policies with imitation learning methods lies in the computational inefficiency. This inefficiency comes from an assumption that underlies these methods, assuming the policy should compute a new action for each state, which is unnecessary and costly. However, we notice the states occurring within K consecutive timesteps differ negligibly and their corresponding actions are extremely similar. Therefore, we challenge this assumption and argue that it is enough to compute an action every K states. With this argument, we propose K-Timestep Adaptive Action Approximation, which replaces the computation of K one-timestep actions approximately with that of one K-timestep action to alleviate the computational inefficiency issue. To demonstrate the theoretical validity of our method, we analyze the errors incurred by the policies learned via the method. The analysis proves these policies can converge to the optimal solution with errors no more than an upper bound dependent on K, revealing the effectiveness of our method. To avoid the difficulty of hyperparameter handcrafting on K, we design a simple but effective auto-hyperparameter tuning strategy. In the proposed strategy, K is added as an extra dimension to the action space of the policy, so it can be tuned adaptively by the policy according to the state without any user intervention. Empirical results on 4 imitation learning tasks show the superiority in computational efficiency of our method, which can effectively reduce new actions to compute in training policies. Weiming Wu, Cong Wang 0007, Yuehu Liu |
IJCNN | 4 |
| 2024 | DDGPnP: Differential degree graph based PnP solution to handle outliers
Zhichao Cui, Zeqi Chen, Chi Zhang 0020, Gaofeng Meng, Yuehu Liu, Xiangmo Zhao |
Comput. Vis. Image Underst. | 5 |
| 2024 | Worst Perception Scenario Search via Recurrent Neural Controller and K-Reciprocal Re-RankingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. A recent theoretical study shows that the generalization on the worst-group of test samples is far more difficult than others. Therefore, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address this, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as the discrete search on the Visual Operation Design Domain (ODD), namely scenario representation, by optimizing LSTM-RNN controller with the worst-performance reward. Moreover, a time-efficient K-reciprocal re-ranking technique is utilized to match the predicted scenario parameters with existing test data. The proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Furthermore, searching performances w.r.t different Visual ODDs are investigated and it is found that visual representations through generative adversarial network contribute to a better performance. Chi Zhang 0020, Xiaoning Ma, Liheng Xu, Haoang Lu, Le Wang 0003, Yuanqi Su, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Multi-Robot Environmental Coverage With a Two-Stage Coordination Strategy via Deep Reinforcement LearningabstractMulti-robot environmental coverage can be widely used in many applications like search and rescue. However, it is challenging to coordinate the robot team for high coverage efficiency. In this paper, we propose a Two-Stage Coordination (TSC) strategy, which consists of a high-level leader module and a low-level action executor. The former provides the robots with the topology and geometry of the environment, which are crucial for robots to learn “where” they should go and avoid invalid coverage. Based on the observed information and the environmental topology, the latter module takes primitive action to reach the sub-goal. To facilitate cooperation among the robots, we aggregate local perception information of neighbors from different hops based on graph neural networks. We compare our method with state-of-the-art multi-robot coverage approaches. Experiments and supporting ablation studies show the superior efficiency, scalability, and generalization of our algorithm especially in unseen style and scale of scenes, and an unseen number of robots. Jiyu Cheng, Hao Zhang 0113, Wei Zhang 0021, Yuehu Liu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Discrepant and Multi-instance Proxies for Unsupervised Person Re-identificationabstractMost recent unsupervised person re-identification methods maintain a cluster uni-proxy for contrastive learning. However, due to the intra-class variance and inter-class similarity, the cluster uni-proxy is prone to be biased and confused with similar classes, resulting in the learned features lacking intra-class compactness and inter-class separation in the embedding space. To completely and accurately represent the information contained in a cluster and learn discriminative features, we propose to maintain discrepant cluster proxies and multi-instance proxies for a cluster. Each cluster proxy focuses on representing a part of the information, and several discrepant proxies collaborate to represent the entire cluster completely. As a complement to the overall representation, multi-instance proxies are used to accurately represent the fine-grained information contained in the instances of the cluster. Based on the proposed discrepant cluster proxies, we construct cluster contrastive loss to use the proxies as hard positive samples to pull instances of a cluster closer and reduce intraclass variance. Meanwhile, instance contrastive loss is constructed by global hard negative sample mining in multiinstance proxies to push away the truly indistinguishable classes and decrease inter-class similarity. Extensive experiments on Market-1501 and MSMT17 demonstrate that the proposed method outperforms state-of-the-art approaches. Chang Zou, Zeqi Chen, Zhichao Cui, Yuehu Liu |
ICCV | 4 |
| 2023 | Reuse Non-Terrain Policies for Learning Terrain-Adaptive Humanoid Locomotion SkillsabstractTerrain-adaptive locomotion skills are prerequisites for humanoid to traverse complex environments including barriers, gaps, caves and etc. Previous imitation learning methods require large amounts of motion sampling for an optimal policy adapted to the new terrain. However, we reveal an interesting fact: the effectiveness of such a costly policy on certain complex terrain is quite similar to that of existing policies pretrained on the flat ground. Inspired by this finding, we present a few-shot imitation learning (FSIL) framework to reuse these pretrained policies as learning primitives for new terrain adaptation. Specifically, a meta-controller, trained with a few control sequences, is proposed to allocate combination weights of each primitive for different terrains. Empirical studies in mainstream problem settings show that our method maintains a high passing rate within a few shots of new terrain while avoiding massive unnecessary data sampling. Further in-depth theoretical analysis shows that, compared to other methods that sample data in new terrain environments hundreds or thousands of times, our method requires only a maximum of 52 shots to achieve optimal control policy for certain terrain adaptation. Yuehu Liu |
ICIP | 3 |
| 2023 | Swin-UNIT: Transformer-based GAN for High-resolution Unpaired Image TranslationabstractThe transformer model has gained a lot of success in various computer vision tasks owing to its capacity of modeling long-range dependencies. However, its application has been limited in the area of high-resolution unpaired image translation using GANs due to the quadratic complexity with the spatial resolution of input features. In this paper, we propose a novel transformer-based GAN for high-resolution unpaired image translation named Swin-UNIT. A two-stage generator is designed which consists of a global style translation (GST) module and a recurrent detail supplement (RDS) module. The GST module focuses on translating low-resolution global features using the ability of self-attention. The RDS module offers quick information propagation from the global features to the detail features at a high resolution using cross-attention. Moreover, we customize a dual-branch discriminator to guide the generator. Extensive experiments demonstrate that our model achieves state-of-the-art results on the unpaired image translation tasks. Yaochen Li, Wenneng Tang, Zhifeng Zhu, Jinhuo Yang, Yuehu Liu |
ACM Multimedia | 6 |
| 2023 | Generating Explanations for Embodied Action Decision from Visual ObservationabstractGetting trust is crucial for embodied agents (such as robots and autonomous vehicles) to collaborate with human beings, especially non-experts. The most direct way for mutual understanding is through natural language explanation. Existing researches consider generating visual explanations for object recognition, while the exploration of explaining embodied decisions remains vacant. In this paper, we study generating action decisions and explanations based on visual observation. Distinct to explanations for recognition, justifying an action needs to show why it's better than other actions. Besides, the understanding of scene structure is required since the agent needs to interact with the environment (e.g. navigation, moving objects). We introduce a new dataset THOR-EAE (Embodied Action Explanation) collected based on AI2-THOR simulator. The dataset consists of over 840,000 egocentric images of indoor embodied observation which are annotated with the optimal action labels and explanation sentences. An explainable decision-making criterion is developed considering scene layout and action attributes for efficient annotation. We propose a graph action justification model, exploiting graph neural networks for obstacle-surroundings relations representation and justifying the actions under the guidance of decision results. Experimental results on THOR-EAE dataset showcase its challenge and the effectiveness of the proposed method. Yuehu Liu, Xinhang Song, Shuqiang Jiang |
ACM Multimedia | 2 |
| 2023 | CaMP: Causal Multi-policy Planning for Interactive Navigation in Multi-room ScenesabstractVisual navigation has been widely studied under the assumption that there may be several clear routes to reach the goal. However, in more practical scenarios such as a house with several messy rooms, there may not. Interactive Navigation (InterNav) considers agents navigating to their goals more effectively with object interactions, posing new challenges of learning interaction dynamics and extra action space. Previous works learn single vision-to-action policy with the guidance of designed representations. However, the causality between actions and outcomes is prone to be confounded when the attributes of obstacles are diverse and hard to measure. Learning policy for long-term action planning in complex scenes also leads to extensive inefficient exploration. In this paper, we introduce a causal diagram of InterNav clarifying the confounding bias caused by obstacles. To address the problem, we propose a multi-policy model that enables the exploration of counterfactual interactions as well as reduces unnecessary exploration. We develop a large-scale dataset containing 600k task episodes in 12k multi-room scenes based on the ProcTHOR simulator and showcase the effectiveness of our method with the evaluations on our dataset. Yuehu Liu, Xinhang Song, Shuqiang Jiang |
NeurIPS | 2 |
| 2023 | Reliable Boundary Samples-Based Proxy Pairs for Unsupervised Person Re-identification
Chang Zou, Zeqi Chen, Yuehu Liu, Chi Zhang 0020 |
PRCV (12) | 3 |
| 2023 | EgoFormer: Transformer-Based Motion Context Learning for Ego-Pose EstimationabstractEgo-pose estimation, i.e. predicting 3D pose of the camera wearer, has an essential value in AR and VR applications. First-person video has an ambiguity in that similar video frames may correspond to totally different body poses because of the invisible body part. However, exploiting the context of a video and establishing a long-term temporal relationship can alleviate this ambiguity. To this end, this paper proposes EgoFormer, a Transformer-based model, to learn the motion context from egocentric videos. Moreover, dynamic features commonly used to characterize first-person video do not provide sufficient temporal information to remove the ambiguity inherent in such videos. Therefore, we present a method that can effectively extract temporal features in first-person videos. Results on real-scene and synthetic datasets show that our method could estimate a sequence of human poses with high accuracy and coherence. Yuehu Liu |
SMC | 4 |
| 2023 | Phase Discriminated Multi-Policy for Visual Room RearrangementabstractEmbodied AI, where the agent learns to accomplish tasks through interaction with its surrounding environment, is drawing increasing attention in the community. As a challenging Embodied AI task, visual room rearrangement aims to restore the initially misplaced objects in a room to the target state. Existing approaches usually use a single policy to learn a mapping from visual observation to action. Those methods may be capable of accomplishing tasks with simple goals such as visual navigation. However, the agent in the rearrangement task has to explore various types of interaction for a long time. Only considering a single policy may easily get stuck in local optimum. In this paper, we propose a Phase Discriminated Multi-Policy (PDMP) model, decomposing the task into specific phases and tackling them with customized policies. In particular, we first introduce the graph representation of object relationships providing scene layout knowledge, which is discriminated to task phases. Then based on the knowledge a hierarchical actor-critic module is proposed to dynamically call the policies capable of navigation or object interaction. Each policy is trained with narrowed action space and dense rewards so that they can better converge and cooperate to reach long-term goals. Comprehensive experiments based on the AI2-THOR platform, show that the proposed model achieves better performance than baselines. Xinhang Song, Yuehu Liu |
SMC | 4 |
| 2023 | Frame-Level Smooth Motion Learning for Human Mesh RecoveryabstractReconstructing accurate and smooth 3D human mesh from a monocular video is still a challenge, due to the temporal consistency requirement of body movements. A frame-level smooth motion learning approach called SLMR model was proposed in this work. Specifically, we first design the temporal encoding with a multi-headed attention mechanism, which captures the global and local temporal context relations of motion to retain more dynamic motion features for improving mesh recovery accuracy. We also propose a probabilistic generative model consisting of a conditional variational autoencoder, which learns the distribution of pose changes in each frame, and solves temporal inconsistency of the body movement. Compared with existing approaches, SLMR can take full advantage of inter-frame motion contexts. Experiments validated the effectiveness and smoothness of the proposed approach for human mesh recovery in the wild. Zhaobing Zhang, Yuehu Liu |
SMC | 2 |
| 2023 | Dual Clustering Co-Teaching With Consistent Sample Mining for Unsupervised Person Re-IdentificationabstractIn unsupervised person Re-ID, peer-teaching strategy leveraging two networks to facilitate training has been proven to be an effective method to deal with the pseudo label noise. However, training two networks with a set of noisy pseudo labels reduces the complementarity of the two networks and results in label noise accumulation. To handle this issue, this paper proposes a novel Dual Clustering Co-teaching (DCCT) approach. DCCT mainly exploits the features extracted by two networks to generate two sets of pseudo labels separately by clustering with different parameters. Each network is trained with the pseudo labels generated by its peer network, which can increase the complementarity of the two networks to reduce the impact of noises. Furthermore, we propose dual clustering with dynamic parameters (DCDP) to make the network adaptive and robust to dynamically changing clustering parameters. Moreover, Consistent Sample Mining (CSM) is proposed to find the samples with unchanged pseudo labels during training for potential noisy sample removal. Extensive experiments demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art unsupervised person Re-ID methods by a considerable margin and surpasses most methods utilizing camera information. Zeqi Chen, Zhichao Cui, Chi Zhang 0020, Jiahuan Zhou, Yuehu Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Proprioception-Driven Wearer Pose Estimation for Egocentric VideoabstractPerceiving proprioception from egocentric video to estimate 3D wearer pose is an attention-grabbing visual task. Yet the invisibility of wearer body and the complex motion modality bring challenges to perceive self-motion from the human visual span. In this work, a data processing framework is designed to convert a raw egocentric video stream into a 3D wearer pose sequence. Critically, a generic and lightweight Self-Perception Excitation (SPE) module is proposed to enhance motion modeling and calibrate spatial correlation in the temporal dimension. Employing ResNet50 embedded with SPE module as a backbone, a two-stream architecture for proprioception representation pipeline is proposed to learn proprioception behaviors from RGB streams and motion streams. Then, the proprioception is incorporated as an additional key control signal in a deep reinforcement learning (DeepRL) based motion imitation policy for estimating the multi-modal wearer pose. By considering proprioception, we indicate for the first time, it is possible to understand the self from an egocentric view and further translate it into a higher understanding of wearer motion. The experimental results demonstrate that the proposed framework is able to outperform the state-of-the-art methods by a large margin on the MoCup dataset and produce highly identifiable proprioception behaviors. Yuehu Liu, Zerun Cai |
ICPR | 2 |
| 2022 | Driver Behavior Decision Making Based on Multi-Action Deep Q Network in Dynamic Traffic Scenes
Yaochen Li, Hujun Liu, Anna Li, Yuehu Liu |
PRCV (1) | 7 |
| 2022 | Orthogonal multi-view tensor-based learning for clustering
Shuangxun Ma, Yuehu Liu, Guangcan Liu, Qinghai Zheng, Chi Zhang 0020 |
Neurocomputing | 2 |
| 2022 | Single-image human mesh reconstruction by parallel spatial feature aggregationabstractAbstract Recovering human mesh from a single image with natural postures is a challenging task in human modeling and animation. Model‐free methods regress the mesh vertices from the input image directly to avoid the 6‐DoF human joint extraction from the 2D image. However, the missing of the global information in spatial feature aggregation of the existing GNNs may result in the undesired deformity and inaccuracy of the recovered human mesh. To address this issue, we propose a parallel‐aggregating network with a novelly designed global layer for spatial feature extracting from random walk normalized matrix. Moreover, the coarse body mesh (head, hand, foot, etc.) provided by the coarsening network can add the human characteristic to the mesh. The local and global spatial features are aggregated to update vertice coordinates following an iterative, coarse‐to‐fine process to obtain an accurate and smooth human mesh. Experiments validated the effectiveness and robustness of the proposed approaches for single‐image human mesh recovery. Yuehu Liu |
Comput. Animat. Virtual Worlds | 2 |
| 2022 | Switching: understanding the class-reversed sampling in tail sample memorization
Chi Zhang 0020, Benyi Hu, Yuhang Liuzhang, Le Wang 0003, Yuehu Liu |
Mach. Learn. | 6 |
| 2022 | Density-Aware Haze Image Synthesis by Self-Supervised Content-Style DisentanglementabstractThe key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method. Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Caption Generation From Road Images for Traffic Scene ModelingabstractIn this traffic-scene-modeling study, we propose an image-captioning network which incorporates element attention into an encoder-decoder mechanism to generate more reasonable scene captions. A visual-relationship-detecting network is also developed to detect the relative positions of object pairs. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image-captioning network is applied to generate the corresponding description of each traffic scene element. The visual-relationship-detecting network is utilized to detect the position relations of all object pairs in the subregion. The static and dynamic traffic elements are appropriately selected and organized to construct a 3D model according to the captions and the position relations. The reconstructed 3D traffic scenes can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max, KITTI and Microsoft’s COCO datasets demonstrate the effectiveness of the proposed framework. Yaochen Li, Yuehu Liu, Jihua Zhu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Adaptive Frequency Hopping Policy for Fast Pose EstimationabstractExisting methods for human pose estimation using motion imitation usually suffer from a mismatch of the frequency between policy and demonstrations: the policy runs at a much higher frequency to enable the agent to get track with the demonstrations, which usually leads to unbearablely long traing time. In this work, we propose an adaptive frequency hopping policy which adaptively ajusts the frequency of policy to accelerate the training process. At the meantime, we design a control policy with muti early termination conditions to bias desired distributions for convergence. Experimental results demonstrate that our method achieves equivalent quantitative quality with a reduction of 50% of training time in comparison to the baselines. Yuehu Liu |
ICIP | 2 |
| 2021 | Multiview spectral clustering via complementary informationabstractAbstract In this article, multiview spectral clustering via complementary information (MSCC) is proposed, in which both the consensus information and the complementary information are explored for multiview clustering. In contrast to most multiview spectral clustering methods, the proposed MSCC considers the differences among multiple views and constructs a similarity matrix for clustering. Furthermore, a convex relaxation is employed and an algorithm that is based on the augmented Lagrange multiplier is proposed for optimizing the objective function of MSCC. In extensive experiments on five real‐world benchmark datasets, our proposed method outperforms two baselines and has significantly improved to several state‐of‐the‐art multiview clustering methods. Shuangxun Ma, Yuehu Liu, Qinghai Zheng, Yaochen Li, Zhichao Cui |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Geometric and semantic analysis of road image sequences for traffic scene construction
Yaochen Li, Yuehu Liu, Yuhui Hong, Jianji Wang 0001 |
Neurocomputing | 3 |
| 2020 | Multi-label X-Ray Imagery Classification via Bottom-Up Attention and Meta Fusion
Benyi Hu, Chi Zhang 0020, Le Wang 0003, Qilin Zhang 0004, Yuehu Liu |
ACCV (6) | 5 |
| 2020 | Calibrank: Effective Lidar-Camera Extrinsic Calibration By Multi-Modal Learning To RankabstractPrecise and online LiDAR-camera extrinsic calibration is one of the prerequisites of multi-modal data fusion for autonomous perception. The existing 6-DoF pose regression networks take majority effort on coarse-to-fine training strategy to gradually approach the global minimum. However, with limited computing resources, the optimal pose parameters seem unreachable. Moreover, recent research on neural network interpretability reveals that learning-based pose regression is nothing but the interpolation with most relevant samples. Motivated by this notion, we propose to solve the calibration problem in a retrieval way. Concretely, the learning-to-rank pipeline is introduced for ranking the top n relevant poses in the gallery set, which is then fused in to the final prediction. To better explore the pose relevance between ground truth samples, we further propose an exponential mapping from parametric space to the relevance space. The superiority of the proposed method is validated and demonstrated in the comparative and ablative experimental analysis. Xiannong Wu, Chi Zhang 0020, Yuehu Liu |
ICIP | 3 |
| 2020 | Caption Generation from Road Images for Traffic Scene ConstructionabstractIn this paper, an image captioning network is proposed for traffic scene modeling, which incorporates element attention into the encoder-decoder mechanism to generate more reasonable scene captions. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image captioning network is applied to generate the corresponding caption of each subregion. The static and dynamic traffic elements are appropriately organized to construct a 3D corridor scene model. The semantic relationships between the traffic elements are specified according to the captions. The constructed 3D scene model can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max and COCO datasets prove the effectiveness of the proposed framework. Yaochen Li, Le Wang 0003, Yuehu Liu |
IV | 5 |
| 2020 | Worst Perception Scenario Search for Autonomous DrivingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. In this paper, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as a discrete search problem. A single layer recurrent neural network with LSTM neurons is employed to predict WPS according to the searching reward, which is optimized by a vanilla policy gradient method. Moreover, to deal with the imbalanced distribution of real traffic scenarios, a KNN-like retrieval is utilized for searching the closest scenario samples. Effective yet efficient, the proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Further experiments reveal that detection networks with structural similarity share the similar WPS. Liheng Xu, Chi Zhang 0020, Yuehu Liu, Le Wang 0003, Li Li 0013 |
IV | 3 |
| 2020 | PyRetri: A PyTorch-based Library for Unsupervised Image Retrieval by Deep Convolutional Neural NetworksabstractDespite significant progress of applying deep learning methods to the field of content-based image retrieval, there has not been a software library that covers these methods in a unified manner. In order to fill this gap, we introduce PyRetri, an open source library for deep learning based unsupervised image retrieval. The library encapsulates the retrieval process in several stages and provides functionality that covers various prominent methods for each stage. The idea underlying its design is to provide a unified platform for deep learning based image retrieval research, with high usability and extensibility. The project source code, with usage examples, sample data and pre-trained models are available at https://github.com/PyRetri/. Benyi Hu, Renjie Song, Xiu-Shen Wei, Yazhou Yao, Xian-Sheng Hua 0001, Yuehu Liu |
ACM Multimedia | 6 |
| 2020 | Intermediate coordinate based pose non-perspective estimation from line correspondencesabstractIn this paper, a non-iterative solution to the non-perspective pose estimation from line correspondences was proposed. Specifically, the proposed method uses an intermediate camera frame and an intermediate world frame, which simplifies the expression of rotation matrix by reducing to the two freedoms from three in the rotation matrix R. Then formulate the pose estimation problem into an optimal problem. Our method solve the parameters of rotation matrix by building the fifteenth-order and fourth-order univariate polynomial. The proposed method can be applied into the pose estimation of the perspective camera. We utilize both the simulated data and real data to conduct the comparative experiments. The experimental results show that the proposed method is comparable or better than existing methods in the aspects of accuracy, stability and efficiency. Yujia Cao, Zhichao Cui, Yuehu Liu, Xiaojun Lv, Kaibei Peng |
MMAsia | 3 |
| 2020 | Coarse-to-fine 3D road model registration for traffic video augmentationabstractThis study addresses the problem of non‐perspective pose estimation from line correspondences in the traffic scenarios. A coarse‐to‐fine 3D road registration method is proposed for this problem in two stages. Firstly, the iterative closest point algorithm is exploited to estimate the pose coarsely. An objective function is then established to incorporate the feature correspondences for refining the coarse pose. Besides, the framework including road registration is employed for traffic video augmentation. The framework begins with the inputs of traffic videos, road information from Geographic Information Systems and 3D models of traffic elements (e.g. vehicles, pedestrians). Subsequently, 3D road model generation and point‐to‐line correspondence establishment are achieved in the preprossessing stage. After road and viewpoint registration, the 3D graphic engine is employed to simulate the traffic scene with the road, viewpoints and traffic elements. The augmented videos are generated by fusing the original frames and newly projected traffic elements. The authors demonstrate the superiority of the proposed registration method by the comparison to state‐of‐the‐arts in both quantitative and qualitative experiments. In addition, the frames of the augmented videos validate the proposed method in the application. Zhichao Cui, Yaochen Li, Chi Zhang 0020, Yuehu Liu, Fuji Ren |
IET Image Process. | 4 |
| 2020 | Spatiotemporal road scene reconstruction using superpixel-based Markov random field
Yaochen Li, Yuehu Liu, Jihua Zhu, Shiqi Ma, Zhenning Niu |
Inf. Sci. | 2 |
| 2020 | A sparse structure for fast circle detection
Yuanqi Su, Bonan Cuan, Yuehu Liu |
Pattern Recognit. | 4 |
| 2019 | Large Scale Traffic Signal Network Optimization - A Paradigm Shift Driven by Big DataabstractTraffic signal is the key method for city traffic control. Existing signal control systems use the loop detector data as the main input which is nearsighted in terms of both space and time. Lacking of effective data collection methods has hindered the development of more sophisticated models. It is therefore very hard to develop an optimization model considering all signals in a region or even a city. In this paper, we will introduce our method for large scale traffic signal optimization, which is the major module of Alibaba's city brain solution. By integrating multiple data sources to sense the whole city's traffic conditions, a layered model is developed based on the divide-and-conquer paradigm to gradually apply different types of data-driven optimization algorithms. It is a paradigm shift to use big data to improve a traditionally closed signal control system, and its effectiveness has been proven in a field test in Shanghai city. Liang Yu 0005, Jinqiang Yu, Maolei Zhang, Yuehu Liu, Wanli Min |
ICDE | 5 |
| 2019 | VIASEG: Visual Information Assisted Lightweight Point Cloud SegmentationabstractRapid and precise point cloud segmentation is one of the prerequisites for real-time and robust autonomous perception and environmental understanding, which requires a balance between speed and accuracy in architecture design. However, recent lightweight architectures, though fast enough, rely on domain adaptation from time-consuming-constructed synthetic dataset and sophisticated post-processing procedure to improve their performance, neglecting the rich visual information acquired by cameras aside from LiDAR sensors. In this paper, such color information is embedded at data-level to boost the performance of real-time point cloud segmentation. Furthermore, a multiscale lightweight fully convolutional network, VIASeg, is proposed based on the newly designed Super Squeeze Residual module and Semantic Connection from higher convolutional layers to lower layers, which improves the performance by feature denoising with high level semantic information. The superiority of the proposed method is validated and demonstrated in the comparative and ablative experimental analysis, while maintaining the real-time characteristic. Zhibin Zhong, Chi Zhang 0020, Yuehu Liu, Ying Wu 0001 |
ICIP | 3 |
| 2019 | Jointly Detecting and Retrieving Vehicles from Road Image Sequences based on CNNabstractIn this paper, a CNN-based vehicle detection and retrieval framework is proposed for the intelligent transportation system. Firstly, the vehicle target is detected from the traffic scene. The proposed object detection method uses a fully convolutional neural network (CNN) based on SqueezeNet, which has the characteristics of real-time, high accuracy and has small model size. Secondly, an intra-class image retrieval method is presented to search vehicles which are similar to the target vehicle in the dataset. The image retrieval results can be used for traffic scenes simulation and modeling. The experiments and comparisons prove the effectiveness of our framework. Yaochen Li, Yuehu Liu, Shanmin Pang, Le Wang 0003, Huihui Huo |
IV | 3 |
| 2019 | Road Scene Layout Reconstruction based on CNN and its Application in Traffic SimulationabstractIn this paper, we propose a road scene prediction framework based on the control points of road boundaries using CNN. Firstly, the image features are extracted and the heatmaps are generated by CNN to locate the control points of road boundaries. The input images are then segmented to specify the scene layout based on the control points. Furthermore, the 3D traffic scene models are constructed. The applications for traffic simulation are then developed. The evaluations and comparisons based on TSD-max dataset prove the effectiveness of the proposed method. Yaochen Li, Yuehu Liu, Zhichao Cui, Chi Zhang 0020 |
IV | 3 |
| 2019 | Homography-based traffic sign localisation and pose estimation from image sequenceabstractThis study proposes a vision‐based method for traffic sign attribute estimation, i.e. 3D position and pose, from image sequences by binocular or monocular cameras. The method starts with acquiring robust feature correspondences based on homography constraints from image pairs. Then the objective function is designed to integrate the feature correspondences to optimise the parameters of the traffic sign plane in the 3D coordinate. Finally, the sign plane is utilised for attribute estimation. In addition, the authors provide an extension for the raw KITTI dataset, which can be utilised for 3D tasks of traffic sign localisation and pose estimation. In the experiments, three popular methods are employed for comparisons based on the publicly available BelgiumTS and KITTI datasets. The results show that the authors’ method based on SIFT and SURF features can locate the traffic signs with a mean error of ∼0.44 and 0.51 m in the BelgiumTS and KITTI datasets, respectively, and estimate the pose with a mean error of ∼14.45° in the KITTI dataset. Zhichao Cui, Yuehu Liu, Fuji Ren |
IET Image Process. | 2 |
| 2019 | Joint Task Difficulties Estimation and Testees Ranking for Intelligence EvaluationabstractIn this paper, we study the testing tasks evaluation and testees ranking problem, in which tasks have different difficulty levels, and testees have different capabilities.We assume that a testee may have a probability to pass a certain task so as to allow certain uncertainty. The goal of this problem is to simultaneously determine the relative difficulty level of each testing task and the relative capability of every testee, purely based on the test outcome. We design two models to solve this problem. The first one assumes that the test outcome follows a certain Bernoulli distribution; while the second one assumes that the test outcome follows a certain Bernoulli distribution with the beta distribution-type a priori knowledge. Then, we form the original problem into likelihood estimation problems and solve them by using coordinate descent algorithms. We show that the beta distribution-type a priori knowledge is needed, when we only carry out a limited number of tests due to time and financial budgets. All these findings are useful to intelligence tests. Finally, we discuss how to extend this statistical learning model for more general cases as well as in a specific case in the field of Computational Social Systems like artificial social cognition evaluation. Chi Zhang 0020, Yuehu Liu, Li Li 0013, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2018 | Fused Discriminative Metric Learning for Low Resolution Pedestrian DetectionabstractLow resolution (LR) is one of the most challenging factor in pedestrian detection. In this paper, we propose a fused discriminative metric learning (F-DML) approach for low resolution pedestrian detection without explicit super resolution. We firstly learn a discriminative high resolution (HR) feature space as target space. Then, an optimal Mahanalobis metric is learned to transform the LR feature space into a new LR classification space, which largely preserves the discriminative structure of the HR feature space. Finally, a weighted K-nearest neighbors classifier is applied in the LR classification space which inherits good discrimination from HR feature space. A new training strategy is proposed to find the fewest and most representative LR-HR exemplars. In addition, we build a new dataset for the evaluation of low resolution pedestrian detection methods. Extensive experimental results demonstrate that the proposed approach performs favorably against the state-of-the-art methods. Xinzhao Li, Yuehu Liu, Zeqi Chen, Jiahuan Zhou, Ying Wu 0001 |
ICIP | 2 |
| 2018 | Dynamic Facial Expression Synthesis Driven by Deformable Semantic PartsabstractDynamic facial expression synthesis has some wild applications in human-computer interaction and virtual reality. The popular data-driven synthesis method like generative adversarial network (GAN) has made a great progress in generating a single face image, but has not well performed for expression sequences. To solve this problem, we design a series of deformable semantic parts to represent facial geometrical movement. And we synthesize the facial appearance by the geometrical driven under the-state-of-art pix2pixHD framework. In order to maintain the person identity among image sequence, we utilize an encoder to constrain the attributes of target face. With the above efforts, our method is capable to synthesize satisfied dynamic facial expression sequences. Nanxue Gong, Yang Yang 0066, Yuehu Liu, Dingdong Liu |
ICPR | 3 |
| 2018 | Traffic Sensory Data Classification by Quantifying Scenario ComplexityabstractFor unmanned ground vehicle (UGV) off-line testing and performance evaluation, massive amount of traffic scenario data is often required. The annotations in current off-line traffic sensory dataset typically include I) types of roadways II) scene types III) specific characteristics that are generally considered challenging for cognitive algorithms. While such annotations are helpful in manual selection of data, they are insufficient for comprehensive and quantitate measurement of per-roadway-segment scenario complexity. To resolve such limitations, we propose a traffic sensory data classification paradigm based on quantifying the scenario complexity for each roadway segment, where such quantification is jointly based on road semantic complexity and traffic element complexity. The road semantic complexity is a proposed measurement of the complexity incurred by the static elements such as curvy roads, intersections, merges and splits, which is predicted with a Support Vector Regression (SVR). The traffic element complexity is a measurement of complexity due to dynamic traffic elements, such as nearby vehicles and pedestrians. Experimental results and a case study verify the efficacy of the proposed method. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004 |
Intelligent Vehicles Symposium | 3 |
| 2018 | Multi-model Traffic Scene Simulation with Road Image Sequences and GIS InformationabstractIn this paper, a new multi-modal traffic scene simulation framework with combined inputs of road image sequences and road information from Geographic Information Systems (GIS) is proposed. The proposed framework contains two major steps, with the first one being a preprocessing step, including 3D road model extraction, camera location and orientation estimation and lane extraction from both GIS and road image sequences. After such preprocessing, the traffic scene reconstruction is reformulated into a 6-degree of freedom (6DoF) pose estimation in the 3D road model. Subsequently, the iterative closest point (ICP) algorithm is exploited for coarse point registration by estimating the pose in the road model. In addition, an objective function is established to incorporate the image features (e.g., lanes) into the road model and to refine the pose estimation. In the experiments with the publicly available KITTI dataset, the proposed method achieves high average Intersection-over-Union (IoU) scores as compared to the ground truth image sequences. Zhichao Cui, Yuehu Liu, Fuji Ren, Qilin Zhang 0004 |
Intelligent Vehicles Symposium | 2 |
| 2018 | A Graded Offline Evaluation Framework for Intelligent Vehicle's Cognitive AbilityabstractCognitive ability evaluation in intelligent vehicles is conventionally evaluated by classical autonomous driving dataset, which lacks comprehensive annotations of driving difficulty. Realistically, different driving conditions require vast different level of cognitive ability, e.g., driving in highly congested traffic is much more challenging than driving on limited access highway; driving in a blizzard/hurricane requires much more robust environmental cognition abilities than driving under ordinary conditions. Different datasets contain different proportions of various driving conditions, rendering intelligent vehicle evaluation susceptible to dataset variations. To overcome such limitations, we propose to first benchmark the driving difficulty with the proposed “Cascaded Tanks Model” and obtain a fine-grained per-segment difficulty rating based on our proposed Semantic Descriptor. With the proposed Graded Offline Evaluation (GOE) framework, it is demonstrated that offline validation of the cognitive abilities in Intelligent Vehicles (IV) is more consistent regardless of dataset choice. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004, Le Wang 0003 |
Intelligent Vehicles Symposium | 2 |
| 2018 | Active target tracking: A simplified view aligning method for binocular camera model
Xinzhao Li, Yuanqi Su, Yuehu Liu, Shaozhuo Zhai, Ying Wu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2018 | Data-Driven State-Increment Statistical Model and Its Application in Autonomous DrivingabstractThe aim of trajectory planning is to generate a feasible, collision-free trajectory to guide an autonomous vehicle from the initial state to the goal state safely. However, it is difficult to guarantee that the trajectory is feasible for the vehicle and the real path of the vehicle is collision-free when the vehicle follows the trajectory. In this paper, a state-increment statistical model (SISM) is proposed to describe the kinodynamic constraints of a vehicle by modeling the controller, the actuator, and the vehicle model jointly. The SISM consists of Gaussian distributions of lateral error increments in all state subspaces which are composed of the curvature radius, the velocity, and the lateral error. It is a data-driven modeling approach that can improve the SISM via increasing the number of samples of the increment-state, which is composed of the state and its corresponding increment of the lateral error. According to the SISM, the experience cost functions are designed to evaluate the trajectories for searching the best one with the lowest cost, and the real path can be predicted directly according to the planned trajectory and the vehicle state. The predicted path can be utilized effectually to evaluate the safety of the vehicle motion. Chao Ma 0024, Jianru Xue, Yuehu Liu, Jing Yang 0014, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | An Iterative Feature-Pair Updating Framework for Rigid Template Matching with OutliersabstractTo deal with the rigid template matching problem in real-world scenarios, we propose a novel iterative feature-pair updating framework which is also robust to high levels of outliers, such as background changing, complex nonrigid deformation and partial occlusion. Given a pair of template image and target image, we first extract a set of corresponding feature-pairs as candidates. Then, we propose a robust objective function under the iterative framework for discriminatively updating these candidates, where the space distance, appearance distance, and the overlapping percentage of feature pairs are integrated simultaneously. Finally, a hierarchical matching strategy is provided with the parameter discussion. Experimental results compared with the-state-of-art methods on public data sets demonstrate the effectiveness of the proposed method. Yang Yang 0066, Qian Kou, Shaoyi Du, Yuehu Liu, Bangyu Wu |
ISM | 5 |
| 2017 | Precise glasses detection algorithm for face with in-plane rotation
Shaoyi Du, Yuehu Liu, Xuetao Zhang 0001, Jianru Xue |
Multim. Syst. | 3 |
| 2016 | A driver fatigue detection method based on multi-sensor signalsabstractFatigue during long-time driving threatens the safety of drivers and transportation. In this paper, we provide an effective method based on multi-sensor signals collected from Kinect2.0 camera and PPG pulse sensor to build a driver fatigue detection system. Unlike most traditional works, we define the transitional process of fatigue and elaborate its effect on training classifiers. The simulation experiments are then designed and 15 groups of data are collected. Our method works in the following steps: 1) feature extraction and fusion, 2) sample labelling and 3) SVM classifier designing. The 10-fold cross-validation accuracy of the classifier is 90.10% and the test accuracy is 83.82%. Experimental results verify that our method to deal with samples in transitional process is universal and more accurate than traditional methods. Moreover, our method based on multi-sensor works better than those dealing with single-sensor. Yuanqi Su, Yuehu Liu, Danchen Zhao |
WACV | 3 |
| 2016 | Iteratively parsing contour fragments for object detection
Yuanqi Su, Yuehu Liu |
Neurocomputing | 3 |
| 2016 | Robust iterative closest point algorithm with bounded rotation angle for 2D registration
Chunjia Zhang, Shaoyi Du, Jianru Xue, Yuehu Liu |
Neurocomputing | 6 |
| 2016 | An efficient depth image-based rendering with depth reliability maps for view synthesis
Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Fast two-cycle curve evolution with narrow perception of background for object tracking and contour refinement
Yaochen Li, Yuanqi Su, Yuehu Liu |
Signal Process. Image Commun. | 3 |
| 2016 | Three-Dimensional Traffic Scenes Simulation From Road Image SequencesabstractIn this paper, we present a novel framework to allow users to tour simulated traffic scenes from the first-person view. Constructing 3-D scenes from road image sequences is in general difficult, due to the intrinsic complexity of dynamic road scenes, which are composed of a drastically moving background, not to mention numerous other surrounding vehicles. With the definitions of the traffic scene models, we first introduce the construction process of the simple traffic scenes. After the detection of road boundaries by a semantic fast two-cycle (FTC) level set method, we generate the control points on road sides to construct the “floor-wall” background scene that is subsequently propagated to each frame. Furthermore, we approach the cluttered traffic scenes through a three-component processing pipeline as follows: 1) traffic elements segmentation; 2) background images inpainting; and 3) traffic scenes construction. The traffic elements in the cluttered images are segmented by the semantic FTC level set method first. A Gaussian mixture model is then employed to inpaint the occluded background utilizing the optical flows. The cluttered traffic scenes can be constructed after the segmentation and inpainting components. The foreground polygons such as vehicles and traffic signs are then modeled. Users can change their viewpoints according to their own interpretations. We present the evaluations of each technical component, followed by our findings from comprehensive user studies, which well demonstrate the effectiveness of the proposed framework in delivering good touring experience to users. Yaochen Li, Yuehu Liu, Yuanqi Su, Gang Hua 0001, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Contour Guided Hierarchical Model for Shape MatchingabstractFor its simplicity and effectiveness, star model is popular in shape matching. However, it suffers from the loose geometric connections among parts. In the paper, we present a novel algorithm that reconsiders these connections and reduces the global matching to a set of interrelated local matching. For the purpose, we divide the shape template into overlapped parts and model the matching through a part-based layered structure that uses the latent variable to constrain parts' deformation. As for inference, each part is used for localizing candidates by the partial matching. Thanks to the contour fragments, the partial matching can be solved via modified dynamic programming. The overlapped regions among parts of the template are then explored to make the candidates of parts meet at their shared points. The process is fulfilled via a refined procedure based on iterative dynamic programming. Results on ETHZ shape and Inria Horse datasets demonstrate the benefits of the proposed algorithm. Yuanqi Su, Yuehu Liu, Bonan Cuan, Nanning Zheng 0001 |
ICCV | 2 |
| 2015 | Fast Two-Cycle level set tracking with narrow perception of backgroundabstractThe problem of tracking foreground objects in a video sequence with moving background remains challenging. In this paper, we propose the Fast Two-Cycle level set method with Narrow band Background (FTCNB) to automatically extract the foreground objects in such video sequences. The level set curve evolution process consists of two successive cycles: one cycle for data dependent term and a second cycle for smoothness regularization. The curve evolution is implemented by computing the signs of region competition terms on two linked lists of contour pixels rather than solving any Partial Differential Equations (PDEs). Maximum A Posterior (MAP) optimization is applied in the FTCNB method for curve refinement with the assistance of optical flows. The comparison with other level set methods demonstrate the tracking accuracy of our method. The tracking speed of the proposed method also outperforms the traditional level set methods. Yaochen Li, Yuanqi Su, Yuehu Liu |
ICME | 3 |
| 2015 | State-statistical model based trajectory-band planning in urban environmentabstractIn the traditional trajectory planning methods, a feasible, collision-free trajectory is generated to guide the vehicle. But generally the vehicle cannot follow the trajectory without tracking deviation because of the vehicle kinematical constraints and the performance of control algorithm. In this paper, State-Statistical Model (SSM) based trajectory-band planning method is proposed to predict the vehicle motion during the vehicle tracks the trajectory. In this method, the statistics of historical states are used to build the SSM which is a normal distribution model of tracking deviation in different segments of curvature radius and velocity. According to the SSM, the inaccessible states of vehicle can be obtained to search the best trajectory and the tracking deviation boundary can be calculated on the trajectory. Then the best trajectory is used as the base line to generate the trajectory-band of which the halfband width is the deviation boundary value. As a result, the trajectory-band can represent the maximum range of vehicle motion accurately. Chao Ma 0024, Jing Yang 0014, Jianru Xue, Yuehu Liu |
Intelligent Vehicles Symposium | 4 |
| 2015 | A line segments extraction based undirected graph from 2D laser scansabstractA novel algorithm to find line segments is proposed from the sequence of points taken by a laser scan, which consists of over segmentation, undirected graph generation and line segments extraction. Firstly, the self-adaptive IEPF method is used to over segment raw points into sub-groups. Secondly, the undirected graph is generated based on merge probability function of sub-groups. Thirdly, according to the main edges, the energy function of undirected graph partition and corresponding minimization method, line segments are extracted. The primary contributions of the work are the main edge, undirected graph and minimization method of energy function, which are more proper to noisy laser scan data. The experimental comparison with 6 state-of-the-art methods, using 2D laser scan data obtained by laser profile sensor and laser range finder, shows better performance and robustness of the new algorithm. Xinzhao Li, Yuehu Liu, Zhenning Niu, Zhichao Cui |
MMSP | 2 |
| 2015 | Autonomous Driving Simulation for Unmanned VehiclesabstractHuman can judge driver's driving ability by observing the vehicle motion in different traffic scenes. Identically, driving behavior can be the main basis for evaluating the performance of an unmanned vehicle in both field test and simulation test. Although simulation test avoids disadvantages of field test, existing simulation systems lack traffic scene data with perception granularity of visual sensors. In order for realizing vehicle-in-loop simulation, simulation technique of driving behaviors must be able to exhibit actual motion of unmanned vehicles. In this paper, we propose an automatic approach of simulating autonomous driving behaviors of vehicles in traffic scene represented by image sequences. Different from general simulation systems, we use actual traffic environment data to build the traffic scene and simulate the driving behaviors. After the proposed method was embedded in scene browser, a typical traffic scene including the intersections was chosen for virtual vehicle to execute the driving tasks of lane change, overtaking, slowing down and stop, right turn and U-Turn. The experimental results show that different driving behaviors of vehicles in typical traffic scene can be exhibited smoothly and realistically. Our method can also be used for generating simulation data of traffic scenes that are difficult to collect. Danchen Zhao, Yuehu Liu, Chi Zhang 0020, Yaochen Li |
WACV | 2 |
| 2015 | Scene text detection method based on the hierarchical modelabstractAs an important step in text‐based information extraction systems, scene text detection has become a popular subject of research in recent years. In this study, the authors present a novel approach to robustly detect texts which are variable in scales, colours, fonts, languages and orientations in scene images. To segment candidate text connected components (CCs) from images, both local contrast and colour consistency are considered in superpixel level. To filter out the non‐text CCs, a hierarchical model is designed. This hierarchical model groups the CCs into three cascaded stages, and is equipped with a well‐designed classifier in each stage. Experimental results on the public ICDAR 2005 dataset and the MSRA‐TD500 dataset show that their approach obtains better performance than other state‐of‐the‐art methods. Yuehu Liu, Zhenhong Jia |
IET Comput. Vis. | 2 |
| 2015 | Double layer multiple task learning for age estimation with insufficient training samples
Yuehu Liu, Nanning Zheng 0001 |
Neurocomputing | 4 |
| 2015 | Video object segmentation by integrating trajectories from points and regions
Zejian Yuan, Yuehu Liu, Nanning Zheng 0001 |
Multim. Tools Appl. | 3 |
| 2015 | Exact solution to median surface problem using 3D graph search and application to parameter space exploration
Zhengwang Wu, Xiaoyi Jiang 0001, Nanning Zheng 0001, Yuehu Liu, Da-Chuan Cheng |
Pattern Recognit. | 4 |
| 2015 | Multi-target tracking by learning local-to-global trajectory models
Jinjun Wang, Zelun Wang, Yihong Gong, Yuehu Liu |
Pattern Recognit. | 5 |
| 2014 | Semantic propagation network with robust spatial context descriptors for multi-class object labeling
Ping Wei 0001, Yuehu Liu, Nanning Zheng 0001, Shaozhuo Zhai |
Neural Comput. Appl. | 2 |
| 2014 | Hybrid constraint SVR for facial age estimation
Lixin Duan, Yuehu Liu |
Signal Process. | 5 |
| 2013 | An iterative parsing approach for contour fragmentsabstractThis paper presents an approach for linking edge points into contour fragments, each of which is an ordered point sequence. In the problem of shape-based object detection, we investigate what characteristics of the contour fragments representation influence the detection performance. It is observed that if interest objects are represented by limited number of contour fragments, with moderate count of noisy points (not belonging to interest objects) brought in, the detection algorithm can reliably obtain interest objects. For this purpose, we propose to iteratively extract a pair of points which preserves the farthest geodesic distance in each connected component of edge map. We conduct experiments on the ETHZ dataset and compare with other typical methods for generating contour fragments. Experimental results demonstrate that the proposed approach outperforms some existing ones in object detection. Yuehu Liu, Yuanqi Su |
ICME | 2 |
| 2013 | Automatic face image annotation based on a single template with constrained warping deformationabstractIn this study, an automatic face image annotation method is proposed by aligning faces with different expressions to an annotated neutral face. This work is useful in reducing tedious manual work for labelling image data in large databases. However, it is challenging because of the appearance variations caused by non‐rigid face deformations under various expressions. Unlike some conventional approaches acquiring sufficient image templates to model the query appearance, only a single given template is necessary for the proposed method. The authors address the problem through dense image alignment. Specifically, image warping in the alignment process is constrained by prior knowledge about facial shape deformation. The proposed method is independent of the appearance model, and is available for unseen faces. In addition, to initialise warping parameters, the authors present a robust patch‐based estimation method. Context information for feature points is carefully modelled to propagate the searching path for local patch matching. The face annotation experiments are performed on some large expressions, with noisy image qualities and in low image resolutions. Comparison results with conventional methods demonstrate the proposed method's superiority on both accuracy and robustness. Yang Yang 0066, Yuehu Liu |
IET Comput. Vis. | 2 |
| 2013 | A Voting Scheme for Partial Object Extraction under Cluttered EnvironmentabstractShape extraction aims to detect and localize objects via the shape information. The paper presents a novel voting scheme that can extract partially occluded objects under cluttered environments using a single shape. It works by jointly figuring out the boundaries and resolving the geometric configurations. To model the missing part lead by occlusion, we discretize the shape template into a set of its subpart, named portions. Our representation of shape template is through a set of portion together with their interconnections. Instead of forming a fully connected network, our interconnections make the portions consistent with the chain along the boundary of shape template. Based on the representation, we formulate an auto-locked objective function that contains both the unary and pairwise terms and balances the effects of missing parts. Min-sum voting scheme with strategy driven by bottom–up information is then proposed to minimize the objective function. Conducted experiments show that proposed algorithm is promising for shape extraction with occlusion and noisy backgrounds and allows the non-rigid deformations. Yuanqi Su, Yuehu Liu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Parameter analysis of fractal image compression and its applications in image sharpening and smoothing
Jianji Wang 0001, Nanning Zheng 0001, Yuehu Liu |
Signal Process. Image Commun. | 3 |
| 2012 | Improved View Synthesis with Depth Reliability MapsabstractAn improved view interpolation scheme with depth reliability maps is presented to enhance the quality of synthesized images in this paper. First of all, the left depth map and the right one can be estimated by a graph cut approach from the input reference images. Then, the predict depth maps are derived by forward mapping, and then a median filter is applied to eliminate small blank points caused by forward mapping. Next, the left predict image and the right one are synthesized by reverse mapping using the reference image with the associated predict depth map. After that, in order to obtain better visual quality, the post-processing operations are developed. On the one hand, the segmentation mask can be derived from the depth reliability maps, which are produced from the depth maps. On the other hand, based on the segmentation mask, a virtual view image is generated by a weighted interpolation scheme from the left and right predict images. Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
DCC | 3 |
| 2012 | Disocclusion using depth reliability map for view synthesisabstractIn this paper, a novel disocclusion scheme is proposed based on the depth reliability maps for virtual view synthesis. One important aspect of the developed method is that the occlusion map is estimated with the depth reliability maps. Therefore, the holes in a virtual view image can be filled in by blending the two predict images based on the occlusion map. The experimental results demonstrate that the addressed method has better performance for view synthesis than the other three conventional methods through the assessment of image quality and algorithm cost. Yi Lai, Xuguang Lan, Yuehu Liu, Nanning Zheng 0001 |
ICASSP | 3 |
| 2012 | Scene text detection with superpixels and hierarchical modelabstractScene text detection is a challenging task for the text-based information extraction systems. We present a novel scene text detection method for this task. The images are over-segmented into meaningful perceptron superpixels, and candidate connected components (CCs)are extracted by combining local contrast and color consistency. The non-text components are then pruned by a hierarchical model consisting of three stages in cascade. Experimental results show that our approach is better than other state-of-the-art methods. Yuehu Liu |
ICIP | 2 |
| 2012 | Video object segmentation by clustering region trajectories
Zejian Yuan, Dapeng Chen, Yuehu Liu, Nanning Zheng 0001 |
ICPR | 4 |
| 2011 | Fractal image coding using SSIMabstractSince Jacquin proposed original fractal image compression technique in 1990, fractal coding method has been developed into various schemes. Traditionally, fractal coding uses mean square error (MSE) to evaluate similarity of image blocks, but the similarity evaluated by MSE usually differs from human visual system (HVS). Compared with MSE, structural similarity (SSIM) is an image measure index which is more appropriate for the HVS. This paper proposes a new fractal coding scheme which uses structural similarity to measure the similarity between image blocks and compute these blocks' coefficients. The experiment results show that the proposed method generates higher quality images for the HVS than MSE scheme. Jianji Wang 0001, Yuehu Liu, Ping Wei 0001, Yaochen Li, Nanning Zheng 0001 |
ICIP | 2 |
| 2011 | A new hybrid method to detect text in natural sceneabstractIn this paper, a new hybrid method for text detection in natural scene is proposed. According the linguistics rules, this algorithm mainly consists of three parts. First, considering both unary property and binary relationship, the conditional random field (CRF) model is introduced for text region detection. Second, connected components (CCs) are extracted by similar stroke width, and filtered coarsely by stroke width analysis. Candidate CCs are then filtered by candidate regions. Finally, text CCs are clustered into words by geometry heuristics. Experiments on the public benchmark ICDAR 2003 dataset show that proposed algorithm can detect text with various font sizes in natural scene. Yuehu Liu, Yuanqi Su |
ICIP | 2 |
| 2011 | 3D facial mesh detection using geometric saliency of surfaceabstractThis paper proposes a 3D facial mesh detection algorithm based on the geometric saliency of surface. Specifically, the geometric saliency of each vertex on 3D triangle mesh is measured by the combination of Gaussian-weighted curvature and spin-image correlation. Salient vertices with similar properties are clustered into regions on the saliency map, and represented as nodes by the graph model. To detect a 3D facial mesh, initialization and registration steps are applied to match each triangle in the graph model with a reference graph, corresponding to a 3D reference facial mesh. Furthermore, the match error between the graph model of the testing 3D mesh and the reference facial mesh is computed to classify face and non-face meshes. Experimental results demonstrate that the proposed algorithm is effective to detect 3D facial meshes and robust to facial expressions and geometric noises. Yaochen Li, Yuehu Liu, Yuanchun Wang 0003, Zhengwang Wu, Yang Yang 0066 |
ICME | 2 |
| 2011 | Expression transfer for facial sketch animation
Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yuanqi Su, Yoshifumi Nishio |
Signal Process. | 3 |
| 2010 | Optimal Trajectory Space Finding for Nonrigid Structure from Motion
Yuanqi Su, Yuehu Liu, Yang Yang 0066 |
ACIVS (1) | 2 |
| 2010 | A Parameterized Representation for the Cartoon Sample Space
Yuehu Liu, Yuanqi Su, Daitao Jia |
MMM | 1 |
| 2010 | Eye synthesis using the eye curve model
Nanning Zheng 0001, Shaoyi Du, Yuehu Liu |
Image Vis. Comput. | 5 |
| 2009 | A Statistical-Structural Constraint Model for Cartoon Face Wrinkle Representation and Generation
Ping Wei 0001, Yuehu Liu, Nanning Zheng 0001, Yang Yang 0066 |
ACCV (3) | 2 |
| 2009 | Example-based performance driven facial shape animationabstractA novel performance driven facial shape animation method is presented for mapping the expressions from the source face to the target face automatically. Unlike the prior expression cloning approaches, the proposed method aims to animate a new target face with the help of real facial expression samples. The basic idea is to learn the shape deformation from samples for target face to generate corresponding expressions. The process consists of two main stages. First of all, source motion vectors are transferred by statistic face model to generate a reasonable expression on the target face. And then, local deformation constraints are proposed to refine the animation results. In the second part, the local deformation characters for each target organ are learned from the samples, which preserve the personality as well as the expression styles. Experimental results on different facial animation demonstrate the feasibility and effectiveness of the proposed method. Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yoshifumi Nishio |
ICME | 3 |
| 2008 | A Method for Deforming-Driven Exaggerated Facial Animation GenerationabstractThis paper proposes a method for automatically generating facial animation with exaggerated features from a facial image. According to facial diversity and identity, an exaggerated face can be determined by the neutral face and the exaggeration effect difference that is represented with the deforming parameters. The proposed method utilizes the central point and the feature vector to represent a facial component and then transforms these feature vectors under the control of deforming parameters to generate the exaggerated facial features. The exaggerated facial animation can be driven to generate by the sequence of the deforming parameters. Experimental results prove that the proposed method has the advantages of simplicity, flexibility and directness, and the generated facial animations are expressive. Ping Wei 0001, Yuehu Liu, Yuanqi Su |
CW | 2 |
| 2007 | Automatic Region-of-Interest Coding in JPEG2000 Based on Morphology Segmentation and LLn Subband Analysis
Fei Wang 0008, Nanning Zheng 0001, Yuehu Liu |
ICIC (3) | 3 |
| 2007 | VLSI Design of a High-Speed and Area-Efficient JPEG2000 EncoderabstractA high-speed VLSI design of an area-efficient JPEG2000 encoder is given. Recursive multilevel 2D discrete wavelet transform (DWT) architecture with dual buffers is proposed to reduce the wavelet coefficients memory to 1/4 tile size, prerate allocation is used to reduce the compressed code memory to 3/4 tile size. A highly pipelined and parallelism implementation of line-based 1-level DWT is proposed using two line-buffers in 5/3 wavelet type and its input speed is up to 2 samples/cycle; code block based address mapping in access wavelet coefficients memory, concurrent state variables generation and multiple parallel and pipeline coding methods are used in the bit plane encoder (BPE) which encodes on average at 40.5 M samples/s at 100 MHz with no memory used; the conditional two-symbol pipeline arithmetic encoder (AE) encodes at 1.3 symbols/cycle. Parallel units in BPE and buffer control between BPE and AE are optimally implemented with low cost without performance loss. Byte representation of rate-distortion slope used reaches a near optimal implementation of post-coding rate distortion in Tier2 with low cost. The compressed file generated by the encoder is fully compatible with ISO/IEC FCD15444-1. The encoder is verified on field-programmable gate array platform with a direct interface to digital video input with tile size 256 times 256 and code block size 32 times 16. The resulting input sampling rate is up to 58 M samples/s when Tier1 operates at 100 MHz. Difference of the peak signal-to-noise ratio of images compressed by our encoder and JasPer is less than 0.2 dB when the compression ratio is greater than 1 bps. Equivalent NAND2 gates synthesized are 90.6 K and on-chip RAM size is 626.75 kb. Unlike other designs the proposed design of JPEG2000 encoder has high compression quality as well as high speed and area-efficiency. Nanning Zheng 0001, Chang Huang, Yuehu Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2006 | A Boosting SVM Chain Learning for Visual Information Retrieval
Zejian Yuan, Yanyun Qu, Yuehu Liu |
ISNN (1) | 4 |
| 2005 | A Cascaded Mixture SVM Classifier for Object Detection
Zejian Yuan, Nanning Zheng 0001, Yuehu Liu |
ISNN (1) | 3 |