EDBT 2026 Demo / reviewers in the wild / expert
Jing Wu 0004
dblp:88/3604-4
· DBLP profile ↗
51ranked-venue papers
11as first author
33since 2021 · last 2025
0000-0001-5123-9861ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-author · 23 since 2021Artificial intelligence and machine learning · 23 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Class Part Parsing Based on Multi-Class BoundariesabstractMulti-class part parsing is a dense prediction task that segments objects into semantic components with multi-level abstractions. Despite its significance, this task remains challenging due to ambiguities at both part and class levels. In this paper, we propose a network that incorporates multi-class boundaries to precisely identify and emphasize the spatial boundaries of part classes, thereby improving segmentation quality. Additionally, we employ a weighted multi-label cross-entropy loss function to ensure balanced and effective learning from all parts. Experimental results validate the effectiveness of the proposed method, demonstrating its ability to enhance baseline performance on benchmark datasets. Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji |
ICIP | 2 |
| 2025 | RGB-D Video Mirror DetectionabstractMirror detection aims to identify mirror areas in a scene, with recent methods either integrating depth information (RGB-D) or making use of temporal information (video). However, utilizing both data is still under-explored due to the lack of a high-quality dataset and an effective method for the RGB-D Video Mirror Detection (DVMD) problem. To the best of our knowledge, this is the first work to address the DVMD problem. To exploit depth and temporal information in mirror segmentation, we first construct a large-scale RGB-D Video Mirror Detection Dataset (DVMD-D), which contains 17977 RGB-D images from 273 diverse videos. We further develop a novel model, named DVMDNet, which can first locate the mirrors based on triple consistencies: local consistency, cross-modality consistency and global consistency, and then refine the mirror boundaries through content discontinuity, taking the temporal information within videos into account. We conduct a comparative study on the DVMD dataset, evaluating 12 state-of-the-art models (including single-image mirror detection, single-image glass detection, RGB-D mirror detection, video shadow detection, video glass detection, and video mirror detection methods). Code is available from https://github.com/UpChen/2025_DVMDNet. Mingchen Xu, Peter Herbert, Yukun Lai, Ze Ji, Jing Wu 0004 |
WACV | 5 |
| 2025 | REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
Comput. Vis. Image Underst. | 8 |
| 2025 | Guest Editorial: Multi-view representation learning for computer visionabstractObject recognition and scene analysis in single-view images may face difficulties such as occlusion and incomplete information, while multi-view learning can address this limitation. When an object or scene is observed from multiple views, information on target objects can be significantly enriched to improve the performance of computer vision tasks. For this reason, multi-view has become one of the important forms of data representation, which leads to the emerging of new research topics on complete or in-complete multi-view learning. Multi-view learning enables the use of multi-source information, nevertheless, the heterogeneous characteristics of data make it difficult to reliably associate information from different views, especially in a complex environment. It remains challenging for tasks to make effective use of the consistent and complementary information between different complete views and to enhance the completeness of potential representation. A wide variety of research is being conducted to explore and discover possible challenges and opportunities to exploit multi-view representation learning for computer vision. The purpose of this Special Issue is to collect high-quality articles on the recent development and trend of multi-view representation learning in computer vision, publish new ideas, theories, solutions and insights on this topic, and showcase their applications. In this Special Issue, we have received 36 papers, all of which underwent peer review. Of the 36 originally submitted papers, 10 have been accepted, which cover a variety of fields, such as person re-identification, gait recognition, 3D object recognition, and behaviour recognition. These accepted papers are mainly divided into three categories. The first category covers the incomplete multi-view data learning theoretics and methods. The papers in this category are of He et al., Kun et al., Fan et al. and Wang et al. The last two categories are both multi-view applications. One of which is 3D-related applications. The papers in this category are of Qi et al. and Sun et al. The other category is about 2D recognition. The papers in this category are of Zhang et al., Huang et al., Zheng et al. and Zhang et al. A brief presentation of each of the paper follows. He et al. present an innovative multi-view subspace clustering method with incomplete graph information. Specifically, they separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. The clustering results on six real-world datasets show that the method outperforms a series of classic incomplete multi-view clustering methods. Kun et al. propose a new method for low-rank-based multi-view subspace clustering based on low-rank correlation analysis. To overcome the limitations of unreliable low-rank structure and imprecise graphs caused by multi-view noise and outliers, they introduce the canonical correlation analysis strategy and a dual regularisation term to characterise the connections between different views adaptively. Experimental results reveal the method's superiority over compared state-of-the-art (SOTA) methods in accuracy, normalised mutual information, and F-score evaluation metrics. Fan et al. address the challenge of partial mapping between the views in multi-view clustering, and propose a self-inferring incomplete multi-view clustering algorithm to explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances. Experimental results show that the method can improve the clustering performance compared with the SOTA methods. Wang et al. propose a semi-paired semi-supervised deep hashing to solve the large-scale multimedia retrieval task. The method is an end-to-end deep neural network model with high-order affinity. To maintain the consistency within the modalities, they introduce a common representation that combines with the labelled information to associate different modalities. Experimental results demonstrate the superior performance of proposed method. Qi et al. propose a double-weighting convolution neural network based on the L2-S grouping mechanism for multi-view 3D object recognition. The goal of the proposed L2-S grouping mechanism is to calculate the discrimination score of views and group views more reasonably. Results of the experiments show that the method can achieve SOTA performance. Sun et al. present a dual-matching method with cross-attention mechanism to address the limitations of matching-based methods caused by a preset fixed disparity range on depth estimation task. To tackle the mismatches on edges and details, they introduce an exquisite module based on left-right consistency. The method is proved to be competitive and effective by experiments conducted under popular benchmarks. Zhang et al. want to answer the following two questions: (1) does a query image with higher resolution than that of the gallery image also affect the pedestrian re-identification performance? If so, and (2) how does it affect performance? So, they propose an end-to-end trainable resolution independent person re-identification network that is composed of a cross-resolution Generative Adversarial Networks and embedding batch normalisation layers. The results demonstrate that the proposed method outperforms the SOTA methods in the pedestrian re-identification task on their expanded benchmark dataset. Huang et al. address the limitation of current gait-based age and gender recognition methods under multi-view scene, and propose an attention-aware spatio–temporal learning framework that employs silhouette sequence as an input to learn essential spatial–temporal gait representation. The proposed method has produced results that outperformed the benchmarks with an Mean Absolute Error of 6.68 years for age estimation and a Correct Classification Rate of 97% for gender classification. Zheng et al. apply deep learning to multi-view classroom behaviour detection. First, they propose an improved detection model based on YOLOv5 to improve the convergence speed of the prediction box. Second, they establish a quantitative evaluation standard for students' classroom attention, and then conduct training and verification by collecting multi-view classroom datasets. Finally, they increase the environment variation in the training model phase to make the model have better generalisation ability. Experiments demonstrate that the method can effectively identify and detect students' behaviours in the classroom from different views. Zhang et al. propose a method for multi-dimensional video anomaly detection, which uses the Object-meta instead of video frames as the input, and the Memory Search Guided Autoencoder with Memory Pools (MSGAE-MP) to reconstruct. The multi-dimensional information carried by the input can be strengthened via Object-meta. The MSGAE-MP construct multi-level memory pools, so as to reconstruct Object-meta in different dimensions. Experiments show that the method is feasible and has achieved excellent results. All of the papers published in this Special Issue show that multi-view representation learning theoretics have developed very fast in recent years. In addition, it is very promising to solve traditional computer vision tasks under multi-view setting, including but not limited to 3D object recognition, person re-identification, gait-based age and gender estimation, and depth estimation. Xin Ning and Chen Wang are responsible for the writing of Proposal and Editorial materials; Jun Zhou is responsible for the processing of articles; and Jing Wu, Lin Gu and Jian Cheng are responsible for the solicitation and publicity of the special issue. Firstly, we would like to thank all the authors for their innovative contributions and all the reviewers for their professional and crucial, yet constructive comments. Also, we wish to express our thanks to Mr Hang Ran, PhD students at Institute of Semiconductors, Chinese Academy of Sciences, for his assistance in this process. Last, we wish to express our gratitude to the editorial team of IET Computer Vision for their support throughout this venture. We hope you enjoy this collection of papers and that the Special Issue can stimulate further research and development in this area. This work is supported by the National Natural Science Foundation of China (Grant no. 61901436). National Natural Science Foundation of China, Grant/Award Number: 61901436. Data sharing is not applicable to this article as no new data were created or analyzed in this study. Xin Ning (SMIEEE) received a B.S. degree in software engineering in 2012, and a Ph.D. degree in electronic circuit and system from the university of Chinese Academy of Sciences, in 2017. He is currently an associate professor with the Laboratory of Artificial Neural Networks and High Speed Circuits, Institute of Semiconductors, Chinese Academy of Sciences. His current research interests include neural networks, intelligent systems and computer vision. He has published as the first or corresponding author in more than 45 papers in journals and refereed conferences. Now he serves as the young associated editor of CAAI Transactions on Intelligent Systems, the guest editor of Elsevier Journal on DISPLAYS. He is also the guest editor of CONNECTION SCIENCE and CONCURR COMP-PRACT E. He was the Website Chair of the IEEE HPBD&IS 2020 and the Publication Chair of the IEEE HPBD&IS 2021. Jun Zhou received a B.S. degree in computer science and a B.E. degree in international business from the Nanjing University of Science and Technology, Nanjing, China, in 1996 and 1998, respectively, an M.S. degree in computer science from Concordia University, Montreal, QC, Canada, in 2002, and a Ph.D. degree in computing science from the University of Alberta, Edmonton, AB, Canada, in 2006. He was a research fellow with the Research School of Computer Science, The Australian National University, Canberra, ACT, Australia, and a researcher with the Canberra Research Laboratory, National Information and Communications Technology Australia, Canberra. In 2012, he joined the School of Information and Communication Technology, Griffith University, Nathan, QLD, Australia, where he is currently a reader. His research interests include pattern recognition, computer vision, and spectral imaging and their applications in remote sensing and environmental informatics. He is the associate editor for the journal of Pattern Recognition and IEEE Trans. on Remote Sensing. Jian Cheng is a professor of Institute of Automation, Chinese Academy of Sciences. He received the B.S. and M.S. degrees in Mathematics from Wuhan University in 1998 and 2001, respectively. After that, he received a Ph.D degree in pattern recognition and intelligent systems from Institute of Automation, Chinese Academy of Sciences in 2004. His current major research interests include deep learning, computer vision, chip design, etc. Jing Wu is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Chen Wang is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Lin Gu received a B.Eng. degree from Shanghai University, Shanghai, China, in 2009, and a Ph.D. degree in computer vision from Australian National University in 2014. After Ph.D. graduation from the Australian National University, he worked as a post-doctoral researcher at A*STAR, Singapore. Then, he was a project researcher with the National Institute of Informatics, Japan, and also a visiting scholar with Kyoto University, Japan. He is currently a research scientist at RIKEN AIP, Japan, and a special researcher with the University of Tokyo, Japan. He is also an in-charge of a Moonshot and an ACT-X Project to improve artificial intelligence by simulating the human brain. His primary research interests lie in machine learning, medical imaging, and computational photography. Xin Ning 0001, Jun Zhou 0001, Jian Cheng 0001, Jing Wu 0004, Chen Wang 0026, Lin Gu 0003 |
IET Comput. Vis. | 4 |
| 2025 | A Survey of Object Goal NavigationabstractObject Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. Though it has drawn increasing attention from researchers in the Embodied AI community, there has not been a contemporary and comprehensive survey of ObjectNav. In this survey, we give an overview of this field by summarizing more than 70 recent papers. First, we give the preliminaries of the ObjectNav: the definition, the simulator, and the metrics. Then, we group the existing works into three categories: 1) end-to-end methods that directly map the observations to actions, 2) modular methods that consist of a mapping module, a policy module, and a path planning module, and 3) zero-shot methods that use zero-shot learning to do navigation. Finally, we summarize the performance of existing works and the main failure modes and discuss the challenges of ObjectNav. This survey would provide comprehensive information for researchers in this field to have a better understanding of ObjectNav.Note to Practitioners—This work was motivated by the increased interest in real-world applications of mobile robots. Object Goal Navigation (ObjectNav), which is an important task in these applications, requires an agent to find an object in an unseen environment. To accomplish that, the agent needs to be equipped with the capability to move in the environment, decide where to go, and recognize the object categories. So far, most works on ObjectNav have been done in a simulation environment. We present an overview of the existing works in ObjectNav and introduce them in three categories. Additionally, we analyze the current performance of ObjectNav and the challenges for future research. This paper provides researchers and practitioners with a comprehensive overview of the developed methods in ObjectNav, which can help them to have a good understanding of this task and develop suitable solutions for applications in the real world. Jing Wu 0004, Ze Ji, Yukun Lai |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Dictionary Based Generative Adversarial Network for Multi-Collection Style TransferabstractMost collection-based style transfer methods require training a separate model for each individual collection of styles, making the extension to multiple collections of styles less flexible. Besides, the existing collection-based methods are also less flexible in extending to new style collections in a continual manner. To address these issues, we propose a novelMultI-Dictionary Generative Adversarial Network framework (MID-GAN)for multi-collection style transfer. Specifically, we design a multi-dictionary architecture within a GAN, with each dictionary consisting of a set of local style codes for a specific style collection. Benefiting from the local style codes used in the dictionary, a stylization module with aligned skip connections is further proposed, which can better preserve both the local details and the overall image structure. The dictionary design allows a flexible extension to new style collections by readily adding new dictionaries and we propose a continual training strategy that can both preserve the style transfer ability of old styles and achieve good transfer results for newly added styles. Extensive experiments are performed to show that the proposed method is better than existing collection-based style transfer methods. We also demonstrate the proposed method can generate diverse meaningful style transfer results of the same style collection. Jing Huo, Shiyin Jin, Jiashen Li, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | RuleExplorer: A Scalable Matrix Visualization for Understanding Tree Ensemble ClassifiersabstractThe high performance of tree ensemble classifiers benefits from a large set of rules, which, in turn, makes the models hard to understand. To improve interpretability, existing methods extract a subset of rules for approximation using model reduction techniques. However, by focusing on the reduced rule set, these methods often lose fidelity and ignore anomalous rules that, despite their infrequency, play crucial roles in real-world applications. This paper introduces a scalable visual analysis method to explain tree ensemble classifiers that contain tens of thousands of rules. The key idea is to address the issue of losing fidelity by adaptively organizing the rules as a hierarchy rather than reducing them. To ensure the inclusion of anomalous rules, we develop an anomaly-biased model reduction method to prioritize these rules at each hierarchical level. Synergized with this hierarchical organization of rules, we develop a matrix-based hierarchical visualization to support exploration at different levels of detail. Our quantitative experiments and case studies demonstrate how our method fosters a deeper understanding of both common and anomalous rules, thereby enhancing interpretability without sacrificing comprehensiveness. Zhen Li 0044, Weikai Yang, Jun Yuan 0003, Jing Wu 0004, Changjian Chen, Yao Ming, Fan Yang 0094, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | GRPSNET: Multi-Class Part Parsing Based on Graph ReasoningabstractMulti-class part parsing is a dense prediction task that decomposes objects into semantic components with multi-level abstractions. Despite the importance of this problem, it remains challenging due to the presence of both part-level and class-level ambiguities. In this paper, we propose GRPSNet network which integrates graph reasoning to capture relationships between parts for part segmentation. These captured relationships help to enhance the recognition and localization of parts. We also propose to exploit the relationships of part boundaries to further enhance the accuracy of part segmentation. The experimental results demonstrate the effectiveness of the proposed method and show that it achieves state-of-the-art performance on the benchmark datasets. Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji |
ICME | 2 |
| 2024 | Fusion of Short-term and Long-term Attention for Video Mirror DetectionabstractTechniques for detecting mirrors from static images have witnessed rapid growth in recent years. However, these methods detect mirrors from single input images. Detecting mirrors from video requires further consideration of temporal consistency between frames. We observe that humans can recognize mirror candidates, from just one or two frames, based on their appearance (e.g. shape, color). However, to ensure that the candidate is indeed a mirror (not a picture or a window), we often need to observe more frames for a global view. This observation motivates us to detect mirrors by fusing appearance features extracted from a short-term attention module and context information extracted from a long-term attention module. To evaluate the performance, we build a challenging benchmark dataset of 19,255 frames from 281 videos. Experimental results demonstrate that our method achieves state-of-the-art performance on the benchmark dataset. Mingchen Xu, Jing Wu 0004, Yukun Lai, Ze Ji |
ICME | 2 |
| 2024 | Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware SamplingabstractDespite impressive results, deep generative models require massive datasets for training, and as dataset size increases, effective evaluation metrics like precision and recall (P&R) become computationally infeasible on commodity hardware. In this paper, we address this challenge by proposing efficient P&R (eP&R) metrics that give almost identical results as the original P&R but with much lower computational costs. Specifically, we identify two redundancies in the original P&R: i) redundancy in ratio computation and ii) redundancy in manifold inside/outside identification. We find both can be effectively removed via hubness-aware sampling, which extracts representative elements from synthetic/real image samples based on their hubness values, i.e., the number of times a sample becomes a k-nearest neighbor to others in the feature space. Thanks to the insensitivity of hubness-aware sampling to exact k-nearest neighbor (k-NN) results, we further improve the efficiency of our eP&R metrics by using approximate k-NN methods. Extensive experiments show that our eP&R matches the original P&R but is far more efficient in time and space. Our code is available at: https://github.com/Byronliang8/Hubness_Precision_Recall Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICML | 2 |
| 2024 | RSMPNet: Relationship Guided Semantic Map PredictionabstractIn semantic navigation, a top-down map with accurate and complete semantic information is vital to subsequent decision-making. However, due to occlusions and limitations of the robot’s field of view (FOV), there are often unobserved areas in the top-down maps. To address this problem, recent works have studied semantic map prediction to complete the top-down maps. In this work, we propose to improve map prediction by integrating relational information. We propose RSMPNet, a relationship-guided semantic map prediction network, which makes use of semantic and spatial relationships to predict unobserved areas from accumulated semantic maps. Specifically, we propose a Relationship Reasoning Layer that includes two modules, namely 1) the Semantic Relationship Graph Reasoning Module (SeGRM) to capture the semantic relationship and 2) the Spatial Relationship Graph Reasoning Module (SpGRM) to utilize the spatial relationship. We also design a semantic relationship enhanced loss to enhance our model to learn semantic relationship information. Experiments show the effectiveness of our proposed network which achieves state-of-the-art performance in semantic map prediction. Our code and dataset are publicly available at https://github.com/jws39/semantic-map-prediction Jing Wu 0004, Ze Ji, Yukun Lai |
WACV | 2 |
| 2024 | Sparse Convolutional Networks for Surface Reconstruction from Noisy Point CloudsabstractReconstructing accurate 3D surfaces from noisy point clouds is a fundamental problem in computer vision. Among different approaches, neural implicit methods that map 3D coordinates to occupancy values benefit from the learning capabilities of deep neural networks and the flexible topology of implicit representations, achieving promising reconstruction results. However, existing methods utilize standard (dense) 3D convolutional neural networks for feature extraction and occupancy prediction, which significantly restricts their capability to reconstruct details. In this paper, we propose a neural implicit method based on sparse convolutions, where features and network calculations only focus on grid points close to the surface to be reconstructed. This allows us to build significantly higher resolution 3D grids and reconstruct high-fidelity details. We further build a 3D residual UNet to extract features which are robust to noise, while ensuring details are retained. A 3D position along with features extracted at the position are fed into the occupancy probability predictor network to obtain occupancy. As features at nearby grid points to the query position may not exist due to the sparse nature, we propose a normalized weight interpolation approach to obtain smooth interpolation with sparse data. Experimental results demonstrate that our method achieves promising results, both qualitatively and quantitatively, outperforming existing methods. Jing Wu 0004, Ze Ji, Yukun Lai |
WACV | 2 |
| 2024 | Towards efficient image and video style transfer via distillation and learnable feature transformation
Jing Huo, Meihao Kong, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Benchmarking visual SLAM methods in mirror environmentsabstractVisual simultaneous localisation and mapping (vSLAM) finds applications for indoor and outdoor navigation that routinely subjects it to visual complexities, particularly mirror reflections. The effect of mirror presence (time visible and its average size in the frame) was hypothesised to impact localisation and mapping performance, with systems using direct techniques expected to perform worse. Thus, a dataset, MirrEnv, of image sequences recorded in mirror environments, was collected, and used to evaluate the performance of existing representative methods. RGBD ORB-SLAM3 and BundleFusion appear to show moderate degradation of absolute trajectory error with increasing mirror duration, whilst the remaining results did not show significantly degraded localisation performance. The mesh maps generated proved to be very inaccurate, with real and virtual reflections colliding in the reconstructions. A discussion is given of the likely sources of error and robustness in mirror environments, outlining future directions for validating and improving vSLAM performance in the presence of planar mirrors. The MirrEnv dataset is available at https://doi.org/10.17035/d.2023.0292477898 . Peter Herbert, Jing Wu 0004, Ze Ji, Yukun Lai |
Comput. Vis. Media | 2 |
| 2024 | GAM: General affordance-based manipulation for contact-rich object disentangling tasksabstractPicking up an entangled object is a difficult manipulation task due to its rich contact dynamics. Most existing solutions fail to produce grasp poses to enable reliable manipulation due to the dependence on simplified assumptions for the motion policies. Grasps generated by these methods tend to drop objects or cause undesired movements of non-grasped objects. To improve such object-disentangling tasks, we propose to extend the concept of reinforcement learning (RL)-based affordance to include arbitrary action consequences and implement a general affordance-based manipulation (GAM) framework. In the GAM, we train an RL agent that uses more fine-grained actions and outperforms previous methods with a smaller chance of dropping objects and making contact with non-grasped hooks. Then, a manipulation affordance prediction (MAP) model is trained to estimate the performances of the RL agent. Finally, the manipulation affordance-based grasp filter (MAGF) selects grasp poses that afford the desired manipulation performances, showing substantial improvements in five challenging hook disentangling tasks in simulation. The experiments show (1) the limitation of TAG generators, (2) the effectiveness of filtering TAGs with predicted manipulation performances based on the general affordance theory, and (3) the importance of avoiding contact with non-grasped objects in contact-rich manipulation. Xintong Yang, Jing Wu 0004, Yukun Lai, Ze Ji |
Neurocomputing | 2 |
| 2024 | Efficient Hierarchical Reinforcement Learning for Mapless Navigation With Predictive Neighbouring Space ScoringabstractSolving reinforcement learning (RL)-based mapless navigation tasks is challenging due to their sparse reward and long decision horizon nature. Hierarchical reinforcement learning (HRL) has the ability to leverage knowledge at different abstract levels and is thus preferred in complex mapless navigation tasks. However, it is computationally expensive and inefficient to learn navigation end-to-end from raw high-dimensional sensor data, such as Lidar or RGB cameras. The use of subgoals based on a compact intermediate representation is therefore preferred for dimension reduction. This work proposes an efficient HRL-based framework to achieve this with a novel scoring method, named Predictive Neighbouring Space Scoring (PNSS). The PNSS model estimates the explorable space for a given position of interest based on the current robot observation. The PNSS values for a few candidate positions around the robot provide a compact and informative state representation for subgoal selection. We study the effects of different candidate position layouts and demonstrate that our layout design facilitates higher performances in longer-range tasks. Moreover, a penalty term is introduced in the reward function for the high-level (HL) policy, so that the subgoal selection process takes the performance of the low-level (LL) policy into consideration. Comprehensive evaluations demonstrate that using the proposed PNSS module consistently improves performances over the use of Lidar only or Lidar and encoded RGB featuresNote to Practitioners—This paper seeks to improve robot mapless navigation capabilities where the robot is expected to navigate to a goal location without knowing the map of the environment. This ability is highly demanded in many applications that require autonomous operations in unstructured environments, including both indoor and outdoor scenarios, involving tasks such as service robots for domestic and public environments, logistics in industrial warehouses, urban search and rescue missions, and disaster relief efforts, where detailed and accurate maps are difficult to obtain in advance. In this work, we focus on reinforcement learning-based mapless navigation. It is known that such methods struggle in complex long-range tasks, e.g. stuck in a local region by multiple objects. Therefore, this paper proposes a novel mapless navigation method inspired by human navigation behaviours. We enable a robot to split a long-range navigation task into multiple segments, by selecting and navigating to short-term goals. These subgoals are selected each time from a number of candidate positions located around the robot. The process stops when the robot reaches the final target location. When selecting a short-term goal, we use a deep neural network to predict the openness around each candidate subgoal position, named the Predictive Neighbouring Space Scoring (PNSS), from raw images and Lidar scans. In addition, we study the effects of different arrangements of candidate subgoal locations and select the optimal one. Experiments conducted in photo-realistic simulation environments demonstrate the effectiveness of our method, showcasing superior performance over baselines. It is worth noting that our agent is only trained in domestic environments using the iGibson simulator. For applications in other environments, additional training in more representative settings specific to corresponding scenarios will be necessary. In the future, our intention is to validate our methods in complex real-world environments and narrow the simulation-to-reality gap for long-horizon navigation tasks. Yan Gao 0021, Jing Wu 0004, Xintong Yang, Ze Ji |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionabstractExisting model evaluation tools mainly focus on evaluating classification models, leaving a gap in evaluating more complex models, such as object detection. In this paper, we develop an open-source visual analysis tool, Uni-Evaluator, to support a unified model evaluation for classification, object detection, and instance segmentation in computer vision. The key idea behind our method is to formulate both discrete and continuous predictions in different tasks as unified probability distributions. Based on these distributions, we develop 1) a matrix-based visualization to provide an overview of model performance; 2) a table visualization to identify the problematic data subsets where the model performs poorly; 3) a grid visualization to display the samples of interest. These visualizations work together to facilitate the model evaluation from a global overview to individual samples. Two case studies demonstrate the effectiveness of Uni-Evaluator in evaluating model performance and making informed improvements. Changjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu 0004, Weikai Yang, Jing Wu 0004, Hang Su 0006, Hanspeter Pfister, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Interactive Reweighting for Mitigating Label Quality IssuesabstractLabel quality issues, such as noisy labels and imbalanced class distributions, have negative effects on model performance. Automatic reweighting methods identify problematic samples with label quality issues by recognizing their negative effects on validation samples and assigning lower weights to them. However, these methods fail to achieve satisfactory performance when the validation samples are of low quality. To tackle this, we develop Reweighter, a visual analysis tool for sample reweighting. The reweighting relationships between validation samples and training samples are modeled as a bipartite graph. Based on this graph, a validation sample improvement method is developed to improve the quality of validation samples. Since the automatic improvement may not always be perfect, a co-cluster-based bipartite graph visualization is developed to illustrate the reweighting relationships and support the interactive adjustments to validation samples and reweighting results. The adjustments are converted into the constraints of the validation sample improvement method to further improve validation samples. We demonstrate the effectiveness of Reweighter in improving reweighting results through quantitative evaluation and two case studies. Weikai Yang, Yukai Guo, Jing Wu 0004, Lan-Zhe Guo, Yufeng Li 0008, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Feature Proliferation - the "Cancer" in StyleGAN and its TreatmentsabstractDespite the success of StyleGAN in image synthesis, the images it synthesizes are not always perfect and the well-known truncation trick has become a standard post-processing technique for StyleGAN to synthesize high quality images. Although effective, it has long been noted that the truncation trick tends to reduce the diversity of synthesized images and unnecessarily sacrifices many distinct image features. To address this issue, in this paper, we first delve into the StyleGAN image synthesis mechanism and discover an important phenomenon, namely Feature Proliferation, which demonstrates how specific features reproduce with forward propagation. Then, we show how the occurrence of Feature Proliferation results in StyleGAN image artifacts. As an analogy, we refer to it as the "cancer" in StyleGAN from its proliferating and malignant nature. Finally, we propose a novel feature rescaling method that identifies and modulates risky features to mitigate feature proliferation. Thanks to our discovery of Feature Proliferation, the proposed feature rescaling method is less destructive and retains more useful image features than the truncation trick, as it is more fine-grained and works in a lower-level feature space rather than a high-level latent space. Experimental results justify the validity of our claims and the effectiveness of the proposed feature rescaling method. Our code is available at https://github.com/songc42/Feature-proliferation. Shuang Song 0012, Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICCV | 3 |
| 2023 | AFPSNet: Multi-Class Part Parsing based on Scaled Attention and Feature FusionabstractMulti-class part parsing is a dense prediction task that seeks to simultaneously detect multiple objects and the semantic parts within these objects in the scene. This problem is important in providing detailed object understanding, but is challenging due to the existence of both class-level and part-level ambiguities. In this paper, we propose to integrate an attention refinement module and a feature fusion module to tackle the part-level ambiguity. The attention refinement module aims to enhance the feature representations by focusing on important features. The feature fusion module aims to improve the fusion operation for different scales of features. We also propose an object-to-part training strategy to tackle the class-level ambiguity, which improves the localization of parts by exploiting prior knowledge of objects. The experimental results demonstrated the effectiveness of the proposed modules and the training strategy, and showed that our proposed method achieved state-of-the-art performance on the benchmark datasets. Njuod Alsudays, Jing Wu 0004, Yukun Lai, Ze Ji |
WACV | 2 |
| 2023 | Interactive Visual Cluster Analysis by Contrastive Dimensionality ReductionabstractWe propose a contrastive dimensionality reduction approach (CDR) for interactive visual cluster analysis. Although dimensionality reduction of high-dimensional data is widely used in visual cluster analysis in conjunction with scatterplots, there are several limitations on effective visual cluster analysis. First, it is non-trivial for an embedding to present clear visual cluster separation when keeping neighborhood structures. Second, as cluster analysis is a subjective task, user steering is required. However, it is also non-trivial to enable interactions in dimensionality reduction. To tackle these problems, we introduce contrastive learning into dimensionality reduction for high-quality embedding. We then redefine the gradient of the loss function to the negative pairs to enhance the visual cluster separation of embedding results. Based on the contrastive learning scheme, we employ link-based interactions to steer embeddings. After that, we implement a prototype visual interface that integrates the proposed algorithms and a set of visualizations. Quantitative experiments demonstrate that CDR outperforms existing techniques in terms of preserving correct neighborhood structures and improving visual cluster separation. The ablation experiment demonstrates the effectiveness of gradient redefinition. The user study verifies that CDR outperforms t-SNE and UMAP in the task of cluster identification. We also showcase two use cases on real-world datasets to present the effectiveness of link-based interactions. Jiazhi Xia, Linquan Huang, Weixing Lin, Xin Zhao 0025, Jing Wu 0004, Yang Chen 0048, Ying Zhao 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Exploring and Exploiting Hubness Priors for High-Quality GAN Latent SamplingabstractDespite the extensive studies on Generative Adversarial Networks (GANs), how to reliably sample high-quality images from their latent spaces remains an under-explored topic. In this paper, we propose a novel GAN latent sampling method by exploring and exploiting the hubness priors of GAN latent distributions. Our key insight is that the high dimensionality of the GAN latent space will inevitably lead to the emergence of hub latents that usually have much larger sampling densities than other latents in the latent space. As a result, these hub latents are better trained and thus contribute more to the synthesis of high-quality images. Unlike the a posterior "cherry-picking", our method is highly efficient as it is an a priori method that identifies high-quality latents before the synthesis of images. Furthermore, we show that the well-known but purely empirical truncation trick is a naive approximation to the central clustering effect of hub latents, which not only uncovers the rationale of the truncation trick, but also indicates the superiority and fundamentality of our method. Extensive experimental results demonstrate the effectiveness of the proposed method. Our code is available at: https://github.com/Byronliang8/HubnessGANSampling. Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICML | 2 |
| 2022 | Hierarchical Reinforcement Learning With Universal Policies for Multistep Robotic ManipulationabstractMultistep tasks, such as block stacking or parts (dis)assembly, are complex for autonomous robotic manipulation. A robotic system for such tasks would need to hierarchically combine motion control at a lower level and symbolic planning at a higher level. Recently, reinforcement learning (RL)-based methods have been shown to handle robotic motion control with better flexibility and generalizability. However, these methods have limited capability to handle such complex tasks involving planning and control with many intermediate steps over a long time horizon. First, current RL systems cannot achieve varied outcomes by planning over intermediate steps (e.g., stacking blocks in different orders). Second, the exploration efficiency of learning multistep tasks is low, especially when rewards are sparse. To address these limitations, we develop a unified hierarchical reinforcement learning framework, named Universal Option Framework (UOF), to enable the agent to learn varied outcomes in multistep tasks. To improve learning efficiency, we train both symbolic planning and kinematic control policies in parallel, aided by two proposed techniques: 1) an auto-adjusting exploration strategy (AAES) at the low level to stabilize the parallel training, and 2) abstract demonstrations at the high level to accelerate convergence. To evaluate its performance, we performed experiments on various multistep block-stacking tasks with blocks of different shapes and combinations and with different degrees of freedom for robot control. The results demonstrate that our method can accomplish multistep manipulation tasks more efficiently and stably, and with significantly less memory consumption. Xintong Yang, Ze Ji, Jing Wu 0004, Yukun Lai, Changyun Wei, Rossitza Setchi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Towards Better Caption Supervision for Object DetectionabstractAs training high-performance object detectors requires expensive bounding box annotations, recent methods resort to free-available image captions. However, detectors trained on caption supervision perform poorly because captions are usually noisy and cannot provide precise location information. To tackle this issue, we present a visual analysis method, which tightly integrates caption supervision with object detection to mutually enhance each other. In particular, object labels are first extracted from captions, which are utilized to train the detectors. Then, the objects detected from images are fed into caption supervision for further improvement. To effectively loop users into the object detection process, a node-link-based set visualization supported by a multi-type relational co-clustering algorithm is developed to explain the relationships between the extracted labels and the images with detected objects. The co-clustering algorithm clusters labels and images simultaneously by utilizing both their representations and their relationships. Quantitative evaluations and a case study are conducted to demonstrate the efficiency and effectiveness of the developed method in improving the performance of object detectors. Changjian Chen, Jing Wu 0004, Shouxing Xiang, Song-Hai Zhang, Qifeng Tang, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | A Unified Understanding of Deep NLP Models for Text ClassificationabstractThe rapid development of deep natural language processing (NLP) models for text classification has led to an urgent need for a unified understanding of these models proposed individually. Existing methods cannot meet the need for understanding different models in one framework due to the lack of a unified measure for explaining both low-level (e.g., words) and high-level (e.g., phrases) features. We have developed a visual analysis tool, DeepNLPVis, to enable a unified understanding of NLP models for text classification. The key idea is a mutual information-based measure, which provides quantitative explanations on how each layer of a model maintains the information of input words in a sample. We model the intra- and inter-word information at each layer measuring the importance of a word to the final prediction as well as the relationships between words, such as the formation of phrases. A multi-level visualization, which consists of a corpus-level, a sample-level, and a word-level visualization, supports the analysis from the overall training set to individual samples. Two case studies on classification tasks and comparison between models demonstrate that DeepNLPVis can help users effectively identify potential problems caused by samples and model architectures and then make informed improvements. Zhen Li 0044, Xiting Wang, Weikai Yang, Jing Wu 0004, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Manifold Alignment for Semantically Aligned Style TransferabstractMost existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of local style patterns, where algorithms are designed to swap similar local features of content and style images. However, the limitation of these existing methods is that they neglect the semantic structure of the content image which may lead to corrupted content structure in the output. In this paper, we make a new assumption that image features from the same semantic region form a manifold and an image with multiple semantic regions follows a multi-manifold distribution. Based on this assumption, the style transfer problem is formulated as aligning two multi-manifold distributions and a Manifold Alignment based Style Transfer (MAST) framework is proposed. The proposed frame-work allows semantically similar regions between the output and the style image share similar style patterns. Moreover, the proposed manifold alignment method is flexible to allow user editing or using semantic segmentation maps as guidance for style transfer. To allow the method to be applicable to photorealistic style transfer, we propose a new adaptive weight skip connection network structure to preserve the content details. Extensive experiments verify the effectiveness of the proposed framework for both artistic and photorealistic style transfer. Code is available at https://github.com/NJUHuoJing/MAST. Jing Huo, Shiyin Jin, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yinghuan Shi, Yang Gao 0001 |
ICCV | 4 |
| 2021 | MLVSNet: Multi-level Voting Siamese Network for 3D Visual TrackingabstractBenefiting from the excellent performance of Siamese-based trackers, huge progress on 2D visual tracking has been achieved. However, 3D visual tracking is still under-explored. Inspired by the idea of Hough voting in 3D object detection, in this paper, we propose a Multi-level Voting Siamese Network (MLVSNet) for 3D visual tracking from outdoor point cloud sequences. To deal with sparsity in outdoor 3D point clouds, we propose to perform Hough voting on multi-level features to get more vote centers and retain more useful information, instead of voting only on the fi-nal level feature as in previous methods. We also design an efficient and lightweight Target-Guided Attention (TGA) module to transfer the target information and highlight the target points in the search area. Moreover, we propose a Vote-cluster Feature Enhancement (VFE) module to exploit the relationships between different vote clusters. Extensive experiments on the 3D tracking benchmark of KITTI dataset demonstrate that our MLVSNet outperforms state-of-the-art methods with significant margins. Code will be available at https://github.com/CodeWZT/MLVSNet. Zhoutao Wang, Qian Xie 0001, Yukun Lai, Jing Wu 0004, Kun Long, Jun Wang 0039 |
ICCV | 4 |
| 2021 | VENet: Voting Enhancement Network for 3D Object DetectionabstractHough voting, as has been demonstrated in VoteNet, is effective for 3D object detection, where voting is a key step. In this paper, we propose a novel VoteNet-based 3D detector with vote enhancement to improve the detection accuracy in cluttered indoor scenes. It addresses the limitations of current voting schemes, i.e., votes from neighboring objects and background have significant negative impacts. Before voting, we replace the classic MLP with the proposed Attentive MLP (AMLP) in the backbone network to get better feature description of seed points. During voting, we design a new vote attraction loss (VALoss) to enforce vote centers to locate closely and compactly to the corresponding object centers. After voting, we then devise a vote weighting module to integrate the foreground/background prediction into the vote aggregation process to enhance the capability of the original VoteNet to handle noise from background voting. The three proposed strategies all contribute to more effective voting and improved performance, resulting in a novel 3D object detector, termed VENet. Experiments show that our method outperforms state-of-the-art methods on benchmark datasets. Ablation studies demonstrate the effectiveness of the proposed components. Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Dening Lu, Mingqiang Wei, Jun Wang 0039 |
ICCV | 3 |
| 2021 | Learning 3D face reconstruction from a single sketch
Jing Wu 0004, Jing Huo, Yukun Lai, Yang Gao 0001 |
Graph. Model. | 2 |
| 2021 | Vote-Based 3D Object Detection with Context Modeling and SOB-3DNMS
Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Kai Xu 0004, Jun Wang 0039 |
Int. J. Comput. Vis. | 3 |
| 2021 | MW-GAN: Multi-Warping GAN for Caricature Generation With Multi-Style Geometric ExaggerationabstractGiven an input face photo, the goal of caricature generation is to produce stylized, exaggerated caricatures that share the same identity as the photo. It requires simultaneous style transfer and shape exaggeration with rich diversity, and meanwhile preserving the identity of the input. To address this challenging problem, we propose a novel framework called Multi-Warping GAN (MW-GAN), including a style network and a geometric network that are designed to conduct style transfer and geometric exaggeration respectively. We bridge the gap between the style/landmark space and their corresponding latent code spaces by a dual way design, so as to generate caricatures with arbitrary styles and geometric exaggeration, which can be specified either through random sampling of latent code or from a given caricature sample. Besides, we apply identity preserving loss to both image space and landmark space, leading to a great improvement in quality of generated caricatures. Experiments show that caricatures generated by MW-GAN have better quality than existing methods. Haodi Hou, Jing Huo, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Analyzing the Noise Robustness of Deep Neural NetworksabstractAdversarial examples, generated by adding small but intentionally imperceptible perturbations to normal examples, can mislead deep neural networks (DNNs) to make incorrect predictions. Although much work has been done on both adversarial attack and defense, a fine-grained understanding of adversarial examples is still lacking. To address this issue, we present a visual analysis method to explain why adversarial examples are misclassified. The key is to compare and analyze the datapaths of both the adversarial and normal examples. A datapath is a group of critical neurons along with their connections. We formulate the datapath extraction as a subset selection problem and solve it by constructing and training a neural network. A multi-level visualization consisting of a network-level visualization of data flows, a layer-level visualization of feature maps, and a neuron-level visualization of learned features, has been designed to help investigate how datapaths of adversarial and normal examples diverge and merge in the prediction process. A quantitative evaluation and a case study were conducted to demonstrate the promise of our method to explain the misclassification of adversarial examples. Kelei Cao, Mengchen Liu, Hang Su 0006, Jing Wu 0004, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Interactive Graph Construction for Graph-Based Semi-Supervised LearningabstractSemi-supervised learning (SSL) provides a way to improve the performance of prediction models (e.g., classifier) via the usage of unlabeled samples. An effective and widely used method is to construct a graph that describes the relationship between labeled and unlabeled samples. Practical experience indicates that graph quality significantly affects the model performance. In this paper, we present a visual analysis method that interactively constructs a high-quality graph for better model performance. In particular, we propose an interactive graph construction method based on the large margin principle. We have developed a river visualization and a hybrid visualization that combines a scatterplot, a node-link diagram, and a bar chart to convey the label propagation of graph-based SSL. Based on the understanding of the propagation, a user can select regions of interest to inspect and modify the graph. We conducted two case studies to showcase how our method facilitates the exploitation of labeled and unlabeled samples for improving model performance. Changjian Chen, Jing Wu 0004, Xiting Wang, Lan-Zhe Guo, Yufeng Li 0008, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | MLCVNet: Multi-Level Context VoteNet for 3D Object DetectionabstractIn this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual information between these objects. Comparatively, we propose Multi-Level Context VoteNet (MLCVNet) to recognize 3D objects correlatively, building on the state-of-the-art VoteNet. We introduce three context modules into the voting and classifying stages of VoteNet to encode contextual information at different levels. Specifically, a Patch-to-Patch Context (PPC) module is employed to capture contextual information between the point patches, before voting for their corresponding object centroid points. Subsequently, an Object-to-Object Context (OOC) module is incorporated before the proposal and classification stage, to capture the contextual information between object candidates. Finally, a Global Scene Context (GSC) module is designed to learn the global scene context. We demonstrate these by capturing contextual information at patch, object and scene levels. Our method is an effective way to promote detection accuracy, achieving new state-of-the-art detection performance on challenging 3D object detection datasets, i.e., SUN RGBD and ScanNet. We also release our code at https://github.com/NUAAXQ/MLCVNet. Qian Xie 0001, Yukun Lai, Jing Wu 0004, Zhoutao Wang, Kai Xu 0004, Jun Wang 0039 |
CVPR | 3 |
| 2020 | Automated Robot-based Large-Scale 3D Surface ImagingabstractThis work develops a robot-based automated 3D imaging system for large-scale surface measurement at high resolution. The system has the advantages of allowing 1) high-resolution 3D surface imaging based on photometric stereo, and 2) automatic stitching of multiple images collected by a robot for large-scale surface measurement. We developed a dome-shaped image acquisition system with 16 individually controlled lights, mounted on a robot (Kuka iiwa lbr). A photometric stereo with a lighting selection mechanism is used for the reconstruction of local surface regions. To allow image stitching for large-scale surface measurement, one challenge arises from the robot arm’s limited encoder precision and accuracy, which is about ±150 µm for its repeatability and even lower for its nominal accuracy. This is unsuitable for the applications of surface metrology or inspection. To compensate the errors introduced by the chained robot arm’s encoders, for image stitching, we experimented with two feature descriptors extracted from the normal and the curvature space respectively, and performed comparative studies with standard feature descriptors from the standard grey-scale intensity space. The normal-based feature descriptor demonstrated advantages of illumination invariance while the curvature-based feature descriptor demonstrated clear advantages of rotation invariance, and feasibility of aligning multiple images with high accuracy. Jingjing Wen, Jing Wu 0004, Ze Ji |
KES | 3 |
| 2019 | Discriminative Features Matter: Multi-layer Bilinear Pooling for Camera Localization
Xiang Wang 0014, Chen Wang 0026, Xiao Bai 0001, Jing Wu 0004, Edwin R. Hancock |
BMVC | 5 |
| 2019 | Computer-assisted Relief Modelling: A Comprehensive SurveyabstractAbstract As an art form between drawing and sculpture, relief has been widely used in a variety of media for signs, narratives, decorations and other purposes. Traditional relief creation relies on both professional skills and artistic expertise, which is extremely time‐consuming. Recently, automatic or semi‐automatic relief modelling from a 3D object or a 2D image has been a subject of interest in computer graphics. Various methods have been proposed to generate reliefs with few user interactions or minor human efforts, while preserving or enhancing the appearance of the input. This survey provides a comprehensive review of the advances in computer‐assisted relief modelling during the past decade. First, we provide an overview of relief types and their art characteristics. Then, we introduce the key techniques of object‐space methods and image‐space methods respectively. Advantages and limitations of each category are discussed in details. We conclude the report by discussing directions for possible future research. Yu-Wei Zhang 0014, Jing Wu 0004, Zhongping Ji, Mingqiang Wei, Caiming Zhang 0001 |
Comput. Graph. Forum | 2 |
| 2019 | InSocialNet: Interactive visual analytics for role - event videosabstractRole–event videos are rich in information but challenging to be understood at the story level. The social roles and behavior patterns of characters largely depend on the interactions among characters and the background events. Understanding them requires analysis of the video contents for a long duration, which is beyond the ability of current algorithms designed for analyzing short-time dynamics. In this paper, we propose InSocialNet, an interactive video analytics tool for analyzing the contents of role–event videos. It automatically and dynamically constructs social networks from role–event videos making use of face and expression recognition, and provides a visual interface for interactive analysis of video contents. Together with social network analysis at the back end, InSocialNet supports users to investigate characters, their relationships, social roles, factions, and events in the input video. We conduct case studies to demonstrate the effectiveness of InSocialNet in assisting the harvest of rich information from role–event videos. We believe the current prototype implementation can be extended to applications beyond movie analysis, e.g., social psychology experiments to help understand crowd social behaviors. Yaohua Pan, Zhibin Niu, Jing Wu 0004, Jiawan Zhang |
Comput. Vis. Media | 3 |
| 2018 | Visual Diagnosis of Tree Boosting MethodsabstractTree boosting, which combines weak learners (typically decision trees) to generate a strong learner, is a highly effective and widely used machine learning method. However, the development of a high performance tree boosting model is a time-consuming process that requires numerous trial-and-error experiments. To tackle this issue, we have developed a visual diagnosis tool, BOOSTVis, to help experts quickly analyze and diagnose the training process of tree boosting. In particular, we have designed a temporal confusion matrix visualization, and combined it with a t-SNE projection and a tree visualization. These visualization components work together to provide a comprehensive overview of a tree boosting model, and enable an effective diagnosis of an unsatisfactory training process. Two case studies that were conducted on the Otto Group Product Classification Challenge dataset demonstrate that BOOSTVis can provide informative feedback and guidance to improve understanding and diagnosis of tree boosting algorithms. Shixia Liu, Jiannan Xiao, Xiting Wang, Jing Wu 0004, Jun Zhu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Improving Shape from Shading with Interactive Tabu Search
Jing Wu 0004, Paul L. Rosin, Xianfang Sun, Ralph R. Martin |
J. Comput. Sci. Technol. | 1 |
| 2014 | Use of non-photorealistic rendering and photometric stereo in making bas-reliefs from photographs
Jing Wu 0004, Ralph R. Martin, Paul L. Rosin, Xianfang Sun, Yukun Lai, Christian Wallraven |
Graph. Model. | 1 |
| 2013 | Making bas-reliefs from photographs of human faces
Jing Wu 0004, Ralph R. Martin, Paul L. Rosin, Xianfang Sun, Frank C. Langbein, Yukun Lai, David Marshall 0001 |
Comput. Aided Des. | 1 |
| 2011 | Supervised relevance maps for increasing the distinctiveness of facial images
Michal Kawulok, Jing Wu 0004, Edwin R. Hancock |
Pattern Recognit. | 2 |
| 2011 | Gender discriminating models from facial surface normals
Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 2010 | Facial gender classification using shape-from-shading
Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
Image Vis. Comput. | 1 |
| 2009 | Semi-supervised Feature Selection for Gender Classification
Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
ACCV (2) | 1 |
| 2009 | Extracting gender discriminating features from facial needle-mapsabstractIn this paper, we show how to extract gender discriminating features from 2.5D facial needle-maps. The standard eigenspace analysis method for non-Euclidean data is principal geodesic analysis (PGA). Based on PGA, we propose a novel supervised weighted PGA method which incorporates local weights into standard PGA to improve gender discriminating capability of the extracted features. The weight map is iteratively optimized from the labeled data, which is different from other gender relevant weights used in the literature. Experimental results illustrate the effectiveness of this method and its successful application to gender classification. Jing Wu 0004, William A. P. Smith, Edwin R. Hancock, Michal Kawulok |
ICIP | 1 |
| 2008 | Gender classification based on facial surface normalsabstractIn this paper, we perform gender classification based on 2.5D facial surface normals (facial needle-maps), and present two novel principal geodesic analysis (PGA) methods, weighted PGA and supervised PGA, to parameterize the facial needle-maps, and compare their performances with PGA for gender classification. Experimental results demonstrate the feasibility of gender classification based on facial needle-maps, and show that incorporating weights or pairwise relationships of labeled data into PGA improves the gender discriminating powers in the leading eigenvectors and the gender classification accuracy. Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
ICPR | 1 |
| 2007 | Gender Classification using Shape from ShadingabstractThe aim in this paper is to show how to use the 2.5D facial surface normals \n(needle-maps) recovered using shape from shading (SFS) to improve \nthe performance of gender classification. We incorporate principal geodesic \nanalysis (PGA) into SFS to guarantee the recovered needle-maps is a possible \nexample defined by a statistical model. Because the recovered facial needlemaps \nsatisfy data-closeness constraint, they not only give the facial shape \ninformation, but also combine the image intensity implicitly. Experiments \nshow that this combination gives better gender classification performance \nthan using facial shape or texture information alone. Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
BMVC | 1 |
| 2007 | Weighted Principal Geodesic Analysis for Facial Gender Classification
Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
CIARP | 1 |
| 2006 | Gender Classification Using Principal Geodesic Analysis and Gaussian Mixture Models
Jing Wu 0004, William A. P. Smith, Edwin R. Hancock |
CIARP | 1 |