Lei Wang 0018

dblp:w/LeiWang18 · DBLP profile ↗
← Back
57ranked-venue papers
16as first author
32since 2021 · last 2026
0000-0001-5990-896XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multifactorial evolutionary algorithm enhanced by symmetry transformation and ridge regression
Guoxing Luo, Yan Wang 0155, Jihua Fan, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001, Xiaoliang Ma 0001
Expert Syst. Appl.5
2026 Phy-HHMR: Physics-Aware Holistic Human Mesh Reconstruction
Binhong Ye, Lei Wang 0018, Baoyu Liu, Xiaoliang Ma 0001, Jun Cheng 0002
IEEE Trans. Hum. Mach. Syst.3
2026 Speech2Blend: A Hybrid Network for Speech-Driven 3-D Facial Animation by Learning Blendshape
abstract
Recent advances in speech-driven facial animation have attracted significant interest across computer graphics, human–computer interaction systems, and immersive virtual reality applications. However, existing methods remain constrained by dependencies on specific reference videos or proprietary face mesh structures, limiting their applicability across diverse production pipelines and reducing compatibility with industry-standard animation workflows. To overcome these fundamental limitations in generalization and deployment flexibility, we propose Speech2Blend—an end-to-end hybrid convolutional-recurrent network that directly learns nonlinear speech-to-blendshape parameter mappings. This novel approach enables markerless speech-driven facial animation generation without restrictive inputs like video references or specialized facial rigs. Trained on the largest available digital human dataset (BEAT) and rigorously evaluated using three benchmark datasets with photorealistic visualization tools, Speech2Blend achieves state-of-the-art performance. It delivers superior audio-visual synchronization through learned temporal dynamics and reduces lip vertex error by 30% compared to existing baseline methods. These advances significantly lower production costs for virtual human speech animation while enabling cross-platform compatibility with common game engines and animation software.
Lei Wang 0018, Gongbin Chen, Feng Liu 0013, Jiaji Wu, Jun Cheng 0002
IEEE Trans. Hum. Mach. Syst.1
2025 Automatic Numbering and Pathological Recognition of Pediatric Teeth Using CNN and Attention Mechanisms
abstract
Preliminary progress has been made in using deep learning networks for tooth segmentation and numbering, as well as pathological identification in dental panoramic images. However, The publicly available datasets specifically for children’s teeth are very scarce. To address this issue, this paper proposes a fully public database of 849 children’s panoramic radiographs. We also introduce two models based on CNN and attention mechanisms: DCD-Net (Dental Classification and Detection Net) and DPD-Net (Dental Pathology Detection Net). The former, when combined with our category refinement model, can automatically segment and number children’s teeth, achieving an [email protected] of 96.4% while significantly reducing the required training images. The latter detects dental pathologies with an [email protected] of 80%.
Hongzhou Zhu, Yuhao Qiu, Shengji Zhu, Lei Wang 0018
ICASSP6
2025 Bidirectional Mixed Augmentation Sample Generation under Dual Perturbations in Semi-supervised Medical Image Segmentation
Yibo Feng, Feng Liu 0013, Zhiyi Shan, Lei Wang 0018, Jun Cheng 0002
PRCV (14)4
2025 Semi-supervised Cephalometric Landmark Detection Using Landmark Contrastive Learning
Zixun Zhan, Xiaoliang Ma 0001, Zhiyi Shan, Shengji Zhu, Lei Wang 0018
PRCV (13)5
2025 LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
Dezhi Zheng, Lei Wang 0018, Liping xiang, Kaijun Deng, Xiaowen Fu, LinLin Shen, Jinbao Wang 0001
PRCV (10)5
2025 A survey of graph neural networks and their industrial applications
Lei Wang 0018, Xiaoliang Ma 0001, Jun Cheng 0002, MengChu Zhou
Neurocomputing2
2025 A Twist Representation and Shape Refinement Method for Human Mesh Recovery
abstract
3D human mesh recovery from single RGB images or monocular videos is a challenging task. The twist representation utilized in existing inverse kinematics-based methods fails to accurately describe the twisting posture when the estimated bone direction is imprecise. Additionally, supervising SMPL shape parameters has the issue of shape estimation overfitting due to limited training data. This often results in compromised bone lengths that subsequently impair the precision of joint positions. To address these issues, we propose a framework that breaks down both human pose and shape into finer components, effectively managing and minimizing errors within each component. The proposed framework integrates two key advancements: the advanced Ortho-Twist and Swing Representation (OTSR) and the Skeleton-Focused Shape Refinement (SFSR). OTSR offers a more sophisticated representation for limb rotations compared to the traditional twist angle and swing representation to enhance the accuracy of twisting posture estimation. SFSR refines the estimated SMPL shape parameters by fitting bone lengths using the estimated joint positions, thereby significantly mitigating shape overfitting and enhancing joint position accuracy in the recovered mesh. We conduct experiments on the Human3.6 M and 3DPW datasets. The results demonstrate the superiority of the proposed framework in both single-image and video scenarios. Additionally, the ablation studies confirm the effectiveness of our proposed modules, and further generalizability experiments demonstrate that our two key advancements can serve as plug-and-play modules to enhance existing methods.
Xiaoyang Hao, Jing Sun 0010, Lei Wang 0018, Jianping Fan 0002
IEEE Trans. Multim.4
2024 AUEditNet: Dual-Branch Facial Action Unit Intensity Manipulation with Implicit Disentanglement
abstract
Facial action unit (AU) intensity plays a pivotal role in quantifying fine-grained expression behaviors, which is an effective condition for facial expression manipulation. How-ever, publicly available datasets containing intensity annotations for multiple AUs remain severely limited, often featuring a restricted number of subjects. This limitation places challenges to the AU intensity manipulation in images due to disentanglement issues, leading researchers to resort to other large datasets with pretrained AU intensity estimators for pseudo labels. In addressing this constraint and fully leveraging manual annotations of AU intensities for precise manipulation, we introduce AUEditNet. Our proposed model achieves impressive intensity manipulation across 12 AUs, trained effectively with only 18 subjects. Utilizing a dual-branch architecture, our approach achieves comprehensive disentanglement of facial attributes and identity without necessitating additional loss functions or implementing with large batch sizes. This approach offers a potential solution to achieve desired facial attribute editing despite the dataset's limited subject count. Our experiments demonstrate AUEdit-Net's superior accuracy in editing AU intensities, affirming its capability in disentangling facial attributes and identity within a limited subject pool. AUEditNet allows conditioning by either intensity values or target images, eliminating the need for constructing AU combinations for specific facial expression synthesis. Moreover, AU intensity estimation, as a downstream task, validates the consistency between real and edited images, confirming the effectiveness of our proposed AU intensity manipulation method.
Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Peng Liu 0039, Ning Bi, Truong Q. Nguyen
CVPR3
2024 Realistic and Visually-Pleasing 3D Generation of Indoor Scenes from a Single Image
Lei Wang 0018, Gongbin Chen, Yuhao Qiu, Jiaji Wu, Jun Cheng 0002
PRCV (6)2
2024 Dental Diagnosis from X-Ray Panoramic Radiography Images: A Dataset and A Hybrid Framework
Gege Shan, Xiaoliang Ma 0001, Xiaojie Bai, Hongzhou Zhu, Shengji Zhu, Lei Wang 0018
PRCV (14)7
2023 Reject Decoding via Language-Vision Models for Text-to-Image Synthesis
abstract
Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Lei Wang 0018, Jun Cheng 0002
AAAI5
2023 ReDirTrans: Latent-to-Latent Translation for Gaze and Head Redirection
abstract
Learning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which successfully redirected gaze directions pre-cisely given the assigned conditions. However, these methods focused on changing gaze directions of the images that only include eyes or restricted ranges of faces with low res-olution (less than$128\times 128$) to largely reduce interference from other attributes such as hairs, which limits application scenarios. To cope with this limitation, we proposed a portable network, called ReDirTrans, achieving latent-to-latent translation for redirecting gaze directions and head orientations in an interpretable manner. ReDirTrans projects input latent vectors into aimed-attribute embed-dings only and redirects these embeddings with assigned pitch and yaw values. Then both the initial and edited embeddings are projected back (deprojected) to the initial latent space as residuals to modify the input latent vec-tors by subtraction and addition, representing old status re-moval and new status addition. The projection of aimed at-tributes only and subtraction-addition operations for status replacement essentially mitigate impacts on other attributes and the distribution of latent vectors. Thus, by combining ReDirTrans with a pretrained fixed e4e-StyleGAN pair, we created ReDirTrans-GAN, which enables accurately redi-recting gaze in full-face images with$1024\times 1024$resolution while preserving other attributes such as identity, expres-sion, and hairstyle. Furthermore, we presented improvements for the downstream learning-based gaze estimation task, using redirected samples as dataset augmentation.
Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Ning Bi, Truong Q. Nguyen
CVPR3
2023 StyleAU: StyleGAN based Facial Action Unit Manipulation for Expression Editing
abstract
Facial expression editing has a wide range of applications, such as emotion detection, human-computer interaction, and social entertainment. However, existing expression editing methods either fail to allow for fine-grained editing, resulting in unnatural and unrealistic facial expressions, or generate artifacts and blurs, leading to poor image quality. In this paper, we propose a novel framework called StyleAU, which is based on StyleGAN and facial action units, to address these problems. Our framework leverages the pre-trained StyleGAN prior knowledge to enable action unit editing of the face in the StyleGAN latent space, allowing precise expression editing. In addition, we use an encoder to extract multi-scale content features to achieve high-fidelity image reconstruction. Our approach qualitatively and quantitatively outperforms competing methods for action unit manipulation and expression editing.
Yanliang Guo, Xianxu Hou, Feng Liu 0013, LinLin Shen, Lei Wang 0018, Zhen Wang 0009, Peng Liu 0039
IJCB5
2023 CCF-Net: A Cascade Center-Based Framework Towards Efficient Human Parts Detection
Kai Ye 0004, Haoqin Ji, Lei Wang 0018, Peng Liu 0039, LinLin Shen
MMM (2)4
2023 Blendshape-Based Migratable Speech-Driven 3D Facial Animation with Overlapping Chunking-Transformer
Jixi Chen, Xiaoliang Ma 0001, Lei Wang 0018, Jun Cheng 0002
PRCV (2)3
2023 Autoencoder and Masked Image Encoding-Based Attentional Pose Network
Long-Hua Hu, Xiaoliang Ma 0001, Lei Wang 0018, Jun Cheng 0002
PRCV (2)4
2023 Enhancing evolutionary multitasking optimization by leveraging inter-task knowledge transfers and improved evolutionary operators
Xiaoliang Ma 0001, Yan Wang 0155, Lei Wang 0018, Yutao Qi
Knowl. Based Syst.6
2023 A Progressive Quadric Graph Convolutional Network for 3D Human Mesh Recovery
abstract
Human mesh recovery from one single image has achieved rapid progress recently, but many methods suffer from the image appearance overfitting since the training data are collected along with accurate 3D annotations in controlled settings of monotonous backgrounds or simple clothes. Some methods regress human mesh vertices from poses to tackle the above problem. However the mesh topologies have not been well exploited, and artifacts are often generated. In this paper, we aim to find an efficient low-cost solution to human mesh reconstruction. To this end, we propose a Progressive Quadric Graph Convolutional Network (PQ-GCN), and design a simple and fast method for 3D human mesh recovery from a single image in the wild. Specifically, we apply quadric-based surface simplification to human meshes and design a progressive graph convolution network, accompanied by mesh feature up-sampling, to deal with the mesh topologies. We carry out a series of studies to validate our method. The results prove that our method achieves superior performance on a challenging in-the-wild dataset, while using 66% fewer parameters than the existing method, Pose2Mesh. Artifacts have also been eliminated and better visual quality has been obtained without any further post-processing and model fitting. Besides, the recovery can be stopped at an earlier stage by adding a decoder head. Consequently, the computational complexity can be reduced greatly.
Lei Wang 0018, Xun-Yu Liu, Xiaoliang Ma 0001, Jiaji Wu, Jun Cheng 0002, MengChu Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2023 Regionwise Generative Adversarial Image Inpainting for Large Missing Areas
abstract
Recently, deep neural networks have achieved promising performance for in-filling large missing regions in image inpainting tasks. They have usually adopted the standard convolutional architecture over the corrupted image, leading to meaningless contents, such as color discrepancy, blur, and other artifacts. Moreover, most inpainting approaches cannot handle well the case of a large contiguous missing area. To address these problems, we propose a generic inpainting framework capable of handling incomplete images with both contiguous and discontiguous large missing areas. We pose this in an adversarial manner, deploying regionwise operations in both the generator and discriminator to separately handle the different types of regions, namely, existing regions and missing ones. Moreover, a correlation loss is introduced to capture the nonlocal correlations between different patches, and thus, guide the generator to obtain more information during inference. With the help of regionwise generative adversarial mechanism, our framework can restore semantically reasonable and visually realistic images for both discontiguous and contiguous large missing areas. Extensive experiments on three widely used datasets for image inpainting task have been conducted, and both qualitative and quantitative experimental results demonstrate that the proposed model significantly outperforms the state-of-the-art approaches, on the large contiguous and discontiguous missing areas.
Yuqing Ma, Xianglong Liu 0001, Shihao Bai, Lei Wang 0018, Aishan Liu, Dacheng Tao, Edwin R. Hancock
IEEE Trans. Cybern.4
2023 Multiobjectivization of Single-Objective Optimization in Evolutionary Computation: A Survey
abstract
Multiobjectivization has emerged as a new promising paradigm to solve single-objective optimization problems (SOPs) in evolutionary computation, where an SOP is transformed into a multiobjective optimization problem (MOP) and solved by an evolutionary algorithm to find the optimal solutions of the original SOP. The transformation of an SOP into an MOP can be done by adding helper-objective(s) into the original objective, decomposing the original objective into multiple subobjectives, or aggregating subobjectives of the original objective into multiple scalar objectives. Multiobjectivization bridges the gap between SOPs and MOPs by transforming an SOP into the counterpart MOP, through which multiobjective optimization methods manage to attain superior solutions of the original SOP. Particularly, using multiobjectivization to solve SOPs can reduce the number of local optima, create new search paths from local optima to global optima, attain more incomparability solutions, and/or improve solution diversity. Since the term "multiobjectivization" was coined by Knowles et al. in 2001, this subject has accumulated plenty of works in the last two decades, yet there is a lack of systematic and comprehensive survey of these efforts. This article presents a comprehensive multifacet survey of the state-of-the-art multiobjectivization methods. Particularly, a new taxonomy of the methods is provided in this article and the advantages, limitations, challenges, theoretical analyses, benchmarks, applications, as well as future directions of the multiobjectivization methods are discussed.
Xiaoliang Ma 0001, Xiaodong Li 0001, Yutao Qi, Lei Wang 0018, Zexuan Zhu 0001
IEEE Trans. Cybern.5
2022 Character animation and retargeting from video streams
abstract
Virtual character animation is widely used in 3D games and virtual reality. Traditional character animation can be achieved through key-frame animation or motion capture technology. These methods have limited applications due to expensive equipments or sophisticated operations. Aiming at a lower-cost solution for this issue, in this paper we propose a method of virtual character animations and retargeting from RGB video streams based on human pose reconstruction. We conduct extensive experiments with different videos and virtual characters, and the resulting character animation is well represented in the virtual scene. The proposed method greatly has reduced the production cost of character animation, which has potential applications in virtual reality.
Gongbin Chen, Lei Wang 0018, Xun-Yu Liu, Long-Hua Hu, Jun Cheng 0002
SMC2
2022 Learning a Contrast Enhancer for Intensity Correction of Remotely Sensed Images
abstract
Low-quality remotely sensed images (RSIs) are not beneficial for the analysis of many activities including agricultural growth, resident migration, forest fire, and etc. Many previous enhancement schemes improve their quality via changing their illumination. However, these approaches often fail in detail and brightness preservation as well as contrast improvement due to that the information from a single image is limited. To address this issue, an enhancement framework, named as global-local enhancement network (GLE-Net), is proposed to correct the intensity via learning extra information from collected training data, including the following three key steps: first, RSIs are decomposed by the discrete wavelet transformation (DWT) method into the low-frequency component and the detail components. Then, the low-frequency component is improved by the global enhancement network while the detail components are enhanced by the local enhancement network in parallel. Finally, the enhanced components are employed to produce high-quality images with the inverse DWT (IDWT) method. The quantitatively and qualitatively comparable experiments on both synthetic and real-world RSIs validate that the proposed GLE-Net method performs well on preserving brightness and fine details, and even outperforms the state-of-the-arts.
Zhenghua Huang, Lei Wang 0018, Qing An, Qin Zhou 0005, Hanyu Hong
IEEE Signal Process. Lett.2
2022 RiFeGAN2: Rich Feature Generation for Text-to-Image Synthesis From Constrained Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual description. The description contains limited information compared with the corresponding image and is ambiguous and abstract, which will complicate the generation and lead to low-quality images. To address this problem, we propose a novel generation text-to-image synthesis method, called RiFeGAN2, to enrich the given description. To improve the enrichment quality while accelerating the enrichment process, RiFeGAN2 exploits a domain-specific constrained model to limit the search scope and then uses an attention-based caption matching model to refine the compatible candidate captions based on constrained prior knowledge. To improve the semantic consistency between the given description and the synthesized results, RiFeGAN2 employs improved SAEMs, SAEM2s, to compact better features of the retrieved captions and effectively emphasize the descriptions via incorporating centre-attention layers. Finally, multi-caption attentional GANs are exploited to synthesize images from those features. Experiments performed on widely-used datasets show that the models can generate vivid images from enriched captions and effectually improve the semantic consistency.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.4
2022 Visual Relationship Detection: A Survey
abstract
Visual relationship detection (VRD) is one newly developed computer vision task, aiming to recognize relations or interactions between objects in an image. It is a further learning task after object recognition, and is important for fully understanding images even the visual world. It has numerous applications, such as image retrieval, machine vision in robotics, visual question answer (VQA), and visual reasoning. However, this problem is difficult since relationships are not definite, and the number of possible relations is much larger than objects. So the complete annotation for visual relationships is much more difficult, making this task hard to learn. Many approaches have been proposed to tackle this problem especially with the development of deep neural networks in recent years. In this survey, we first introduce the background of visual relations. Then, we present categorization and frameworks of deep learning models for visual relationship detection. The high-level applications, benchmark datasets, as well as empirical analysis are also introduced for comprehensive understanding of this task.
Jun Cheng 0002, Lei Wang 0018, Jiaji Wu, Xiping Hu, Gwanggil Jeon, Dacheng Tao, MengChu Zhou
IEEE Trans. Cybern.2
2022 Enhanced Multifactorial Evolutionary Algorithm With Meme Helper-Tasks
abstract
Evolutionary multitasking (EMT) is an emerging research direction in the field of evolutionary computation. EMT solves multiple optimization tasks simultaneously using evolutionary algorithms with the aim to improve the solution for each task via intertask knowledge transfer. The effectiveness of intertask knowledge transfer is the key to the success of EMT. The multifactorial evolutionary algorithm (MFEA) represents one of the most widely used implementation paradigms of EMT. However, it tends to suffer from noneffective or even negative knowledge transfer. To address this issue and improve the performance of MFEA, we incorporate a prior-knowledge-based multiobjectivization via decomposition (MVD) into MFEA to construct strongly related meme helper-tasks. In the proposed method, MVD creates a related multiobjective optimization problem for each component task based on the corresponding problem structure or decision variable grouping to enhance positive intertask knowledge transfer. MVD can reduce the number of local optima and increase population diversity. Comparative experiments on the widely used test problems demonstrate that the constructed meme helper-tasks can utilize the prior knowledge of the target problems to improve the performance of MFEA.
Xiaoliang Ma 0001, Jian Yin 0004, Anmin Zhu, Xiaodong Li 0001, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001
IEEE Trans. Cybern.6
2022 Image Hallucination From Attribute Pairs
abstract
Recent image-generation methods have demonstrated that realistic images can be produced from captions. Despite the promising results achieved, existing caption-based generation methods confront a dilemma. On the one hand, the image generator should be provided with sufficient details for realistic hallucination, meaning that longer sentences with rich content are preferred, but on the other hand, the generator is meanwhile fragile to long sentences due to their complex semantics and syntax like long-range dependencies and the combinatorial explosion of object visual features. Toward alleviating this dilemma, a novel approach is proposed in this article to hallucinate images from attribute pairs, which can be extracted from natural language processing (NLP) toolsets in the presence of complex semantics and syntax. Attribute pairs, therefore, enable our image generator to tackle long sentences handily and alleviate the combinatorial explosion, and at the same time, allow us to enlarge the training dataset and to produce hallucinations from randomly combined attribute pairs at ease. Experiments on widely used datasets demonstrate that the proposed approach yields results superior to the state of the art.
Fuxiang Wu, Jun Cheng 0002, Xinchao Wang, Lei Wang 0018, Dapeng Tao
IEEE Trans. Cybern.4
2022 Merged Differential Grouping for Large-Scale Global Optimization
abstract
The divide-and-conquer strategy has been widely used in cooperative co-evolutionary algorithms to deal with large-scale global optimization problems, where a target problem is decomposed into a set of lower-dimensional and tractable subproblems to reduce the problem complexity. However, such a strategy usually demands a large number of function evaluations to obtain an accurate variable grouping. To address this issue, a merged differential grouping (MDG) method is proposed in this article based on the subset–subset interaction and binary search. In the proposed method, each variable is first identified as either a separable variable or a nonseparable variable. Afterward, all separable variables are put into the same subset, and the nonseparable variables are divided into multiple subsets using a binary-tree-based iterative merging method. With the proposed algorithm, the computational complexity of interaction detection is reduced to$O(\max \{n,n_{ns}\times \log _{2} k\})$, where$n$,$n_{ns}(\leq n)$, and$k( < n)$indicate the numbers of decision variables, nonseparable variables, and subsets of nonseparable variables, respectively. The experimental results on benchmark problems show that MDG is very competitive with the other state-of-the-art methods in terms of efficiency and accuracy of problem decomposition.
Xiaoliang Ma 0001, Xiaodong Li 0001, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001
IEEE Trans. Evol. Comput.4
2021 An Adaptive Multi-objective Multifactorial Evolutionary Algorithm Based on Mixture Gaussian Distribution
abstract
In recent decades, multi-objective multifactorial evolutionary algorithm (MOMFEA) has become a very promising research direction. How to achieve effective knowledge transfer between similar tasks is the key issue to affect the performance of the algorithm. In this paper, an adaptive MOMFEA (AMOMFEA) is proposed by exploiting the mixture Gaussian distribution of the population distributions of related tasks to help solve the target task. Wasserstein distance is used to measure the inter-task relevance in that the weight coefficient in the mixture distribution is proportional to the inter-task relevance. Experimental results on benchmark problems validate the effectiveness and efficiency of the proposed method in comparison with MOMFEA and NSGA-II.
Mengfan Xu, Zexuan Zhu 0001, Yutao Qi, Lei Wang 0018, Xiaoliang Ma 0001
CEC4
2021 Visual relationship detection with recurrent attention and negative sampling
Lei Wang 0018, Peizhen Lin, Jun Cheng 0002, Feng Liu 0013, Xiaoliang Ma 0001, Jian Yin 0004
Neurocomputing1
2021 Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image Synthesis
abstract
Due to the development of Generative Adversarial Networks (GANs), significant progress has been achieved in text-to-image synthesis task. However, most previous works have only focus on learning the semantic consistency between paired images and sentences, without exploring the semantic correlation between different yet related sentences that describe the same image, which leads to significant visual variation among the synthesized images. Accordingly, in this article, we propose a new method for text-to-image synthesis, dubbed Multi-sentence Auxiliary Generative Adversarial Networks (MA-GAN); this approach not only improves the generation quality but also guarantees the generation similarity of related sentences by exploring the semantic correlation between different sentences describing the same image. More specifically, we propose a Single-sentence Generation and Multi-sentence Discrimination (SGMD) module that explores the semantic correlation between multiple related sentences in order to reduce the variation between their generated images and enhance the reliability of the generated results. Moreover, a Progressive Negative Sample Selection mechanism (PNSS) is designed to mine more suitable negative samples for training, which can effectively promote detailed discrimination ability in the generative model and facilitate the generation of more fine-grained results. Extensive experiments on Oxford-102 and CUB datasets reveal that our MA-GAN significantly outperforms the state-of-the-art methods.
Yanhua Yang, Lei Wang 0018, De Xie, Cheng Deng 0002, Dacheng Tao
IEEE Trans. Image Process.2
2020 RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual sequence, which usually contains limited information compared with the corresponding image and so is ambiguous and abstractive. The limited textual information only describes a scene partly, which will complicate the generation with complementing the other details implicitly and lead to low-quality images. To address this problem, we propose a novel rich feature generating text-to-image synthesis, called RiFeGAN, to enrich the given description. In order to provide additional visual details and avoid conflicting, RiFeGAN exploits an attention-based caption matching model to select and refine the compatible candidate captions from prior knowledge. Given enriched captions, RiFeGAN uses self-attentional embedding mixtures to extract features across them effectually and handle the diverging features further. Then it exploits multi-captions attentional generative adversarial networks to synthesize images from those features. The experiments conducted on widely-used datasets show that the models can generate images from enriched captions effectually and improve the results significantly.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
CVPR4
2020 Phase-Sensitive Model for Temporal Action Proposal Generation
abstract
Temporal action proposal generation is an important and challenging task, aiming to localize the position where an action or event may occur in an untrimmed video. In this paper, we propose an efficient and end-to-end framework to generate temporal action proposals, named Phase-Sensitive Model (PSM), which fully understands all phases of temporal information. In particular, the PSM consists two modules: Boundary Phase Classification (BPC) and Action Phase Classification (APC). The BPC aims to provide two temporal boundary phase confidence maps by rich local information, while the APC is designed to generate an action phase confidence map by global features. Moreover, we introduce a new method boundary probability calculation to get the final score. Our experiments on ActivityNet-1.3 show a significant improvement with remarkable efficiency and generalizability.
Ziliang Ren, Lei Wang 0018, Jun Cheng 0002
HealthCom4
2019 Coupled CycleGAN: Unsupervised Hashing Network for Cross-Modal Retrieval
abstract
In recent years, hashing has attracted more and more attention owing to its superior capacity of low storage cost and high query efficiency in large-scale cross-modal retrieval. Benefiting from deep leaning, continuously compelling results in cross-modal retrieval community have been achieved. However, existing deep cross-modal hashing methods either rely on amounts of labeled information or have no ability to learn an accuracy correlation between different modalities. In this paper, we proposed Unsupervised coupled Cycle generative adversarial Hashing networks (UCH), for cross-modal retrieval, where outer-cycle network is used to learn powerful common representation, and inner-cycle network is explained to generate reliable hash codes. Specifically, our proposed UCH seamlessly couples these two networks with generative adversarial mechanism, which can be optimized simultaneously to learn representation and hash codes. Extensive experiments on three popular benchmark datasets show that the proposed UCH outperforms the state-of-the-art unsupervised cross-modal hashing methods.
Chao Li 0033, Cheng Deng 0002, Lei Wang 0018, De Xie, Xianglong Liu 0001
AAAI3
2019 Collect and Select: Semantic Alignment Metric Learning for Few-Shot Learning
abstract
Few-shot learning aims to learn latent patterns from few training examples and has shown promises in practice. However, directly calculating the distances between the query image and support image in existing methods may cause ambiguity because dominant objects can locate anywhere on images. To address this issue, this paper proposes a Semantic Alignment Metric Learning (SAML) method for few-shot learning that aligns the semantically relevant dominant objects through a "collect-and-select'' strategy. Specifically, we first calculate a relation matrix (RM) to "`collect" the distances of each local region pairs of the 3D tensor extracted from a query image and the mean tensor of the support images. Then, the attention technique is adapted to "select" the semantically relevant pairs and put more weights on them. Afterwards, a multi-layer perceptron (MLP) is utilized to map the reweighted RMs to their corresponding similarity scores. Theoretical analysis demonstrates the generalization ability of SAML and gives a theoretical guarantee. Empirical results demonstrate that semantic alignment is achieved. Extensive experiments on benchmark datasets validate the strengths of the proposed approach and demonstrate that SAML significantly outperforms the current state-of-the-art methods. The source code is available at https://github.com/haofusheng/SAML.
Fusheng Hao, Fengxiang He, Jun Cheng 0002, Lei Wang 0018, Jianzhong Cao, Dacheng Tao
ICCV4
2019 Coarse-to-Fine Image Inpainting via Region-wise Convolutions and Non-Local Correlation
abstract
Recently deep neural networks have achieved promising performance for filling large missing regions in image inpainting tasks. They usually adopted the standard convolutional architecture over the corrupted image, where the same convolution filters try to restore the diverse information on both existing and missing regions, and meanwhile ignores the long-distance correlation among the regions. Only relying on the surrounding areas inevitably leads to meaningless contents and artifacts, such as color discrepancy and blur. To address these problems, we first propose region-wise convolutions to locally deal with the different types of regions, which can help exactly reconstruct existing regions and roughly infer the missing ones from existing regions at the same time. Then, a non-local operation is introduced to globally model the correlation among different regions, promising visual consistency between missing and existing regions. Finally, we integrate the region-wise convolutions and non-local correlation in a coarse-to-fine framework to restore semantically reasonable and visually realistic images. Extensive experiments on three widely-used datasets for image inpainting tasks have been conducted, and both qualitative and quantitative experimental results demonstrate that the proposed model significantly outperforms the state-of-the-art approaches, especially for the large irregular missing regions.
Yuqing Ma, Xianglong Liu 0001, Shihao Bai, Lei Wang 0018, Dailan He, Aishan Liu
IJCAI4
2019 Single-Image De-Raining With Feature-Supervised Generative Adversarial Network
abstract
De-raining, which aims at rain-steak removal from images, is a practical task in computer vision. However, it is difficult due to its ill-posed nature. In this letter, we propose a deep neural network architecture, feature-supervised generative adversarial network (FS-GAN) for single-image rain removal. Its main idea is to train a generative adversarial network (GAN) for which the supervision from ground truth is imposed on different layers of the generator network. We design a feature-supervised generator, a discriminator, an optimization target, as well as the detailed structure of FS-GAN. Experiments show that the proposed FS-GAN achieves better performance than state-of-the-art de-raining methods on both synthetic and real-world images in terms of quantitative and visual quality.
Lei Wang 0018, Fuxiang Wu, Jun Cheng 0002, MengChu Zhou
IEEE Signal Process. Lett.2
2018 Ensemble One-Dimensional Convolution Neural Networks for Skeleton-Based Action Recognition
abstract
This letter proposes an ensemble neural network (Ensem-NN) for skeleton-based action recognition. The Ensem-NN is introduced based on the idea of ensemble learning, “two heads are better than one.” According to the property of skeleton sequences, we design one-dimensional convolution neural network with residual structure asBase-Net. From entirety to local, from focus to motion, we designed four different subnets based on theBase-Netto extract diverse features. The first subnet is aTwo-stream Entirety Net, which performs on the entirety skeleton and explores both temporal and spatial features. The second is aBody-part Net, which can extract fine-grained spatial and temporal features. The third is anAttention Net, in which a channel-wised attention mechanism can learn important frames and feature channels.Frame-difference Net, as the fourth subnet, aims at exploring motion features. Finally, the four subnets are fused as one ensemble network. Experimental results show that the proposed Ensem-NN performs better than state-of-the-art methods on three widely used datasets.
Yangyang Xu 0004, Jun Cheng 0002, Lei Wang 0018, Haiying Xia, Feng Liu 0013, Dapeng Tao
IEEE Signal Process. Lett.3
2017 PolSAR image compression based on online sparse K-SVD dictionary learning
Jing Bai 0003, Lei Wang 0018, Licheng Jiao
Multim. Tools Appl.3
2016 Locally estimated heterogeneity property and its fuzzy filter application for deinterlacing
Gwanggil Jeon, Marco Anisetti, Lei Wang 0018, Ernesto Damiani
Inf. Sci.3
2016 Robust object representation by boosting-like deep learning architecture
Lei Wang 0018, Baochang Zhang 0001, Jungong Han, LinLin Shen, Chengshan Qian
Signal Process. Image Commun.1
2015 2D to 3D conversion with motion-type adaptive depth estimation
Cheolkon Jung, Lei Wang 0018, Licheng Jiao
Multim. Syst.2
2015 Hyperspectral image compression based on lapped transform and Tucker decomposition
Lei Wang 0018, Jing Bai 0003, Jiaji Wu, Gwanggil Jeon
Signal Process. Image Commun.1
2015 Bayer Pattern CFA Demosaicking Based on Multi-Directional Weighted Interpolation and Guided Filter
abstract
In this letter, we proposed a new framework for color image demosaicking by using different strategies on green (G) and red/blue (R/B) components. Firstly, for G component, the missing samples are estimated by eight-direction weighted interpolation via exploiting spatial and spectral correlations of neighboring pixels. The G plane can be well reconstructed by considering the joint contribution of pre-estimations along eight interpolation directions with different weighting factors. Secondly, we estimate R/B components using guided filter with the reconstructed G plane as guidance image. Simulation results verify that, the proposed framework performs better than state-of-the-art demosaicking methods in term of color peak signal-to-noise ratio (CPSNR) and feature similarity index measure (FSIM), as well as higher visual quality.
Lei Wang 0018, Gwanggil Jeon
IEEE Signal Process. Lett.1
2014 Example-Based Video Stereolization With Foreground Segmentation and Depth Propagation
abstract
With advances in 3DTV technology, video stereolization has attracted much attention in recent years. Although video stereolization can enrich stereoscopic 3D contents, it is hard to create good depth maps from monocular 2D videos. In this paper, we propose an automatic example-based video stereolization method with foreground segmentation and depth propagation, called EBVS. To consider both performance and computational complexity, we separately estimate depth maps according to the key and non-key frames. In the key frames, we first estimate an initial depth map based on examples from the RGB-D training data set, then refine it to preserve boundaries of foreground objects. In the non-key frames, we propagate the depth map of the key frame using motion compensation, and generate depth maps. Finally, we employ depth-image-based-rendering (DIBR) to generate stereoscopic views from 2D videos and their depth maps. Extensive experiments verify that the proposed EBVS produces visually pleasing and realistic stereoscopic 3D views from 2D videos.
Lei Wang 0018, Cheolkon Jung
IEEE Trans. Multim.1
2011 Shape-Adaptive Reversible Integer Lapped Transform for Lossy-to-Lossless ROI Coding of Remote Sensing Two-Dimensional Images
abstract
In this letter, we propose a shape-adaptive (SA) reversible integer lapped transform (SA-RLT) method. The new method can deal with arbitrarily shaped image areas while guaranteeing completely reversible integer-to-integer transform. Based on SA-RLT and object-based set partitioned embedded block coder, a new region-of-interest (ROI) compression scheme is designed for 2-D remote sensing images. Numerical experiments reveal that SA-RLT performs better than integer SA discrete wavelet transform, and the new ROI compression scheme performs comparably even better than the JPEG2000-ROI scheme. Advantages in hardware implementation have been preserved by SA-RLT, such as parallel processing and low memory requirement.
Licheng Jiao, Lei Wang 0018, Jiaji Wu, Jing Bai 0003, Shuang Wang 0001, Biao Hou
IEEE Geosci. Remote. Sens. Lett.2
2010 Lossy-to-lossless image compression based on multiplier-less reversible integer time domain lapped transform
Lei Wang 0018, Licheng Jiao, Jiaji Wu, Guangming Shi, Yanjun Gong
Signal Process. Image Commun.1
2009 3D medical image compression based on multiplierless low-complexity RKLT and shape-adaptive wavelet transform
abstract
A multiplierless low complexity reversible integer Karhunen-Loe¿ve transform (Low-RKLT) is proposed based on matrix factorization. Conventional methods based on KLT suffer from high computational complexity and unability of applying in lossless medical image compression. To solve the two problems, multiplierless Low-RKLT is investigated using multi-lifting in this paper. Combined with ROI coding method, we have proposed a progressive lossy-to-lossless ROI compression method for three dimensional (3D) medical images with high performance. In our proposed method Low-RKLT is used for the inter-frame decorrelation after SA-DWT in the spatial domain. Simulation results show that, the proposed method performs much better in both lossless and lossy compression than 3D-DWT-based method.
Lei Wang 0018, Jiaji Wu, Licheng Jiao, Guangming Shi
ICIP1
2009 Lossy-to-Lossless Hyperspectral Image Compression Based on Multiplierless Reversible Integer TDLT/KLT
abstract
We proposed a new transform scheme of multiplierless reversible time-domain lapped transform and Karhunen-Loeve transform (RTDLT/KLT) for lossy-to-lossless hyperspectral image compression. Instead of applying discrete wavelet transform (DWT) in the spatial domain, RTDLT is applied for decorrelation. RTDLT can be achieved by existing discrete cosine transform and pre- and postfilters, while the reversible transform is guaranteed by a matrix factorization method. In the spectral direction, reversible integer low-complexity KLT is used for decorrelation. Owing to completely reversible transform, the proposed method can realize progressive lossy-to-lossless compression from a single embedded code-stream file. Numerical experiments on benchmark images show that the proposed transform scheme performs better than 5/3DWT-based methods in both lossy and lossless compressions, comparable with the optimal 9/7DWT-FloatKLT-based lossy compression method.
Lei Wang 0018, Jiaji Wu, Licheng Jiao, Guangming Shi
IEEE Geosci. Remote. Sens. Lett.1
2008 Lossy to lossless image compression based on reversible integer DCT
abstract
A progressive image compression scheme is investigated using reversible integer discrete cosine transform (RDCT) which is derived from the matrix factorization theory. Previous techniques based on DCT suffer from bad performance in lossy image compression compared with wavelet image codec. And lossless compression methods such as IntDCT, I2I-DCT and so on could not compare with JPEG-LS or integer discrete wavelet transform (DWT) based codec. In this paper, lossy to lossless image compression can be implemented by our proposed scheme which consists of RDCT, coefficients reorganization, bit plane encoding, and reversible integer pre- and post-filters. Simulation results show that our method is competitive against JPEG-LS and JPEG2000 in lossless compression. Moreover, our method outperforms JPEG2000 (reversible 5/3 filter) for lossy compression, and the performance is even comparable with JPEG2000 which adopted irreversible 9/7 floating-point filter (9/7F filter).
Lei Wang 0018, Jiaji Wu, Licheng Jiao, Li Zhang 0004, Guangming Shi
ICIP1
2006 A Novel Model of Artificial Immune Network and Simulations on Its Dynamics
Lei Wang 0018, Yinling Nie, Weike Nie, Licheng Jiao
ISNN (1)1
2005 A Novel Classifier with the Immune-Training Based Wavelet Neural Network
Lei Wang 0018, Yinling Nie, Weike Nie, Licheng Jiao
ISNN (2)1
2001 An Immune Neural Network Used for Classification
abstract
Based on an analysis of immune phenomena in nature and utilizing performances of ANN, a novel network mode, i.e., an immune neural network (INN), is proposed which integrates the immune mechanism and the function of neural information processing. The learning algorithm of an INN selects an excitation function and adaptive algorithm of the network. This model makes it easy for a user to utilize directly the characteristic information of a problem and to simplify the original structure by adjusting the excitation function with prior knowledge, improving the working efficiency and searching accuracy. A theoretical analysis and a simulation test for the twin-spiral problem show that, compared with an artificial neural network, INN is not only effective but also feasible. INN can simplify the structure of the existent model and show good working performance.
Lei Wang 0018, Licheng Jiao
ICDM1
2000 A novel genetic algorithim based on immunity
abstract
A novel optimal algorithm, immune genetic algorithm (IGA), is proposed based on the theory of immunity in biology, which constructs an immune operator accomplished by two steps, a vaccination and an immune selection. The detail processes of realizing IGA are presented. The methods of selecting vaccines and constructing an immune operator are also given. IGA is illustrated to be able to restrain the degenerate phenomenon evidently during the evolutionary process with examples of TSP, improve the searching capability and efficiency, therefore increase the convergent speed greatly.
Lei Wang 0018, Licheng Jiao
ISCAS1
2000 A novel genetic algorithm based on immunity
abstract
A novel algorithm, the immune genetic algorithm (IGA), is proposed based on the theory of immunity in biology which mainly constructs an immune operator accomplished by two steps: 1) a vaccination and 2) an immune selection. IGA proves theoretically convergent with probability 1. Strategies and methods of selecting vaccines and constructing an immune operator are also given. IGA is illustrated to be able to restrain the degenerate phenomenon effectively during the evolutionary process with examples of TSP, and can improve the searching ability and adaptability, greatly increase the convergence rate.
Licheng Jiao, Lei Wang 0018
IEEE Trans. Syst. Man Cybern. Part A2
1999 The immune evolutionary algorithm
abstract
A novel algorithm, the immune evolutionary algorithm (IEA), is proposed based on immune theory in biology, which constructs an immune operator accomplished by two steps: a vaccination and an immune selection. IEA is shown to converge to the global optimum with probability 1. Strategies of selecting vaccines and methods of constructing an immune operator are also given, with an example of TSP. A simulation test on the 75-city TSP shows that IEA can not only restrain the degenerate phenomenon during the evolutionary process, but can also improve the searching ability and the adaptability greatly, therefore increasing the convergent speed.
Lei Wang 0018, Licheng Jiao
KES1