VLDB 2026 Research / reviewers in the wild / expert
Yi Wang 0037
dblp:17/221-37
· DBLP profile ↗
51ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0001-5721-776XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 8 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiEn: Hierarchical ensemble learning for semi-supervised medical image segmentation
Long Ma 0002, Xinwei Xue, Chengpei Xu, Weimin Wang 0007, Yi Wang 0037 |
Neurocomputing | 7 |
| 2026 | Model-aware ellipse detection via parametric correlation learningabstractEllipse detection presents a significant challenge in computer vision and pattern recognition, often hindered by traditional parameter regression methods that fail to account for the unique geometric characteristics and complex parameter interactions of ellipses. These limitations frequently result in imprecise detections, notably with small or partially occluded ellipses. To overcome these challenges, we propose EDNet, a novel ellipse detection network that exploits the geometric properties of ellipses, thus moving beyond the reliance on internal textures. EDNet improves ellipse detection by refining the loss function to better capture the relationship between the error and each parameters during training. It features a LoG-like Edge Detection Module (LEDM) and an Edge Guided Module (EGM) for precise boundary extraction and multi-scale feature enhancement. Additionally, an auxiliary component estimates ellipse vertices, boosting accuracy for occluded ellipses. Experimental results on two wildly-used benchmark datasets demonstrate that EDNet achieves significant improvements, with an average detection accuracy increase of 6% and 10% over leading state-of-the-art models. • We concentrate on the geometric characteristics of ellipse detection via Edge Detection Module and Edge Guided Module. • We design an auxiliary head for the estimation of four ellipse vertices, invoking additional feature attention on these pivotal points. • We establish the relations between the error and geometric characteristics of the ellipse by a model-aware loss function. Qi Jia 0001, Zezheng Liu, Yu Liu 0012, Yi Wang 0037, Xinwei Xue, Weimin Wang 0007 |
Signal Process. | 4 |
| 2026 | Reliable Multi-Prototypical Contrastive Learning for Semi-Supervised Heterogeneous and Multi-Organ Medical Image SegmentationabstractAccurate multi-organ segmentation across heterogeneous medical images is pivotal for real-world surgical navigation. The scarcity of annotation constitutes a well-established consensus in the field, prompting semi-supervised learning to emerge as a prominent solution. However, two critical bottlenecks persist in clinical translation: (1) inter-class feature ambiguity, and (2) high multi-source sample heterogeneity. To tackle these bottlenecks, we propose TP-Net, a semi-supervised framework for multi-organ segmentation in heterogeneous medical images, which innovatively integrates reliable multi-prototype contrastive learning. Specifically, we propose a heterogeneous prototype dynamic evolution mechanism that self-adaptively models fragmented intra-class distributions in multi-source data. Then, to mitigate prototype shift, we first devise an uncertainty-aware cross-domain alignment strategy grounded in the smoothness assumption, which constructs reliable prototypes by propagating reliable pixel prediction distributions from labeled to unlabeled domains. Furthermore, this is synergized with contrastive separation that enforces feature proximity to class-matched prototypes in the embedding space, effectively resolving inter-class ambiguity by minimizing overlap between adjacent organ clusters via prototype repulsion. Experimental results on two public datasets and one in-house dataset prove that the proposed method achieves state-of-the-art performance. Further, clinical validation with our self-developed surgical navigation system demonstrated the clinical viability of the proposed method. The code associated with this work will be made publicly available at https://github.com/JIESHUREN330/TP-Net/tree/main. Xiangjun Yang, Jieshu Ren, Dongpei Liu, Yi Wang 0037, Zhihui Wang 0001, Bin Liu 0040 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | 2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT ImagesabstractClassifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instance selection. Thus, they cannot capture weak visual cues related to PD-L1 status on CT images. To address this, we propose a 2.5D top-K ranked multiple instance learning method. We design a 2.5D instance feature encoder, which takes advantage of knowledge from a 2D pre-trained model and has trainable parameters to learn information for 3D instances. In addition, we design a top-K ranked multiple instance learning strategy, which fully exploits the bag-level labels to select representative instances to eliminate the effect of atypical instances and guide the network to learn effective information. We demonstrate that our method can not only outperform state-of-the-art MIL methods on the PD-L1 status classification but also generalize well on a COVID-19 classification task. Huadong Liu, Yongcen Li, Xinchen Ye, Hongkai Wang 0002, Yi Wang 0037, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu |
ICASSP | 8 |
| 2025 | EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001 |
MICCAI (13) | 6 |
| 2025 | NuSegDG: Integration of heterogeneous space and Gaussian kernel for domain-generalized nuclei segmentation
Zhenye Lou, Qing Xu 0014, Zekun Jiang, Xiangjian He, Chenxin Li, Zhen Chen 0013, Yi Wang 0037, Maggie M. He, Wenting Duan |
Knowl. Based Syst. | 7 |
| 2025 | Enhancing video salient object detection via SAM-based multimodal energy prompting
Yi Wang 0037, Feng Hou, Li-li Liu |
Pattern Anal. Appl. | 2 |
| 2025 | Knowledge-sharing hierarchical memory fusion network for scribble-supervised video salient object detection
Feng Hou, Yi Wang 0037, Guangzhu Chen, Ruili Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2025 | Wireless Charging for Uncertain Location NodesabstractBenefiting from Wireless Power Transfer (WPT) technology, Wireless Rechargeable Sensor Networks (WRSNs) effectively address the lifetime bottleneck of sensor nodes, enabling them to work perpetually. Most state-of-the-art studies assume that all WRSNs’ information is known or precise in advance. However, sensor nodes may be deployed randomly in a large-scale area, and some critical information (such as node location) may be unavailable or difficult to obtain precisely. In this work, we eliminate the effect of uncertain or imprecise node location and formalize theMaximizingChargingEnergy utility for uncertain location nodesproblem (i.e., MCE problem). With magnetic resonance coupling and beamforming technologies, we propose a novel node localization method to determine precise node location information. In addition, we present a reinforcement learning framework and a charging path scheduling method to maximize charging energy. To validate the effectiveness of our proposed scheme in real-world scenarios, we conduct test-bed experiments. The results demonstrate that our approach significantly improves charging efficiency by an average of 20.9% in a large-scale network, even when the locations of sensors are entirely unknown. Chi Lin 0001, Shibo Hao, Yi Wang 0037, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Impossible Trinity in Underwater Optical Wireless CommunicationabstractUnderwater Optical Wireless Communication (UOWC) is considered a promising approach, offering the potential for flexible and high-speed communication under the surface of the water. However, the interdependent relationship among three key performance elements, namely communication distance, bit error rate, and communication rate, has been largely overlooked. This oversight impedes the complete utilization of the system performance. In this work, we innovatively introduce a “UOWC Impossible Trinity” model and theorems to establish relationships among the three key performance elements, which clarify the inherent constraints within UOWC system optimization. Moreover, we formulate the Underwater Optical Communication Trade-offs (UOCT) Problem to maximize communication performance. Furthermore, we provide feasible non-dominated solution sets, considering the constraints of real environments and user demands of specific scenarios. Our model has been validated by extensive simulations, demonstrating that our approach not only clarifies fundamental limitations of UOWC systems, but also provides practical guidelines for designing and optimizing the systems. Our approach has been experimentally validated with an impressive accuracy of over 95%, surpassing conventional models, which not only enhances the understanding of UOWC system optimization but also validates the existence of inherent trade-offs. Furthermore, our approach demonstrates a significant increase in communication distance, outperforming traditional methods by more than 20%. Chi Lin 0001, Yi Wang 0037, Yu Sun 0077, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001 |
ICNP | 3 |
| 2024 | Novelty Detection Based Discriminative Multiple Instance Feature Mining to Classify NSCLC PD-L1 Status on HE-Stained Histopathological Images
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yi Wang 0037, Hongkai Wang 0002, Dingpin Huang, Fangyi Xu, Yi Gan, Yuan Tu, Hongjie Hu |
MICCAI (4) | 6 |
| 2024 | Multimodal Energy Prompting for Video Salient Object Detection
Feng Hou, Yi Wang 0037 |
MMAsia | 3 |
| 2024 | MSTMENet: Multi-Scale Spatio-Temporal Mapping and Evolution Network for Video Deraining
Fengqi Li, Mengchao Guo, Renxuan Xiong, Donglei Yang, Yi Wang 0037, Fengqiang Xu |
MMAsia | 5 |
| 2024 | Structured Bipartite Graph Ensemble Clustering
Chen Wang 0108, Feng Hou, Yi Wang 0037, Ruili Wang 0001 |
MMAsia | 3 |
| 2024 | CSUNet: Contour-Sensitive Underwater Salient Object Detection
Yi Wang 0037, Shijun Yan, Tianzhu Wang, Zhihan Wang, Weirong Sun, Yu Zhao 0054, Xinwei Xue |
MMAsia | 2 |
| 2024 | Convolutional transformer network for fine-grained action recognition
Yujun Ma, Ruili Wang 0001, Ming Zong, Wanting Ji, Yi Wang 0037 |
Neurocomputing | 5 |
| 2024 | IENet: inheritance enhancement network for video salient object detection
Yi Wang 0037, Feng Hou, Ruili Wang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | WBNet: Weakly-supervised salient object detection via scribble and pseudo-background priorsabstractWeakly supervised salient object detection (WSOD) methods endeavor to boost sparse labels to get more salient cues in various ways. Among them, an effective approach is using pseudo labels from multiple unsupervised self-learning methods, but inaccurate and inconsistent pseudo labels could ultimately lead to detection performance degradation. To tackle this problem, we develop a new multi-source WSOD framework, WBNet, that can effectively utilize pseudo-background (non-salient region) labels combined with scribble labels to obtain more accurate salient features. We first design a comprehensive salient pseudo-mask generator from multiple self-learning features. Then, we pioneer the exploration of generating salient pseudo-labels via point-prompted and box-prompted Segment-Anything Models (SAM). Then, WBNet leverages a pixel-level Feature Aggregation Module (FAM), a mask-level Transformer-decoder (TFD), and an auxiliary Boundary Prediction Module (EPM) with a hybrid loss function to handle complex saliency detection tasks. Comprehensively evaluated with state-of-the-art methods on five widely used datasets, the proposed method significantly improves saliency detection performance. The code and results are publicly available at https://github.com/yiwangtz/WBNet. Yi Wang 0037, Ruili Wang 0001, Xiangjian He, Chi Lin 0001, Tianzhu Wang, Qi Jia 0001, Xin Fan 0001 |
Pattern Recognit. | 1 |
| 2024 | A Handwriting Recognition System With WiFiabstractHandwriting recognition systems are a convenient and alternative way of writing in the air with fingers rather than typing on keyboards. However, existing recognition systems are limited by their low accuracy and the requirement to wear dedicated devices. To address these issues, we propose WiWrite, an accurate contactless handwriting recognition system that allows users to write in the air without wearing any wearable devices. Specifically, we employ a novelCSI division schemeto process the noisy raw WiFi channel state information (CSI), which stabilizes the CSI phase and reduces noise in CSI amplitude. To automatically retain low noise data for identification in the LOS scenario, we propose a self-paced dense convolutional network (SPDCN), which is a self-paced loss function based on a modified convolutional neural network coupled with a dense convolutional network. Furthermore, to achieve accurate handwriting recognition in the NLOS scenario, we combine ADOA and PCA algorithms to remove location-induced interference and extract action features. Comprehensive experiments show the merits of WiWrite, revealing that the recognition accuracy for the same-size input and different-size input are 93.6% and 89.0%, respectively. Moreover, WiWrite can achieve accurate recognition regardless of environment and target diversity in LOS and NLOS scenarios. Chi Lin 0001, Asfandeyar Ahmad, Rongsheng Qu, Yi Wang 0037, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | AirWrite: An Aerial Handwriting Trajectory Tracking and Recognition System With mmWaveabstractIn the field of human-computer interaction (HCI), handwriting trajectory tracking and recognition have attracted significant attention due to their wide range of applications. However, many existing approaches rely on handheld devices and are highly susceptible to factors such as environmental conditions, location, and writing style. To overcome these limitations, we propose AirWrite, a novel contactless aerial system for handwriting trajectory tracking and recognition using mmWave technology. We introduce a signal clipping method based on the Doppler effect caused by user actions to accurately remove non-handwriting signals in the time domain. Additionally, we analyze power variations within the signal frequency interval to determine the handwriting frequency and employ a band-pass filter to eliminate dynamic environmental noise effectively. Through extensive experiments, we demonstrate that AirWrite can precisely track handwriting trajectories in noisy environments regardless of distance, angle, handwriting speed, character size, or in the presence of obstacles. Furthermore, we present an effective handwritten character recognition method for AirWrite that recognizes alphabets, numbers, and words. AirWrite can achieve an average accuracy of over 96% with only a 34 KB small dataset within 0.15 s for recognition. Chi Lin 0001, Zhouhe Sun, Asfandeyar Ahmad, Xinxin Fan, Yi Wang 0037, Lei Wang 0005, Xin Fan 0001, Guowei Wu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | A rotation robust shape transformer for cartoon character recognition
Qi Jia 0001, Yi Wang 0037, Xin Fan 0001, Haibin Ling, Longin Jan Latecki |
Vis. Comput. | 3 |
| 2023 | Pixels, Regions, and Objects: Multiple Enhancement for Salient Object DetectionabstractSalient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet. Yi Wang 0037, Ruili Wang 0001, Xin Fan 0001, Tianzhu Wang, Xiangjian He |
CVPR | 1 |
| 2023 | Spatial frequency enhanced salient object detection
Yi Wang 0037, Tianzhu Wang, Ruili Wang 0001 |
Inf. Sci. | 2 |
| 2023 | Near Optimal Charging Schedule for 3-D Wireless Rechargeable Sensor NetworksabstractWireless rechargeable sensor networks (WRSNs) have become a hot research issue owing to the breakthrough of wireless power transfer (WPT) technology. Previous theoretical schemes are mostly designed for 2-D networks, and few of them are tailored for 3-D scenarios, making them not suitable for wide adoptions in practical applications. In this paper, we address the issue of how to serve a 3-D WRSN with an unmanned aerial vehicle (UAV). Our main concern is to maximize the charged energy for sensors supplied by the UAV, which has energy constraints. We respectively develop a spatial discretization scheme to construct a finite feasible set of charging spots for the UAV in a 3-D environment and a temporal discretization scheme to determine the appropriate charging duration for each charging spot. Then, we reduce the problem into a submodular maximization problem with routing constraints and present a cost-efficient algorithm (CEA) with a provable approximation ratio to solve it. Finally, test-bed experiments are conducted to show the feasibility of our schemes in practical scenarios. Extensive simulations are taken to verify the superior performance of our algorithm in charged energy and robustness. The charged energy of our scheme outperforms other competing methods by at least$18.2\%$. Chi Lin 0001, Wei Yang 0039, Haipeng Dai 0001, Teng Li 0003, Yi Wang 0037, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 5 |
| 2022 | Best of Both Worlds: See and Understand Clearly in the DarkabstractRecently, with the development of intelligent technology, the perception of low-light scenes has been gaining widespread attention. However, existing techniques usually focus on only one task (e.g., enhancement) and lose sight of the others (e.g., detection), making it difficult to perform all of them well at the same time. To overcome this limitation, we propose a new method that can handle visual quality enhancement and semantic-related tasks (e.g., detection, segmentation) simultaneously in a unified framework. Specifically, we build a cascaded architecture to meet the task requirements. To better enhance the entanglement in both tasks and achieve mutual guidance, we develop a new contrastive-alternative learning strategy for learning the model parameters, to largely improve the representational capacity of the cascaded architecture. Notably, the contrastive learning mechanism establishes the communication between two objective tasks in essence, which actually extends the capability of contrastive learning to some extent. Finally, extensive experiments are performed to fully validate the advantages of our method over other state-of-the-art works in enhancement, detection, and segmentation. A series of analytical evaluations are also conducted to reveal our effectiveness. The code is available at https://github.com/k914/contrastive-alternative-learning. Xinwei Xue, Long Ma 0002, Yi Wang 0037, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 4 |
| 2022 | Function Representation Based Analytic Shape Hollowing Optimization
Shengfa Wang, Baojun Li, Yi Wang 0037, Zhongxuan Luo, Ligang Liu 0001 |
Comput. Aided Des. | 4 |
| 2022 | Proactive and intelligent evaluation of big data queries in edge clouds with materialized views
Qiufen Xia, Lizhen Zhou, Wenhao Ren, Yi Wang 0037 |
Comput. Networks | 4 |
| 2021 | Temporal Rain Decomposition with Spatial Structure Guidance for Video DerainingabstractRecently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In this work, we propose a multi-frame deraining network with temporal rain decomposition and spatial structure guidance to more effectively accomplish video deraining. A learnable decomposition method is defined to learn the distribution of rain, where the location map acts on a single-frame deraining block. We construct a multi-frame fusion module with a detailed guidance map to integrate temporal and spatial information. Many evaluated experiments demonstrate that our algorithm performs favorably on video deraining tasks compared with other methods. The elaborate ablation study in terms of network architecture fully indicates the effectiveness of our network. Xinwei Xue, Ying Ding 0006, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001 |
ICASSP | 4 |
| 2021 | Searching Frame-Recurrent Attentive Deformable Network for Real-Time Video DerainingabstractVideo deraining has become an issue of great interest since rain streaks inevitably affect video quality. Most of the existing works focus on heuristically designing the network architecture to integrate available information derived from the temporal dimension. However, their inferences take a long time, so that the practicability is somewhat ignored. To solve this problem, we develop a real-time video deraining network in a frame-recurrent manner. It includes a fast attentive deformable alignment module and an automatically-discovered spatial-temporal reconstruction module. In which, the alignment is composed of a single newly-built deformable convolution under the channel attention mechanism to keep the accurate motion consistency and reduce time-consuming by a wide margin. The reconstruction part for the first time introduces the architecture search technique for video deraining to automatically discover a high-effective architecture by designing an effective and compact search space. Experimental results demonstrate remarkable superiority both in computational efficiency and actual performance compared to other state-of-the-art approaches. Xinwei Xue, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001 |
ICME | 4 |
| 2021 | Underwater Species Detection using Channel Sharpening AttentionabstractWith the continuous exploration of marine resources, underwater artificial intelligent robots play an increasingly important role in the fish industry. However, the detection of underwater objects is a very challenging problem due to the irregular movement of underwater objects, the occlusion of sand and rocks, the diversity of water illumination, and the poor visibility and low color contrast in the underwater environment. In this article, we first propose a real-world underwater object detection dataset (UODD), which covers more than 3K images of the most common aquatic products. Then we propose Channel Sharpening Attention Module (CSAM) as a plug-and-play module to further fuse high-level image information, providing the network with the privilege of selecting feature maps. Fusion of original images through CSAM can improve the accuracy of detecting small and medium objects, thereby improving the overall detection accuracy. We also use Water-Net as a preprocessing method to remove the haze and color cast in complex underwater scenes, which shows a satisfactory detection result on small-sized objects. In addition, we use the class weighted loss as the training loss, which can accurately describe the relationship between classification and precision of bounding boxes of targets, and the loss function converges faster during the training process. Experimental results show that the proposed method reaches a maximum AP of 50.1%, outperforming other traditional and state-of-the-art detectors. In addition, our model only needs an average inference time of 25.4 ms per image, which is quite fast and might suit the real-time scenario. Lihao Jiang, Yi Wang 0037, Qi Jia 0001, Shengwei Xu, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Xinwei Xue, Ruili Wang 0001 |
ACM Multimedia | 2 |
| 2021 | Novelty Detection and Online Learning for Chunk Data StreamsabstractDatastream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while incrementally updating the model efficiently and stably, especially for high-dimensional and/or large-scale data streams. This paper proposes an efficient framework for novelty detection and incremental learning for unlabeled chunk data streams. First, an accurate factorization-free kernel discriminative analysis (FKDA-X) is put forward through solving a linear system in the kernel space. FKDA-X produces a Reproducing Kernel Hilbert Space (RKHS), in which unlabeled chunk data can be detected and classified by multiple known-classes in a single decision model with a deterministic classification boundary. Moreover, based on FKDA-X, two optimal methods FKDA-CX and FKDA-C are proposed. FKDA-CX uses the micro-cluster centers of original data as the input to achieve excellent performance in novelty detection. FKDA-C and incremental FKDA-C (IFKDA-C) using the class centers of original data as their input have extremely fast speed in online learning. Theoretical analysis and experimental validation on under-sampled and large-scale real-world datasets demonstrate that the proposed algorithms make it possible to learn unlabeled chunk data streams with significantly lower computational costs and comparable accuracies than the state-of-the-art approaches. Yi Wang 0037, Xiangjian He, Xin Fan 0001, Chi Lin 0001, Fengqi Li, Tianzhu Wang, Zhongxuan Luo, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Joint Luminance and Chrominance Learning for Underwater Image EnhancementabstractRecently, learning-based works have been widely-investigated to enhance underwater images. However, interactions between various degradation factors (e.g., color distortion and haze effects) inevitably cause negative interference during the inference phase. Thus, these works cannot fully remove degraded factors. To address this problem, we propose a novel Joint Luminance and Chrominance Learning Network (JLCL-Net). Concretely, we reformulate the task as luminance reconstruction (for haze removal), and chrominance correction (for color correction) sub-tasks by separating the luminance and chrominance (i.e., color appearance) of the underwater images. In this way, we successfully realize the disentanglement in degraded factors to avoid introducing interference. We specify the reconstruction by integrating the atmospheric scattering model, which endows the adaptive dehazing ability over different scenarios. The correction learns to compensate for color by a simple network to reverse the color attenuation process. To this end, we obtain our JLCL-Net. To better train it, we design a new multi-stage cross-space training strategy, which progressively updates the network parameters to enlarge the network potentiality. Extensive evaluations are presented to fully verify our superiority against other methods. Xinwei Xue, Zhenhua Hao, Long Ma 0002, Yi Wang 0037, Risheng Liu |
IEEE Signal Process. Lett. | 4 |
| 2021 | DENet: A Universal Network for Counting Crowd With Varying Densities and ScalesabstractCounting people or objects with significantly varying scales and densities has attracted much interest from the research community and yet it remains an open problem. In this paper, we propose a simple but efficient and effective network, named DENet, which is composed of two components,i.e., a detection network (DNet) and an encoder-decoder estimation network (ENet). We first run the DNet on the input image to detect and count individuals who can be segmented clearly. Then, the ENet is utilized to estimate the density maps of the remaining areas, typically with low resolution and high densities where individuals cannot be detected. For this purpose, we propose a modified Xception network as the encoder for feature extraction and a combination of dilated convolution and transposed convolution as the decoder. When evaluated on the ShanghaiTech Part A, UCF and WorldExpo’10 datasets, our DENet has achieved lower Mean Absolute Error (MAE) than those of the state-of-the-art methods. Lei Liu 0036, Jie Jiang 0005, Wenjing Jia, Saeed Amirgholipour Kasmani, Yi Wang 0037, Michelle Zeibots, Xiangjian He |
IEEE Trans. Multim. | 5 |
| 2020 | BG-Net: Boundary-Guided Network for Lung Segmentation on Clinical CT ImagesabstractLung segmentation on CT images is a crucial step for a computer-aided diagnosis system of lung diseases. The existing deep learning based lung segmentation methods are less efficient to segment lungs on clinical CT images, especially that the segmentation on lung boundaries is not accurate enough due to complex pulmonary opacities in practical clinics. In this paper, we propose a boundary-guided network (BG-Net) to address this problem. It contains two auxiliary branches that seperately segment lungs and extract the lung boundaries, and an aggregation branch that efficiently exploits lung boundary cues to guide the network for more accurate lung segmentation on clinical CT images. We evaluate the proposed method on a private dataset collected from the Osaka university hospital and four public datasets including StructSeg [1], HUG [2], VESSEL12 [3], and a Novel Coronavirus 2019 (COVID-19) dataset [4]. Experimental results show that the proposed method can segment lungs more accurately and outperform several other deep learning based methods. Rui Xu 0002, Yi Wang 0037, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001, Shoji Kido, Noriyuki Tomiyama |
ICPR | 2 |
| 2019 | A lightweight methodology of 3D printed objects utilizing multi-scale porous structures
Jiangbei Hu, Shengfa Wang, Yi Wang 0037, Fengqi Li, Zhongxuan Luo |
Vis. Comput. | 3 |
| 2018 | Fast Factorization-free Kernel Learning for Unlabeled Chunk Data StreamsabstractData stream analysis aims at extracting discriminative information for classification from continuously incoming samples. It is extremely challenging to detect novel data while updating the model in an efficient and stable fashion, especially for the chunk data. This paper proposes a fast factorization-free kernel learning method to unify novelty detection and incremental learning for unlabeled chunk data streams in one framework. The proposed method constructs a joint reproducing kernel Hilbert space from known class centers by solving a linear system in kernel space. Naturally, unlabeled data can be detected and classified among multi-classes by a single decision model. And projecting samples into the discriminative feature space turns out to be the product of two small-sized kernel matrices without needing such time-consuming factorization like QR-decomposition or singular value decomposition. Moreover, the insertion of a novel class can be treated as the addition of a new orthogonal basis to the existing feature space, resulting in fast and stable updating schemes. Both theoretical analysis and experimental validation on real-world datasets demonstrate that the proposed methods learn chunk data streams with significantly lower computational costs and comparable or superior accuracy than the state of the art. Yi Wang 0037, Nan Xue 0004, Xin Fan 0001, Jiebo Luo 0001, Risheng Liu, Zhongxuan Luo |
IJCAI | 1 |
| 2017 | Fast Online Incremental Learning on Mixture Streaming DataabstractThe explosion of streaming data poses challenges to feature learning methods including linear discriminant analysis (LDA). Many existing LDA algorithms are not efficient enough to incrementally update with samples that sequentially arrive in various manners. First, we propose a new fast batch LDA (FLDA/QR) learning algorithm that uses the cluster centers to solve a lower triangular system that is optimized by the Cholesky-factorization. To take advantage of the intrinsically incremental mechanism of the matrix, we further develop an exact incremental algorithm (IFLDA/QR). The Gram-Schmidt process with reorthogonalization in IFLDA/QR significantly saves the space and time expenses compared with the rank-one QR-updating of most existing methods. IFLDA/QR is able to handle streaming data containing 1) new labeled samples in the existing classes, 2) samples of an entirely new (novel) class, and more significantly, 3) a chunk of examples mixed with those in 1) and 2). Both theoretical analysis and numerical experiments have demonstrated much lower space and time costs (2~10 times faster) than the state of the art, with comparable classification accuracy. Yi Wang 0037, Xin Fan 0001, Zhongxuan Luo, Tianzhu Wang, Maomao Min, Jiebo Luo 0001 |
AAAI | 1 |
| 2017 | Incremental zero-shot learning based on attributes for image classificationabstractInstead of assuming a closed-world environment comprising a fixed number of objects, modern pattern recognition systems need to recognize outliers, identify anomalies, or discover entirely new objects, which is known as zero-shot object recognition. However, many existing zero-shot learning methods are not efficient enough to incrementally update themselves with new samples mixed with known or novel class labels. In this paper, we propose an incremental zero-shot learning framework (IIAP/QR) based on indirect-attribute-prediction (IAP) model. Firstly, a fast incremental classifier based on null space based linear discriminant analysis with QR-updating (NLDA/QR) is put forward, which can solve small-sample-size (SSS) problem and unequal-sample-size (USS) problem that usually occur in incremental learning using the centroid of each class as input. Then with the probabilistic inference of Class-Attribute layer and Attribute-Zero shot classification layer, IIAP/QR model can efficiently update itself for the insertion of both new samples to the existing class and totally novel classes with comparable recognition accuracy for zero-shot object recognition. Nan Xue 0004, Yi Wang 0037, Xin Fan 0001, Maomao Min |
ICIP | 2 |
| 2017 | Two-Layer Gaussian Process Regression With Example Selection for Image DehazingabstractResearchers have devoted great efforts to image dehazing with prior assumptions in the past decade. Recently developed example-based approaches typically lack elegant models for the hazy process and meanwhile demand synthetic hazy images by manual selection. The priors from observations, and those trained from synthetic images cannot always reflect true structural information of natural images in practice. In this paper, we present a learning model for haze removal by using two-layer Gaussian process regression (GPR). By using training examples, the two-layer GPR establishes a direct relationship from the input image to the depth-dependent transmission, and learns local image priors to further improve the estimation. We also provide a systematic scheme to automatically collect suitable training pairs, which works for both simulated examples and images of natural scenes. Both qualitative and quantitative comparisons on real-world and synthetic hazy images demonstrate the effectiveness of the proposed approach, especially for white or bright objects and heavy haze regions in which traditional methods may fail. Xin Fan 0001, Yi Wang 0037, Xianxuan Tang, Renjie Gao, Zhongxuan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | An efficient mesh-based face beautifier on mobile devices
Xin Fan 0001, Yuyao Feng, Yi Wang 0037, Shengfa Wang, Zhongxuan Luo |
Neurocomputing | 4 |
| 2016 | Haze editing with natural transmission
Xin Fan 0001, Yi Wang 0037, Renjie Gao, Zhongxuan Luo |
Vis. Comput. | 2 |
| 2015 | A camera motion histogram descriptor for video shot classification
Muhammad Abul Hasan, Min Xu 0001, Xiangjian He, Yi Wang 0037 |
Multim. Tools Appl. | 4 |
| 2015 | Online gesture-based interaction with visual oriental characters based on manifold learning
Yi Wang 0037, Xin Fan 0001, Xiangjian He, Qi Jia 0001, Renjie Gao |
Signal Process. | 1 |
| 2015 | Local N-Ary Pattern and Its Extension for Texture ClassificationabstractTexture image classification is important in computer vision research. To effectively capture texture patterns, a distinctive feature such as a local binary pattern (LBP) is needed. An LBP is robust against monotonic and gray-scale variations and it computes quickly. Its robustness and speed advantage have made it popular in various texture analysis applications. However, an LBP is sensitive to noise, particularly smooth weak illumination gradients in near-uniform regions. To mitigate the effect of noise and increase distinctiveness, a local ternary pattern (LTP) is proposed. Compared with a binary coding LBP, an LTP adopts ternary coding. As a result, an LTP can better tolerate noise and is significantly more distinctive. These advantages of an LTP effectively improve its classification accuracy. However, the potential of ternary coding is not fully explored in LTPs because a ternary pattern is split into a pair of binary patterns. In this paper, to fully explore the distinctiveness in the local pattern, the feature extraction process is formulated as an integer decomposition problem, which is a generalized version of the Bachet de Meziriac weight problem (BMWP). Following this generalization, a local n-ary pattern (LNP) is proposed, for which the LBP is a special case parametrized under n = 2. The LTP is not a special case of the LNP. Both LBP and LTP are used as benchmark methods to evaluate LNPs performance due to their well-recognized success. In addition, a rotation-invariant and uniform LNP is also proposed and compared with a rotation-invariant and uniform LBP. The proposed LNP achieves significantly improved texture classification accuracy compared with the LBP and also demonstrates considerable improvement over the LTP. Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Jie Yang 0002, Yi Wang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Real-time estimation of hand gestures based on manifold learning from monocular videos
Yi Wang 0037, Zhongxuan Luo, Xin Fan 0001, Yunzhen Wu |
Multim. Tools Appl. | 1 |
| 2014 | Localizing relevant frames in web videos using topic model and relevance filtering
Lei Yi, Bin Liu 0040, Yi Wang 0037 |
Mach. Vis. Appl. | 4 |
| 2012 | Looking into the world on Google Maps with view direction estimated photos
Jinhui Tang 0001, Yi Wang 0037, Bin Liu 0001 |
Neurocomputing | 3 |
| 2006 | Off-Line Motion Description for Fast Video Stream Generation in MPEG-4 AVC/H.264abstractThe rate-distortion optimal mode decision as well as motion estimation adopted in H.264 brings a big challenge to real-time encoding and transcoding due to the high computation complexity. In this paper, we propose a hierarchical motion description model to present the motion data of each macroblock (MB) from coarsely to finely. A preprocessing approach is developed to estimate the motion data for each MB at each quality level with regard to its reference quality, its adjacent MBs and the target bit-rate. The resulting motion data can be coded and stored as metadata in a media file or a stream. Moreover, we propose a method to readily extract the specific motion data from the model for each MB at given bit-rates. Experimental results have shown the effectiveness of our proposed motion description model in terms of coding efficiency as well as fast bit-rate adaptation in comparison with that of H.264 Yi Wang 0037, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Houqiang Li, Zhengkai Liu |
ICME | 1 |
| 2006 | An attention based spatial adaptation scheme for H.264 videos on mobilesabstractWhen browsing videos in mobile devices, people often feel that resolution greatly affects their perceptual experience in the limited screen size. In this paper, an attention based spatial video adaptation scheme is proposed to overcome display constraints by producing the region of interest. According to the size of the target display, we automatically detect and crop the informative region in each frame to generate a smooth sequence. To avoid costly fully encoding operations, we employ a set of transcoding techniques based on the H.264 standard. Experimental results show that this approach not only improves the perceptual quality but also saves the bandwidth and computation, especially for the videos which are not well edited Yi Wang 0037, Xin Fan 0001, Houqiang Li, Zhengkai Liu, Mingjing Li |
MMM | 1 |
| 2006 | An Attention Based Spatial Adaptation Scheme for H.264 Videos on MobilesabstractWith the growing popularity of personal digital assistant devices and smart phones, consumers have become increasingly enthusiastic to watching videos from these mobile devices. However, when browsing videos in mobiles, users often feel that the display resolution greatly affects their perceptual experience with the limited screen size. In this paper, an attention based spatial video adaptation scheme is proposed to overcome the display constraints by producing and displaying the region of interest. According to the size of the target display, we automatically detect and crop the informative region in each frame to generate a smooth sequence. To avoid costly full encoding operations, we develop a set of transcoding techniques based on the H.264 standard. Experimental results show that this approach not only improves the perceptual quality but also saves the bandwidth and computation, especially for the videos which have not been well edited. Yi Wang 0037, Houqiang Li, Xin Fan 0001, Chang Wen Chen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | A novel video coding scheme for mobile devicesabstractIn this paper, we propose a novel video coding profile for the multimedia applications oriented for mobile wireless communication. Because mobile devices generally have limited computational capability and constrained power consumption, we have considered jointly both video coding efficiency and implementation feasibility when developing the proposed video coding scheme. We achieve a good trade-off between coding efficiency and complexity by minimizing the video reconstruction distortion caused by adopting reduced complexity algorithms. A number of experiments have been conducted and the results have shown the efficiency of the proposed profile. A demo implemented on the Nokia 6600 cell phone demonstrates the feasibility of video coding scheme for mobile devices in wireless communication applications. Yi Wang 0037, Houqiang Li, Chang Wen Chen |
MUM | 1 |