VLDB 2026 Research / reviewers in the wild / expert
Wang Yuan
dblp:11/3629
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-9422-777XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OrientTongue: an oriented and attention-enhanced framework for fine-grained tongue diagnosisabstractTongue diagnosis in Traditional Chinese Medicine contains rich clinical information, yet conventional visual assessment remains subjective and poorly standardized. To address key challenges in automated tongue image analysis—small-scale targets, low-contrast lesions, background interference, and strong directional variations—we propose an enhanced YOLOv8-based end-to-end detection framework. The model integrates Bottleneck Transformer (BoT3) modules in the backbone to strengthen global dependency modeling, and inserts Convolutional Block Attention Module (CBAM) attention in the detection head to improve feature discrimination under noisy and low-contrast conditions. To better capture irregular and oriented tongue features, we adopt oriented bounding boxes and design an adaptive NMS strategy that adjusts suppression thresholds based on object scale, improving both small-lesion recall and large-object precision. Experiments on the Tongue-det dataset covering seven clinically relevant tongue phenotypes show an [email protected] of 0.581, with most categories achieving sample-level F1 scores above 0.75. Ablation studies confirm consistent performance gains from each component, especially for subtle and low-contrast features such as rotten coating. Overall, the framework enhances accuracy, robustness, and interpretability, providing a promising pathway toward objective and intelligent tongue diagnosis Tao Jiang 0032, Wang Yuan, Liping Tu, Ji Cui, Lizhuang Ma, Jiatuo Xu |
Expert Syst. Appl. | 2 |
| 2025 | Learning Pulse Image with Deep Dynamic Frequency Network for Cardiovascular Diseases Diagnosis
Ji Cui, Litai Pang, Shiju Zhao, Zhengyuan Peng, Lingzhi Zeng, Tao Jiang 0032, Mengchen Liang, Jinlian Huang, Wang Yuan, Xin Tan 0002, Lizhuang Ma, Jiatuo Xu |
Vis. Comput. | 10 |
| 2024 | Boosting Data Center Performance via Intelligently Managed Multi-backend Disaggregated MemoryabstractExisting disaggregated memory (DM) systems face a problem of underutilized far memory bandwidth, which greatly limits the data throughput when processing data-intensive applications. Specifically, prior works all target runtime design for a single PCIe-based secondary memory device (i.e., single-backend far memory) with low data bandwidth and high system overhead. In this work, we take the first step to realize a well-crafted, multi-backend DM system with scale-out far memory paths. We propose xDM, a novel DM management scheme that can dynamically build and implicitly select appropriate far memory access paths. As part of xDM, we devise a smart far memory configuration strategy that can further optimize bandwidth usage effectiveness by tuning a wide set of key parameters based on synthesized information of application page data. Our design shows up to $3.9 \times$ data swap performance speedup, $2.8 \times$ data throughput increase, and $5.1 \times$ data center task throughput improvement compared with state-of-the-art works. Jing Wang 0055, Hanzhang Yang, Chao Li 0009, Yiming Zhuansun, Wang Yuan, Xiaofeng Hou, Minyi Guo, Yang Hu 0001, Yaqian Zhao |
SC | 5 |
| 2023 | Self-supervised Contrastive Feature Refinement for Few-Shot Class-Incremental Learning
Shengjin Ma, Wang Yuan, Xin Tan 0002, Zhizhong Zhang 0001, Lizhuang Ma |
CAD/Graphics | 2 |
| 2023 | Whole-Body Control of an Autonomous Mobile Manipulator Using Model Predictive Control and Adaptive Fuzzy TechniqueabstractWhole-body control (WBC) has emerged as an important framework in manipulation for mobile manipulators. However, most existing WBC frameworks require known dynamics. Considering whole-body manipulation and optimization with unknown dynamics, this article presents the WBC of a nonholonomic mobile manipulator using model predictive control (MPC) and fuzzy logic system. First, by constructing a dynamics-based feedback linearized robotic multi-input-multi-output (MIMO) system, an MPC-based WBC strategy is proposed for mobile manipulator. Such a strategy can provide the optimal control inputs with the specified optimization index and constraints. Thereafter, a primal-dual neural network effectively addresses the constrained quadratic programming (QP) problem over a finite receding horizon brought by the MPC. Then, in order to convert the intermediate control signals into the optimal control torques that can be executed by actuators, an adaptive FLS is employed to approximate the unknown dynamics. The novel elements of the current design control approach refer to the dynamics-based feedback linearized robotic MIMO system and the combination of an MPC module with an adaptive fuzzy controller. Finally, the trajectory tracking experiments performed on a mobile dual-arm robot demonstrate the effectiveness of the proposed method. Wang Yuan, Yong-Hua Liu, Chun-Yi Su, Feng Zhao 0004 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | Task-Level Self-Supervision for Cross-Domain Few-Shot LearningabstractLearning with limited labeled data is a long-standing problem. Among various solutions, episodic training progres-sively classifies a series of few-shot tasks and thereby is as-sumed to be beneficial for improving the model’s generalization ability. However, recent studies show that it is eveninferior to the baseline model when facing domain shift between base and novel classes. To tackle this problem, we pro-pose a domain-independent task-level self-supervised (TL-SS) method for cross-domain few-shot learning.TL-SS strategy promotes the general idea of label-based instance-levelsupervision to task-level self-supervision by augmenting mul-tiple views of tasks. Two regularizations on task consistencyand correlation metric are introduced to remarkably stabi-lize the training process and endow the generalization ability into the prediction model. We also propose a high-order associated encoder (HAE) being adaptive to various tasks.By utilizing 3D convolution module, HAE is able to generate proper parameters and enables the encoder to flexibly toany unseen tasks. Two modules complement each other andshow great promotion against state-of-the-art methods experimentally. Finally, we design a generalized task-agnostic test,where our intriguing findings highlight the need to re-think the generalization ability of existing few-shot approaches. Wang Yuan, Zhizhong Zhang 0001, Cong Wang 0039, Yuan Xie 0006, Lizhuang Ma |
AAAI | 1 |
| 2022 | Mutually Reinforcing Structure with Proposal Contrastive Consistency for Few-Shot Object Detection
TianXue Ma, Mingwei Bi, Jian Zhang 0079, Wang Yuan, Zhizhong Zhang 0001, Yuan Xie 0006, Shouhong Ding, Lizhuang Ma |
ECCV (20) | 4 |
| 2022 | Spoof Face Detection Via Semi-Supervised Adversarial TrainingabstractFace spoofing causes severe security threats in face recognition systems. The previous anti-spoofing mainly focused on supervised techniques, typically with either binary or auxiliary supervision. Most of them have to ‘see’ both spoofing face data and live face data during training to realize the task of face anti-spoofing. In this paper, we propose a semi-supervised adversarial learning framework for spoof face detection, which largely relaxes the supervision condition. To capture the underlying structure of live face data in latent representation space, we propose to train the live face data only, with a convolutional Encoder-Decoder network acting as a Generator, and a second convolutional network serving as a Discriminator. The generator and discriminator are trained by competing with each other while collaborating to understand the live faces. Since the spoof face detection is video-based (i.e., temporal information), we intuitively take the optical flow maps converted from consecutive video frames as input. Our approach is free of the spoof faces, thus being robust and general to different types of face spoofing (even unknown spoofing). Experiments on cross-dataset tests show that our semi-supervised method achieves better or comparable results to state-of-the-art supervised techniques. We also conduct ablation studies for the proposed method. Chengwei Chen, Yaping Jing, Xuequan Lu, Wang Yuan, Lizhuang Ma |
IJCNN | 4 |
| 2021 | Cross-Modality Graph Neural Network For Few-Shot LearningabstractFew-shot learning, which attempts to predict unlabeled samples with only a few labeled samples, has drawn more and more attention. Though recent works have achieved promising progress, none of them have noticed to establish consistency among episodes, leading to the ambiguity in latent embedding space. In this paper, we propose a novel Cross-Modality Graph Neural Network (CMGNN) to uncover the associations among episodes for consistent global embedding. Since the semantic information induced from NLP is relatively fixed compared to visual information space, we leverage it to construct meta nodes for each category to guide the corresponding visual feature learning through GNN. Moreover, to ensure global embedding, a distance loss function is designed to force the visual nodes closer to their associated meta nodes to a greater extent. Extensive experiments and ablation studies on four benchmark datasets show its superiority over many SOTA comparison methods. Shubao Liu, Yuan Xie 0006, Wang Yuan, Lizhuang Ma |
ICME | 3 |
| 2021 | Both Comparison and Induction are Indispensable for Cross-Domain Few-Shot LearningabstractFew-shot learning (FSL), aiming to extract new knowledge from very small amount of labeled samples, has attracted noticeable attentions recently. However, most of existing methods often fail when facing huge domain shift between seen and unseen classes. We think this should be attributed to the episode strategy which ignore utilizing support samples to induct the test classes. So in this paper, for the first time, we propose a bilevel episode strategy (BL-ES) to train a inductive graph network (IGN) that learn to both comparison and induction. Specifically, first, outer episodes in BL-ES simulate the cross-domain few-shot tasks constantly, while inner episodes learn to drive IGN to induct the common features of test classes. Then, the propsoed IGN captures the correlation among all samples to update meta points of each category in induction module. Finally, we introduce a geometrical constraint term utilizing meta points into the training loss, to update the nodes and edges in feature space. This way improves the robustness of training process. Extensive experiments show that our framework outperforms the state-of-the-art FSL alternatives, and are more suitable for real-world applications. Wang Yuan, TianXue Ma, Yuan Xie 0006, Zhizhong Zhang 0001, Lizhuang Ma |
ICME | 1 |
| 2021 | Learn from Concepts: Towards the Purified Memory for Few-shot LearningabstractHuman beings have a great generalization ability to recognize a novel category by only seeing a few number of samples. This is because humans possess the ability to learn from the concepts that already exist in our minds. However, many existing few-shot approaches fail in addressing such a fundamental problem, {\it i.e.,} how to utilize the knowledge learned in the past to improve the prediction for the new task. In this paper, we present a novel purified memory mechanism that simulates the recognition process of human beings. This new memory updating scheme enables the model to purify the information from semantic labels and progressively learn consistent, stable, and expressive concepts when episodes are trained one by one. On its basis, a Graph Augmentation Module (GAM) is introduced to aggregate these concepts and knowledge learned from new tasks via a graph neural network, making the prediction more accurate. Generally, our approach is model-agnostic and computing efficient with negligible memory cost. Extensive experiments performed on several benchmarks demonstrate the proposed method can consistently outperform a vast number of state-of-the-art few-shot learning methods. Xuncheng Liu, Shaohui Lin, Yanyun Qu, Lizhuang Ma, Wang Yuan, Zhizhong Zhang 0001, Yuan Xie 0006 |
IJCAI | 6 |
| 2021 | Multisensor-Based Navigation and Control of a Mobile Service RobotabstractService robot navigation must take the humans into account explicitly so as to produce motion behaviors that reflect its social awareness. Generally, the navigation problems of mobile service robot can be summarized to three aspects: 1) human detection; 2) robot real-time localization; and 3) robot motion planning. The purpose of this paper is to provide a feasible strategy to integrate these three aspects to achieve a conscious, safe, accurate, robust, and efficient navigation. We first introduce the human detection system for recognition of human gesture using a weighted dynamic time warping (DTW) with kinematic constraints. Thus, by interpreting the human body language through gesture recognition, robot motion behaviors like heading to the assigned position or following people can be activated. Then, for the robot localization, a simultaneous localization and mapping (SLAM) method based on artificial and natural landmark recognition is employed to provide absolute position feedback in real time. For the motion planning, a novel quadrupole potential field (QPF) method is proposed to plan collision-free trajectories, adequately considering the nonholomic constraint of the mobile robot system. Then, a robust kinematic controller is designed for trajectory tracking to account for slip disturbances. Such a design automatically merges path finding, trajectory generation, and trajectory tracking in a closed-loop fashion, achieving simultaneous motion planning for obstacle avoidance and feedback stabilization to a desired position and orientation even in the presence of slippage. Finally, experiments prove the effectiveness and feasibility of the proposed strategy, showing a good navigation performance on mobile service robot. Wang Yuan, Zhijun Li 0001, Chun-Yi Su |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Residual Attention Network for Wavelet Domain Super-ResolutionabstractSingle-image super-resolution plays an important role in computer vision area. However, previous works using convolutional neural networks perform badly when reconstructing high frequency details, result in over-smooth and lacking of textural information in the output. At the same time, super-resolution computation always relays on convolutional neural networks with huge depth, which is super tricky to train and use. In this paper, we propose a novel network with better textural details in wavelet domain, which is composed of a feature extract layer, residual channel attention groups (RCAG) and a residual up-sampling layer based on inverse discrete wavelet transform. Channel attention and spatial attention layers are inserted into residual channel and spatial attention blocks (RCSAB), enhancing the learning of high frequency information with attention maps. Composed of a chain of RCSAB and a channel attention layer with short skip connection, RCAG is good at catching long-term high frequency information. Then the feature mapping component is composed of a chain of RCAG. Experiment shows that our method performs better than state-of-the-art methods on benchmark datasets in different scales. Jing Liu 0031, Yuan Xie 0006, Wang Yuan, Lizhuang Ma |
ICASSP | 4 |
| 2019 | Object-Level Salience Detection by Progressively Enhanced Network
Wang Yuan, Xin Tan 0002, Chengwei Chen, Shouhong Ding, Lizhuang Ma |
ICANN (3) | 1 |
| 2018 | Neural-Dynamic Optimization-Based Model Predictive Control for Tracking and Formation of Nonholonomic Multirobot SystemsabstractIn this paper, a neural-dynamic optimization-based nonlinear model predictive control (NMPC) is developed for the multiple nonholonomic mobile robots formation. First, a model-based monocular vision method is developed to obtain the location information of the leader. Then, a separation-bearing-orientation scheme (SBOS) control strategy is proposed. During the formation motion, the leader robot is controlled to track the desired trajectory and the desired leader-follower relationship can be maintained through the SBOS method. Finally, the model predictive control (MPC) is utilized to maintain the desired leader-follower relationship. To solve the MPC generated constrained quadratic programming problem, the neural-dynamic optimization approach is used to search for the global optimal solution. Compared to other existing formation control approaches, the proposed solution is that the NMPC scheme exploit prime-dual neural network for online optimization. Finally, by using several actual mobile robots, the effectiveness of the proposed approach has been verified through the experimental studies. Zhijun Li 0001, Wang Yuan, Fan Ke, Xiaoli Chu, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | A Skin Segmentation Algorithm Based on Stacked AutoencodersabstractA good skin detector that is capable of capturing skin tones under different conditions is important for human-machine interaction applications. In a general situation, skin detectors, such as skin probability maps or Gaussian mixture models, achieve acceptable skin segmentation results. However, the false positive rate increases significantly when the skin tones are in shadow or when skin-like background objects are under similar illumination. In this paper, we propose a novel skin feature learning algorithm based on stacked autoencoders, which are deep neural networks. To overcome the problems encountered in skin segmentation that are caused by different ethnicities and varying illumination conditions, the stacked autoencoders are utilized to learn more discriminative representations of the skin area in both the RGB color space and the HSV color space. Unlike traditional machine learning methods, instead of predicting each pixel individually, our algorithm utilizes blocks to learn the representations and detect the skin areas. The algorithm exploits the learning ability of deep neural networks to learn high-level representations of skin tones. Experiments on test images show that the proposed algorithm achieves acceptable results on several publicly available data sets. To reduce the difficulty of detecting skin pixels in these data sets, the ground truths of these data sets are commonly focused on foreground skin area detection. Our skin detector is also able to detect background areas, as shown in our experiments. You Lei, Wang Yuan, Hongpeng Wang 0002, Wenhu You, Wu Bo |
IEEE Trans. Multim. | 2 |