VLDB 2026 Research / reviewers in the wild / expert
Wenwu Yang
dblp:27/6241
· DBLP profile ↗
22ranked-venue papers
12as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerabstractExisting multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single-person pose estimation. This design relies on heuristic operations such as tracking, RoI cropping, and non-maximum suppression, limiting both accuracy and efficiency. In this paper, we present a fully end-to-end framework for multi-person 2D pose estimation in videos, effectively eliminating heuristic operations. A key challenge is to associate individuals across frames under complex and overlapping temporal trajectories. To address this, we introduce a novel Pose-Aware Video transformEr Network (PAVE-Net), which features a spatial encoder to model intra-frame relations and a spatiotemporal pose decoder to capture global dependencies across frames. To achieve accurate temporal association, we propose a pose-aware attention mechanism that enables each pose query to selectively aggregate features corresponding to the same individual across consecutive frames. Additionally, we explicitly model spatiotemporal dependencies among pose keypoints to improve accuracy. Notably, our approach is the first end-to-end method for multi-frame 2D human pose estimation. Extensive experiments show that PAVE-Net substantially outperforms prior image-based end-to-end methods, achieving a 6.0 mAP improvement on PoseTrack2017, and delivers accuracy competitive with state-of-the-art two-stage video-based approaches, while offering significant gains in efficiency. Yonghui Yu, Jiahang Cai, Xun Wang 0007, Wenwu Yang |
AAAI | 4 |
| 2026 | Multi-Distribution Knowledge Distillation With Wasserstein Distance for Efficient 3D Human ReconstructionabstractRecently, transformer-based Human Mesh Recovery (HMR) from monocular images has achieved remarkable progress by effectively modeling long-range dependencies among body parts. However, the large-scale transformer architectures that drive these state-of-the-art results impose prohibitive computational and storage demands, severely limiting their deployment in real-time or resource-constrained scenarios. Existing knowledge distillation methods offer a potential solution but typically rely on a single distribution, either final outputs or intermediate features, thus failing to capture the diverse and complementary representations inherent in transformer-based HMR models, such as attention patterns. To address these limitations, we propose Multi-dIstribution kNowledge Distillation (MIND), a framework tailored for transformer-based HMR that transfers knowledge from multiple complementary distributions: final outputs for high-level task knowledge, intermediate features for visual representation learning, and internal attention heatmaps to preserve spatial focus patterns. Furthermore, motivated by the inherent conceptual similarity between Generative Adversarial Networks (GANs) and knowledge distillation, we introduce a Wasserstein distillation strategy that leverages the Wasserstein-GAN framework to robustly align these distributions between the teacher and student models. Extensive experiments on Human3.6M, 3DPW, and COCO demonstrate that MIND not only surpasses existing distillation baselines, reducing PA-MPJPE by 3.1mmon 3DPW, but also matches or exceeds the state-of-the-art teacher model HMR2.0 [1] while using only 17.8% of its parameters and 13.7% of its computation cost. Xun Wang 0007, Wenwu Yang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | TwinPose: Person-Specific Subspaces for Multi-View 3D Pose EstimationabstractFollowing the success of deep neural networks in 2D pose estimation, reconstruction-based approaches have significantly advanced multi-person 3D pose estimation from sparse multi-view images. These methods typically detect 2D poses independently in each view and then associate them for 3D reconstruction. However, despite strong progress, recent state-of-the-art methods still face critical limitations: 1) They often depend on global optimization over a large and complex set of multi-view 2D joints to jointly infer 3D poses for all individuals, making the process highly complex and prone to suboptimal solutions; 2) Their tight coupling with the bottom-up detector OpenPose hinders the use of more advanced top-down or single-stage 2D pose estimators and restricts the integration of richer instance-level cues learned by these models. To address these limitations, we propose TwinPose, a novel framework that alleviates the complexity of global pose inference by optimizing within person-specific 3D pose subspaces, while fully supporting diverse 2D pose detectors and effectively leveraging pose-instance cues. The key idea is to introduce a twin pose — a 3D counterpart of each 2D pose — that inherits its instance representation and aggregates geometrically consistent 2D joints from other views. All twin poses are unified in a common 3D space, where those belonging to the same individual naturally share a number of bones. This structural property enables association by counting shared bones, forming person-specific subspaces from which each individual's 3D pose can be inferred independently in an efficient and robust manner. Extensive experiments demonstrate that TwinPose achieves state-of-the-art performance in both accuracy and efficiency across multiple public and proprietary datasets. Importantly, it is fully detector-agnostic, allowing seamless integration with current and future advances in 2D pose estimation while remaining highly robust to noisy or imperfect 2D predictions. Project page with code and additional resources: https://github.com/zgspose/TwinPose Wenwu Yang, Tianyi He, Jiwei Ding, Xun Wang 0007, Kun Zhou 0001 |
ACM Trans. Graph. | 1 |
| 2025 | Context-Aware Academic Emotion Dataset and BenchmarkabstractAcademic emotion analysis plays a crucial role in evaluating students' engagement and cognitive states during the learning process. This paper addresses the challenge of automatically recognizing academic emotions through facial expressions in real-world learning environments. While significant progress has been made in facial expression recognition for basic emotions, academic emotion recognition remains underexplored, largely due to the scarcity of publicly available datasets. To bridge this gap, we introduce RAER, a novel dataset comprising approximately 2,700 video clips collected from around 140 students in diverse, natural learning contexts such as classrooms, libraries, laboratories, and dormitories, covering both classroom sessions and individual study. Each clip was annotated independently by approximately ten annotators using two distinct sets of academic emotion labels with varying granularity, enhancing annotation consistency and reliability. To our knowledge, RAER is the first dataset capturing diverse natural learning scenarios. Observing that annotators naturally consider context cues-such as whether a student is looking at a phone or reading a book-alongside facial expressions, we propose CLIP-CAER (CLIP-based Context-aware Academic Emotion Recognition). Our method utilizes learnable text prompts within the vision-language model CLIP to effectively integrate facial expression and context cues from videos. Experimental results demonstrate that CLIP-CAER substantially outperforms state-of-the-art video-based facial expression recognition methods, which are primarily designed for basic emotions, emphasizing the crucial role of context in accurately recognizing academic emotions. Project page: https://zgsfer.github.io/CAER Luming Zhao, Jingwen Xuan, Jiamin Lou, Yonghui Yu, Wenwu Yang |
ICCV | 5 |
| 2024 | Video-Based Human Pose Regression via Decoupled Space-Time AggregationabstractBy leveraging temporal dependency in video se-quences, multi-frame human pose estimation algorithms have demonstrated remarkable results in complicated sit-uations, such as occlusion, motion blur, and video defocus. These algorithms are predominantly based on heatmaps, re-sulting in high computation and storage requirements per frame, which limits their flexibility and real-time application in video scenarios, particularly on edge devices. In this paper, we develop an efficient and effective video-based hu-man pose regression method, which bypasses intermediate representations such as heatmaps and instead directly maps the input to the output joint coordinates. Despite the inher-ent spatial correlation among adjacent joints of the human pose, the temporal trajectory of each individual joint ex-hibits relative independence. In light of this, we propose a novel Decoupled Space-Time Aggregation network (DSTA) to separately capture the spatial contexts between adja-cent joints and the temporal cues of each individual joint, thereby avoiding the conflation of spatiotemporal dimensions. Concretely, DSTA learns a dedicated feature token for each joint to facilitate the modeling of their spatiotemporal dependencies. With the proposed joint-wise local-awareness attention mechanism, our method is capable of efficiently and flexibly utilizing the spatial dependency of adjacent joints and the temporal dependency of each joint itself. Extensive experiments demonstrate the superiority of our method. Compared to previous regression-based single-frame human pose estimation methods, DSTA significantly enhances performance, achieving an 8.9 mAP improvement on PoseTrack2017. Furthermore, our approach either sur-passes or is on par with the state-of-the-art heatmap-based multi-frame human pose estimation methods. Project page: https://github.com/zgspose/DSTA. Jijie He, Wenwu Yang |
CVPR | 2 |
| 2024 | Empirical Study of Move Smart Contract Security: Introducing MoveScan for Enhanced AnalysisabstractMove, a programming language for smart contracts, stands out for its focus on security. However, the practical security efficacy of Move contracts remains an open question. This work conducts the first comprehensive empirical study on the security of Move contracts. Our initial step involves collaborating with a security company to manually audit 652 contracts from 92 Move projects. This process reveals eight types of defects, with half previously unreported. These defects present potential security risks, cause functional flaws, mislead users, or waste computational resources. To further evaluate the prevalence of these defects in real-world Move contracts, we present MoveScan, an automated analysis framework that translates bytecode into an intermediate representation (IR), extracts essential meta-information, and detects all eight defect types. By leveraging MoveScan, we uncover 97,028 defects across all 37,302 deployed contracts in the Aptos and Sui blockchains, indicating a high prevalence of defects. Experimental results demonstrate that the precision of MoveScan reaches 98.85%, with an average project analysis time of merely 5.45 milliseconds. This surpasses previous state-of-the-art tools MoveLint, which exhibits an accuracy of 87.50% with an average project analysis time of 71.72 milliseconds, and Move Prover, which has a recall rate of 6.02% and requires manual intervention. Our research also yields new observations and insights that aid in developing more secure Move contracts. Shuwei Song, Jiachi Chen, Ting Chen 0002, Xiapu Luo, Wenwu Yang, Leqing Wang, Feng Luo 0009, Zheyuan He |
ISSTA | 6 |
| 2024 | Multi-threshold deep metric learning for facial expression recognitionabstractFeature representations generated through triplet-based deep metric learning offer significant advantages for facial expression recognition (FER). Each threshold in triplet loss inherently shapes a distinct distribution of inter-class variations, leading to unique representations of expression features. Nonetheless, pinpointing the optimal threshold for triplet loss presents a formidable challenge, as the ideal threshold varies not only across different datasets but also among classes within the same dataset. In this paper, we propose a novel multi-threshold deep metric learning approach that bypasses the complex process of threshold validation and markedly improves the effectiveness in creating expression feature representations. Instead of choosing a single optimal threshold from a valid range, we comprehensively sample thresholds throughout this range, which ensures that the representation characteristics exhibited by the thresholds within this spectrum are fully captured and utilized for enhancing FER. Specifically, we segment the embedding layer of the deep metric learning network into multiple slices, with each slice representing a specific threshold sample. We subsequently train these embedding slices in an end-to-end fashion, applying triplet loss at its associated threshold to each slice, which results in a collection of unique expression features corresponding to each embedding slice. Moreover, we identify the issue that the traditional triplet loss may struggle to converge when employing the widely-used Batch Hard strategy for mining informative triplets, and introduce a novel loss termed dual triplet loss to address it. Extensive evaluations demonstrate the superior performance of the proposed approach on both posed and spontaneous facial expression datasets. Wenwu Yang, Jinyi Yu, Tuo Chen, Zhenguang Liu, Xun Wang 0007, Jianbing Shen |
Pattern Recognit. | 1 |
| 2024 | Learning 3D Face Reconstruction From the Cycle-Consistency of Dynamic FacesabstractReconstructinga 3D face from a single image is a crucial task in numerous multimedia applications. Face images with ground-truth 3D face shapes are scarce, so unsupervised deep learning methods, which rely primarily on the free supervision signal derived from the visual disparity between the input image and the rendered counterpart of the predicted 3D face, have proven superior for reconstructing 3D faces. However, it is challenging for such techniques to decouple the dynamic 3D face properties such as pose or expression from a single 2D image, especially when similar local visual appearance changes can be caused by both pose and expression motion, resulting in imprecise 3D face reconstruction. In this article, a novel cycle-consistency in dynamic 3D face characteristics is introduced as a free supervisory signal for learning accurate 3D face shapes from unlabeled facial images. The main idea of cycle-consistency is to explicitly inject the head pose or facial expression variation between video frames into a face image, and then to extract and reverse the injected variation in order to reconstruct the face image to its original state. In our model, a CNN network with multiple branches is proposed to disentangle 3D face properties like identity, expression, pose, and texture from 2D facial images, one branch for each 3D face property. During training, our model learns to completely decouple the dynamic 3D face properties (pose and expression) to be useful for performing cycle-consistent face reconstruction. Extensive experiments demonstrate the superiority of our approach. On the challenging AFLW2000-3D, MICC Florence, and NoW datasets, our method outperforms or is on par with the state of the art. Wenwu Yang, Yeqing Zhao, Bailin Yang, Jianbing Shen |
IEEE Trans. Multim. | 1 |
| 2023 | Demystifying DeFi MEV Activities in Flashbots BundleabstractDecentralized Finance, mushrooming in permissionless blockchains, has attracted a recent surge in popularity. Due to the transparency of permissionless blockchains, opportunistic traders can compete to earn revenue by extracting Miner Extractable Value (MEV), which undermines both the consensus security and efficiency of blockchain systems. The Flashbots bundle mechanism further aggravates the MEV competition because it empowers opportunistic traders with the capability of designing more sophisticated MEV extraction. In this paper, we conduct the first systematic study on DeFi MEV activities in Flashbots bundle by developing ActLifter, a novel automated tool for accurately identifying DeFi actions in transactions of each bundle, and ActCluster, a new approach that leverages iterative clustering to facilitate us to discover known/unknown DeFi MEV activities. Extensive experimental results show that ActLifter can achieve nearly 100% precision and recall in DeFi action identification, significantly outperforming state-of-the-art techniques. Moreover, with the help of ActCluster, we obtain many new observations and discover 17 new kinds of DeFi MEV activities, which occur in 53.12% of bundles but have not been reported in existing studies. Zihao Li 0001, Jianfeng Li 0006, Zheyuan He, Xiapu Luo, Ting Wang 0006, Xiaoze Ni, Wenwu Yang, Ting Chen 0002 |
CCS | 7 |
| 2019 | CR-Morph: Controllable Rigid Morphing for 2D Animation
Wenwu Yang, Jing Hua 0001, Kun-Yang Yao |
J. Comput. Sci. Technol. | 1 |
| 2019 | Building anatomically realistic jaw kinematics model from data
Wenwu Yang, Nathan Marshak, Daniel Sýkora, Srikumar Ramalingam, Ladislav Kavan |
Vis. Comput. | 1 |
| 2018 | FTP-SC: Fuzzy Topology Preserving Stroke CorrespondenceabstractAbstract Stroke correspondence construction is a precondition for vectorized 2D animation inbetweening and remains a challenging problem. This paper introduces the FTP‐SC, a fuzzy topology preserving stroke correspondence technique, which is accurate and provides the user more effective control on the correspondence result than previous matching approaches. The method employs a two‐stage scheme to progressively establish the stroke correspondence construction between the keyframes. In the first stage, the stroke correspondences with high confidence are constructed by enforcing the preservation of the so‐called “fuzzy topology” which encodes intrinsic connectivity among the neighboring strokes. Starting with the high‐confidence correspondences, the second stage performs a greedy matching algorithm to generate a full correspondence between the strokes. Experimental results show that the FTP‐SC outperforms the existing approaches and can establish the stroke correspondence with a reasonable amount of user interaction even for keyframes with large geometric and spatial variations between strokes. Wenwu Yang, Seah Hock Soon, Hong-Ze Liew, Daniel Sýkora |
Comput. Graph. Forum | 1 |
| 2018 | Context-Aware Computer Aided InbetweeningabstractThis paper presents a context-aware computer aided inbetweening (CACAI) technique that interpolates planar strokes to generate inbetween frames from a given set of key frames. The inbetweening is context-aware in the sense that not only the stroke's shape but also the context (i.e., the neighborhood of a stroke) in which a stroke appears are taken into account for the stroke correspondence and interpolation. Given a pair of successive key frames, the CACAI automatically constructs the stroke correspondence between them by exploiting the context coherence between the corresponding strokes. Meanwhile, the construction algorithm is able to incorporate the user's interaction with ease and allows the user more effective control over the correspondence process than existing stroke matching techniques. With a one-to-one stroke correspondence, the CACAI interpolates the shape and context between the corresponding strokes for the generation of intermediate frames. In the interpolation sequence, both the shape of individual strokes and the spatial layout between them are well retained such that the feature characteristics and visual appearance of the objects in the key frames can be fully preserved even when complex motions are involved in these objects. We have developed a prototype system to demonstrate the ease of use and effectiveness of the CACAI. Wenwu Yang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Topology-aware moving least square deformation for 2D charactersabstractAbstract Deformation method based on moving least squares (MLS) allows the user to manipulate 2D characters using either sets of points or line segments in real time. However, the traditional MLS deformation spreads the deformation of the controls with respect to the spatial distance, but oblivious to the shape topology, which would possibly lead to distortion. In this paper, we present a topology‐aware MLS deformation approach for 2D characters. First, a Laplace equation is solved to obtain a set of weights, which are called harmonic weights. Then, the MLS deformation is performed by using the harmonic weights as the deformation influence of the user‐specified controls. Finally, the possible distortion in the traditional MLS deformation can be effectively avoided, as the harmonic weights spread the deformation of the controls in a localized and topology‐aware way. In addition, a simple but effective area‐preserving variant of MLS deformation is proposed, which is suitable for the editing of incompressible objects. Copyright © 2015 John Wiley & Sons, Ltd. Xun Wang 0007, Wenwu Yang, Wangbin Kou, Bailin Yang |
Comput. Animat. Virtual Worlds | 2 |
| 2015 | Recognition of Low-Resolution Logos in Vehicle Images Based on Statistical Random Sparse DistributionabstractTraditional image recognition approaches can achieve high performance only when the images have high resolution and superior quality. A new vehicle logo recognition (VLR) method is proposed to treat low-resolution and poor-quality images captured from urban crossings in intelligent transport system, and the proposed approach is based on statistical random sparse distribution (SRSD) feature and multiscale scanning. The SRSD feature is a novel feature representation strategy that uses the correlation between random sparsely sampled pixel pairs as an image feature and describes the distribution of a grayscale image statistically. Multiscale scanning is a creative classification algorithm that locates and classifies a logo integrally, which alleviates the effect of propagation errors in traditional methods by processing the location and classification separately. Experiments show an overall recognition rate of 97.21% for a set of 3370 vehicle images, which showed that the proposed algorithm outperforms classical VLR methods for low-resolution and inferior quality images and is very suitable for on-site supervision in ITSs. Haoyu Peng, Xun Wang 0007, Huiyan Wang 0002, Wenwu Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2014 | Part-to-part morphing for planar curves
Wenwu Yang, Xun Wang 0007 |
Vis. Comput. | 1 |
| 2013 | Shape-aware skeletal deformation for 2D characters
Xun Wang 0007, Wenwu Yang, Haoyu Peng |
Vis. Comput. | 2 |
| 2012 | One new model based predictive torque control algorithm for doubly salient permanent magnet synchronous machinesabstractThe doubly salient permanent-magnet synchronous machine (DSPMSM) is a new type of brushless machine with permanent magnet locating in its stator pole. Compared with other traditional PMSMs, it can offer advantages of high power/torque density, simple mechanical structure and wide speed range for high speed cruising, which is attractive to the applications of wind energy, plug-in hybrid electrical vehicle, etc. However, due to the nature of salient poles in both the stator and rotor, the DSPMSM suffers from severe torque and flux ripples for its variable magnetic circuits and equivalent air gap length. The conventional switching-table-based direct torque control (DTC) receives increasing attention for its merits of quick dynamic response, strong robustness and simple control structure. However, during the conventional DTC algorithm, large ripple of both torque and air gap flux often occurs for its hysteresis control based on Bang-Bang modification principle. This paper presents one improved strategy to reduce the torque ripple of DSPMSM drive system by the help of model based predictive torque control (MPTC), which is an improved algorithm in the base of conventional DTC. Similar as the traditional MPTC strategy, the new algorithm still requires one completely decoupling control scheme. By selecting the best voltage vector to satisfy the demands of torque and flux, the new method can obviously reduce both torque and flux ripples. Comprehensive simulation results are finally presented to validate theoretical analysis, and further experiments will be available in the near future. Wei Xu 0006, Wenwu Yang, Xinghuo Yu 0001, Jinwei He |
IECON | 2 |
| 2012 | Structure Preserving Manipulation and Interpolation for Multi-element 2D ShapesabstractAbstract This paper presents a method that generates natural and intuitive deformations via direct manipulation and smooth interpolation for multi‐element 2D shapes. Observing that the structural relationships between different parts of a multi‐element 2D shape are important for capturing its feature semantics, we introduce a simple structure called a feature frame to represent such relationships. A constrained optimization is solved for shape manipulation to find optimal deformed shapes under user‐specified handle constraints. Based on the feature frame, local feature preservation and structural relationship maintenance are directly encoded into the objective function. Beyond deforming a given multi‐element 2D shape into a new one at each key frame, our method can automatically generate a sequence of natural intermediate deformations by interpolating the shapes between the key frames. The method is computationally efficient, allowing real‐time manipulation and interpolation, as well as generating natural and visually plausible results. Wenwu Yang, Jieqing Feng, Xun Wang 0007 |
Comput. Graph. Forum | 1 |
| 2009 | 2D shape morphing via automatic feature matching and hierarchical interpolation
Wenwu Yang, Jieqing Feng |
Comput. Graph. | 1 |
| 2009 | 2D shape manipulation via topology-aware rigid gridabstractAbstract This paper presents a new method which allows user to manipulate a two‐dimensional shape in an intuitive and flexible way. The shape is discretized as a regular grid. User places handles on the grid and manipulates the shape by moving the handles to the desired positions. To meet the constraints of the user's manipulation, the grid is then deformed in an as‐rigid‐as‐possible way. However, this straightforward approach tends to produce unnatural deformations when the grid resolution is not high enough to capture the topological structure of the shape. In the proposed method, the regular grid is trimmed and only the cells that are inside the fatty regions of the shape are preserved, namely “interior grid.” When user manipulates the shape, the interior grid and the shape boundary curve are deformed with minimum distortions. To make the deformations of the interior grid and the boundary curve consistent, a junction energy is introduced. In this way, the unnatural deformation effects could be effectively removed and the physically plausible results can be obtained. Meanwhile, the proposed approach provides user an intuitive and simple way to adjust the shape global and local stiffnesses. The deformation is formulated as an energy minimization problem. The energy function is non‐quadratic and could be efficiently solved using an iterative solver with the fast summation technique that exploits the interior grid and boundary curve regularities. In addition, the method could be easily extended to manipulate curves and stick figures. Experimental results demonstrate the capability and flexibility of the new method. Copyright © 2009 John Wiley & Sons, Ltd. Wenwu Yang, Jieqing Feng |
Comput. Animat. Virtual Worlds | 1 |
| 2008 | Shape deformation with tunable stiffness
Wenwu Yang, Jieqing Feng, Xiaogang Jin 0001 |
Vis. Comput. | 1 |