EDBT 2026 Demo / reviewers in the wild / expert
Yuxin Mao
dblp:74/6497
· DBLP profile ↗
44ranked-venue papers
16as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Spatial Decay for Vision TransformersabstractVision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce Spatial Decay Transformer (SDT), featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers. Yuxin Mao, Zhen Qin 0003, Jinxing Zhou, Bin Fan 0002, Jing Zhang 0052, Yiran Zhong, Yuchao Dai |
AAAI | 1 |
| 2026 | CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationabstractThe Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting cross-modal salient anchors, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a Mutual Event Agreement Evaluation module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a Cross-modal Salient Anchor Identification module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an Anchor-based Temporal Propagation module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance. Jinxing Zhou, Yanghao Zhou, Yuxin Mao, Zhangling Duan, Dan Guo 0001 |
AAAI | 4 |
| 2026 | Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual AdaptationabstractMainstream research in audio-visual learning has focused on designing task-specific expert models, primarily implemented through sophisticated multimodal fusion approaches. Recently, a few efforts have aimed to develop more task-independent or universal audiovisual embedding networks, encoding advanced representations for use in various audiovisual downstream tasks. This is typically achieved by fine-tuning large pretrained transformers, such as Swin-V2-L and HTS-AT, in a parameter-efficient manner through techniques such as tuning only a few adapter layers inserted into the pretrained transformer backbone. Although these methods are parameter-efficient, they suffer from significant training memory consumption due to gradient backpropagation through the deep transformer backbones, which limits accessibility for researchers with constrained computational resources. In this paper, we present Meta-Token Learning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-visual tasks. Instead of sequentially modifying the output feature distribution of the transformer backbone, Mettle utilizes a lightweight Layer-Centric Distillation (LCD) module to distill in parallel the intact audio or visual features embedded by each transformer layer into compact meta-tokens. This distillation process considers both pretrained knowledge preservation and task-specific adaptation. The obtained meta-tokens can be directly applied to classification tasks, such as audio-visual event localization and audio-visual video parsing. To further support fine-grained segmentation tasks, such as audio-visual segmentation, we introduce a Meta-Token Injection (MTI) module, which utilizes the audio and visual meta-tokens distilled from the top transformer layer to guide feature adaptation in earlier layers. Extensive experiments on multiple audiovisual benchmarks demonstrate that our method significantly reduces memory usage and training time while maintaining parameter efficiency and competitive accuracy. Jinxing Zhou, Zhihui Li 0001, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li 0001, Yuxin Mao, Mingfei Han 0002, Xiaojun Chang, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | AS-Level Topology Inference for Path-Aware NetworkingabstractAs the frequency of cyber attacks (e.g., DDoS) continues to rise, it becomes increasingly crucial to trace malicious packets and identify which autonomous systems (ASes) they originate from and traverse through. Path-aware networking (PAN) architectures, forwarding packets based on in-packet path identifiers (PIDs), enable end-hosts to gain insights into the AS-level topology and facilitate AS-level path traceback. However, the study on AS-level topology inference and path traceback under PAN has been overlooked, primarily due to the diversity of path identification. Some PAN architectures (e.g., SCION) use deterministic PIDs to identify inter-domain paths, while others (e.g., CoLoR) adopt prefix-deterministic PIDs. In the latter, PID-Prefix is used to identify inter-domain paths, and PID-Suffix is reserved for additional security functions, making the inference problem complicated. This paper focuses on PAN with prefix-deterministic PIDs and investigates how to use in-packet PIDs to infer the AS-level topology. Our goal is to construct a tree-like AS-level topology with the corresponding PID-Prefixes based on in-packet PIDs. We first introduce two macro metrics (i.e., completion rate and precision rate) and two micro metrics (i.e., traceback division and traceback aggregation) to evaluate the inference accuracy. Furthermore, we propose an Alternating Expanding and Checking (AEC) algorithm for AS-level topology inference. AEC algorithm constructs the topology by iteratively expanding and checking the current topology. The expanding phase performs one-hop PID-Prefix inference and generates new child ASes, while the checking phase verifies the consistency of PID sequences with the current topology from the perspective of each new child AS. Experiments based on empirical Internet topology of 201 ASes show that the accuracy of AEC reaches 99.6% when the observer receives only 30 packets from each AS. Yuxin Mao, Hongbin Luo, Zhiyuan Wang 0004, Shan Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Tri-Ergon: Fine-Grained Video-to-Audio Generation with Multi-Modal Conditions and LUFS ControlabstractVideo-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal conditions. To overcome these limitations, we introduce Tri-Ergon, a diffusion-based V2A model that incorporates textual, auditory, and pixel-level visual prompts to enable detailed and semantically rich audio synthesis. Additionally, we introduce Loudness Units relative to Full Scale (LUFS) embedding, which allows for precise manual control of the loudness changes over time for individual audio channels, enabling our model to effectively address the intricate correlation of video and audio in real-world Foley workflows. Tri-Ergon is capable of creating 44.1 kHz high-fidelity stereo audio clips of varying lengths up to 60 seconds, which significantly outperforms existing state-of-the-art V2A methods that typically generate mono audio for a fixed duration. Bingliang Li, Fengyu Yang 0005, Yuxin Mao, Qingwen Ye, Yiran Zhong |
AAAI | 3 |
| 2025 | Towards Open-Vocabulary Audio-Visual Event LocalizationabstractThe Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible. Most research in this field assumes a closed-set setting, which restricts these models’ ability to handle test data containing event categories absent (unseen) during training. Recently, a few studies have explored AVEL in an open-set setting, enabling the recognition of unseen events as "unknown", but without providing category-specific semantics. In this paper, we advance the field by introducing the Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) problem, which requires localizing audio-visual events and predicting explicit categories for both seen and unseen test data at inference. To address this new task, we propose the OV-AVEBench dataset, comprising 24,800 videos across 67 real-life audiovisual scenes (seen:unseen = 46:21), each with manual segment-level annotation. We also establish three evaluation metrics for this task. Moreover, we investigate two baseline approaches, one training-free and one using a further fine-tuning paradigm. Specifically, we utilize the unified multimodal space from the pretrained ImageBind model to extract audio, visual, and textual (event classes) features. The training-free baseline then determines predictions by comparing the consistency of audio-text and visual-text feature similarities. The fine-tuning baseline incorporates lightweight temporal layers to encode temporal relations within the audio and visual modalities, using OVAVEBench training data for model fine-tuning. We evaluate these baselines on the proposed OV-AVEBench dataset and discuss potential directions for future work in this new field. Jinxing Zhou, Dan Guo 0001, Ruohao Guo, Yuxin Mao, Yiran Zhong, Xiaojun Chang, Meng Wang 0001 |
CVPR | 4 |
| 2025 | Generative Transformer for Accurate and Reliable Salient Object DetectionabstractWe explore the impact of transformers on accurate and reliable salient object detection. For accuracy, we integrate the transformer with a deterministic model and delineate its advantages in structural modeling. Regarding reliability, we address the transformer’s tendency to produce overly confident, incorrect predictions. To gauge reliability implicitly, we introduce a latent variable model within the transformer framework, termed the inferential generative adversarial network (iGAN). The stochastic nature of the latent variable facilitates the estimation of predictive uncertainty, which serves as an auxiliary measure of the model’s prediction reliability. Different from the conventional GAN, which defines the distribution of the latent variable as fixed standard normal distribution$\mathcal {N}(0,\mathbf {I})$. The proposed iGAN infers the latent variable by gradient-based Markov Chain Monte Carlo (MCMC), namely Langevin dynamics, leading to an input-dependent latent variable model. We apply our proposed iGAN to fully supervised salient object detection, explaining that iGAN within the transformer framework leads to both accurate and reliable salient object detection. The source code and experimental results are publicly available via our project page:https://npucvr.github.io/TransformerSOD. Yuxin Mao, Jing Zhang 0052, Zhexiong Wan, Aixuan Li, Yunqiu Lv, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Achieving Packet Traceback by Inferring AS-Level Topology Based on Cryptographic Path Identifiers
Hongbin Luo, Shan Zhang 0001, Yuxin Mao, Zhiyuan Wang 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Contrastive Conditional Latent Diffusion for Audio-Visual SegmentationabstractAudio-visual Segmentation (AVS) is conceptualized as a conditional generation task, where audio is considered as the conditional variable for segmenting the sound producer(s). In this case, audio should be extensively explored to maximize its contribution for the final segmentation task. We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio and the final segmentation map is modeled to guarantee the strong correlation between them. To achieve semantic-correlated representation learning, our framework incorporates a latent diffusion model. The diffusion model learns the conditional generation process of the ground-truth segmentation map, resulting in ground-truth aware inference during the denoising process at the test stage. As our model is conditional, it is vital to ensure that the conditional variable contributes to the model output. We thus extensively model the contribution of the audio signal by minimizing the density ratio between the conditional probability of the multimodal data, e.g. conditioned on the audio-visual data, and that of the unimodal data, e.g. conditioned on the audio data only. In this way, our latent diffusion model via density ratio optimization explicitly maximizes the contribution of audio for AVS, which can then be achieved with contrastive learning as a constraint, where the diffusion part serves as the main objective to achieve maximum likelihood estimation, and the density ratio optimization part imposes the constraint. By adopting this latent diffusion model via contrastive learning, we effectively enhance the contribution of audio for AVS. The effectiveness of our solution is validated through experimental results on the benchmark dataset. Code and results are online via our project page: https://github.com/OpenNLPLab/DiffusionAVS. Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yunqiu Lv, Dong Li 0033, Yiran Zhong, Yuchao Dai |
IEEE Trans. Image Process. | 1 |
| 2024 | Improving Audio-Visual Segmentation with Bidirectional GenerationabstractThe aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the contribution of each modality is implicitly or explicitly modeled. Nevertheless, the interconnections between different modalities tend to be overlooked in audio-visual modeling. In this paper, inspired by the human ability to mentally simulate the sound of an object and its visual appearance, we introduce a bidirectional generation framework. This framework establishes robust correlations between an object's visual characteristics and its associated sound, thereby enhancing the performance of AVS. To achieve this, we employ a visual-to-audio projection component that reconstructs audio features from object segmentation masks and minimizes reconstruction errors. Moreover, recognizing that many sounds are linked to object movements, we introduce an implicit volumetric motion estimation module to handle temporal dynamics that may be challenging to capture using conventional optical flow methods. To showcase the effectiveness of our approach, we conduct comprehensive experiments and analyses on the widely recognized AVSBench benchmark. As a result, we establish a new state-of-the-art performance level in the AVS benchmark, particularly excelling in the challenging MS3 subset which involves segmenting multiple sound sources. Code is released in: https://github.com/OpenNLPLab/AVS-bidirectional. Dawei Hao, Yuxin Mao, Xiaodong Han, Yuchao Dai, Yiran Zhong |
AAAI | 2 |
| 2024 | Label-Anticipated Event Disentanglement for Audio-Visual Video Parsing
Jinxing Zhou, Dan Guo 0001, Yuxin Mao, Yiran Zhong, Xiaojun Chang, Meng Wang 0001 |
ECCV (10) | 3 |
| 2024 | TAVGBench: Benchmarking Text to Audible-Video GenerationabstractThe Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field, we have developed a comprehensive Text to Audible-Video Generation Benchmark (TAVGBench), which contains over 1.7 million clips with a total duration of 11.8 thousand hours. We propose an automatic annotation pipeline to ensure each audible video has detailed descriptions for both its audio and video contents. We also introduce the Audio-Visual Harmoni score (AVHScore) to provide a quantitative measure of the alignment between the generated audio and video modalities. Additionally, we present a baseline model for TAVG called TAVDiffusion, which uses a two-stream latent diffusion model to provide a fundamental starting point for further research in this area. We achieve the alignment of audio and video by employing cross-attention and contrastive learning. Through extensive experiments and evaluations on TAVGBench, we demonstrate the effectiveness of our proposed model under both conventional metrics and our proposed metrics. The dataset and code can be found on this page https://npucvr.github.io/TAVGBench/ and on github https://github.com/OpenNLPLab/TAVGBench. Yuxin Mao, Xuyang Shen, Jing Zhang 0052, Zhen Qin 0003, Jinxing Zhou, Mochu Xiang, Yiran Zhong, Yuchao Dai |
ACM Multimedia | 1 |
| 2024 | Mutual Information Regularization for Weakly-Supervised RGB-D Salient Object DetectionabstractIn this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective multimodal representation learning via inter-modal mutual information regularization. In particular, following the principle of disentangled representation learning, we introduce a mutual information upper bound with a mutual information minimization regularizer to encourage the disentangled representation of each modality for salient object detection. Based on our multimodal representation learning framework, we introduce an asymmetric feature extractor for our multimodal data, which is proven more effective than the conventional symmetric backbone setting. We also introduce multimodal variational auto-encoder as stochastic prediction refinement techniques, which takes pseudo labels from the first training stage as supervision and generates refined prediction. Experimental results on benchmark RGB-D salient object detection datasets verify both effectiveness of our explicit multimodal disentangled representation learning method and the stochastic prediction refinement strategy, achieving comparable performance with the state-of-the-art fully supervised models. Our code and data are available at:https://npucvr.github.io/MIRV/. Aixuan Li, Yuxin Mao, Jing Zhang 0052, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Decomposed Guided Dynamic Filters for Efficient RGB-Guided Depth CompletionabstractRGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic filters, which generate spatially-variant depth-wise separable convolutional filters from RGB features to guide depth features, have been proven to be effective in this task. However, the dynamically generated filters require massive model parameters, computational costs and memory footprints when the number of feature channels is large. In this paper, we propose to decompose the guided dynamic filters into a spatially-shared component multiplied by content-adaptive adaptors at each spatial location. Based on the proposed idea, we introduce two decomposition schemes$\mathcal {A}$and$\mathcal {B}$, which decompose the filters by splitting the filter structure and using spatial-wise attention, respectively. The decomposed filters not only maintain the favorable properties of guided dynamic filters as being content-dependent and spatially-variant, but also reduce model parameters and hardware costs, as the learned adaptors are decoupled with the number of feature channels. Extensive experimental results demonstrate that the methods using our schemes outperform state-of-the-art methods on the KITTI dataset, and rank 1st and 2nd on the KITTI benchmark at the time of submission. Meanwhile, they also achieve comparable performance on the NYUv2 dataset. In addition, our proposed methods are general and could be employed as plug-and-play feature fusion blocks in other multi-modal fusion tasks such as RGB-D salient object detection. Yuxin Mao, Qi Liu 0054, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Joint Appearance and Motion Learning for Efficient Rolling Shutter CorrectionabstractRolling shutter correction (RSC) is becoming increasingly popular for RS cameras that are widely used in commercial and industrial applications. Despite the promising performance, existing RSC methods typically employ a two-stage network structure that ignores intrinsic infor-mation interactions and hinders fast inference. In this pa-per, we propose a single-stage encoder-decoder-based network, named JAMNet, for efficient RSC. It first extracts pyramid features from consecutive RS inputs, and then simultaneously refines the two complementary information (i.e., global shutter appearance and undistortion motion field) to achieve mutual promotion in a joint learning de-coder. To inject sufficient motion cues for guiding joint learning, we introduce a transformer-based motion embed-ding module and propose to pass hidden states across pyra-mid levels. Moreover, we present a new data augmentation strategy “vertical flip + inverse order” to release the potential of the RSC datasets. Experiments on various benchmarks show that our approach surpasses the state-of-the-art methods by a large margin, especially with a 4.7 dB PSNR leap on real-world RSC. Code is available at https://github.com/GitCVfb/JAMNet. Bin Fan 0002, Yuxin Mao, Yuchao Dai, Zhexiong Wan, Qi Liu 0054 |
CVPR | 2 |
| 2023 | Multimodal Variational Auto-encoder based Audio-Visual SegmentationabstractWe propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies, where models are trained to fit the discrete samples in the dataset. With a limited and less diverse dataset, the resulting performance is usually unsatisfactory. In contrast, we address this problem from an effective representation learning perspective, aiming to model the contribution of each modality explicitly. Specifically, we find that audio contains critical category information of the sound producers, and visual data provides candidate sound producer(s). Their shared information corresponds to the target sound producer(s) shown in the visual data. In this case, cross-modal shared representation learning is especially important for AVS. To achieve this, our ECMVAE factorizes the representations of each modality with a modality-shared representation and a modality-specific representation. An orthogonality constraint is applied between the shared and specific representations to maintain the exclusive attribute of the factorized latent code. Further, a mutual information maximization regularizer is introduced to achieve extensive exploration of each modality. Quantitative and qualitative evaluations on the AVSBench demonstrate the effectiveness of our approach, leading to a new state-of-the-art for AVS, with a 3.84 mIOU performance leap on the challenging MS3 subset for multiple sound source segmentation. Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yiran Zhong, Yuchao Dai |
ICCV | 1 |
| 2023 | RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow EstimationabstractRecently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiDAR sensors adopt a frame-based data acquisition mechanism, their performance is limited by the fixed low sampling rates, especially in highly-dynamic scenes. By contrast, the event camera can asynchronously capture the intensity changes with a very high temporal resolution, providing complementary dynamic information of the observed scenes. In this paper, we incorporate RGB images, Point clouds and Events for joint optical flow and scene flow estimation with our proposed multi-stage multimodal fusion model, RPEFlow. First, we present an attention fusion module with a cross-attention mechanism to implicitly explore the internal cross-modal correlation for 2D and 3D branches, respectively. Second, we introduce a mutual information regularization term to explicitly model the complementary information of three modalities for effective multimodal feature learning. We also contribute a new synthetic dataset to advocate further research. Experiments on both synthetic and real datasets show that our model outperforms the existing state-of-theart by a wide margin. Code and dataset is available at https://npucvr.github.io/RPEFlow. Zhexiong Wan, Yuxin Mao, Jing Zhang 0052, Yuchao Dai |
ICCV | 2 |
| 2023 | Continuous Parametric Optical FlowabstractIn this paper, we present continuous parametric optical flow, a parametric representation of dense and continuous motion over arbitrary time interval. In contrast to existing discrete-time representations (i.e., flow in between consecutive frames), this new representation transforms the frame-to-frame pixel correspondences to dense continuous flow. In particular, we present a temporal-parametric model that employs B-splines to fit point trajectories using a limited number of frames. To further improve the stability and robustness of the trajectories, we also add an encoder with a neural ordinary differential equation (NODE) to represent features associated with specific times. We also contribute a synthetic dataset and introduce two evaluation perspectives to measure the accuracy and robustness of continuous flow estimation. Benefiting from the combination of explicit parametric modeling and implicit feature optimization, our model focuses on motion continuity and outperforms the flow-based and point-tracking approaches for fitting long-term and variable sequences. Jianqin Luo, Zhexiong Wan, Yuxin Mao, Bo Li 0090, Yuchao Dai |
NeurIPS | 3 |
| 2023 | BTD: An effective business-related hot topic detection scheme in professional social networks
Lujie Zhou, Yuxin Mao, Naixue Xiong |
Inf. Sci. | 2 |
| 2023 | MUNet: Motion uncertainty-aware semi-supervised video object segmentation
Jiadai Sun, Yuxin Mao, Yuchao Dai, Yiran Zhong |
Pattern Recognit. | 2 |
| 2023 | Deep Idempotent Network for Efficient Single Image Blind DeblurringabstractSingle image blind deblurring is highly ill-posed as neither the latent sharp image nor the blur kernel is known. Even though considerable progress has been made, several major difficulties remain for blind deblurring, including the trade-off between high-performance deblurring and real-time processing. Besides, we observe that current single image blind deblurring networks cannot further improve or stabilize the performance but significantly degrades the performance when re-deblurring is repeatedly applied. This implies the limitation of these networks in modeling an ideal deblurring process. In this work, we make two contributions to tackle the above difficulties: (1) We introduce the idempotent constraint into the deblurring framework and present a deep idempotent network to achieve improved blind non-uniform deblurring performance with stable re-deblurring. (2) We propose a simple yet efficient deblurring network with lightweight encoder-decoder units and a recurrent structure that can deblur images in a progressive residual fashion. Extensive experiments on synthetic and realistic datasets prove the superiority of our proposed framework. Remarkably, our proposed network is nearly$6.5\times $smaller and$6.4\times $faster than the state-of-the-art while achieving comparable high performance. Yuxin Mao, Zhexiong Wan, Yuchao Dai, Xin Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Learning Dense and Continuous Optical Flow From an Event CameraabstractEvent cameras such as DAVIS can simultaneously output high temporal resolution events and low frame-rate intensity images, which own great potential in capturing scene motion, such as optical flow estimation. Most of the existing optical flow estimation methods are based on two consecutive image frames and can only estimate discrete flow at a fixed time interval. Previous work has shown that continuous flow estimation can be achieved by changing the quantities or time intervals of events. However, they are difficult to estimate reliable dense flow, especially in the regions without any triggered events. In this paper, we propose a novel deep learning-based dense and continuous optical flow estimation framework from a single image with event streams, which facilitates the accurate perception of high-speed motion. Specifically, we first propose an event-image fusion and correlation module to effectively exploit the internal motion from two different modalities of data. Then we propose an iterative update network structure with bidirectional training for optical flow prediction. Therefore, our model can estimate reliable dense flow as two-frame-based methods, as well as estimate temporal continuous flow as event-based methods. Extensive experimental results on both synthetic and real captured datasets demonstrate that our model outperforms existing event-based state-of-the-art methods and our designed baselines for accurate dense and continuous optical flow estimation. Zhexiong Wan, Yuchao Dai, Yuxin Mao |
IEEE Trans. Image Process. | 3 |
| 2020 | Automatic obstacle avoidance of quadrotor UAV via CNN-based learning
Xi Dai, Yuxin Mao, Tianpeng Huang, Na Qin 0001, Deqing Huang, Yanan Li 0001 |
Neurocomputing | 2 |
| 2018 | Generating hybrid interior structure for 3D printing
Yuxin Mao, Lifang Wu, Dong-Ming Yan 0001, Jianwei Guo 0003, Chang Wen Chen, Baoquan Chen |
Comput. Aided Geom. Des. | 1 |
| 2016 | A game-based incentive model for service cooperation in VANETsabstractSummary Because of the highly dynamic topology and the unstable service status of nodes, services in vehicular ad hoc networks (VANETs) are not always reliable enough for users. Nodes in such a VANET incline to be selfish, which will even enhance this situation. In this work, we present an incentive model for VANETs to support more reliable services in network. We model the situation of service request and response in VANETs by using game theory. We consider the competitive and cooperative relationship between the nodes to formulate the game for VANETs. A contribution measurement is given in order to encourage cooperation during the game. Nodes are encouraged to provide more services to their neighbors in order to acquire more services from other nodes in our model. We also conduct a simulation for the proposed model and give detailed analysis in this work. From the results of the simulation, we argue that we could enable a VANET to support more reliable service by configuring suitable parameters for it. Copyright © 2014 John Wiley & Sons, Ltd. Yuxin Mao, Ping Zhu 0007, Guiyi Wei, Mohammad Mehedi Hassan, M. Anwar Hossain 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | A Game Theoretical Model for Energy-Aware DTN Routing in MANETs with Nodes' Selfishness
Yuxin Mao, Ping Zhu 0007 |
Mob. Networks Appl. | 1 |
| 2013 | Cooperation Dynamics on Collaborative Social Networks of Heterogeneous PopulationabstractIn collaborative social networks (CSNs), autonomous individuals cooperate for their common reciprocity interests. The intrinsic heterogeneity of individuals' capability and willingness makes significant impact on the promotion of cooperation rate. In this paper, we propose a two-phase Heterogeneous Public Goods Game (HPGG) model to study the cooperation dynamics in CSNs. We introduce two factors to represent the heterogeneity of individual behaviors and the benefit-to-cost enhancement of population, respectively. Based on HPGG CSN model, we quantitatively investigate the relationship between cooperation rate and individuals' heterogeneous behaviors from an evolutionary game perspective. Simulations on the population structure of scale-free networks show the evolution of cooperation in CSNs has no-trivial dependence on the individuals' heterogeneous behaviors. Compared with standard PGG and single-phase heterogeneous PGG, HPGG provides a more precise mechanism to promote cooperation rate of CSNs. Finally, data traces collected from real experiments also demonstrate the preciseness of HPGG in formulating the cooperation dynamics on CSNs. Guiyi Wei, Ping Zhu 0007, Athanasios V. Vasilakos, Yuxin Mao, Jun Luo 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2012 | Personalized Services Recommendation Based on Context-Aware QoS PredictionabstractWith the increase of published Web services, it has become a great challenge to recommend service consumers the best services with regard to the quality of services (QoS). Collaborative filtering is often employed to predict the QoS of a specific service to a certain consumer. However, in existing collaborative filtering based service recommendation approaches, the context under which consumers submit a recommendation request is seldom taken into account when filtering similar recommenders and their corresponding experience. In this paper, we propose a new method dubbed CASR (Context-Aware Services Recommendation) by referring to previous service invocation experiences under similar context with the current consumer, which is of great importance in the personalized service recommendation system. First, the proposed algorithm clusters the service invocation records according to the similarity on context properties and selects the cluster that is most similar to the context of current consumer. Then it predicts the QoS of an unused service for current consumer based on the filtered recommendation records by Bayesian inference. Experimental results demonstrate that the proposed approach can significantly improve the accuracy of QoS prediction and service recommendation. Li Kuang, Yingjie Xia, Yuxin Mao |
ICWS | 3 |
| 2009 | Subontology-Based Resource Management for Web-Based e-LearningabstractRecent advances in Web and information technologies have resulted in many e-learning resources. There is an emerging requirement to manage and reuse relevant resources together to achieve on-demand e-learning in the Web. Ontologies have become a key technology for enabling semantic-driven resource management. We argue that to meet the requirements of semantic-based resource management for Web-based e-learning, one should go beyond using domain ontologies statically. In this paper, we provide a semantic mapping mechanism to integrate e-learning databases by using ontology semantics. Heterogeneous e-learning databases can be integrated under a mediated ontology. Taking into account the locality of resource reuse, we propose to represent context-specific portions from the whole ontology as sub-ontologies. We present a sub-ontology-based approach for resource reuse by using an evolutionary algorithm. We also conduct simulation experiments to evaluate the approach with a traditional Chinese medicine e-learning scenario and obtain promising results. Zhaohui Wu 0001, Yuxin Mao, Huajun Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Semantic Web Development for Traditional Chinese Medicine
Zhaohui Wu 0001, Tong Yu 0003, Huajun Chen, Xiaohong Jiang 0002, Chunying Zhou, Yu Zhang 0008, Yuxin Mao, Yi Feng 0004, Aining Yin |
AAAI | 7 |
| 2008 | Information retrieval and knowledge discovery on the semantic web of traditional chinese medicineabstractWe conduct the first systematical adoption of the Semantic Web solution in the integration, management, and utilization of TCM information and knowledge resources. As the results, the largest TCM Semantic Web ontology is engineered as the uniform knowledge representation mechanism; the ontology-based query and search engine is deployed, mapping legacy and heterogeneous relational databases to the Semantic Web layer for query and search across database boundaries; the first global herb-drug interaction network is mapped through semantic integration, and the semantic graph mining methodology is implemented for discovering and interpreting interesting patterns from this network. The platform and underlying methodology are proved effective in TCM-related drug usage, discovery, and safety analysis. Zhaohui Wu 0001, Tong Yu 0003, Huajun Chen, Xiaohong Jiang 0002, Yi Feng 0004, Yuxin Mao, Jingming Tang, Chunying Zhou |
WWW | 6 |
| 2008 | Dynamic sub-ontology evolution for traditional Chinese medicine web ontology
Yuxin Mao, Zhaohui Wu 0001, Wenya Tian, Xiaohong Jiang 0002, William Kwok-Wai Cheung |
J. Biomed. Informatics | 1 |
| 2007 | Towards Semantic e-Science for Traditional Chinese MedicineabstractBACKGROUND: Recent advances in Web and information technologies with the increasing decentralization of organizational structures have resulted in massive amounts of information resources and domain-specific services in Traditional Chinese Medicine. The massive volume and diversity of information and services available have made it difficult to achieve seamless and interoperable e-Science for knowledge-intensive disciplines like TCM. Therefore, information integration and service coordination are two major challenges in e-Science for TCM. We still lack sophisticated approaches to integrate scientific data and services for TCM e-Science. RESULTS: We present a comprehensive approach to build dynamic and extendable e-Science applications for knowledge-intensive disciplines like TCM based on semantic and knowledge-based techniques. The semantic e-Science infrastructure for TCM supports large-scale database integration and service coordination in a virtual organization. We use domain ontologies to integrate TCM database resources and services in a semantic cyberspace and deliver a semantically superior experience including browsing, searching, querying and knowledge discovering to users. We have developed a collection of semantic-based toolkits to facilitate TCM scientists and researchers in information sharing and collaborative research. CONCLUSION: Semantic and knowledge-based techniques are suitable to knowledge-intensive disciplines like TCM. It's possible to build on-demand e-Science system for TCM based on existing semantic and knowledge-based techniques. The presented approach in the paper integrates heterogeneous distributed TCM databases and services, and provides scientists with semantically superior experience to support collaborative research in TCM discipline. Huajun Chen, Yuxin Mao, Xiaoqing Zheng, Yi Feng 0004, Shuiguang Deng, Aining Yin, Chunying Zhou, Jingming Tang, Xiaohong Jiang 0002, Zhaohui Wu 0001 |
BMC Bioinform. | 2 |
| 2006 | Concept Map Model for Web Ontology Exploration
Yuxin Mao, Zhaohui Wu 0001, Huajun Chen, Xiaoqing Zheng |
APWeb | 1 |
| 2006 | RDF/RDFS-based Relational Database IntegrationabstractWe study the problem of answering queries through a RDF/RDFS ontology, given a set of view-based mappings between one or more relational schemas and this target ontology. Particularly, we consider a set of RDFS semantic constraints such as rdfs:subClassof, rdfs:subPropertyof, rdfs:domain, and rdfs:range, which are present in RDF model but neither XML nor relational models. We formally define the query semantics in such an integration scenario, and design a novel query rewriting algorithm to implement the semantics. On our approach, we highlight the important role played by RDF Blank Node in representing incomplete semantics of relational data. A set of semantic tools supporting relational data integration by RDF are also introduced. The approach have been used to integrate 70 relational databases at China Academy of Traditional Chinese Medicine. Huajun Chen, Zhaohui Wu 0001, Yuxin Mao |
ICDE | 4 |
| 2006 | Towards a Semantic Web of Relational Databases: A Practical Semantic Toolkit and an In-Use Case from Traditional Chinese Medicine
Huajun Chen, Yuxin Mao, Jinmin Tang, Chunying Zhou, Aining Yin, Zhaohui Wu 0001 |
ISWC | 4 |
| 2006 | DartGrid: a semantic infrastructure for building database Grid applicationsabstractAbstract In the presence of a Database Grid where a huge number of highly diverse, widely distributed, autonomously managed databases can be involved in a sharing cycle, database tools and middleware should be well suited for schema mediation and query processing in a semantically meaningful way. In this paper, an implemented system called DartGrid is presented. DartGrid is intended to provide a semantic infrastructure for building database grid applications. We explore the essential and fundamental roles played by Resource Description Framework (RDF) semantics for database grids and implement a set of semantically enabled tools and grid services such as semantic browser, semantic mapping tools, ontology service, semantic query service and semantic registration service. We propose an RDF‐View‐based approach for relational schema mediation and describe the view‐based semantic query rewriting algorithm implemented in DartGrid. DartGrid has been used to build a real database grid application for Traditional Chinese Medicine in China. Copyright © 2006 John Wiley & Sons, Ltd. Huajun Chen, Zhaohui Wu 0001, Yuxin Mao, Guozhou Zheng |
Concurr. Comput. Pract. Exp. | 3 |
| 2006 | Dynamic Query Optimization Approach for Semantic Database Grid
Xiaoqing Zheng, Huajun Chen, Zhaohui Wu 0001, Yuxin Mao |
J. Comput. Sci. Technol. | 4 |
| 2005 | Interactive Semantic-Based Visualization Environment for Traditional Chinese Medicine Information
Yuxin Mao, Zhaohui Wu 0001, Huajun Chen, Yumeng Ye |
APWeb | 1 |
| 2005 | Rewriting Queries Using Views for RDF-Based Relational IntegrationabstractWe study the problem of answering queries through a target RDF ontology, given a set of view-based mappings between one or more source relational schemas and this target ontology. We design a novel query rewriting algorithm that can efficiently rewrites a RDF query into relational queries using a set of RDF views. With our approach, we highlight the important role played by RDF blank node in representing incomplete semantics of relational data when integrating them using RDF. Huajun Chen, Zhaohui Wu 0001, Yuxin Mao |
ICTAI | 3 |
| 2005 | Sub-Ontology Evolution for Service Composition with Application to Distributed E-LearningabstractIn order to meet the on-demand requirement of service composition at a large scale, one should go beyond the use of static domain ontologies but allow different focused aspects of the ontologies to be distributed as sub-ontologies. To demonstrate the feasibility of the sub-ontology idea, we suggest a possible implementation using the semantic Web technology and apply it to distributed e-learning. Yuxin Mao, William Kwok-Wai Cheung, Zhaohui Wu 0001, Jiming Liu 0001 |
ICTAI | 1 |
| 2005 | Context-Based Web Ontology Service for TCM Information SharingabstractWeb ontologies as the foundation of the semantic Web were proposed to solve the problem of integrating and sharing heterogeneous information resources in the Web. Massive amount of domain specific ontologies have been constructed and published in different domains on the Web. However, several limitations make existent ontologies not suitable for high-level and large-scale applications. In this paper, we described a context-based ontology service with a large-scale traditional Chinese medicine (TCM) ontology, which provides clients with an interactive interface and intelligent inter-operations to assist users in sharing and exploiting large-scale TCM information and can be used as a semantic view for domain specific problem solving. Yuxin Mao |
ICWS | 1 |
| 2005 | RDF-Based Ontology View for Relational Schema Mediation in Semantic Web
Huajun Chen, Zhaohui Wu 0001, Yuxin Mao |
KES (2) | 3 |
| 2005 | An Interactive Visual Model for Web Ontologies
Yuxin Mao, Zhaohui Wu 0001, Huajun Chen, Xiaoqing Zheng |
KES (2) | 1 |