EDBT 2026 Demo / reviewers in the wild / expert
Yanhua Yang
dblp:123/2397
· DBLP profile ↗
44ranked-venue papers
11as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposing Prompts, Composing Actions: A Multi-Granularity Prompting Approach for Incremental Action LearningabstractContinual learning for action recognition is a critical capability for next-generation Extended Reality (XR) systems. Yet it faces a severe real-world challenge: strict user privacy that prohibits data rehearsal. While recent prompt-based continual learning methods show promise, we argue their core 'flat,' single-granularity design fundamentally misaligns with the complexity of human actions. This monolithic architecture fails to model the inherent hierarchical structure and overlooks standard action primitives shared across tasks, resulting in suboptimal performance and hindered knowledge transfer. To overcome this limitation, we propose DPCA, a novel spatio-temporal continual learning framework with multi-granularity adaptive prompting. DPCA learns three synergistic components to resolve this mismatch. First, the task-specific prompter employs a multi-granularity query system to capture the unique, compositional semantics of each action. Second, the task-agnostic prompter learns a globally shared vocabulary of ``action primitives," providing a stable and generalizable knowledge base to mitigate catastrophic forgetting. Finally, we introduce a Dissimilarity Attention Rectification at each granularity level, leveraging a reverse attention mechanism to model class-agnostic background information and effectively alleviating overfitting. The synergy between these components enables robust model adaptation without requiring access to past data. Rigorous experiments on multiple large-scale benchmarks (including NTU RGB+D), under a strict rehearsal-free, few-shot protocol, confirm that DPCA establishes a new state-of-the-art. This advance paves the way for the realization of truly adaptive and privacy-respecting XR systems. Xinyi Cheng, Jiexi Yan, Yanhua Yang |
AAAI | 5 |
| 2026 | Mix-QSAM2: Mixed-Precision Quantization for High Fidelity Segmentation in Resource Constrained ScenariosabstractThe Segment Anything Model 2 (SAM2) has established a new benchmark for high-precision image and video segmentation, offering significant potential for a wide range of computer vision tasks. Despite its impressive performance, the model's substantial computational and memory requirements present a significant obstacle to its practical deployment on resource-constrained devices. In this paper, we introduce a novel framework for optimizing SAM2 through two synergistic, importance-driven strategies: quantization and memory management. Specifically, an Importance-driven Mixed-Precision Quantization scheme, which analyzes the sensitivity of each layer using a Weight-Activation Importance Score, is employed to enable a targeted bit-width assignment, preserving model accuracy by keeping critical layers at higher precision. Then, the Selective Importance-driven Synthesis (SIS) mechanism is proposed to address the inefficient accumulation of redundant data in the memory bank. SIS intelligently compresses the memory by identifying the most contextually similar historical frames and synthesizing them into a single, representative feature, thereby preserving informational diversity while enhancing temporal context understanding. Extensive experiments on the COCO and SA-V benchmarks validate our approach, showing that our optimized model consistently outperforms state-of-the-art quantization methods. Our work provides a principled framework for the co-design of quantization and dynamic memory management, offering a practical path toward deploying powerful video segmentation models in real-world applications. Yuzhe Duan, Xuanxuan Ren, Guizhe Dong, Xu Yang 0019, Yanhua Yang |
AAAI | 5 |
| 2026 | Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset DistillationabstractDataset distillation has achieved remarkable progress as an effective approach for data compression. However, real-world data often comes from diverse domains, leading to potential mismatches between the domains of synthesized images and those of the evaluation set. Existing methods primarily assume domain alignment between them, which limits their generalization ability in the above cross-domain scenarios. In this paper, we aim to ensure that images synthesized from known domains maintain robust performance on unseen domains and propose a novel framework called Channel-masked Asymmetric Distribution Matching (CADM). During asymmetric distribution matching, domain-sensitive channels of real data are selectively masked at different layers to extract domain-invariant features that guide synthetic data optimization. To further improve synthetic data representation, we introduce a class-focused domain-agnostic regularization to capture class-relevant knowledge while ignoring domain-specific information. Experiments show that our method produces domain-robust synthetic data and substantially improves generalization performance on unseen domains. Jiexi Yan, Guangtao Lyu, Erkun Yang, Guihai Chen, Yanhua Yang |
AAAI | 7 |
| 2026 | Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If CalibratedabstractDespite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world. Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002 |
AAAI | 6 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Counterfactual Risk Minimization for Out-of-Distribution GeneralizationabstractThe out-of-distribution (OOD) property in data is deemed as one main challenge hindering the generalization ability of machine learning algorithms. However, the underlying reasons for this property remain an intriguing and open question that has yet to be fully understood. In this paper, we seek to enhance our understanding of the OOD phenomenon by framing it as a problem of distribution shift and addressing it through two complementary causal perspectives. The first is a generative causal view that elucidates the data generation process. We introduce a novel three-dimensional coordinate system to represent three fundamental distribution shifts, illustrating their role in various OOD generalization problems. The second is an anti-causal view that focuses on the model learning process. We develop an effective approach dubbed Counterfactual Risk Minimization (CRM) to address arbitrary distribution shifts in a unified framework. Additionally, we introduce a new multi-domain visual recognition dataset called CONA to facilitate further exploration of OOD generalization. We conduct evaluations of CRM alongside several state-of-the-art competitors on four benchmark datasets under the three distribution shifts. The results not only affirm CRM's superiority but also shed light on potential future directions. Code and data: https://github.com/muliyangm/CRM. Yanhua Yang, Muli Yang, Henan Wang, Cheng Deng 0002, Hongyuan Zhu 0002 |
IEEE Trans. Image Process. | 1 |
| 2026 | Future-Trend-Aware Filter-Based PD-MRAC Method for Quadrotors With Unknown Strong DisturbancesabstractRobust flight in complex and windy environments is critical for both single and multiple quadrotors. Existing methods either learn disturbance model at high computational cost or use error-based adaptive control with a speed-stability trade-off that makes tuning difficult. To address these issues, this paper proposes a future-trend-aware filter-based PD-MRAC (Proportional-Derivative Model Reference Adaptive Control) for single quadrotor and a distributed PD-MRAC for multiple quadrotor formation. By embedding a trend-aware derivative term in the adaptive update laws, the controller obtains anticipatory information about the error evolution, enabling rapid adaptation while mitigating oscillations. For more disturbance-sensitive multi-quadrotors, we design a robust distributed protocol under a directed graph, improving resilience to disturbances. The approach maintains low computational cost and supports fast adaptive updates. Extensive simulations and real-world experiments validate improvement. For single quadrotor, RMSE reduced by around 57% versus the baselines and by around 12% versus the DJI Mavic 2. For multi-quadrotors, formation results show enhanced robustness in simulation and effective real-world indoor/outdoor experiments under strong winds. Our project page is athttps://xiongtao-shi.github.io/PD-MRAC/. Yanhua Yang, Chenxin Yu, Xiongtao Shi, Changchun Hua, James Lam, Youmin Gong, Jie Mei 0002 |
IEEE Trans. Robotics | 1 |
| 2026 | High-Speed AAV-Assisted OTFS-Enabled Intelligent Data Collection in Large-Scale Wireless Sensor NetworksabstractSixth-generation (6G) communication emphasizes the deep integration of sensing, communication, and computing to support intelligent and rapid-response networks. Autonomous aerial vehicles (AAVs), known for their superior flexibility, terrain adaptability, and low deployment costs, are promising candidates for data collection in large-scale wireless sensor networks (WSNs). However, many existing AAVs-assisted data collection studies assume that AAVs operate at relatively low speeds and incorporate hovering time during data collection. In such scenarios, the AAVs inevitably require longer flying durations and consume much energy. Additionally, they often neglect the impact of the Doppler effect during the data collection. In most general and realistic scenarios, AAVs typically fly at high speeds without the need to hover, and the Doppler effect highly impacts the communication between sensor nodes (SNs) and AAVs. To address this, we propose a data collection framework that leverages orthogonal time frequency space (OTFS) modulation and non-orthogonal multiple access (NOMA) to mitigate the Doppler-induced interference in the up-link. We formulate an AAV-assisted data collection efficiency maximization problem by jointly considering AAV energy consumption, and the SNs’ uploading rates and bit error rates (BERs). Given the NP-hard nature of this problem, we design a three-step solution: the first two steps employ heuristic algorithms and the third step integrates a bi-directional long short-term memory (BiLSTM) for intelligent AAV symbol detection. Simulation results validate the superiority of our proposed solution. Jiujia Yin, Xilong Liu, Nirwan Ansari, Yanhua Yang |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Detecting Open World Objects via Partial Attribute AssignmentabstractDespite being trained on massive data, today’s vision foundation models still fall short in detecting open world objects. Apart from recognizing known objects from training, a successful Open World Object Detection (OWOD) system must also be able to detect unknown objects never seen before, without confusing them with the backgrounds. Unlike prevailing prior works that rely on probability models to learn "objectness", we focus on learning fine-grained, class-agnostic attributes, allowing the detection of both known and unknown objects in an explainable manner. In this paper, we propose Partial Attribute Assignment (PASS), aiming to automatically select and optimize a small, relevant subset of attributes from a large attribute pool. Specifically, we model attribute selection as a Partial Optimal Transport (POT) problem between known visual objects and the attribute pool, in which more relevant attributes signify more transported mass. PASS follows a curriculum schedule that progressively selects and optimizes a targeted subset of attributes during training, promoting stability and accuracy. Our method enjoys end-to-end optimization by minimizing the POT distance and the classification loss on known visual objects, demonstrating high training efficiency and superior OWOD performance among extensive experimental evaluations.‡ Muli Yang, Gabriel James Goenawan, Huaiyuan Qin, Xi Peng 0001, Yanhua Yang, Hongyuan Zhu 0002 |
CVPR | 6 |
| 2025 | A Novel Efficient Lightweight Multi-scale Network for Apple Leaf Disease Identification
Sheng Pei, Yanhua Yang |
ICIC (5) | 2 |
| 2025 | Q-MiniSAM2: A Quantization-based Benchmark for Resource-Efficient Video SegmentationabstractSegment Anything Model 2 (SAM2) is a new-generation, high-precision model for image and video segmentation, offering extensive application prospects across numerous computer vision fields. However, as a large-scale model, its huge memory demands and expansive computing costs pose challenges for practical deployment. This paper presents Q-MiniSAM2, an efficient Quantization-based segmentation benchmark tailored to optimize SAM2 by Minimizing memory consumption and accelerating computations. We begin with applying Post-Training Quantization (PTQ) to SAM2, requiring only a relatively small dataset for network calibration, thereby eliminating the need for retraining. Building upon PTQ, we further introduce a Hierarchy-based Video Quantization method to enhance the model’s capacity to capture video semantics and temporal correlations across different time scales. Furthermore, we observe that SAM2’s memory overhead is predominantly concentrated on processing historical frames, and the redundant cross-attention computations significantly increase memory and computational costs due to the imperceptible change of the short time intervals between these frames. To tackle this issue, an Adaptive Mutual-KV mechanism is proposed to mitigate excessive cross-attention by leveraging inter-frame similarities. Comprehensive experiments demonstrate that the proposed approach achieves superior performance compared to state-of-the-art methods, underscoring its potential for efficient and scalable video segmentation. Xuanxuan Ren, Xu Yang 0019, Yanhua Yang |
IJCAI | 5 |
| 2025 | Risk-Aware Informative Path Planning for Information Gathering of a 3D SurfaceabstractSurface information acquisition by robots faces challenges such as sensor uncertainty, limited resources, and dynamic environment, all of which often lead to reduced collection efficiency and accuracy. To address these issues, this paper proposes a Risk-aware Informative Path Planning (RIPP) framework. The framework is capable of adaptively selecting the target region according to the mutual information between the expected detection viewpoints. The uncertainty risk caused by noisy sensing can be effectively managed using the Conditional Value at Risk (CVaR)-based method. Therefore, a CVaR-based Greedy Algorithm (CGA) is proposed to select the optimal set of inspection viewpoints. To further enhance information acquisition efficiency, the drone’s path is optimized using a novel Adaptive Fractional Particle Swarm Optimization (AFPSO) algorithm. This approach enables the drone to autonomously select trajectories rich in high-value information. This framework is evaluated in the context of 3D surface temperature inspection of large storage tanks. Simulation and experimental results show that RIPP framework significantly reduces information reconstruction errors in tank surface inspections by demonstrating clear advantages over existing methods. The effectiveness and feasibility of RIPP framework in surface information acquisition task are verified, which provides a new solution for efficient monitoring in complex environment. Mengfei Xu, Yang Chen 0032, Mian Hu, Yanhua Yang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Consensus of multiagent systems via a distributed event-triggered intermittent control
Yawen Zhou, Yanhua Yang, Yufeng Zhou 0003, Li Chai 0008 |
Inf. Sci. | 2 |
| 2024 | Multi-Relational Deep Hashing for Cross-Modal SearchabstractDeep cross-modal hashing retrieval has recently made significant progress. However, existing methods generally learn hash functions with pairwise or triplet supervisions, which involves learning the relevant information by splicing partial similarity between data pairs; notably, this approach only captures the data similarity locally and incompletely, resulting in sub-optimal retrieval performance. In this paper, we propose a novel Multi-Relational Deep Hashing (MRDH) approach, which can fully bridge the modality gap by comprehensively modeling the similarity relationship between data in different modalities. In more detail, to investigate the inter-modal relationships, we constrain the consistency of cross-modal pairwise similarities to maintain the semantic similarity across modalities. Moreover, to further capture complete similarity information, we design a new similarity metric, which we term cross-modal global similarity, by encouraging hash codes of similar data pairs from different modalities to approach a common center and hash codes for dissimilar pairs to converge to different centers. Adopting this approach enables our model to generate more discriminative hash codes. Extensive experiments on three benchmark datasets demonstrate the superiority of our method on cross-modal hashing retrieval. Erkun Yang, Yanhua Yang, Cheng Deng 0002 |
IEEE Trans. Image Process. | 3 |
| 2024 | Implicit Compositional Generative Network for Length-Variable Co-Speech Gesture SynthesisabstractCo-speech gesture synthesis is a practical yet challenging task that aims to generate body motion sequences in line with speech audio. Most of the existing methods can only generate the gesture sequence with a fixed number of frames, which does not satisfy the high-quality requirement of the virtual speech video in real-world applications. In this paper, we propose a novel Implicit Compositional Generative Network (ICGN) for length-variable co-speech gesture synthesis. In ICGN, the implicit neural representation is captured and optimized for a whole gesture sequence of arbitrary length with temporal embeddings. Moreover, to enforce the synthesized gestures more realistic and consistent, we compositionally generate the gesture sequence through a well-designed asymmetric two-stream network that effectively captures and utilizes the rich correlations between speech audio and human body motions. In this way, the coarse and fine-grained gestures are synthesized, respectively, according to the corresponding content-aware and emotion-aware audio components. Extensive experiments on four widely-used benchmarks demonstrate that the proposed method renders realistic human gestures and achieves the superior performance against several state-of-the-art methods. Jiexi Yan, Yanhua Yang, Cheng Deng 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Dual-Stream Contrastive Learning for Compositional Zero-Shot RecognitionabstractCompositional Zero-Shot learning (CZSL) requires recognizing unseen attribute-object compositions using observed visual primitives attributes and objects in a training set, which is a critical capacity for learning systems because the long tail of new combinations dominates the distribution in the real world. However, CZSL is a challenging problem because learning systems tend to learn the dependencies between objects and attributes, which is not conducive to composition classification, and incorrect dependencies will mislead the classification of new combinations of known attributes and objects. This paper primarily introduces a novel yet effective dual-stream contrastive learning method with two main objectives: making the learned representations discriminative and transferring knowledge more efficiently from seen to unseen compositions. Specifically, we generate positive and negative pairs based on the similarity of different concepts (attributes and objects), independently capturing the discriminative representations of concepts. Meanwhile, unlike existing contrastive methods that select negative samples randomly, we construct confusable compositional representations as the negatives to explore the intrinsic relevance between attributes and objects, which can improve the generalization from seen to unseen compositions. Experimental results on two benchmarks show that the proposed method outperforms the state-of-the-arts. Yanhua Yang, Xu Yang 0019, Cheng Deng 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | CrossFormer: Cross-Modal Representation Learning via Heterogeneous Graph TransformerabstractTransformers have been recognized as powerful tools for various cross-modal tasks due to their superior ability to perform representation learning through self-attention. Existing transformer-based cross-modal models can be categorized into single-stream and dual-stream ones. By performing fine-grained interaction with self-attention on the cross-modal concatenated features, the former can simultaneously learn intra- and inter-modal correlations. However, this simple concatenation treats the inputs of different modalities equally; as a result, the heterogeneous differences between modalities are ignored, leading to a modality gap. The latter process the inputs of different modalities separately, then perform cross-modal interaction on the subsequently fused networks, resulting in a failure to integrate the fine-grained correlations of both intra- and inter-modality in a uniform module. To this end, we propose an effective heterogeneous graph transformer for dual-stream cross-modal representation learning, named CrossFormer, which constructs a heterogeneous graph as a bridge to achieve fine-grained intra- and inter-modal interaction on a dual-stream network. Specifically, we first represent multi-modal data with a heterogeneous graph, then develop a dual-positional encoding strategy that enables the heterogeneous graph to obtain the relative positional information. Finally, a dual-stream self-attention is performed on the heterogeneous graph, bridging the gap between modalities and effectively capturing fine-grained intra- and inter-modal interactions simultaneously. Extensive experiments on various cross-modal tasks demonstrate the superiority of our method. Erkun Yang, Cheng Deng 0002, Yanhua Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Fully Distributed Event-Triggered Consensus of MIMO MASs With Parametric Uncertainties and External DisturbancesabstractThis article studies the consensus problem of a class of multi-input–multi-output (MIMO) multiagent systems (MASs) subject to parametric uncertainties and external disturbances via a fully distributed model reference adaptive event-triggered control (MRA-ETC) protocol. Incorporate both the MRA and ETC, a reference model using the predicted relative state information initialized by intermittently collected state information of neighbors as input and a self-contained adaptive estimator of uncertainties are constructed, where the transmitted information is asynchronous and intermittent. We consider both the cases of matched and unmatched external disturbances, where each agent is assigned a reference model to track. Asymptotic consensus is achieved for the case of matched disturbances by using the sliding-mode control, while for the case of unmatched disturbances, the uniformly ultimately bound consensus is obtained via the adaptive$\sigma$-modification technique. The communication resources have been significantly saved via the proposed event-triggered mechanism (ETM) with strictly excluding the Zeno behavior. Moreover, with the help of the designed adaptive control gains for the reference models, the consensus algorithm can be implemented in a fully distributed fashion without employing any global information of the MASs. Finally, the simulation examples are illustrated to show the correctness of the proposed control schemes. Yanhua Yang, Jie Mei 0002, Ai-Guo Wu 0001, Guangfu Ma |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Learning with Diversity: Self-Expanded Equalization for Better Generalized Deep Metric LearningabstractExploring good generalization ability is essential in deep metric learning (DML). Most existing DML methods focus on improving the model robustness against category shift to keep the performance on unseen categories. However, in addition to category shift, domain shift also widely exists in real-world scenarios. Therefore, learning better generalization ability for the DML model is still a challenging yet realistic problem. In this paper, we propose a new self-expanded equalization (SEE) method to effectively generalize the DML model to both unseen categories and domains. Specifically, we take a ‘min-max’ strategy combined with a proxy-based loss to adaptively augment diverse out-of-distribution samples that vastly expand the span of original training data. To take full advantage of the implicit cross-domain relations between source and augmented samples, we introduce a domain-aware equalization module to induce the domain-invariant distance metric by regularizing the feature distribution in the metric space. Extensive experiments on two benchmarks and a large-scale multi-domain dataset demonstrate the superiority of our SEE over the existing DML methods. Jiexi Yan, Zhihui Yin, Erkun Yang, Yanhua Yang, Heng Huang 0001 |
ICCV | 4 |
| 2023 | Frequency domain regularization for iterative adversarial attacks
Tengjiao Li, Maosen Li, Yanhua Yang, Cheng Deng 0002 |
Pattern Recognit. | 3 |
| 2023 | Adaptive Bias-Aware Feature Generation for Generalized Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to recognize unseen classes that never appear during training. Recently, generative adversarial networks (GANs) have been introduced to convert ZSL into a supervised learning problem by synthesizing unseen visual features. However, since unseen classes are never experienced for the generator during training, the synthesized unseen visual features often become heavily biased towards seen classes, or sometimes there is even no meaningful class that can be assigned to them. This is known as thebias problem. In this paper, we propose a novel method, namely Adaptive Bias-Aware GAN (ABA-GAN), to alleviate generating biased visual features. For this purpose, we build a semantic adversarial network to regularize the feature generator. Specifically, an adaptive adversarial loss is proposed to constrain the feature distributions, which avoids the generation of meaningless visual features. Meanwhile, a domain divider is presented to explicitly distinguish synthesized visual features between seen and unseen domains, such that the bias towards seen classes can be alleviated. Moreover, we propose a novel metric named bias score (BS) to explicitly quantify the degree of the strong bias. Extensive experiments on four widely used benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art approaches under both ZSL and GZSL protocols. Yanhua Yang, Xiaozhe Zhang, Muli Yang, Cheng Deng 0002 |
IEEE Trans. Multim. | 1 |
| 2022 | Learning Universal Adversarial Perturbation by Adversarial ExampleabstractDeep learning models have shown to be susceptible to universal adversarial perturbation (UAP), which has aroused wide concerns in the community. Compared with the conventional adversarial attacks that generate adversarial samples at the instance level, UAP can fool the target model for different instances with only a single perturbation, enabling us to evaluate the robustness of the model from a more effective and accurate perspective. The existing universal attack methods fail to exploit the differences and connections between the instance and universal levels to produce dominant perturbations. To address this challenge, we propose a new universal attack method that unifies instance-specific and universal attacks from a feature perspective to generate a more dominant UAP. Specifically, we reformulate the UAP generation task as a minimax optimization problem and then utilize the instance-specific attack method to solve the minimization problem thereby obtaining better training data for generating UAP. At the same time, we also introduce a consistency regularizer to explore the relationship between training data, thus further improving the dominance of the generated UAP. Furthermore, our method is generic with no additional assumptions about the training data and hence can be applied to both data-dependent (supervised) and data-independent (unsupervised) manners. Extensive experiments demonstrate that the proposed method improves the performance by a significant margin over the existing methods in both data-dependent and data-independent settings. Code is available at https://github.com/lisenxd/AT-UAP. Maosen Li, Yanhua Yang, Xu Yang 0019, Heng Huang 0001 |
AAAI | 2 |
| 2022 | Progressive Self-Attention Network with Unsymmetrical Positional Encoding for Sequential RecommendationabstractIn real-world recommendation systems, the preferences of users are often affected by long-term constant interests and short-term temporal needs. The recently proposed Transformer-based models have proved superior in the sequential recommendation, modeling temporal dynamics globally via the remarkable self-attention mechanism. However, all equivalent item-item interactions in original self-attention are cumbersome, failing to capture the drifting of users' local preferences, which contain abundant short-term patterns. In this paper, we propose a novel interpretable convolutional self-attention, which efficiently captures both short- and long-term patterns with a progressive attention distribution. Specifically, a down-sampling convolution module is proposed to segment the overall long behavior sequence into a series of local subsequences. Accordingly, the segments are interacted with each item in the self-attention layer to produce locality-aware contextual representations, during which the quadratic complexity in original self-attention is reduced to nearly linear complexity. Moreover, to further enhance the robust feature learning in the context of Transformers, an unsymmetrical positional encoding strategy is carefully designed. Extensive experiments are carried out on real-world datasets, \eg ML-1M, Amazon Books, and Yelp, indicating that the proposed method outperforms the state-of-the-art methods w.r.t. both effectiveness and efficiency. Yuehua Zhu, Shaohua Jiang, Muli Yang, Yanhua Yang, Leon Wenliang Zhong |
SIGIR | 5 |
| 2021 | Understanding and Improving Early Stopping for Learning with Noisy LabelsabstractThe memorization effect of deep neural network (DNN) plays a pivotal role in many state-of-the-art label-noise learning methods. To exploit this property, the early stopping trick, which stops the optimization at the early stage of training, is usually adopted. Current methods generally decide the early stopping point by considering a DNN as a whole. However, a DNN can be considered as a composition of a series of layers, and we find that the latter layers in a DNN are much more sensitive to label noise, while their former counterparts are quite robust. Therefore, selecting a stopping point for the whole network may make different DNN layers antagonistically affect each other, thus degrading the final performance. In this paper, we propose to separate a DNN into different parts and progressively train them to address this problem. Instead of the early stopping which trains a whole DNN all at once, we initially train former DNN layers by optimizing the DNN with a relatively large number of epochs. During training, we progressively train the latter DNN layers by using a smaller number of epochs with the preceding layers fixed to counteract the impact of noisy labels. We term the proposed method as progressive early stopping (PES). Despite its simplicity, compared with the traditional early stopping, PES can help to obtain more promising and stable results. Furthermore, by combining PES with existing approaches on noisy label training, we achieve state-of-the-art performance on image classification benchmarks. The code is made public at https://github.com/tmllab/PES. Yingbin Bai, Erkun Yang, Bo Han 0003, Yanhua Yang, Yinian Mao, Gang Niu 0001, Tongliang Liu |
NeurIPS | 4 |
| 2021 | Multi-models and dual-sampling periods quality prediction with time-dimensional K-means and state transition-LSTM network
Xiongtao Shi, Yonggang Li 0002, Yanhua Yang, Bei Sun, Fang Qi |
Inf. Sci. | 3 |
| 2021 | Bipartite consensus of double-integrator multi-agent systems with nonuniform communication time delays
Wenfeng Hu, Yanhua Yang, Guo Chen 0002, Min Meng 0003 |
Neural Comput. Appl. | 2 |
| 2021 | Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image SynthesisabstractDue to the development of Generative Adversarial Networks (GANs), significant progress has been achieved in text-to-image synthesis task. However, most previous works have only focus on learning the semantic consistency between paired images and sentences, without exploring the semantic correlation between different yet related sentences that describe the same image, which leads to significant visual variation among the synthesized images. Accordingly, in this article, we propose a new method for text-to-image synthesis, dubbed Multi-sentence Auxiliary Generative Adversarial Networks (MA-GAN); this approach not only improves the generation quality but also guarantees the generation similarity of related sentences by exploring the semantic correlation between different sentences describing the same image. More specifically, we propose a Single-sentence Generation and Multi-sentence Discrimination (SGMD) module that explores the semantic correlation between multiple related sentences in order to reduce the variation between their generated images and enhance the reliability of the generated results. Moreover, a Progressive Negative Sample Selection mechanism (PNSS) is designed to mine more suitable negative samples for training, which can effectively promote detailed discrimination ability in the generative model and facilitate the generation of more fine-grained results. Extensive experiments on Oxford-102 and CUB datasets reveal that our MA-GAN significantly outperforms the state-of-the-art methods. Yanhua Yang, Lei Wang 0018, De Xie, Cheng Deng 0002, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2020 | Progressive Domain-Independent Feature Decomposition Network for Zero-Shot Sketch-Based Image RetrievalabstractZero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a specific cross-modal retrieval task for searching natural images given free-hand sketches under the zero-shot scenario. Most existing methods solve this problem by simultaneously projecting visual features and semantic supervision into a low-dimensional common space for efficient retrieval. However, such low-dimensional projection destroys the completeness of semantic knowledge in original semantic space, so that it is unable to transfer useful knowledge well when learning semantic features from different modalities. Moreover, the domain information and semantic information are entangled in visual features, which is not conducive for cross-modal matching since it will hinder the reduction of domain gap between sketch and image. In this paper, we propose a Progressive Domain-independent Feature Decomposition (PDFD) network for ZS-SBIR. Specifically, with the supervision of original semantic knowledge, PDFD decomposes visual features into domain features and semantic ones, and then the semantic features are projected into common space as retrieval features for ZS-SBIR. The progressive projection strategy maintains strong semantic supervision. Besides, to guarantee the retrieval features to capture clean and complete semantic information, the cross-reconstruction loss is introduced to encourage that any combinations of retrieval features and domain features can reconstruct the visual features. Extensive experiments demonstrate the superiority of our PDFD over state-of-the-art competitors. Xinxun Xu, Muli Yang, Yanhua Yang, Hao Wang 0062 |
IJCAI | 3 |
| 2020 | Rotating consensus control of double-integrator multi-agent systems with event-based communication
Rui Ding 0017, Wenfeng Hu, Yanhua Yang |
Sci. China Inf. Sci. | 3 |
| 2020 | Robust Cumulative Crowdsourcing Framework Using New Incentive Payment Function and Joint Aggregation ModelabstractIn recent years, crowdsourcing has gained tremendous attention in the machine learning community due to the increasing demand for labeled data. However, the labels collected by crowdsourcing are usually unreliable and noisy. This issue is mainly caused by: 1) nonflexible data collection mechanisms; 2) nonincentive payment functions; and 3) inexpert crowd workers. We propose a new robust crowdsourcing framework as a comprehensive solution for all these challenging problems. Our unified framework consists of three novel components. First, we introduce a new flexible data collection mechanism based on the cumulative voting system, allowing crowd workers to express their confidence for each option in multi-choice questions. Second, we design a novel payment function regarding the settings of our data collection mechanism. The payment function is theoretically proved to be incentive-compatible, encouraging crowd workers to disclose truthfully their beliefs to get the maximum payment. Third, we propose efficient aggregation models, which are compatible with both single-option and multi-option crowd labels. We define a new aggregation model, called simplex constrained majority voting (SCMV), and enhance it by using the probabilistic generative model. Furthermore, fast optimization algorithms are derived for the proposed aggregation models. Experimental results indicate higher quality for the crowd labels collected by our proposed mechanism without increasing the cost. Our aggregation models also outperform the state-of-the-art models on multiple crowdsourcing data sets in terms of accuracy and convergence speed. Kamran Ghasedi Dizaji, Hongchang Gao, Yanhua Yang, Heng Huang 0001, Cheng Deng 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Deep Asymmetric Metric Learning via Rich Relationship MiningabstractLearning effective distance metric between data has gained increasing popularity, for its promising performance on various tasks, such as face verification, zero-shot learning, and image retrieval. A major line of researches employs hard data mining, which makes efforts on searching a subset of significant data. However, hard data mining based approaches only rely on a small percentage of data, which is apt to overfitting. This motivates us to propose a novel framework, named deep asymmetric metric learning via rich relationship mining (DAMLRRM), to mine rich relationship under satisfying sampling size. DAMLRRM constructs two asymmetric data streams that are differently structured and of unequal length. The asymmetric structure enables the two data streams to interlace each other, which allows for the informative comparison between new data pairs over iterations. To improve the generalization ability, we further relax the constraint on the intra-class relationship. Rather than greedily connecting all possible positive pairs, DAMLRRM builds a minimum-cost spanning tree within each category to ensure the formation of a connected region. As such there exists at least one direct or indirect path between arbitrary positive pairs to bridge intra-class relevance. Extensive experimental results on three benchmark datasets including CUB-200-2011, Cars196, and Stanford Online Products show that DAMLRRM effectively boosts the performance of existing deep metric learning approaches. Yanhua Yang, Cheng Deng 0002, Feng Zheng 0001 |
CVPR | 2 |
| 2019 | Zero-shot Metric LearningabstractIn this work, we tackle the zero-shot metric learning problem and propose a novel method abbreviated as ZSML, with the purpose to learn a distance metric that measures the similarity of unseen categories (even unseen datasets). ZSML achieves strong transferability by capturing multi-nonlinear yet continuous relation among data. It is motivated by two facts: 1) relations can be essentially described from various perspectives; and 2) traditional binary supervision is insufficient to represent continuous visual similarity. Specifically, we first reformulate a collection of specific-shaped convolutional kernels to combine data pairs and generate multiple relation vectors. Furthermore, we design a new cross-update regression loss to discover continuous similarity. Extensive experiments including intra-dataset transfer and inter-dataset transfer on four benchmark datasets demonstrate that ZSML can achieve state-of-the-art performance. Huanhuan Cao, Yanhua Yang, Erkun Yang, Cheng Deng 0002 |
IJCAI | 3 |
| 2019 | Adaptive graph weighting for multi-view dimensionality reduction
Yanhua Yang, Cheng Deng 0002, Feiping Nie 0001 |
Signal Process. | 2 |
| 2018 | Unsupervised Deep Generative Adversarial Hashing NetworkabstractUnsupervised deep hash functions have not shown satisfactory improvements against their shallow alternatives, and usually require supervised pretraining to avoid overfitting. In this paper, we propose a new deep unsupervised hashing function, called HashGAN, which efficiently obtains binary representation of input images without any supervised pretraining. HashGAN consists of three networks, a generator, a discriminator and an encoder. By sharing the parameters of the encoder and discriminator, we benefit from the adversarial loss as a data-dependent regularization in training our deep hash function. Moreover, a novel hashing loss function is introduced for real images, which results in minimum entropy, uniform frequency, consistent and independent hash bits. Furthermore, we employ a collaborative loss in training our model, enforcing similar random inputs and hash bits for synthesized images. In our experiments, HashGAN outperforms the previous unsupervised hash functions in image retrieval and achieves the state-of-the-art performance in image clustering on benchmark datasets. We also provide an ablation study, showing the contribution of each component in our loss function. Kamran Ghasedi Dizaji, Feng Zheng 0001, Najmeh Sadoughi, Yanhua Yang, Cheng Deng 0002, Heng Huang 0001 |
CVPR | 4 |
| 2018 | Centralized Ranking Loss with Weakly Supervised Localization for Fine-Grained Object RetrievalabstractFine-grained object retrieval has attracted extensive research focus recently. Its state-of-the-art schemesare typically based upon convolutional neural network (CNN) features. Despite the extensive progress, two issues remain open. On one hand, the deep features are coarsely extracted at image level rather than precisely at object level, which are interrupted by background clutters. On the other hand, training CNN features with a standard triplet loss is time consuming and incapable to learn discriminative features. In this paper, we present a novel fine-grained object retrieval scheme that conquers these issues in a unified framework. Firstly, we introduce a novel centralized ranking loss (CRL), which achieves a very efficient (1,000times training speedup comparing to the triplet loss) and discriminative feature learning by a ?centralized? global pooling. Secondly, a weakly supervised attractive feature extraction is proposed, which segments object contours with top-down saliency. Consequently, the contours are integrated into the CNN response map to precisely extract features ?within? the target object. Interestingly, we have discovered that the combination of CRL and weakly supervised learning can reinforce each other. We evaluate the performance ofthe proposed scheme on widely-used benchmarks including CUB200-2011 and CARS196. We havereported significant gains over the state-of-the-art schemes, e.g., 5.4% over SCDA [Wei et al., 2017]on CARS196, and 3.7% on CUB200-2011. Xiawu Zheng, Rongrong Ji, Xiaoshuai Sun, Yongjian Wu 0001, Feiyue Huang, Yanhua Yang |
IJCAI | 6 |
| 2018 | Joint Generative-Discriminative Aggregation Model for Multi-Option Crowd LabelsabstractAlthough some crowdsourcing aggregation models have been introduced to aggregate noisy crowd labels, these models mostly consider single-option (i.e. discrete) crowd labels as the input variables, and are not compatible with multi-option (i.e. non-deterministic) crowd data. In this paper, we propose a novel joint generative-discriminative aggregation model, which is able to efficiently deal with both single-option and multi-option crowd labels. Considering the confidence of workers for each option as the input data, we first introduce a new discriminative aggregation model, called Constrained Weighted Majority Voting (CWMVL1), which improves the performance of majority voting method. CWMVL1 considers flexible reliability parameters for crowd workers, employs L1-norm loss function to deal with noisy crowd data, and includes optimization constraints to have probabilistic outputs. We prove that our object is convex, and derive an efficient optimization algorithm. Moreover, we integrate the discriminative CWMVL1 model with a generative model, resulting in a powerful joint aggregation model. Combination of these sub-models is obtained in a probabilistic framework rather than a heuristic way. For our joint model, we derive an efficient optimization algorithm, which alternates between updating the parameters and estimating the potential true labels. Experimental results indicate that the proposed aggregation models achieve superior or competitive results in comparison with the state-of-the-art models on single-option and multi-option crowd datasets, while having faster convergence rates and more reliable predictions. Kamran Ghasedi Dizaji, Yanhua Yang, Heng Huang 0001 |
WSDM | 2 |
| 2018 | Compressed multi-scale feature fusion network for single image super-resolution
Xinxia Fan, Yanhua Yang, Cheng Deng 0002, Jie Xu 0012, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2017 | An improved cooperative control method of DC microgrid based on nearest neighbors communicationabstractDC (Direct Current) microgrid has gained more attention caused by the development of distributed generations and DC loads. The control objectives for DC microgrid are voltage regulation and current sharing. In order to achieve these objectives, consensus based control methods with communication have been extensively studied recently. It is a common strategy to modify the voltage set point of local controller according to the average voltage and current in the system. However, there are some limitations for this method, such as specific communication structure and scalability limitation. This paper proposes an improved cooperative control method. The structure of local controller is simplified and one voltage correction item is generated by the cooperative controller to adjust the reference voltage for local controller. In addition, a novel method in calculating the average voltage and current difference of the multi-node system is applied to reduce the effect of communication delay and increase the scalability. Finally, a four-converter system with ring communication structure is simulated based on Matlab/Simulink to verify the performance of control method. Yanhua Yang, Muhammad Mansoor Khan, Jianyang Yu |
IECON | 1 |
| 2017 | Exploring hybrid spatio-temporal convolutional networks for human action recognition
Hao Wang 0062, Yanhua Yang, Erkun Yang, Cheng Deng 0002 |
Multim. Tools Appl. | 2 |
| 2017 | Latent Max-Margin Multitask Learning With Skelets for 3-D Action RecognitionabstractRecent emergence of low-cost and easy-operating depth cameras has reinvigorated the research in skeleton-based human action recognition. However, most existing approaches overlook the intrinsic interdependencies between skeleton joints and action classes, thus suffering from unsatisfactory recognition performance. In this paper, a novel latent max-margin multitask learning model is proposed for 3-D action recognition. Specifically, we exploit skelets as the mid-level granularity of joints to describe actions. We then apply the learning model to capture the correlations between the latent skelets and action classes each of which accounts for a task. By leveraging structured sparsity inducing regularization, the common information belonging to the same class can be discovered from the latent skelets, while the private information across different classes can also be preserved. The proposed model is evaluated on three challenging action data sets captured by depth cameras. Experimental results show that our model consistently achieves superior performance over recent state-of-the-art approaches. Yanhua Yang, Cheng Deng 0002, Dapeng Tao, Shaoting Zhang 0001, Wei Liu 0005, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Discriminative Multi-instance Multitask Learning for 3D Action RecognitionabstractAs the prosperity of low-cost and easy-operating depth cameras, skeleton-based human action recognition has been extensively studied recently. However, most of the existing methods partially consider that all 3D joints of a human skeleton are identical. Actually, these 3D joints exhibit diverse responses to different action classes, and some joint configurations are more discriminative to distinguish a certain action. In this paper, we propose a discriminative multi-instance multitask learning (MIMTL) framework to discover the intrinsic relationship between joint configurations and action classes. First, a set of discriminative and informative joint configurations for the corresponding action class is captured in multi-instance learning model by regarding the action and the joint configurations as a bag and its instances, respectively. Then, a multitask learning model with group structure constraints is exploited to further reveal the intrinsic relationship between the joint configurations and different action classes. We conduct extensive evaluations of MIMTL using three benchmark 3D action recognition datasets. Experimental results show that our proposed MIMTL framework performs favorably compared with several state-of-the-art approaches. Yanhua Yang, Cheng Deng 0002, Shangqian Gao, Wei Liu 0005, Dapeng Tao, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Multi-task human action recognition via exploring super-category
Yanhua Yang, Ruishan Liu, Cheng Deng 0002, Xinbo Gao 0001 |
Signal Process. | 1 |
| 2013 | Analysis and prediction of jitter of internet one-way time-delay for teleoperation systemsabstractInternet random time delay constitutes a major challenge for Internet-based teleoperation systems. Since the uncertain time delay may degrade the system performance, and even lead to instability. Although large-scale research works have been conducted on the understanding, testing, and analysis of Internet round trip time delay (RTT), it is not appropriate to use RTT in control methods design. It is well known that control commands arrive at the slave site in an aperiodic manner as a result of Internet random time delay. The subsequent control command will terminate the execution of the current command. Consequently, we will get precise information for slave system once the time delay jitter can be known in advance. This paper proposes a novel research idea for Internet-based teleoperation system from the point of view of delay jitter prediction. Statistical properties of the time delay jitter are investigated. Furthermore, the sparse multivariate linear regression method is used to give prediction on Internet time delay jitter. Simulation results demonstrate that sparse multivariate linear regressive method gives a precise prediction, which indicates that the proposed control method based on Internet time delay jitter has a broad prospect in teleoperation systems. Jianning Hua, Yanhua Yang |
INDIN | 3 |
| 2012 | Generalized predictive control for space teleoperation systems with long time-varying delaysabstractPrior researches of generalized predictive control (GPC) in teleoperation systems have mainly considered short transmitted time delays or single degree of freedom (DOF) manipulators of master and slave. This paper presents a GPC strategy for space teleoperation systems in which the master and slave manipulators are both multi-DOF and the communication network brings long time-varying delays. The nonlinear dynamics of the multi-DOF slave manipulator is linearized and described by a linear state-space equation. Meanwhile, a nonlinear compensator is used to compensate the nonlinear parts of the slave. Then, based on the equation, a state-space model based GPC controller is designed on the master side to stabilize the system and make the slave manipulator track the master position and velocity no matter whether the manipulator contacts with the environment or not. Finally, a simulation example is given to illustrate the effectiveness of the proposed method. Yanhua Yang, Fangping Yang, Jianning Hua |
SMC | 1 |