Youtian Du

dblp:34/3813 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-1714-3433ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Learning to Parse and Reconstruct: Bidirectional Modeling of Question-to-Program Mapping
abstract
Neuro-symbolic learning has emerged as a promising paradigm for interpretable visual reasoning, where mapping natural language questions to executable programs plays a central role. However, most existing methods focus exclusively on the forward program generation from questions while overlooking the reverse process of reconstructing questions from programs. In this paper, we propose BiPaR (Bidirectional Parsing and Reconstruction), a Transformer-based framework that jointly models both program parsing and question reconstruction within a unified architecture. Unlike previous approaches that only perform forward parsing, BiPaR introduces reverse program-to-question reconstruction as a powerful auxiliary signal, which improves program generation quality and accelerates convergence, particularly under limited supervision. We further provide a theoretical analysis showing how reverse reconstruction facilitates faster optimization during training. The bidirectional modeling makes BiPaR well-suited for both supervised and semi-supervised learning scenarios. We present two architectural variants: BiPaR-Full, which employs encoder-decoder Transformers for both modules; and BiPaR-DOnly, a lightweight variant that employs a decoder-only structure for question reconstruction, reducing model complexity. Experiments on CLEVR and a GQA subset demonstrate that BiPaR significantly outperforms standard Transformer baselines. Furthermore, in the semi-supervised learning setting, BiPaR achieves notable improvements by leveraging additional questions without program annotations.
Zeying Duan, Youtian Du, Yuanlin Chang
AAAI2
2025 IVAC-$\mathbf {P^{2}L}$: Leveraging Irregular Repetition Priors for Improving Video Action Counting
abstract
The quantification of repetitive actions in videos, a task commonly referred to as Video Action Counting (VAC), is a critical challenge in understanding and analyzing content in sports, fitness, and daily activities. Traditional approaches to VAC have largely overlooked the nuanced irregularities inherent in action repetitions, such as interruptions and variable lengths between cycles. Addressing this gap, our study introduces a novel perspective on VAC, focusing on Irregular Video Action Counting (IVAC), which emphasizes the importance of modeling the irregular repetition priors present in video content. We conceptualize these priors through two key aspects:Inter-cycle ConsistencyandCycle-interval Inconsistency. Inter-cycle Consistency ensures that thespatiotemporalrepresentations across all cycle segments in a videoremainhomogeneous, thereby reflecting the uniformity of actions betweendifferent cycle segments. In contrast, Cycle-interval Inconsistency mandates a clear semantic distinction between the representations of cycle segments and intervals, acknowledging the inherent dissimilarities in content. To effectively encapsulate these priors, we introduce a novel methodology consisting of consistency and inconsistency modules, underpinned by a tailored pull-push loss ($\mathrm {P^{2}~L}$) mechanism. This approach employs a pull loss to enhance the cohesion among cycle segment features and a push loss to distinctly differentiate between cycle and interval segment features. Empirical evaluations on the RepCount dataset illustrate that our IVAC-$\mathrm {P^{2}~L}$model sets a new benchmark in state-of-the-art performance for the VAC task. Moreover, our model demonstrates adaptability and generalization across diverse video content, achieving superior performance on two additional datasets, UCFRep and Countix, without necessitating dataset-specific fine-tuning. These findings not only validate the effectiveness of our approach in addressing the complexities of irregular repetitions in videos but also open new avenues for future research in video understanding and analysis.
Zhi-Qi Cheng, Youtian Du, Lei Zhang 0006
IEEE Trans. Multim.3
2024 Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
abstract
Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robustness. Furthermore, current datasets may not provide a precise diagnostic for these methods. To tackle these challenges, firstly, we propose a novel dataset, *MUSIC-AVQA-R*, crafted in two steps: rephrasing questions within the test split of a public dataset (*MUSIC-AVQA*) and subsequently introducing distribution shifts to split questions. The former leads to a large, diverse test space, while the latter results in a comprehensive robustness evaluation on rare, frequent, and overall questions. Secondly, we propose a robust architecture that utilizes a multifaceted cycle collaborative debiasing strategy to overcome bias learning. Experimental results show that this architecture achieves state-of-the-art performance on MUSIC-AVQA-R, notably obtaining a significant improvement of 9.32\%. Extensive ablation experiments are conducted on the two datasets mentioned to analyze the component effectiveness within the debiasing strategy. Additionally, we highlight the limited robustness of existing multi-modal QA methods through the evaluation on our dataset. We also conduct experiments combining various baselines with our proposed strategy on two datasets to verify its plug-and-play capability. Our dataset and code are available at <https://github.com/reml-group/MUSIC-AVQA-R>.
Jie Ma 0001, Pinghui Wang, Wangchun Sun, Lingyun Song, Hongbin Pei, Jun Liu 0002, Youtian Du
NeurIPS8
2024 Human-Machine Collaborative Reinforcement Learning for Power Line Flow Regulation
abstract
The complexity and uncertainty in power systems leads to a great challenge for controlling the power grid using traditional manual adjustment methods. Reinforcement learning is a promising data-driven paradigm to address control issues in power grids. This article presents a novel human–machine collaborative (HMC) framework for line flow control. We formulate the collaboration between humans and machines as an extended Markov decision process (MDP) and introduce a human–machine collaborative reinforcement learning (HMC-RL) approach, which comprises a routing module, a machine dispatching module and an HMC dispatching module. The routing module determines whether the power system should be operated by the machine or through human–machine collaboration. The machine dispatching module predicts a machine dispatching action for regulating line flow, while the HMC module predicts an HMC dispatching action with human assistance. Experimental results conducted on the IEEE 39-bus and IEEE 118-bus systems demonstrate that our HMC-RL approach can significantly improve the performance of regulation compared to the machine dispatching policy. Specifically, HMC-RL achieves a 40.03% performance improvement on the IEEE 118-bus system, with 25.8% of human participation.
Youtian Du, Yuanlin Chang, Yanhao Huang
IEEE Trans. Ind. Informatics2
2023 Improving weakly supervised phrase grounding via visual representation contextualization with contrastive learning
Youtian Du, Suzan Verberne, Fons J. Verbeek
Appl. Intell.2
2023 Fine-grained label learning in object detection with weak supervision of captions
Youtian Du, Suzan Verberne, Fons J. Verbeek
Multim. Tools Appl.2
2023 One-Stage Visual Relationship Referring With Transformers and Adaptive Message Passing
abstract
There exist a variety of visual relationships among entities in an image. Given a relationship query $\langle subject, predicate, object \rangle $ , the task of visual relationship referring (VRR) aims to disambiguate instances of the same entity category and simultaneously localize the subject and object entities in an image. Previous works of VRR can be generally categorized into one-stage and multi-stage methods. The former ones directly localize a pair of entities from the image but they suffer from low prediction accuracy, while the latter ones perform better but they are indirect to localize only a couple of entities by pre-generating a rich amount of candidate proposals. In this paper, we formulate the task of VRR as an end-to-end bounding box regression problem and propose a novel one-stage approach, called VRR-TAMP, by effectively integrating Transformers and an adaptive message passing mechanism. First, visual relationship queries and images are respectively encoded to generate the basic modality-specific embeddings, which are then fed into a cross-modal Transformer encoder to produce the joint representation. Second, to obtain the specific representation of each entity, we introduce an adaptive message passing mechanism and design an entity-specific information distiller SR-GMP, which refers to a gated message passing (GMP) module that works on the joint representation learned from a single learnable token. The GMP module adaptively distills the final representation of an entity by incorporating the contextual cues regarding the predicate and the other entity. Experiments on VRD and Visual Genome datasets demonstrate that our approach significantly outperforms its one-stage competitors and achieves competitive results with the state-of-the-art multi-stage methods.
Youtian Du, Yabin Zhang 0001, Shuai Li 0014, Lei Zhang 0006
IEEE Trans. Image Process.2
2022 Common quantitative characteristics of music melodies - pursuing the constrained entropy maximization casually in composition
Nan Nan, Xiaohong Guan, Yixin Wang 0007, Youtian Du
Sci. China Inf. Sci.4
2021 Leveraging attention-based visual clue extraction for image classification
abstract
Abstract Deep learning‐based approaches have made considerable progress in image classification tasks, but most of the approaches lack interpretability, especially in revealing the decisive information causing the categorization of images. This paper seeks to answer the question of what clues encode the discriminative visual information between image categories and can help improve the classification performance. To this end, an attention‐based clue extraction network (ACENet) is introduced to mine the decisive local visual information for image classification. ACENet constructs a clue‐attention mechanism, that is global‐local attention, between the image and visual clue proposals extracted from it and then introduces a contrastive loss defined over the achieved discrete attention distribution to increase the discriminability of clue proposals. The loss encourages considerable attention to be devoted to discriminative clue proposals, that is those similar within the same category and dissimilar across categories. The experimental results for the Negative Web Image (NWI) dataset and the public ImageNet2012 dataset demonstrate that ACENet can extract true clues to improve the image classification performance and outperforms the baselines and the state‐of‐the‐art methods.
Yunbo Cui, Youtian Du, Chang Su 0002
IET Image Process.2
2021 Learning Fundamental Visual Concepts Based on Evolved Multi-Edge Concept Graph
abstract
In general, visual media comprises a set of elements of basic semantics, named fundamental visual concepts, that may not be semantically decomposed, such as objects, scenes and actions. This paper proposes a dynamic learning framework for fundamental visual concept learning from image-textual description paired data based on an evolved multi-edge concept graph (EMCG). First, we construct a multi-edge concept graph to represent the relationships between visual concept instances, in which we introduce two types of edges named visual edges and semantic edges to describe the connection strength in terms of visual appearance and semantic content. Second, we evolve the graph by updating connection strength based on the predicted results of concept learning. Finally, we present a growth algorithm for the multi-edge concept graph to handle cross-dataset concept learning. Driven by the predictions, the multi-edge concept graph can dynamically evolve over time by adjusting the connection strength to adapt better to the observations. In addition, our approach can be considered a weakly-supervised learning algorithm since no labeled concepts are employed for learning. Experimental results demonstrate that evolution can significantly improve the learning of fundamental visual concepts by$\text{14.2}\%$,$\text{7.9}\%$and$\text{12.7}\%$in terms of F1-score for the MSRC, VOC2012 and MSCOCO datasets, respectively, and that the proposed EMCG approach largely outperforms the compared approaches.
Youtian Du, Guangxun Zhang, Zhongmin Cai, Chang Su 0002
IEEE Trans. Multim.2
2020 Harmonics Based Representation in Clarinet Tone Quality Evaluation
abstract
Music tone quality evaluation is generally performed by experts. It could be subjective and short of consistency and fairness as well as time-consuming. In this paper we present a new method for identifying the clarinet reed quality by evaluating tone quality based on the harmonic structure and energy distribution. We first decouple the quality of reed and clarinet pipe based on the acoustic harmonics, and discover that the reed quality is strongly relevant to the even parts of the harmonics. Then we construct a features set consisting of the even harmonic envelope and the energy distribution of harmonics in spectrum. The annotated clarinet audio data are recorded from 3 levels of performers and the tone quality is classified by machine learning. The results show that our new method for identifying low and medium high tones significantly outperforms previous methods.
Yixin Wang 0007, Xiaohong Guan, Youtian Du, Nan Nan
ICASSP3
2020 Gender recognition using motion data from multiple smart devices
Youtian Du, Zhongmin Cai
Expert Syst. Appl.2
2020 Kernel-Based Mixture Mapping for Image and Text Association
abstract
Modeling the relationship between multimodal media, including images, videos, and text, can reduce the gap between the modalities and promote cross-media retrieval, image annotation, etc. In this paper, we propose a new approach called kernel-based mixture mapping (KMM) to model the semantic correlations between web images and text. With this approach, we first construct latent high-dimensional feature spaces based on kernel theory to address the nonlinearity of both the data distributions in the input spaces and the cross-model correlation. Second, we present a probabilistic neighborhood model to describe the spatial locality of semantics by assuming that proximate examples in feature spaces generally have the same semantics and a conditional model to describe cross-modal conditional dependency. Finally, we build a probabilistic mixture model to jointly model the spatial locality of semantics and the conditional dependency between different modalities. By combining nonlinear transformation and probabilistic models, KMM can address the nonlinearity of cross-modal correlation, the complexity of semantic distributions at the global scale, and the continuity of semantic distributions at the local scale. We present a hybrid optimization algorithm to find the solution of KMM based on expectation-maximization and subgradient ascent; this algorithm avoids estimating the parameters of KMM in high-dimensional feature space and is proved to converge to an (local) optimal solution. We demonstrate the performance of KMM using four public datasets. The experimental results show that our approach outperforms the compared methods when modeling the relationships between images and text.
Youtian Du, Yunbo Cui, Chang Su 0002
IEEE Trans. Multim.1
2019 Fundamental Visual Concept Learning From Correlated Images and Text
abstract
Heterogeneous web media consists of many visual concepts, such as objects, scenes and activities, that cannot be semantically decomposed. The task of learning fundamental visual concepts (FVCs) plays an important role in automatically understanding the elements that compose all visual media, as well as in applications of retrieval, annotation, etc. In this paper, we formulate the problem of FVC learning and propose an approach to this problem called neighboring concept distributing (NCD). Our approach models all data using a concept graph, which considers the visual patches in images as nodes and generates the inter-image edges between visual patches in different images and the intra-image edges between visual patches in the same image. The NCD approach distributes semantic information from images to visual patches based on measurements over the concept graph, including fitness, distinctiveness, smoothness and sparseness, without any pre-trained concept detectors or classifiers. We analyze the learnability of the proposed approach and find that, under some conditions, all concepts can be correctly learned with an arbitrarily high probability as the size of the data increases.We demonstrate the performance of the NCD approach using three public datasets. Experimental results show that our approach outperforms state-of-the-art approaches when learning visual concepts from correlated media.
Youtian Du, Yunbo Cui
IEEE Trans. Image Process.1
2018 Toward capturing heterogeneity for inferring diffusion networks: A mixed diffusion pattern model
Chang Su 0002, Xiaohong Guan, Youtian Du, Minhua Zhang
Knowl. Based Syst.3
2017 Wikipedia-Based Entity Semantifying in Open Information Extraction
abstract
In the recent years, Open Information Extraction (OIE), an unsupervised strategy which extracts open-domain facts of knowledge from massive heterogeneous text corpora, has achieved impressive improvements. However, the facts (generally represented by a triple) extracted by OIE systems are in lack of clear semantics and then difficult for computer systems to understand. In this paper, we present a new method to semantify the facts by mapping the string arguments in the triples to the corresponding real-world entities based on the existing knowledge base Wikipedia. First, for each query of string argument, we consider a set of its most likely mapping entities and assign each candidate a fused prior probability. Then we calculate the graph-based similarity between candidates as the contextual evidence by propagating semantics on the neighborhood graph of candidates. Finally, we transform the mapping task into an optimization problem and find the maximum a posteriori (MAP) mapping by combining the prior information and contextual evidence through Bayes' theorem. Due to the fusion of multiple cues and the semantics propagation over the graph, our approach improves the performance of the entity semantifying. Experimental results demonstrate the effectiveness of our approach.
Qiuhao Lu, Youtian Du
ICDAR2
2015 Learning Semantic Correlation of Web Images and Text with Mixture of Local Linear Mappings
abstract
This paper proposes a new approach, called mixture of local linear mappings (MLLM), to the modeling of semantic correlation between web images and text. We consider that close examples generally represent a uniform concept and can be supposed to be locally transformed based on a linear mapping into the feature space of another modality. Thus, we use a mixture of local linear transformations, each local component being constrained by a neighborhood model into a finite local space, instead of a more complex nonlinear one. To handle the sparseness of data representation, we introduce the constraints of sparseness and non-negativeness into the approach. MLLM is with good interpretability due to its explicit closed form and concept-related local components, and it avoids the determination of capacity that is often considered for nonlinear transformations. Experimental results demonstrate the effectiveness of the proposed approach.
Youtian Du
ACM Multimedia1
2015 Semantic Correlation Mining between Images and Texts with Global Semantics and Local Mapping
Jiao Xue, Youtian Du, Hanbing Shui
MMM (2)2
2013 Maximizing topic propagation driven by multiple user nodes in micro-blogging
abstract
This work investigates the maximization of topic propagation jointly driven by multiple user nodes in micro-blogging. In this paper, we propose a new method to find a set of user nodes that jointly propagate topics approximately the most widely. First, we obtain multiple nodes with strong influence; Second, we exactly compute the breadth of information spread driven by a single node based on probabilistic models; Finally, we analyze the information propagation jointly driven by multiple nodes and derive an approximately optimal set of driving nodes. We find that the breadth of information propagation jointly driven by multiple nodes is approximately linear with both the breadth of information propagation of single driving nodes and the strength of tie among them, which indicates that selecting the optimal driving nodes needs to consider the link information among them as well as the ability of each node. Experimental results demonstrate the effectiveness of our method.
Chang Su 0002, Youtian Du, Xiaohong Guan, Chenhe Wu
LCN2
2013 Multi-view semi-supervised web image classification via co-graph
Youtian Du, Qian Li 0024, Zhongmin Cai, Xiaohong Guan
Neurocomputing1
2013 Tagging photos using users' vocabularies
Xueming Qian, Youtian Du, Xingsong Hou
Neurocomputing4
2013 Video content categorization using the double decomposition
Youtian Du, Feng Chen 0007, Wenli Xu, Xueming Qian
Multim. Tools Appl.1
2013 User Authentication Through Mouse Dynamics
abstract
Behavior-based user authentication with pointing devices, such as mice or touchpads, has been gaining attention. As an emerging behavioral biometric, mouse dynamics aims to address the authentication problem by verifying computer users on the basis of their mouse operating styles. This paper presents a simple and efficient user authentication approach based on a fixed mouse-operation task. For each sample of the mouse-operation task, both traditional holistic features and newly defined procedural features are extracted for accurate and fine-grained characterization of a user's unique mouse behavior. Distance-measurement and eigenspace-transformation techniques are applied to obtain feature components for efficiently representing the original mouse feature space. Then a one-class learning algorithm is employed in the distance-based feature eigenspace for the authentication task. The approach is evaluated on a dataset of 5550 mouse-operation samples from 37 subjects. Extensive experimental results are included to demonstrate the efficacy of the proposed approach, which achieves a false-acceptance rate of 8.74%, and a false-rejection rate of 7.69% with a corresponding authentication time of 11.8 seconds. Two additional experiments are provided to compare the current approach with other approaches in the literature. Our dataset is publicly available to facilitate future research.
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan, Youtian Du, Roy A. Maxion
IEEE Trans. Inf. Forensics Secur.4
2010 Enhancing Web Page Classification via Local Co-training
abstract
In this paper we propose a new multi-view semi-supervised learning algorithm called Local Co-Training(LCT). The proposed algorithm employs a set of local models with vector outputs to model the relations among examples in a local region on each view, and iteratively refines the dominant local models (i.e. the local models related to the unlabeled examples chosen for enriching the training set) using unlabeled examples by the co-training process. Compared with previous co-training style algorithms, local co-training has two advantages: firstly, it has higher classification precision by introducing local learning; secondly, only the dominant local models need to be updated, which significantly decreases the computational load. Experiments on WebKB and Cora datasets demonstrate that LCT algorithm can effectively exploit unlabeled data to improve the performance of web page classification.
Youtian Du, Xiaohong Guan, Zhongmin Cai
ICPR1
2008 Activity recognition through multi-scale motion detail analysis
Youtian Du, Feng Chen 0007, Wenli Xu
Neurocomputing1
2008 Hierarchical group process representation in multi-agent activity recognition
Feng Chen 0007, Wenli Xu, Youtian Du
Signal Process. Image Commun.4
2007 Human Interaction Representation and Recognition Through Motion Decomposition
abstract
Human action recognition is one of the most important problems in video content analysis and computer vision. In this letter, we propose a novel framework of human interaction recognition through motion decomposition. Interactions contain not only motions corresponding to each person but also motion details on different scales. Hence, we decompose an interaction into multiple interacting stochastic processes in the above two aspects. Under the framework, we present a Coupled Hierarchical Durational-State Dynamic Bayesian Network (CHDS-DBN) to model interactions by modeling the multiple stochastic processes. The effectiveness of the approach is demonstrated by experiments of two-person interaction recognition.
Youtian Du, Feng Chen 0007, Wenli Xu
IEEE Signal Process. Lett.1