Junqing Yu

dblp:79/1515 · DBLP profile ↗
← Back
65ranked-venue papers
4as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 2 first-author · 24 since 2021Artificial intelligence and machine learning · 27 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning
abstract
Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decision-making in high-stakes domains such as mathematical reasoning and legal judgment.In this study, we present a systematic analysis of logical reasoning under controlled increases in logical complexity, and reveal a previously unrecognized phenomenon, which we term Logical Phase Transitions: rather than degrading smoothly, logical reasoning performance remains stable within a regime but collapses abruptly beyond a critical logical depth, mirroring physical phase transitions such as water freezing beyond a critical temperature threshold.Building on this insight, we propose Neuro-Symbolic Curriculum Tuning, a principled framework that adaptively aligns natural language with logical symbols to establish a shared representation, and reshapes training dynamics around phase-transition boundaries to progressively strengthen reasoning at increasing logical depths.Experiments on five benchmarks show that our approach effectively mitigates logical reasoning collapse at high complexity, yielding average accuracy gains of +1.26 in naive prompting and +3.95 in CoT, while improving generalization to unseen logical compositions.Code and data are available at: https://github.com/ AI4SS/Logical-Phase-Transitions.
Xinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu, Zikai Song
ACL (1)4
2026 Semantic-Aware Logical Reasoning via a Semiotic Framework
abstract
Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang, Zikai Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034, Zikai Song
ACL (1)5
2026 EvoCap: Enhancing video captioning via self-evolving video-LLMs with knowledge consolidation
Yangliu Hu, Minye Wu, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
Knowl. Based Syst.3
2025 Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
abstract
A recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recover normal patterns exclusively, thus reporting abnormal patterns as outliers. Yet, existing attempts neglect the various formations of anomaly and predict normal samples at the feature level regardless that abnormal objects in surveillance videos are often relatively small. To address this, a novel patch-based diffusion model is proposed, specifically engineered to capture fine-grained local information. We further observe that anomalies in videos manifest themselves as deviations in both appearance and motion. Therefore, we argue that a comprehensive solution must consider both of these aspects simultaneously to achieve accurate frame prediction. To address this, we introduce innovative motion and appearance conditions that are seamlessly integrated into our patch diffusion model. These conditions are designed to guide the model in generating coherent and contextually appropriate predictions for both semantic content and motion relations. Experimental results on four challenging video anomaly detection datasets empirically substantiate the efficacy of our proposed approach, demonstrating that it consistently outperforms most existing methods in detecting abnormal behaviors.
Hang Zhou 0010, Jiale Cai, Yuteng Ye, Yonghui Feng, Chenxing Gao, Junqing Yu, Zikai Song, Wei Yang 0034
AAAI6
2025 Temporal Coherent Object Flow for Multi-Object Tracking
abstract
Multi-object tracking is a challenging vision task that requires simultaneous reasoning about object detection and object association. Conventional solutions use frame as the basic unit and typically rely on a motion predictor that exploits the appearance features to associate detected candidates, leading to insufficient adaptability to long-term associations. In this study, we propose a section-based multi-object tracking approach that integrates a temporal coherent Object Flow Tracker (OFTrack), capable of achieving simultaneous multi-frame tracking by treating multiple consecutive frames as the basic processing unit, denoted as a “section”. Our OFTrack boosts the optical flow to the object flow by employing object perception and section-based motion estimation strategies. Object perception adopts object-aware sampling and scale-aware correlation to enable precise target discrimination. Motion estimation models the correlation of different objects in multi-frames via specialized temporal-spatial attention to achieve robust association in very long videos. Additionally, to address the oscillation of unpredictable trajectories in multi-frame estimation, we have designed temporal coherent enhancement including the trajectory masking pre-training and the smoothing constraint on trajectory curves. Comprehensive experiments on several widely used benchmarks demonstrate the superior performance of our approach.
Zikai Song, Run Luo, Lintao Ma, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang 0034
AAAI6
2025 SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
abstract
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their fine-grained video understanding abilities. Hence we propose two key contributions: (1) Self-Supervised Fragment Fine-Tuning (SF2T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs’ performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF2T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
CVPR5
2025 Ref-GS: Directional Factorization for 2D Gaussian Splatting
abstract
In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting [8], which enables photorealistic view-dependent appearance rendering and precise geometry recovery. Ref-GS builds upon the deferred rendering of Gaussian splatting and applies directional encoding to the deferred-rendered surface, effectively reducing the ambiguity between orientation and viewing angle. Next, we introduce a spherical Mip-grid to capture varying levels of surface roughness, enabling roughness-aware Gaussian shading. Additionally, we propose a simple yet efficient geometry-lighting factorization that connects geometry and lighting via the vector outer product, significantly reducing renderer overhead when integrating volumetric attributes. Our method achieves superior photorealistic rendering for a range of open-world scenes while also accurately recovering geometry. See our interactive project page.
Youjia Zhang, Anpei Chen, Yumin Wan, Zikai Song, Junqing Yu, Yawei Luo, Wei Yang 0034
CVPR5
2025 CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation
abstract
Segmentation of brain structures from MRI is crucial for evaluating brain morphology, yet existing CNN and transformer-based methods struggle to delineate complex structures accurately. While current diffusion models have shown promise in image segmentation, they are inadequate when applied directly to brain MRI due to neglecting anatomical information. To address this, we propose Collaborative Anatomy Diffusion (CA-Diff), a framework integrating spatial anatomical features to enhance segmentation accuracy of the diffusion model. Specifically, we introduce distance field as an auxiliary anatomical condition to provide global spatial context, alongside a collaborative diffusion process to model its joint distribution with anatomical structures, enabling effective utilization of anatomical features for segmentation. Furthermore, we introduce a consistency loss to refine relationships between the distance field and anatomical structures and design a time adapted channel attention module to enhance the U-Net feature fusion procedure. Extensive experiments show that CA-Diff outperforms state-of-the-art (SOTA) methods.
Qilong Xing, Zikai Song, Yuteng Ye, Yuke Chen, Youjia Zhang, Na Feng, Junqing Yu, Wei Yang 0034
ICME7
2025 Optimized View and Geometry Distillation from Multi-view Diffuser
abstract
Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and over-smoothing in extracted geometry persist. Previous methods integrate multi-view consistency modules or impose additional supervisory to enhance view consistency while compromising on the flexibility of camera positioning and limiting the versatility of view synthesis. In this study, we consider the radiance field optimized during geometry extraction as a more rigid consistency prior, compared to volume and ray aggregation used in previous works. We further identify and rectify a critical bias in the traditional radiance field optimization process through score distillation from a multi-view diffuser. We introduce an Unbiased Score Distillation (USD) that utilizes unconditioned noises from a 2D diffusion model, greatly refining the radiance field fidelity. We leverage the rendered views from the optimized radiance field as the basis and develop a two-step specialization process of a 2D diffusion model, which is adept at conducting object-specific denoising and generating high-quality multi-view images. Finally, we recover faithful geometry and texture directly from the refined multi-view images. Empirical evaluations demonstrate that our optimized geometry and view distillation technique generates comparable results to the state-of-the-art models trained on extensive datasets, all while maintaining freedom in camera positioning. Source code of our work is publicly available at: https://youjiazhang.github.io/USD/.
Youjia Zhang, Zikai Song, Junqing Yu, Yawei Luo, Wei Yang 0034
IJCAI3
2025 Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients
Qilong Xing, Zikai Song, Bingxin Gong, Junqing Yu, Wei Yang 0034
MICCAI (15)5
2025 MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation
Qilong Xing, Zikai Song, Youjia Zhang, Na Feng, Junqing Yu, Wei Yang 0034
MICCAI (5)5
2025 MVP: Winning Solution to SMP Challenge 2025 Video Track
abstract
Social media platforms serve as central hubs for content dissemination, opinion expression, and public engagement across diverse modalities. Accurately predicting the popularity of social media videos enables valuable applications in content recommendation, trend detection, and audience engagement. In this paper, we present Multimodal Video Predictor (MVP), our winning solution to the Video Track of the SMP Challenge 2025. MVP constructs expressive post representations by integrating deep video features extracted from pretrained models with user metadata and contextual information. The framework applies systematic preprocessing techniques, including log-transformations and outlier removal, to improve model robustness. A gradient-boosted regression model is trained to capture complex patterns across modalities. Our approach ranked first in the official evaluation of the Video Track, demonstrating its effectiveness and reliability for multimodal video popularity prediction on social platforms. The source code is available at https://github.com/yllhwa/SMPDVideo.
Liliang Ye, Yunyao Zhang, Yafeng Wu, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang 0034, Zikai Song
ACM Multimedia5
2025 TimeJudge: empowering video-LLMs as zero-shot judges for temporal consistency in video captions
abstract
Video large language models (video-LLMs) have demonstrated impressive capabilities in multimodal understanding, but their potential as zero-shot evaluators for temporal consistency in video captions remains underexplored. Existing methods notably underperform in detecting critical temporal errors, such as missing, hallucinated, or misordered actions. To address this gap, we introduce two key contributions. (1) TimeJudge: a novel zero-shot framework that recasts temporal error detection as answering calibrated binary question pairs. It incorporates modality-sensitive confidence calibration and uses consistency-weighted voting for robust prediction aggregation. (2) TEDBench: a rigorously constructed benchmark featuring videos across four distinct complexity levels, specifically designed with fine-grained temporal error annotations to evaluate video-LLM performance on this task. Through a comprehensive evaluation of multiple state-of-the-art video-LLMs on TEDBench, we demonstrate that TimeJudge consistently yields substantial gains in terms of recall and F1-score without requiring any task-specific fine-tuning. Our approach provides a generalizable, scalable, and training-free solution for enhancing the temporal error detection capabilities of video-LLMs.
Yangliu Hu, Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
Frontiers Inf. Technol. Electron. Eng.3
2025 Sensitivity Seismic Attenuation Estimation Using the Sublimated Centroid Frequency Shift Method
abstract
Seismic attenuation is an important reservoir characteristic, as it is highly sensitive to changes in rock pore fluids. The sensitivity of attenuation makes it a valuable parameter for monitoring reservoir dynamics. This paper presents an advanced method for quality factor (Q) estimation and attenuation difference (AD) estimation, which builds upon the frequency-independent centroid frequency shift (FiCFS) and divergence-based CFS methods. In contrast to conventionalQ-value estimation techniques, our approach provides frequency independence, enhanced noise immunity and heightened sensitivity to minor attenuation variations. This method addresses the issues associated with centroid frequency discrepancies and simplifies the calculations by eliminating the necessity for traveltime differences. The superiority of the noise immunity and amplification effect of AD estimation is demonstrated by synthetic data experiments, which also facilitate easier monitoring of reservoir changes. The efficacy of this method is validated through the CO2injection monitoring at the Frio II site, which demonstrates its ability to detect fluid changes within the reservoir. The in-depth analysis of variance, divergence, and centroid frequency difference reveals the direct impact of these disparate methodologies and substantiates the sublimation of the novel approach.
Junqing Yu, Xiangyang Cao, Huijian Li
IEEE Trans. Geosci. Remote. Sens.1
2024 Attacking Transformers with Feature Diversity Adversarial Perturbation
abstract
Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit low transferability to other structures and tasks. In this paper, we present a label-free white-box attack approach for ViT-based models that exhibits strong transferability to various black-box models, including most ViT variants, CNNs, and MLPs, even for models developed for other modalities. Our inspiration comes from the feature collapse phenomenon in ViTs, where the critical attention mechanism overly depends on the low-frequency component of features, causing the features in middle-to-end layers to become increasingly similar and eventually collapse. We propose the feature diversity attacker to naturally accelerate this process and achieve remarkable performance and transferability.
Chenxing Gao, Hang Zhou 0010, Junqing Yu, Yuteng Ye, Jiale Cai, Junle Wang, Wei Yang 0034
AAAI3
2024 AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion
abstract
Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion. Recent advances in diffusion models have enabled significant progress in human motion synthesis. However, existing methods struggle to handle text inputs that describe complex or long motions. In this paper, we propose the Adaptable Motion Diffusion (AMD) model, which leverages a Large Language Model (LLM) to parse the input text into a sequence of concise and interpretable anatomical scripts that correspond to the target motion. This process exploits the LLM’s ability to provide anatomical guidance for complex motion synthesis. We then devise a two-branch fusion scheme that balances the influence of the input text and the anatomical scripts on the inverse diffusion process, which adaptively ensures the semantic fidelity and diversity of the synthesized motion. Our method can effectively handle texts with complex or long motion descriptions, where existing methods often fail. Experiments on datasets with relatively more complex motions, such as CLCD1 and CLCD2, demonstrate that our AMD significantly outperforms existing state-of-the-art models.
Beibei Jing, Youjia Zhang, Zikai Song, Junqing Yu, Wei Yang 0034
AAAI4
2024 Progressive Text-to-Image Diffusion with Soft Latent Direction
abstract
In spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose enduring challenges. This paper introduces an innovative progressive synthesis and editing operation that systematically incorporates entities into the target image, ensuring their adherence to spatial and relational constraints at each sequential step. Our key insight stems from the observation that while a pre-trained text-to-image diffusion model adeptly handles one or two entities, it often falters when dealing with a greater number. To address this limitation, we propose harnessing the capabilities of a Large Language Model (LLM) to decompose intricate and protracted text descriptions into coherent directives adhering to stringent formats. To facilitate the execution of directives involving distinct semantic operations—namely insertion, editing, and erasing—we formulate the Stimulus, Response, and Fusion (SRF) framework. Within this framework, latent regions are gently stimulated in alignment with each operation, followed by the fusion of the responsive latent components to achieve cohesive entity manipulation. Our proposed framework yields notable advancements in object synthesis, particularly when confronted with intricate and lengthy textual inputs. Consequently, it establishes a new benchmark for text-to-image generation tasks, further elevating the field's performance standards.
Yuteng Ye, Jiale Cai, Hang Zhou 0010, Guanwen Li, Youjia Zhang, Zikai Song, Chenxing Gao, Junqing Yu, Wei Yang 0034
AAAI8
2024 Dynamic Feature Pruning and Consolidation for Occluded Person Re-identification
abstract
Occluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as human body key points and semantic segmentations, which easily fail in the presence of heavy occlusion and other humans as occluders. In this paper, we propose a feature pruning and consolidation (FPC) framework to circumvent explicit human structure parsing. The framework mainly consists of a sparse encoder, a multi-view feature mathcing module, and a feature consolidation decoder. Specifically, the sparse encoder drops less important image tokens, mostly related to background noise and occluders, solely based on correlation within the class token attention. Subsequently, the matching stage relies on the preserved tokens produced by the sparse encoder to identify k-nearest neighbors in the gallery by measuring the image and patch-level combined similarity. Finally, we use the feature consolidation module to compensate pruned features using identified neighbors for recovering essential information while disregarding disturbance from noise and occlusion. Experimental results demonstrate the effectiveness of our proposed framework on occluded, partial, and holistic Re-ID datasets. In particular, our method outperforms state-of-the-art results by at least 8.6% mAP and 6.0% Rank-1 accuracy on the challenging Occluded-Duke dataset.
Yuteng Ye, Hang Zhou 0010, Jiale Cai, Chenxing Gao, Youjia Zhang, Junle Wang, Qiang Hu 0003, Junqing Yu, Wei Yang 0034
AAAI8
2024 Agnostic Feature Compression with Semantic Guided Channel Importance Analysis
abstract
Distributing the computational workload of neural networks across cloud servers and local devices is an effective strategy for deploying resource-intensive deep models to edge devices. Therefore, compressing the deep features without compromising performance is crucial for saving server storage and transmission bandwidth. However, existing feature compression approaches are model or task specific and require training from scratch. In this paper, we propose a general and efficient framework for compressing deep features without requiring any prior knowledge of the semantics or task of the features. Our key observation is that different parts of the the feature map have different importance levels for a specific task. We can apply compression operation to a deeper degree for less irrelevant parts to achieve a high compression rate, while preserving the performance by applying a lower compression ratio to the more important parts. Focusing on this idea, we use the activation map generated by GradCAM [1] to classify each deep feature channel into essential and peripheral categories. To improve classification accuracy, we utilise semantic segmentations to provide natural boundaries for scoring each semantic channel. Peripheral channels are compressed using binary compression to achieve a high compression rate, while essential channels are compressed using mask compression. To effectively separate the essential and peripheral channels for a given input feature, we adopt a data-driven approach to identify the essential channels from datasets. Experimental results demonstrate that our method outperforms the state-of-the-art feature compression methods and is generalizable for various deep models.
Wei Yang 0034, Junqing Yu, Zikai Song
ICME3
2024 Autogenic Language Embedding for Coherent Point Tracking
abstract
Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on temporal modeling techniques to improve local feature similarity, often overlooking the valuable semantic consistency inherent in tracked points. In this paper, we introduce a novel approach leveraging language embeddings to enhance the coherence of frame-wise visual features related to the same object. Our proposed method, termed autogenic language embedding for visual feature enhancement, strengthens point correspondence in long-term sequences. Unlike existing visual-language schemes, our approach learns text embeddings from visual features through a dedicated mapping network, enabling seamless adaptation to various tracking tasks without explicit text annotations. Additionally, we introduce a consistency decoder that efficiently integrates text tokens into visual features with minimal computational overhead. Through enhanced visual consistency, our approach significantly improves tracking trajectories in lengthy videos with substantial appearance variations. Extensive experiments on widely-used tracking benchmarks demonstrate the superior performance of our method, showcasing notable enhancements compared to trackers relying solely on visual cues. The code will be available at https://github.com/SkyeSong38/ALTrack.
Zikai Song, Run Luo, Lintao Ma, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
ACM Multimedia5
2024 Coupled Mamba: Enhanced Multimodal Fusion with Coupled State Space Model
abstract
The essence of multi-modal fusion lies in exploiting the complementary information inherent in diverse modalities.However, most prevalent fusion methods rely on traditional neural architectures and are inadequately equipped to capture the dynamics of interactions across modalities, particularly in presence of complex intra- and inter-modality correlations.Recent advancements in State Space Models (SSMs), notably exemplified by the Mamba model, have emerged as promising contenders. Particularly, its state evolving process implies stronger modality fusion paradigm, making multi-modal fusion on SSMs an appealing direction. However, fusing multiple modalities is challenging for SSMs due to its hardware-aware parallelism designs. To this end, this paper proposes the Coupled SSM model, for coupling state chains of multiple modalities while maintaining independence of intra-modality state processes. Specifically, in our coupled scheme, we devise an inter-modal hidden states transition scheme, in which the current state is dependent on the states of its own chain and that of the neighbouring chains at the previous time-step. To fully comply with the hardware-aware parallelism, we obtain the global convolution kernel by deriving the state equation while introducing the historical state.Extensive experiments on CMU-MOSEI, CH-SIMS, CH-SIMSV2 through multi-domain input verify the effectiveness of our model compared to current state-of-the-art methods, improved F1-Score by 0.4%, 0.9%, and 2.3% on the three datasets respectively, 49% faster inference and 83.7% GPU memory save. The results demonstrate that Coupled Mamba model is capable of enhanced multi-modal fusion.
Wenbing Li, Hang Zhou 0010, Junqing Yu, Zikai Song, Wei Yang 0034
NeurIPS3
2023 Compact Transformer Tracker with Correlative Masked Modeling
abstract
Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with the well-known attention mechanism. Most recent advances focus on exploring attention mechanism variants for better information aggregation. We find these schemes are equivalent to or even just a subset of the basic self-attention mechanism. In this paper, we prove that the vanilla self-attention structure is sufficient for information aggregation, and structural adaption is unnecessary. The key is not the attention structure, but how to extract the discriminative feature for tracking and enhance the communication between the target and search image. Based on this finding, we adopt the basic vision transformer (ViT) architecture as our main tracker and concatenate the template and search image for feature embedding. To guide the encoder to capture the invariant feature for tracking, we attach a lightweight correlative masked decoder which reconstructs the original template and search image from the corresponding masked tokens. The correlative masked decoder serves as a plugin for the compact transformer tracker and is skipped in inference. Our compact tracker uses the most simple structure which only consists of a ViT backbone and a box head, and can run at 40 fps. Extensive experiments show the proposed compact transform tracker outperforms existing approaches, including advanced attention variants, and demonstrates the sufficiency of self-attention in tracking tasks. Our method achieves state-of-the-art performance on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-10k benchmarks. Our project is available at https://github.com/HUSTDML/CTTrack.
Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
AAAI3
2023 Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection
abstract
Learning discriminative features for effectively separating abnormal events from normality is crucial for weakly supervised video anomaly detection (WS-VAD) tasks. Existing approaches, both video and segment level label oriented, mainly focus on extracting representations for anomaly data while neglecting the implication of normal data. We observe that such a scheme is sub-optimal, i.e., for better distinguishing anomaly one needs to understand what is a normal state, and may yield a higher false alarm rate. To address this issue, we propose an Uncertainty Regulated Dual Memory Units (UR-DMU) model to learn both the representations of normal data and discriminative features of abnormal data. To be specific, inspired by the traditional global and local structure on graph convolutional networks, we introduce a Global and Local Multi-Head Self Attention (GL-MHSA) module for the Transformer network to obtain more expressive embeddings for capturing associations in videos. Then, we use two memory banks, one additional abnormal memory for tackling hard samples, to store and separate abnormal and normal prototypes and maximize the margins between the two representations. Finally, we propose an uncertainty learning scheme to learn the normal data latent space, that is robust to noise from camera switching, object changing, scene transforming, etc. Extensive experiments on XD-Violence and UCF-Crime datasets demonstrate that our method outperforms the state-of-the-art methods by a sizable margin.
Hang Zhou 0010, Junqing Yu, Wei Yang 0034
AAAI2
2023 NeReF: Neural Refractive Field for Fluid Surface Reconstruction and Rendering
abstract
We present a novel Neural Refractive Field (NeReF) to recover wavefront of transparent fluids by simultaneously estimating the surface position and normal of the fluid front. Unlike prior arts that treat the reconstruction target as a single layer of the surface, NeReF is specifically formulated to recover a volumetric normal field with its corresponding density field. A query ray will be refracted by NeReF according to its accumulated refractive point and normal, and we employ the correspondences and uniqueness of refracted ray for NeReF optimization. We show NeReF, as a global optimization scheme, can more robustly tackle refraction distortions detrimental to traditional methods for correspondence matching. Furthermore, the continuous NeReF representation of wavefront enables view synthesis as well as normal integration. We validate our approach on both synthetic and real data and show it is particularly suitable for sparse multi-view acquisition. We hence build a small light field array and experiment on various surface shapes to demonstrate high fidelity NeReF reconstruction.
Wei Yang 0034, Junming Cao, Qiang Hu 0003, Lan Xu 0003, Junqing Yu, Jingyi Yu 0001
ICCP6
2023 NeMF: Inverse Volume Rendering with Neural Microflake Field
abstract
Recovering the physical attributes of an object’s appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Recent approaches adopt the emerging implicit scene representations and have shown impressive results. However, they unanimously adopt a surface-based representation, and hence can not well handle scenes with very complex geometry, translucent object and etc. In this paper, we propose to conduct inverse volume rendering, in contrast to surface-based, by representing a scene using microflake volume, which assumes the space is filled with infinite small flakes and light reflects or scatters at each spatial location according to microflake distributions. We further adopt the coordinate networks to implicitly encode the microflake volume, and develop a differentiable microflake volume renderer to train the network in an end-to-end way in principle. Our NeMF enables effective recovery of appearance attributes for highly complex geometry and scattering object, enables high-quality relighting, material editing, and especially simulates volume rendering effects, such as scattering, which is infeasible for surface-based approaches. Our data and code are available at: https://github.com/YoujiaZhang/NeMF.
Youjia Zhang, Teng Xu 0008, Junqing Yu, Yuteng Ye, Yanqing Jing, Junle Wang, Jingyi Yu 0001, Wei Yang 0034
ICCV3
2023 Resource scheduling techniques in cloud from a view of coordination: a holistic survey
abstract
Nowadays, the management of resource contention in shared cloud remains a pending problem. The evolution and deployment of new application paradigms (e.g., deep learning training and microservices) and custom hardware (e.g., graphics processing unit (GPU) and tensor processing unit (TPU)) have posed new challenges in resource management system design. Current solutions tend to trade cluster efficiency for guaranteed application performance, e.g., resource over-allocation, leaving a lot of resources underutilized. Overcoming this dilemma is not easy, because different components across the software stack are involved. Nevertheless, massive efforts have been devoted to seeking effective performance isolation and highly efficient resource scheduling. The goal of this paper is to systematically cover related aspects to deliver the techniques from the coordination perspective, and to identify the corresponding trends they indicate. Briefly, four topics are involved. First, isolation mechanisms deployed at different levels (micro-architecture, system, and virtualization levels) are reviewed, including GPU multitasking methods. Second, resource scheduling techniques within an individual machine and at the cluster level are investigated, respectively. Particularly, GPU scheduling for deep learning applications is described in detail. Third, adaptive resource management including the latest microservice-related research is thoroughly explored. Finally, future research directions are discussed in the light of advanced work. We hope that this review paper will help researchers establish a global view of the landscape of resource management techniques in shared cloud, and see technology trends more clearly.
Yuzhao Wang, Junqing Yu, Zhibin Yu 0001
Frontiers Inf. Technol. Electron. Eng.2
2023 A multi-scale multi-level deep descriptor with saliency for image retrieval
Zebin Wu 0002, Junqing Yu
Multim. Tools Appl.2
2022 Transformer Tracking with Cyclic Shifting Window Attention
abstract
Transformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel-to-pixel attention strategy on flattened image features and unavoidably ignore the integrity of ob-jects. In this paper, we propose a new transformer ar-chitecture with multi-scale cyclic shifting window attention for visual object tracking, elevating the attention from pixel to window level. The cross-window multi-scale at-tention has the advantage of aggregating attention at dif-ferent scales and generates the best fine-scale match for the target object. Furthermore, the cyclic shifting strat-egy brings greater accuracy by expanding the window sam-ples with positional information, and at the same time saves huge amounts of computational power by removing redun-dant calculations. Extensive experiments demonstrate the superior performance of our method, which also sets the new state-of-the-art records on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-lOk benchmarks. Our project is available at https://github.com/SkyeSong38/CSWinTT.
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
CVPR2
2022 Category-Level Adversarial Adaptation for Semantic Segmentation Using Purified Features
abstract
We target the problem named unsupervised domain adaptive semantic segmentation. A key in this campaign consists in reducing the domain shift, so that a classifier based on labeled data from one domain can generalize well to other domains. With the advancement of adversarial learning method, recent works prefer the strategy of aligning the marginal distribution in the feature spaces for minimizing the domain discrepancy. However, based on the observance in experiments, only focusing on aligning global marginal distribution but ignoring the local joint distribution alignment fails to be the optimal choice. Other than that, the noisy factors existing in the feature spaces, which are not relevant to the target task, entangle with the domain invariant factors improperly and make the domain distribution alignment more difficult. To address those problems, we introduce two new modules, Significance-aware Information Bottleneck (SIB) and Category-level alignment (CLA), to construct a purified embedding-based category-level adversarial network. As the name suggests, our designed network, CLAN, can not only disentangle the noisy factors and suppress their influences for target tasks but also utilize those purified features to conduct a more delicate level domain calibration, i.e., global marginal distribution and local joint distribution alignment simultaneously. In three domain adaptation tasks, i.e., GTA5 → Cityscapes, SYNTHIA → Cityscapes and Cross Season, we validate that our proposed method matches the state of the art in segmentation accuracy.
Yawei Luo, Ping Liu 0004, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Para: Harvesting CPU time fragments in Big Data Analytics
abstract
Modern data analytics typically run tasks on statically reserved resources (e.g., CPU and memory), which is prone to over-provision to guarantee the Quality of Service (QoS), leading to a large amount of resource time fragments. As a result, the resource utilization of a data analytics cluster is severely under-utilized. Workload co-location on shared resources has been substantially studied, but they are unaware the sizes of resource time fragments, making them hard to improve the resource utilization and guarantee QoS at the same time. In this paper, we propose Para, an event-driven scheduling mechanism, to harvest the CPU time fragments in co-located big data analytic workloads. Para innovates three techniques: 1) identifying the Idle CPU Time Window (ICTW) associated with each CPU core by capturing the task-switch event; 2) designing a runtime communication mechanism between each task execution of a workload and the underlying resource management system; 3) designing a pull-based scheduler to schedule a workload to run in the ICTW of another workload. We implement Para based on Apache Mesos and Spark. And the experimental results show that Para improves the CPU utilization by 44% and 30% on average relative to the original Mesos and enhanced Mesos under Spark's dynamic mode (MSDM), respectively. Moreover, Para increases the averaged task throughput of Mesos and MSDM by 4.8x and 1.7x, respectively, while guaranteeing the execution time of the primary applications.
Yuzhao Wang, Hongliang Qu, Junqing Yu, Zhibin Yu 0001
CLOUD3
2021 Distractor-Aware Tracker with a Domain-Special Optimized Benchmark for Soccer Player Tracking
abstract
Player tracking in broadcast soccer videos has received widespread attention in the field of sports video analysis, however, we note that there is not a suitable tracking algorithm specifically for soccer video, and the existing benchmarks used for soccer player tracking cover few scenarios with low difficulties. From the observation of the soccer scene that interference and occlusion are knotty problems because the distractors are extremely similar to the targets, a distractor-aware player tracking algorithm and a high-quality benchmark for soccer play tracking (BSPT) have been presented. The distractor-aware player tracking algorithm is able to perceive semantic information about distracting players in the background by similarity judgment, the semantic distractor-aware information is encoded into a context vector and is constantly updated as the objects move through a video sequence. Distractor-aware information is then appended to the tracking result of the baseline tracker to improve the intra-class discriminative power. BSPT contains a total of 120 sequences with rich annotations. Each sequence covers 8 specialized frame-level attributes from soccer scenarios and the player occlusion situations are finely divided into 4 categories for a more comprehensive comparison. In the experimental section, the performance of our algorithm and the other 14 compared trackers are evaluated on BSPT with detailed analysis. Experimental results reveal the effectiveness of the proposed distractor-aware model especially under the attribute of occlusion. The BSPT benchmark and raw experimental results are available on the project page at http://media.hust.edu.cn/BSPT.htm.
Zikai Song, Zhiwen Wan, Junqing Yu, Yi-Ping Phoebe Chen
ICMR5
2021 Shot Boundary Detection Through Multi-stage Deep Convolution Neural Network
Tingting Wang 0003, Na Feng, Junqing Yu, Yunfeng He, Yangliu Hu, Yi-Ping Phoebe Chen
MMM (1)3
2021 TSSBV: A Conflict-Free Flow Rule Management Algorithm in SDN Switches
abstract
Software-Defined Network (SDN) radically changes the network architecture in a mobile broadband network by decoupling the network logic from the underlying forwarding devices. The application in 5G systems will face the challenge of detecting conflict policy provided by the SDN controller. Multiple active network functions in 5G systems with the same priority will potentially trigger conflicts among policies with overlapped flow space, causing the flow table explosion. Focusing on the problem that state-of-art flow rule management methods may cause high consumption of ternary content addressable memory (TCAM) during the process of conflict detection and resolution, we propose TSSBV. Our TSSBV algorithm groups the flow entries with the same prefix length together and decreases the redundant bit in vectors. Also, we present techniques for global prioritization of flow rules in SDN switches and provide dynamic priorities for unassisted resolution of these conflicts. We demonstrate the effectiveness and feasibility of TSSBV in a simulated SDN networks through generating synthetic flows that reflect the characteristics of the flows used in real OpenFlow scenarios. Both the emulations and results confirm that TSSBV is more efficient than state-of-art methods in terms of less initialization time and lower conflict search time.
Qizhao Zhou, Junqing Yu
VTC Spring2
2021 A visibility-based surface reconstruction method on the GPU
Hailong Pan, Keyang Luo, Yawei Luo, Junqing Yu
Comput. Aided Geom. Des.5
2021 A dynamic and lightweight framework to secure source addresses in the SDN-based networks
Qizhao Zhou, Junqing Yu
Comput. Networks2
2020 Fine-Grain Level Sports Video Search Engine
Zikai Song, Junqing Yu, Hengyou Cai, Yangliu Hu, Yi-Ping Phoebe Chen
MMM (1)2
2020 Im-OFDP: An Improved OpenFlow-based Topology Discovery Protocol for Software Defined Network
Yongpu Gu, Junqing Yu
Networking3
2020 Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation
abstract
We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be available when learning to adapt. This setting is realistic but more challenging, in which conventional adaptation approaches are prone to failure due to the scarce of unlabeled target data. To this end, we propose a novel Adversarial Style Mining approach, which combines the style transfer module and task-specific module into an adversarial manner. Specifically, the style transfer module iteratively searches for harder stylized images around the one-shot target sample according to the current learning state, leading the task model to explore the potential styles that are difficult to solve in the almost unseen target domain, thus boosting the adaptation performance in a data-scarce scenario. The adversarial learning framework makes the style transfer module and task-specific module benefit each other during the competition. Extensive experiments on both cross-domain classification and segmentation benchmarks verify that ASM achieves state-of-the-art adaptation performance under the challenging one-shot setting.
Yawei Luo, Ping Liu 0004, Junqing Yu, Yi Yang 0001
NeurIPS4
2020 SSET: a dataset for shot segmentation, event detection, player tracking in soccer videos
Na Feng, Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Yizhu Zhao, Yunfeng He
Multim. Tools Appl.3
2020 Every node counts: Self-ensembling graph convolutional networks for semi-supervised learning
Yawei Luo, Rongrong Ji, Junqing Yu, Ping Liu 0004, Yi Yang 0001
Pattern Recognit.4
2019 Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation
abstract
We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distributions of the two domains to be similar. A popular strategy is to align the marginal distribution in the feature space through adversarial learning. However, this global alignment strategy does not consider the local category-level feature distribution. A possible consequence of the global movement is that some categories which are originally well aligned between the source and target may be incorrectly mapped. To address this problem, this paper introduces a category-level adversarial network, aiming to enforce local semantic consistency during the trend of global alignment. Our idea is to take a close look at the category-level data distribution and align each class with an adaptive adversarial loss. Specifically, we reduce the weight of the adversarial loss for category-level aligned features while increasing the adversarial force for those poorly aligned. In this process, we decide how well a feature is category-level aligned between source and target by a co-training approach. In two domain adaptation tasks, i.e., GTA5 → Cityscapes and SYNTHIA → Cityscapes, we validate that the proposed method matches the state of the art in segmentation accuracy.
Yawei Luo, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
CVPR4
2019 Significance-Aware Information Bottleneck for Domain Adaptive Semantic Segmentation
abstract
For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image classification, but usually fails in semantic segmentation tasks in which the latent representations are overcomplex. In this work, we equip the adversarial network with a “significance-aware information bottleneck (SIB)”, to address the above problem. The new network structure, called SIBAN, enables a significance-aware feature purification before the adversarial adaptation, which eases the feature alignment and stabilizes the adversarial training course. In two domain adaptation tasks, i.e., GTA5 → Cityscapes and SYNTHIA → Cityscapes, we validate that the proposed method can yield leading results compared with other feature-space alternatives. Moreover, SIBAN can even match the state-of-the-art output-space methods in segmentation accuracy, while the latter are often considered to be better choices for domain adaptive segmentation task.
Yawei Luo, Ping Liu 0004, Junqing Yu, Yi Yang 0001
ICCV4
2019 Soccer Video Event Detection Based on Deep Learning
Junqing Yu, Aiping Lei, Yangliu Hu
MMM (2)1
2019 Vector quantization: a review
abstract
Vector quantization (VQ) is a very effective way to save bandwidth and storage for speech coding and image coding. Traditional vector quantization methods can be divided into mainly seven types, tree-structured VQ, direct sum VQ, Cartesian product VQ, lattice VQ, classified VQ, feedback VQ, and fuzzy VQ, according to their codebook generation procedures. Over the past decade, quantization-based approximate nearest neighbor (ANN) search has been developing very fast and many methods have emerged for searching images with binary codes in the memory for large-scale datasets. Their most impressive characteristics are the use of multiple codebooks. This leads to the appearance of two kinds of codebook: the linear combination codebook and the joint codebook. This may be a trend for the future. However, these methods are just finding a balance among speed, accuracy, and memory consumption for ANN search, and sometimes one of these three suffers. So, finding a vector quantization method that can strike a balance between speed and accuracy and consume moderately sized memory, is still a problem requiring study.
Zebin Wu 0002, Junqing Yu
Frontiers Inf. Technol. Electron. Eng.2
2019 A multi-level descriptor using ultra-deep feature for image retrieval
Zebin Wu 0002, Junqing Yu
Multim. Tools Appl.2
2018 Macro-Micro Adversarial Network for Human Parsing
Yawei Luo, Zhedong Zheng, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
ECCV (9)5
2017 Optimized residual vector quantization for efficient approximate nearest neighbor search
Liefu Ai, Junqing Yu, Zebin Wu 0002, Yunfeng He
Multim. Syst.2
2017 Soccer Video Event Annotation by Synchronization of Attack-Defense Clips and Match Reports With Coarse-Grained Time Information
abstract
The annotation of significant events within soccer videos is a fundamental step in content-based soccer video retrieval. In this paper, we propose a soccer video annotation approach based on semantic matching with coarse time constraints, where video events and external text information (match reports) are synchronized using their semantic correspondence in the temporal sequence. Unlike the state-of-the-art soccer video analysis methods that assume that the time of an event's occurrence is given precisely by the external text information, this paper considers the problem of annotating soccer videos using match reports with coarse-grained time information. The contributions of the proposed approach are: 1) we propose a more generalized approach that synchronizes video events with text descriptions using high-level semantics with coarse time constraints, rather than assuming that the timestamp is given exactly in the text description; 2) the detection of event boundaries is improved by attack-defense transition analysis (ADTA); 3) a robust and fast center circle detection algorithm is proposed for the classification of soccer field zones and ADTA; and 4) unlike conventional audio-based whistle detection, we propose a novel Hough-transform-based algorithm from the perspective of image processing. This allows the game start time to be detected, and further helps the synchronization of video and text events. The experimental results conducted on a large number of soccer videos validate the effectiveness of the proposed approach.
Zengkai Wang, Junqing Yu, Yunfeng He
IEEE Trans. Circuits Syst. Video Technol.2
2016 Accurate localization for mobile device using a multi-planar city model
abstract
This paper presents a novel method for estimating the unknown 6DOF pose of a mobile device. The method is based on matching between the mobile image and the virtual city model which is merely composed of 3D points on planar building facade. The main contributions of this paper are as follows: firstly, we design a new plane generation strategy which fuses the 3D model points, photo homography and the orientation of the buildings together within RANSAC framework. Secondly, we propose a novel energy-based method which can parallel solve the mobile poses as well as the best 2D-3D matches. Thirdly, a client/server mode is established to support a speedy localization experience on the mobile devices. To the best of our knowledge, this is the first implementation that uses such a multi-planar model to accurately locate the mobile device in a scene of city scale. Experiment shows that the localization performance becomes faster and more robust comparing with other methods when adding the multi-planar information of the buildings to localization algorithm.
Yawei Luo, Hailong Pan, Yuesong Wang 0001, Junqing Yu
ICPR5
2016 Dense 3D reconstruction combining depth and RGB information
Hailong Pan, Yawei Luo, Liya Duan, Liu Yi, Yizhu Zhao, Junqing Yu
Neurocomputing8
2015 Fast terrain mapping from low altitude digital imagery
Yawei Luo, Benchang Wei, Hailong Pan, Junqing Yu
Neurocomputing5
2015 Wide area localization and tracking on camera phones for mobile augmented reality systems
Benchang Wei, Liya Duan, Junqing Yu, Tan Mao
Multim. Syst.4
2014 A scalable flow rule translation implementation for software defined security
abstract
Software defined networking brings many possibilities to network security, one of the most important security challenge it can help with is the possibility to make network traffic pass through specific security devices, in other words, determine where to deploy these devices logically. However, most researches focus on high level policy and interaction framework but ignored how to translate them to low-level OpenFlow rules with scalability. We analyze different actions used in common security scenarios and resource constraints of physical switch. Based on them, we propose a rule translation implementation which can optimize the resource consumption according to different actions by selecting forward path dynamically.
Junqing Yu
APNOMS4
2014 Affection arousal based highlight extraction for soccer video
Zengkai Wang, Junqing Yu, Yunfeng He
Multim. Tools Appl.2
2013 StreamTMC: Stream compilation for tiled multi-core architectures
Haitao Wei, Mingkang Qin, Junqing Yu, Dongrui Fan, Guang R. Gao
J. Parallel Distributed Comput.4
2013 High-dimensional indexing technologies for large scale content-based image retrieval: a review
abstract
The boom of Internet and multimedia technology leads to the explosion of multimedia information, especially image, which has created an urgent need of quickly retrieving similar and interested images from huge image collections. The content-based high-dimensional indexing mechanism holds the key to achieving this goal by efficiently organizing the content of images and storing them in computer memory. In the past decades, many important developments in high-dimensional image indexing technologies have occurred to cope with the ‘curse of dimensionality’. The high-dimensional indexing mechanisms can mainly be divided into three categories: tree-based index, hashing-based index, and visual words based inverted index. In this paper we review the technologies with respect to these three categories of mechanisms, and make several recommendations for future research issues.
Liefu Ai, Junqing Yu, Yunfeng He
J. Zhejiang Univ. Sci. C2
2013 On-Device Mobile Visual Location Recognition by Integrating Vision and Inertial Sensors
abstract
This paper deals with the problem of city scale on-device mobile visual location recognition by fusing the inertial sensors and computer vision techniques. The main contributions are as follows: Firstly, we design an efficient vector quantization strategy by combining the Transform Coding (TC) and Residual Vector Quantization (RVQ). Our method can compress a visual descriptor into only several bytes while providing reasonable searching accuracy, which makes the managing of city scale image database directly on mobile devices come true. Secondly, we integrate the information from inertial sensors into the Vector of Locally Aggregated Descriptors (VLAD) generation and image similarity evaluation processes. Our method is not only fast enough for on-device implementation, but it also can improve the location recognition accuracy obviously. Thirdly, we also release a set of 1.295 million geo-tagged street view images with the information from inertial sensors, as well as a difficult set of query images. These resources can be used as a new benchmark to facilitate further research in the area. Experimental results prove the validity of the proposed methods for on-device mobile visual location recognition applications.
Yunfeng He, Juan Gao, Jianzhong Yang, Junqing Yu
IEEE Trans. Multim.5
2012 Software Pipelining for Stream Programs on Resource Constrained Multicore Architectures
abstract
Stream programming model has been productively applied to a number of important application domains. Software pipelining is an important code scheduling technique for stream programs. However, the multicore evolution has presented a new dimension of challenges: that is how to orchestrate the best software pipelining schedule in the face of resource constrained architectures (e.g., number of cores, available memory, and bandwidth)? In this paper, we proposed a new solution methodology to address the problem above. Our main contributions include the following. A unified Integer Linear Programming (ILP) formulation has been proposed that combines the requirement of both rate-optimal software pipelining and the minimization of intercore communication overhead. Next, an extended formulation has been proposed to formulate the schedule under memory size constrained systems. It orchestrates the rate-optimal software pipelining execution for stream programs with strict memory, processor cores, and communication constraints. A solution testbed has been implemented for the proposed problem formulations. This has been realized by extending the Brook programming environment with our software pipelining support-named DFBrook. An experimental study has been conducted to verify the effectiveness of the proposed solutions.
Haitao Wei, Junqing Yu, Huafei Yu, Mingkang Qin, Guang R. Gao
IEEE Trans. Parallel Distributed Syst.2
2010 Minimizing communication in rate-optimal software pipelining for stream programs
abstract
Stream programming model has been productively applied to a number of important applications domains. Software pipelining is an important code scheduling technique for stream programs. However, the multi-/many-core evolution has presented a new dimension of challenges: that is while searching a best software pipelining schedule how to ensure the communications between processing cores are also minimized? In this paper, we proposed a new solution methodology to address the above problem. Our main contributions include the following. A unified formulation has been proposed that combines the requirement of both rate-optimal software pipelining and the minimization of inter-core communication overhead. This formulation has been developed based on a synchronized dataflow graph model, and is expressed as an integer linear programming problem. A solution testbed has been implemented for the proposed problem formulation on the IBM Cell architecture. This has been realized by extending the Brook stream programming environment with our software pipelining support -- named DFBrook. An experimental study has been conducted to verify the effectiveness of the proposed solution. And a comparison of other scheduling methods has demon-strated the performance superiority of our proposed method.
Haitao Wei, Junqing Yu, Huafei Yu, Guang R. Gao
CGO2
2009 An improved valence-arousal emotion space for video affective content representation and recognition
abstract
To understand video affective content automatically, the primary task is to transform the abstract concept of emotion into the form which can be handled by the computer easily. An improved V-A emotion space is proposed to address this problem. It unifies the discrete and dimensional emotion model by introducing the typical fuzzy emotion subspace. Fuzzy C-mean clustering (FCM) algorithm is adopted to divide the V-A emotion space into the subspaces and Gaussian mixture model (GMM) is used to determine their membership functions. Based on the proposed emotion space, the maximum membership principle and the threshold principle are introduced to represent and recognize video affective content. A video affective content database is created to validate the proposed model. The experimental results show that the improved emotion space can be used as a solution to represent and recognize video affective content.
Junqing Yu, Xiaoqiang Hu
ICME2
2007 Video Affective Content Representation and Recognition Using Video Affective Tree and Hidden Markov Models
Junqing Yu
ACII2
2007 Video Affective Content Recognition Based on Genetic Algorithm Combined HMM
Junqing Yu
ICEC2
2007 Video Processing and Retrieval on Cell Processor Architecture
Junqing Yu, Haitao Wei
ICEC1
2005 Content-Based News Video Mining
Junqing Yu, Yunfeng He, Shijun Li 0001
ADMA1
2005 Ontology-Based HTML to XML Conversion
Shijun Li 0001, Weijie Ou, Junqing Yu
WAIM3