EDBT 2026 Demo / reviewers in the wild / expert
Jinhai Yang 0001
dblp:98/2303
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0003-1101-6705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder EstimationabstractAdaptive video streaming requires efficient bitrate ladder construction to meet heterogeneous network conditions and end-user demands. Per-title encoding optimization typically traverses numerous encoding parameters to search the Pareto-optimal operating points for each video. Recently, researchers have attempted to predict the content-optimized bitrate ladder for pre-encoding overhead reduction [1] . However, as shown in Fig. 1 , current methods actually estimate the optimal encoding parameters that lie on the Pareto front and thus still require subsequent pre-encodings. Jinhai Yang 0001, Mengxi Guo, Shijie Zhao 0001, Li Zhang 0006 |
DCC | 1 |
| 2024 | Joint Intra & Inter-Grained Reasoning: A New Look Into Semantic Consistency of Image-Text RetrievalabstractMultimodal understanding aims at constructing semantic correlations among modalities of data while performing various downstream tasks. As one of the primary multimodal downstream tasks, image-text retrieval imposes a high demand on semantic alignment because of the independent expression paradigms of images and text. Existing methods mainly construct a joint embedding space at a single granularity level (either global or local). However, such single reasoning paradigms lack granularity interaction, resulting in semantic inconsistency and cross-domain catastrophes. To address these issues, we design a novel Joint Intra and Inter-grained Network (JIIGNet), focusing on not only intra- but also inter-grained interaction between modalities by combining scene information (global) with region-level (local) instances. Specifically, we simultaneously initiate three specific alignment modules, i.e., global-grained, local-grained, and cross-grained alignment modules, followed by Triplet Attention Refinement to better refine the fused embedding at the alignment-level with proper self and cross attention. For different scenarios, a Style Adaptation Head is further designed to smartly accommodate different samples. We validate JIIGNet through extensive experiments conducted on two widely used datasets: Flickr-30 K and MS-COCO, demonstrating the effectiveness of our proposed method. Renjie Pan 0001, Hua Yang 0001, Cunyan Li, Jinhai Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Psychology-Guided Environment Aware Network for Discovering Social Interaction Groups from VideosabstractSocial interaction is a common phenomenon in human societies. Different from discovering groups based on the similarity of individuals’ actions, social interaction focuses more on the mutual influence between people. Although people can easily judge whether or not there are social interactions in a real-world scene, it is difficult for an intelligent system to discover social interactions. Initiating and concluding social interactions are greatly influenced by an individual’s social cognition and the surrounding environment, which are closely related to psychology. Thus, converting the psychological factors that impact social interactions into quantifiable visual representations and creating a model for interaction relationships poses a significant challenge. To this end, we propose a Psychology-Guided Environment Aware Network (PEAN) that models social interaction among people in videos using supervised learning. Specifically, we divide the surrounding environment into scene-aware visual-based and human-aware visual-based descriptions. For the scene-aware visual clue, we utilize 3D features as global visual representations. For the human-aware visual clue, we consider instance-based location and behaviour-related visual representations to map human-centred interaction elements in social psychology: distance, openness, and orientation. In addition, we design an environment aware mechanism to integrate features from visual clues, with a Transformer to explore the relation between individuals and construct pairwise interaction strength features. The interaction intensity matrix reflecting the mutual nature of the interaction is obtained by processing the interaction strength features with the interaction discovery module. An interaction constrained loss function composed of interaction critical loss function and smoothFβloss function is proposed to optimize the whole framework to improve the distinction of the interaction matrix and alleviate class imbalance caused by pairwise interaction sparsity. Given the diversity of real-world interactions, we collect a new dataset named Social Basketball Activity Dataset (Soical-BAD), covering complex social interactions. Our method achieves the best performance among social-CAD, social-BAD, and their combined dataset named Video Social Interaction Dataset (VSID). Jinhai Yang 0001, Hua Yang 0001, Renjie Pan 0001, Pingrui Lai, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Self-Asymmetric Invertible Network for Compression-Aware Image RescalingabstractHigh-resolution (HR) images are usually downscaled to low-resolution (LR) ones for better display and afterward upscaled back to the original size to recover details. Recent work in image rescaling formulates downscaling and upscaling as a unified task and learns a bijective mapping between HR and LR via invertible networks. However, in real-world applications (e.g., social media), most images are compressed for transmission. Lossy compression will lead to irreversible information loss on LR images, hence damaging the inverse upscaling procedure and degrading the reconstruction accuracy. In this paper, we propose the Self-Asymmetric Invertible Network (SAIN) for compression-aware image rescaling. To tackle the distribution shift, we first develop an end-to-end asymmetric framework with two separate bijective mappings for high-quality and compressed LR images, respectively. Then, based on empirical analysis of this framework, we model the distribution of the lost information (including downscaling and compression) using isotropic Gaussian mixtures and propose the Enhanced Invertible Block to derive high-quality/compressed LR images in one forward pass. Besides, we design a set of losses to regularize the learned LR images and enhance the invertibility. Extensive experiments demonstrate the consistent improvements of SAIN across various image rescaling datasets in terms of both quantitative and qualitative evaluation under standard image compression formats (i.e., JPEG and WebP). Code is available at https://github.com/yang-jin-hai/SAIN. Jinhai Yang 0001, Mengxi Guo, Shijie Zhao 0001, Li Zhang 0006 |
AAAI | 1 |
| 2023 | Sine: Similarity-Regularized Intra-Class Exploitation for Cross-Granularity Few-Shot LearningabstractFew-shot learning aims for rapid adaptation with few samples. Recently, cross-granularity few-shot learning has emerged as a promising research area, where models observe coarse labels but target fine-grained recognition among novel classes. As coarse supervision tends to eliminate feature discrimination among underlying sub-classes, existing methods commonly utilize self-supervision as a complement to explore intra-class variation. However, current methods suffer from an intrinsic conflict between contrastive learning and coarse supervision. In this paper, we locate the root cause of the intrinsic conflict. Then, we resolve it by exploiting the similarity among augmented views while ignoring the unreasonable constraint between negative pairs. Besides, we decouple contrastive learning and coarse supervision into parallel branches to better regularize the latent space. Albeit simple, our approach consistently outperforms state-of-the-art methods across different benchmarks. Jinhai Yang 0001, Hua Yang 0001 |
ICASSP | 1 |
| 2021 | PRN: Psychology-Inspired Relation Network for Detecting Social Interaction Groups from Single Images
Jinhai Yang 0001, Hua Yang 0001, Guangtao Zhai |
BMVC | 2 |
| 2021 | MPASNET: Motion Prior-Aware Siamese Network For Unsupervised Deep Crowd Segmentation In Video ScenesabstractCrowd segmentation is a fundamental task serving as the basis of crowded scene analysis, and it is highly desirable to obtain refined pixel-level segmentation maps. However, it remains a challenging problem, as existing approaches either require dense pixel-level annotations to train deep learning models or merely produce rough segmentation maps from optical or particle flows with physical models. In this paper, we propose the Motion Prior-Aware Siamese Network (MPASNET) for unsupervised crowd semantic segmentation. This model not only eliminates the need for annotation but also yields high-quality segmentation maps. Specially, we first analyze the coherent motion patterns across the frames and then apply a circular region merging strategy on the collective particles to generate pseudo-labels. Moreover, we equip MPASNET with siamese branches for augmentation-invariant regularization and siamese feature aggregation. Experiments over benchmark datasets indicate that our model outperforms the state-of-the-arts by more than 12% in terms of mIoU. Jinhai Yang 0001, Hua Yang 0001 |
ICIP | 1 |
| 2021 | Towards Cross-Granularity Few-Shot Learning: Coarse-to-Fine Pseudo-Labeling with Visual-Semantic Meta-EmbeddingabstractFew-shot learning aims at rapidly adapting to novel categories with only a handful of samples at test time, which has been predominantly tackled with the idea of meta-learning. However, meta-learning approaches essentially learn across a variety of few-shot tasks and thus still require large-scale training data with fine-grained supervision to derive a generalized model, thereby involving prohibitive annotation cost. In this paper, we advance the few-shot classification paradigm towards a more challenging scenario, i.e, cross-granularity few-shot classification, where the model observes only coarse labels during training while is expected to perform fine-grained classification during testing. This task largely relieves the annotation cost since fine-grained labeling usually requires strong domain-specific expertise. To bridge the cross-granularity gap, we approximate the fine-grained data distribution by greedy clustering of each coarse-class into pseudo-fine-classes according to the similarity of image embeddings. We then propose a meta-embedder that jointly optimizes the visual- and semantic-discrimination, in both instance-wise and coarse class-wise, to obtain a good feature space for this coarse-to-fine pseudo-labeling process. Extensive experiments and ablation studies are conducted to demonstrate the effectiveness and robustness of our approach on three representative datasets. Jinhai Yang 0001, Hua Yang 0001, Lin Chen 0019 |
ACM Multimedia | 1 |