EDBT 2026 Demo / reviewers in the wild / expert
Jinhui Yi
dblp:276/5919
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language ModelsabstractWe present an efficient encoder-free approach for video-language understanding that achieves competitive performance while significantly reducing computational overhead. Current video-language models typically rely on heavyweight image encoders (300M-1.1B parameters) or video encoders (1B-1.4B parameters), creating a substantial computational burden when processing multi-frame videos. Our method introduces a novel Spatio-Temporal Alignment Block (STAB) that directly processes video inputs without requiring pre-trained encoders while using only 45M parameters for visual processing - at least a 6.5× reduction compared to traditional approaches. The STAB architecture combines Local Spatio-Temporal Encoding for fine-grained feature extraction, efficient spatial downsampling through learned attention and separate mechanisms for modeling frame-level and video-level relationships. Our model achieves comparable or superior performance to encoder-based approaches for open-ended video question answering on standard benchmarks. The fine-grained video question-answering evaluation demonstrates our model’s effectiveness, outperforming the encoder-based approaches Video-ChatGPT and Video-LLaVA in key aspects like correctness and temporal understanding. Extensive ablation studies validate our architectural choices and demonstrate the effectiveness of our spatio-temporal modeling approach while achieving 3-4× faster processing speeds than previous methods. Code is available at https://jh-yi.github.io/Video-Panda. Jinhui Yi, Syed Talal Wasim, Yanan Luo, Muzammal Naseer, Juergen Gall |
CVPR | 1 |
| 2025 | PeT-KeyStAtion: Parameter-efficient Transformer with Keypoint-guided Spatial-temporal Aggregation for Video-based Person Re-identificationabstractVideo-based Person Re-identification (ReID) is crucial in visual surveillance, focusing on matching video snippets of individuals across multiple non-overlapping cameras. Existing methods either conduct ReID at the image level without leveraging temporal information, or employ complex temporal information aggregation techniques, which results in substantial network size and reduced performance efficiency. Recent advances in Vision Transformer (ViT) architectures leverage diverse large-scale datasets alongside sophisticated architectures to achieve enhanced fine-grained feature discrimination. To fully explore the potential of ViT architectures without adding substantial additional modules for video-based ReID, we propose PeT-KeyStAtion: a Parameter-efficient Transformer with Keypoint- guided Spatial-temporal Aggregation using a Spatial-Temporal and Keypoint (STK) Module with lightweight adapters. Our framework effectively captures and aggregates spatial, temporal, and keypoint information with only 11% of the parameters compared to full fine-tuning. Extensive experiments show that our method outperforms state-of-the-art baselines on MARS and iLIDS-VID, and achieves promising performance on LS-VID. Xingan Ma, Jinhui Yi, Juergen Gall |
ICASSP | 2 |
| 2024 | MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
Jinhui Yi, Yanan Luo, Marion Deichmann, Gabriel Schaaf, Juergen Gall |
BMVC | 1 |
| 2024 | Learning to Estimate Package Delivery Time in Mixed Imbalanced Delivery and Pickup Logistics ServicesabstractAccurately estimating package delivery time is essential to the logistics industry, which enables reasonable work allocation and on-time service guarantee. This becomes even more necessary in mixed logistics scenarios where couriers handle a high volume of delivery and a smaller number of pickup simultaneously. However, most of the related works treat the pickup and delivery patterns on couriers' decision behavior equally, neglecting that the pickup has a greater impact on couriers' decision-making compared to the delivery due to its tighter time constraints. In such context, we have three main challenges: 1) multiple spatiotemporal factors are intricately interconnected, significantly affecting couriers' delivery behavior; 2) pickups have stricter time requirements but are limited in number, making it challenging to model their effects on couriers' delivery process; 3) couriers' spatial mobility patterns are critical determinants of their delivery behavior, but have been insufficiently explored. To deal with these, we propose TransPDT, a Transformer-based multi-task package delivery time prediction model. We first employ the Transformer encoder architecture to capture the spatio-temporal dependencies of couriers' historical travel routes and pending package sets. Then we design the pattern memory to learn the patterns of pickup in the imbalanced dataset via attention mechanism. We also set the route prediction as an auxiliary task of delivery time prediction, and incorporate the prior courier spatial movement regularities in prediction. Extensive experiments on real industry-scale datasets demonstrate the superiority of our method. A system based on TransPDT is deployed internally in JD Logistics to track more than 2000 couriers handling hundreds of thousands of packages per day in Beijing, and the average daily delivery timely rate of deployed stations is 0.68% higher than the non-deployed stations. Jinhui Yi, Huan Yan 0003, Haotian Wang 0008, Yong Li 0008 |
SIGSPATIAL/GIS | 1 |
| 2024 | Rethinking Temporal Self-Similarity For Repetitive Action CountingabstractCounting repetitive actions in long untrimmed videos is a challenging task that has many applications such as rehabilitation. State-of-the-art methods predict action counts by first generating a temporal self-similarity matrix (TSM) from the sampled frames and then feeding the matrix to a predictor network. The self-similarity matrix, however, is not an optimal input to a network since it discards too much information from the frame-wise embeddings. We thus rethink how a TSM can be utilized for counting repetitive actions and propose a framework that learns embeddings and predicts action start probabilities at full temporal resolution. The number of repeated actions is then inferred from the action start probabilities. In contrast to current approaches that have the TSM as an intermediate representation, we propose a novel loss based on a generated reference TSM, which enforces that the self-similarity of the learned frame-wise embeddings is consistent with the self-similarity of repeated actions. The proposed framework achieves state-of-the-art results on three datasets, i.e., RepCount, UCFRep, and Countix. Yanan Luo, Jinhui Yi, Yazan Abu Farha, Moritz Wolter, Juergen Gall |
ICIP | 2 |
| 2024 | A Multistar Topological Features Guided Point-Like Space Target Detection FrameworkabstractPoint-like space target detection (PSTD) is a critical technology for space situation awareness. However, it grapples with challenges arising from intense noise, stray light, and stars resembling targets, making it difficult for existing methods to differentiate between point-like targets and stars. To address these issues, we propose a PSTD framework guided by local target continuity, global background volatility, and multistar topological features. First, we tailor a background estimation approach to minimize noise and stray light while preserving target details. Then, we introduce a novel K-vector-based star angular distance (KV-SAD) voting strategy, which uses multistar topological features to accurately identify key points. Finally, using these identified key points, we apply a fast linear attitude estimator (FLAE) to create a star template matrix and set up redundant windows for separating similarly shaped targets from stars. Extensive experiments across four space image sequences achieved an average AUC($P_{d}, \tau $) of 93.36%, which is 12.79% higher than the best baseline method, showing the best balance between detection rate and false alarm rate. This significant improvement underscores the strong generalization of our novel PSTD framework to diverse and challenging space scenarios. Bingbing Dan, Hongfeng Long, Jinhui Yi, Enhai Liu, Yuebo Ma, Rujin Zhao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | RCCNet: A Spatial-Temporal Neural Network Model for Logistics Delivery Timely Rate PredictionabstractIn logistics service, the delivery timely rate is a key experience indicator, which is highly essential to the competitive advantage of express companies. Prediction on it enables intervention on couriers with low predicted results in advance, thus ensuring employee productivity and customer satisfaction. Currently, few related works focus on couriers’ level delivery timely rate prediction, and there are complex spatial correlations between couriers and road districts in the express scenario, which makes traditional real-time prediction approaches hard to utilize. To deal with this, we propose a deep spatial-temporal neural network, RCCNet to model spatial-temporal correlations. Specifically, we adopt Node2vec, which can encode the road network-based graph directly to capture spatial correlations between road districts. Further, we calculate couriers’ historical time-series similarity to build a graph and employ graph convolutional networks to capture the correlation between couriers. We also leverage historical sequential information with long short-term memory networks. We conduct experiments with real-world express datasets. Compared with other competitive baseline methods widely used in industry, the experiment results demonstrate its superior performance over multiple baselines. Jinhui Yi, Huan Yan 0003, Haotian Wang 0008, Yong Li 0008 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | DeepSTA: A Spatial-Temporal Attention Network for Logistics Delivery Timely Rate Prediction in Anomaly ConditionsabstractPrediction of couriers' delivery timely rates in advance is essential to the logistics industry, enabling companies to take preemptive measures to ensure the normal operation of delivery services. This becomes even more critical during anomaly conditions like the epidemic outbreak, during which couriers' delivery timely rate will decline markedly and fluctuates significantly. Existing studies pay less attention to the logistics scenario. Moreover, many works focusing on prediction tasks in anomaly scenarios fail to explicitly model abnormal events, e.g., treating external factors equally with other features, resulting in great information loss. Further, since some anomalous events occur infrequently, traditional data-driven methods perform poorly in these scenarios. To deal with them, we propose a deep spatial-temporal attention model, named DeepSTA. To be specific, to avoid information loss, we design an anomaly spatio-temporal learning module that employs a recurrent neural network to model incident information. Additionally, we utilize Node2vec to model correlations between road districts, and adopt graph neural networks and long short-term memory to capture the spatial-temporal dependencies of couriers. To tackle the issue of insufficient training data in abnormal circumstances, we propose an anomaly pattern attention module that adopts a memory network for couriers' anomaly feature patterns storage via attention mechanisms. The experiments on real-world logistics datasets during the COVID-19 outbreak in 2022 show the model outperforms the best baselines by 12.11% in MAE and 13.71% in MSE, demonstrating its superior performance over multiple competitive baselines. Jinhui Yi, Huan Yan 0003, Haotian Wang 0008, Yong Li 0008 |
CIKM | 1 |
| 2023 | DAS: Efficient Street View Image Sampling for Urban PredictionabstractStreet view data is one of the most common data sources for urban prediction tasks, such as estimating socioeconomic status, sensing physical urban changes, and identifying urban villages. Typical research in this field consists of two steps: acquiring a dataset with a street view image sampling algorithm and designing a prediction algorithm for urban prediction tasks. However, most of the previous research focuses on the prediction algorithms, leaving the sampling algorithms underexplored. To fill this gap, we set out to investigate how different street view image sampling algorithms affect the performance of the follow-up tasks and develop an effective street view image sampling algorithm for urban prediction. Through a comprehensive analysis of the performance of different sampling algorithms in three of the most common urban prediction tasks, including commercial activeness prediction, urban liveliness prediction, and urban population prediction, we provide solid empirical evidence that the sampling algorithm significantly affects the performance of the prediction model. Specifically, the performance differences of different sampling algorithms can reach over 25%. Further, we revealed that the sampling step size and the sampling quality are two important factors that affect the performance of a sampling algorithm, while the sampling angle has little influence. Inspired by our analysis results, we propose an effective street view image sampling algorithm, DAS, which contains a denoising module and an adaptive sampling module. It can dynamically adjust the sampling step size to adapt to the optimal size for each region and get rid of the impact of noise images in the meantime. Experiments on three large-scale datasets demonstrate its superior performance over multiple state-of-the-art baselines, and further ablation study shows the effectiveness of each module. Finally, through a thorough discussion of our findings and experimental results, we provide insights into the street view image sampling algorithm design, and we call for more researches in this blank area. Guozhen Zhang 0001, Jinhui Yi, Yong Li 0008, Depeng Jin |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Spatial-Temporal Consistency Network for Low-Latency Trajectory ForecastingabstractTrajectory forecasting is a crucial step for autonomous vehicles and mobile robots in order to navigate and interact safely. In order to handle the spatial interactions between objects, graph-based approaches have been proposed. These methods, however, model motion on a frame-to-frame basis and do not provide a strong temporal model. To overcome this limitation, we propose a compact model called Spatial-Temporal Consistency Network (STC-Net). In STC-Net, dilated temporal convolutions are introduced to model long-range dependencies along each trajectory for better temporal modeling while graph convolutions are employed to model the spatial interaction among different trajectories. Furthermore, we propose a feature-wise convolution to generate the predicted trajectories in one pass and refine the forecast trajectories together with the reconstructed observed trajectories. We demonstrate that STC-Net generates spatially and temporally consistent trajectories and outperforms other graph-based methods. Since STC-Net requires only 0.7k parameters and forecasts the future with a latency of only 1.3ms, it advances the state-of-the-art and satisfies the requirements for realistic applications. Shijie Li 0006, Yanying Zhou, Jinhui Yi, Juergen Gall |
ICCV | 3 |