EDBT 2026 Demo / reviewers in the wild / expert
Jun-Yan He
dblp:173/3747
· DBLP profile ↗
36ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0002-6628-6924ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal DiffusionabstractCurrent text-to-image models face challenges in visual text rendering: text encoders like CLIP and T5 lack glyph-level understanding and often struggle to distinguish between the specific words to be rendered and their intended semantic meaning within prompts. In addition, inconsistencies between the base model and its plugins further compromise the quality of synthesized images. In this paper, we enhance the existing text-to-image method by addressing the following aspects: (1) Text-Glyph Alignmentin a Visual Question Answering (VQA) manner to enable glyph understanding for the text encoder. This involves establishing an explicit alignment between the representations of the glyphs and their detailed attribute descriptions, which boosts the model's ability to capture fine-grained visual features of the text. (2) Accurate and harmony visual text rendering: integrating pre-aligned glyph-visual embeddings with semantic text tokens through the Multimodal Diffusion Transformer(MMDiT) synchronously, ensuring coherent feature alignment and enhancing both the robustness and fidelity of visual text rendering. (3) Image Aesthetic Refinement: leveraging a multisource data training strategy that incorporates diverse, high-quality image-text pairs from various domains, exposing the model to extensive linguistic and visual diversity while maintaining superior aesthetic quality throughout training. Our experiments demonstrate that the proposed approach significantly outperforms the existing state-of-the-art method. Lishuai Gao, Jun-Yan He, Yingsen Zeng, Xiaoming Wei |
AAAI | 2 |
| 2025 | POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchabstractAchieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs three key contributions: (1) Pseudo-range multilateration is utilized to correct heatmap errors, improving landmark localization accuracy. By integrating multiple anchor points, it reduces the impact of individual heatmap inaccuracies, leading to robust overall positioning. (2) To enhance the pseudo-range accuracy of selected anchor points, a new loss function, named multilateration anchor loss, is proposed. This loss function enhances the accuracy of the distance map, mitigates the risk of local optima, and ensures optimal solutions. (3) A single-step parallel computation algorithm is introduced, boosting computational efficiency and reducing processing time. Extensive evaluations across five benchmark datasets demonstrate that POPoS consistently outperforms existing methods, particularly excelling in low-resolution heatmaps scenarios with minimal computational overhead. These advantages make POPoS as a highly efficient and accurate tool for FLD, with broad applicability in real-world scenarios. Chong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu 0001, Xian-Sheng Hua 0001 |
AAAI | 2 |
| 2025 | UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal PromptsabstractEmotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module(EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. To facilitate reproducibility and further research, we have made our code publicly available at https://github.com/KTTRCDL/UMETTS. Zhi-Qi Cheng, Jun-Yan He, Junyao Chen, Xiaomao Fan, Xiaojiang Peng, Alex Hauptmann 0001 |
ICASSP | 3 |
| 2025 | Dual-Rate Dynamic Teacher for Source-Free Domain Adaptive Object Detection
Qi He 0007, Xiao Wu 0001, Jun-Yan He, Shuai Li 0014 |
ICCV | 3 |
| 2025 | MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt SynthesisabstractMetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the creation of customizable WordArt, ranging from semantic enhancements to intricate textural elements. A central feedback mechanism leverages insights from both multimodal models and user evaluations, enabling iterative refinement of design parameters. Through this iterative process, MetaDesigner dynamically adjusts hyperparameters to align with user-defined stylistic and thematic preferences, consistently delivering WordArt that excels in visual quality and contextual resonance. Empirical evaluations underscore the system's versatility and effectiveness across diverse WordArt applications, yielding outputs that are both aesthetically compelling and context-sensitive. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Jingdong Sun, Qi He 0007, Wangmeng Xiang, Jin-Peng Lan, Xianhui Lin, Kang Zhu, Bin Luo 0008, Yifeng Geng, Xuansong Xie, Alex Hauptmann 0001 |
ICLR | 1 |
| 2025 | Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose EstimationabstractAccurate 3D Human Pose Estimation (HPE) in video sequences demands both precision and a robust architectural framework. Building upon the recent success of transformers in computer vision, we introduce the Refined Temporal Pyramidal Compression-and-Amplification (RTPCA) transformer, an approach that tackles a critical issue in current transformer-based methods: the underutilization of intra-and inter-block relations through attention mechanisms. The cornerstone of our approach is the meticulously designed Temporal Pyramidal Compression-and-Amplification (TPCA) module, which ingeniously leverages a temporal pyramid paradigm to significantly enhance multi-scale key and value representations from intra-block attention. Recognizing that focusing solely on individual modules while overlooking their interconnections can limit performance, we introduce the Cross-Layer Refinement (XLR) module. This carefully crafted component is designed to amplify inter-block communication by linking keys and values across adjacent blocks, creating more coherent attention patterns. The seamless integration of TPCA and XLR results in a powerful synergy that facilitates a rich semantic representation through the dynamic interaction of queries, keys, and values. This synergistic approach enables the RTPCA transformer to achieve remarkable performance on leading benchmarks, such as Human3.6M, HumanEva-I, and MPI-INF-3DHP, with only a small computational overhead. We demonstrate the effectiveness of the RTPCA transformer through extensive experiments and comparisons with state-of-the-art methods. The source code is available at https://github.com/hbing-l/RTPCA.git. Zhi-Qi Cheng, Wangmeng Xiang, Jun-Yan He, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ICME | 4 |
| 2025 | DualEnhance: External Multimodal Foundation Models Guidance and Internal Fast-Slow Teacher RegulationabstractSource-Free Domain Adaptive Object Detection addresses cross-domain detection on an unlabeled target domain without accessing source data. Existing methods implement self-training with Mean Teacher but are bottlenecked by error accumulation from noisy pseudo-labels generated via recursive teacher-student updates. This issue is handled through the proposed dual enhancements: (1) External Guidance via Multimodal Foundation Models (FMs); (2) Internal Regulation through Fast-Slow Teacher. First, despite FMs' multimodal comprehension, their semantic misalignment with a specific task introduces noise during adaptation. Bidirectional Distillation mitigates this by calibrating the FM using task-specific knowledge transferred from the source detector. The aligned cross-modal knowledge then propagates through high-quality pseudo-label generation. Second, the conventional Mean Teacher suffers from plasticity-stability dilemma, where rapid adaptation corrupts historical knowledge. Fast-Slow Teacher introduces dual-velocity knowledge consolidation: The Fast Teacher dynamically captures emerging domain features, while the Slow Teacher preserves stable historical knowledge and periodically resets the Fast Teacher, establishing an error-correcting dynamic equilibrium. Experiments show our method achieves significant improvements over SOTA. Qi He 0007, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Zhaoquan Yuan |
ACM Multimedia | 3 |
| 2025 | GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph LayoutsabstractText logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (e.g., text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences. Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Chenyang Li 0007, Jin-Peng Lan, Jun-Yan He, Bin Luo 0008, Yifeng Geng |
ACM Multimedia | 8 |
| 2025 | DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming PerceptionabstractThe advancement of autonomous driving systems hinges on the ability to achieve low-latency and high-accuracy perception. To address this critical need, this paper introduces Dynamic Routering Network (DyRoNet), a low-rank enhanced dynamic routing framework designed for streaming perception in autonomous driving systems. DyRoNet integrates a suite of pre-trained branch networks, each meticulously fine-tuned to function under distinct environmental conditions. At its core, the framework offers a speed router module, developed to assess and route input data to the most suitable branch for processing. This approach not only addresses the inherent limitations of conventional models in adapting to diverse driving conditions but also ensures the balance between performance and efficiency. Extensive experimental evaluations demonstrating the adaptability of DyRoNet to diverse branch selection strategies, resulting in significant performance enhancements across different scenarios. This work not only establishes a new benchmark for streaming perception but also provides valuable engineering insights for future work.11Project: https://tastevision.github.io/DyRoNet/ Xiang Huang 0004, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li 0007, Wangmeng Xiang, Baigui Sun |
WACV | 3 |
| 2025 | UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
Zhi-Qi Cheng, Gabriel Moreira, Jiawen Zhu 0003, Jingdong Sun, Bukun Ren, Jun-Yan He, Qi Dai 0001, Xian-Sheng Hua 0001 |
WACV | 7 |
| 2025 | Exploring Dynamic Transformer for Efficient Object TrackingabstractThe speed-precision tradeoff is a critical problem in visual object tracking, as it typically requires low latency and is deployed on resource-constrained platforms. Existing solutions for efficient tracking primarily focus on lightweight backbones or modules, which, however, come at a sacrifice in precision. In this article, inspired by dynamic network routing, we propose DyTrack, a dynamic transformer framework for efficient tracking. Real-world tracking scenarios exhibit varying levels of complexity. We argue that a simple network is sufficient for easy video frames, while more computational resources should be assigned to difficult ones. DyTrack automatically learns to configure proper reasoning routes for different inputs, thereby improving the utilization of the available computational budget and achieving higher performance at the same running speed. We formulate instance-specific tracking as a sequential decision problem and incorporate terminating branches to intermediate layers of the model. Furthermore, we propose a feature recycling mechanism to maximize computational efficiency by reusing the outputs of predecessors. Additionally, a target-aware self-distillation strategy is designed to enhance the discriminating capabilities of early-stage predictions by mimicking the representation patterns of the deep model. Extensive experiments demonstrate that DyTrack achieves promising speed-precision tradeoffs with only a single model. For instance, DyTrack obtains 64.9% area under the curve (AUC) on LaSOT with a speed of 256 fps. Jiawen Zhu 0003, Xin Chen 0032, Haiwen Diao, Shuai Li 0014, Jun-Yan He, Chenyang Li 0007, Bin Luo 0008, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Person in Uniforms Re-IdentificationabstractPerson in Uniforms Re-identification (PU-ReID) is an emerging computer vision task for various intelligent video surveillance applications. PU-ReID is much understudied due to the absence of large-scale annotated datasets, also this task is extremely challenging because many individuals captured in surveillance videos wear same clothing, introducing significant interference for retrieval tasks owing to the high visual similarity of outfits and subtle differences among individuals. This research initiates the exploration of person in uniforms re-identification, a novel and challenging task tailored for real industrial scenarios. To address these issues, a novel framework is proposed for PU-ReID, which aims to reduce the visual impact of similar uniforms and learn the unique cues derived from human parts and detailed visual features. Specifically, several novel techniques are built in this study: first, a uniform feature separation method with orthogonal constraints is proposed to extract non-uniform features. Second, multi-view subspace feature alignment is introduced to integrate soft-biometrics including optics-related visual features, contextual information of human parts, and cloth-invariant biometric features. In addition, to close the gap between academic research and real-world settings, a new person in uniforms ReID dataset named PU-151 is constructed, which consists of 151 gas station employees in uniforms from 1,488 videos. At last, extensive experiments conducted on five datasets demonstrate that the proposed approach significantly outperforms the state-of-the-art methods. This advancement can drive further developments in re-identification and person search technologies. Chong-Yang Xiang, Xiao Wu 0001, Jun-Yan He, Zhaoquan Yuan, Tingquan He |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionabstractMultimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose AnyRef, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation. Code and models are available at https://github.com/jwh97nn/AnyRef Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo 0008, Xuansong Xie |
CVPR | 5 |
| 2024 | AnyText: Multilingual Visual Text Generation and EditingabstractDiffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated image, as synthesized text often contains blurred, unreadable, or incorrect characters, making visual text generation one of the most challenging issues in this field. To address this issue, we introduce AnyText, a diffusion-based multilingual visual text generation and editing model, that focuses on rendering accurate and coherent text in the image. AnyText comprises a diffusion pipeline with two primary elements: an auxiliary latent module and a text embedding module. The former uses inputs like text glyph, position, and masked image to generate latent features for text generation or editing. The latter employs an OCR model for encoding stroke data as embeddings, which blend with image caption embeddings from the tokenizer to generate texts that seamlessly integrate with the background. We employed text-control diffusion loss and text perceptual loss for training to further enhance writing accuracy. AnyText can write characters in multiple languages, to the best of our knowledge, this is the first work to address multilingual visual text generation. It is worth mentioning that AnyText can be plugged into existing diffusion models from the community for rendering or editing text accurately. After conducting extensive evaluation experiments, our method has outperformed all other approaches by a significant margin. Additionally, we contribute the first large-scale multilingual text images dataset, AnyWord-3M, containing 3 million image-text pairs with OCR annotations in multiple languages. Based on AnyWord-3M dataset, we propose AnyText-benchmark for the evaluation of visual text generation accuracy and quality. Our project will be open-sourced soon to improve and promote the development of text generation technology. Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, Xuansong Xie |
ICLR | 3 |
| 2024 | DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVsabstractExisting nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable vision system. To address this, we propose a novel architecture called Darkness Clue-Prompted Tracking (DCPT) that achieves robust UAV tracking at night by efficiently learning to generate darkness clue prompts. Without a separate enhancer, DCPT directly encodes anti-dark capabilities into prompts using a darkness clue prompter (DCP). Specifically, DCP iteratively learns emphasizing and undermining projections for darkness clues. It then injects these learned visual prompts into a daytime tracker with fixed parameters across transformer layers. Moreover, a gated feature aggregation mechanism enables adaptive fusion between prompts and between prompts and the base model. Extensive experiments show state-of-the-art performance for DCPT on multiple dark scenario benchmarks. The unified end-to-end learning of enhancement and tracking in DCPT enables a more trainable system. The darkness clue prompting efficiently injects anti-dark knowledge without extra modules. Code is available at https://github.com/bearyi26/DCPT. Jiawen Zhu 0003, Huayi Tang, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Shihao Qiu, Shengming Li, Huchuan Lu |
ICRA | 4 |
| 2024 | Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningabstractAccurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling.
However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset. Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001 |
NeurIPS | 3 |
| 2024 | Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human InteractionsabstractVision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-and-Language Navigation (HA-VLN), extending traditional VLN by incorporating dynamic human activities and relaxing key assumptions. We propose the Human-Aware 3D (HA3D) simulator, which combines dynamic human activities with the Matterport3D dataset, and the Human-Aware Room-to-Room (HA-R2R) dataset, extending R2R with human activity descriptions. To tackle HA-VLN challenges, we present the Expert-Supervised Cross-Modal (VLN-CM) and Non-Expert-Supervised Decision Transformer (VLN-DT) agents, utilizing cross-modal fusion and diverse training strategies for effective navigation in dynamic human environments. A comprehensive evaluation, including metrics considering human activities, and systematic analysis of HA-VLN's unique challenges, underscores the need for further research to enhance HA-VLN agents' real-world robustness and adaptability. Ultimately, this work provides benchmarks and insights for future research on embodied AI and Sim2Real transfer, paving the way for more realistic and applicable VLN systems in human-populated environments. Zhi-Qi Cheng, Yifei Dong 0002, Yuxuan Zhou 0004, Jun-Yan He, Qi Dai 0001, Teruko Mitamura, Alex Hauptmann 0001 |
NeurIPS | 6 |
| 2023 | Optimal Proposal Learning for Deployable End-to-End Pedestrian DetectionabstractEnd-to-end pedestrian detection focuses on training a pedestrian detection model via discarding the Non-Maximum Suppression (NMS) post-processing. Though a few methods have been explored, most of them still suffer from longer training time and more complex deployment, which cannot be deployed in the actual industrial applications. In this paper, we intend to bridge this gap and propose an Optimal Proposal Learning (OPL) framework for deployable end-to-end pedestrian detection. Specifically, we achieve this goal by using CNN-based light detector and introducing two novel modules, including a Coarse-to-Fine (C2F) learning strategy for proposing precise positive proposals for the Ground-Truth (GT) instances by reducing the ambiguity of sample assignment/output in training/testing respectively, and a Completed Proposal Network (CPN) for producing extra information compensation to further recall the hard pedestrian samples. Extensive experiments are conducted on CrowdHuman, TJU-Ped and Caltech, and the results show that our proposed OPL method significantly outperforms the competing methods. Binghui Chen, Jun-Yan He, Yifeng Geng, Xuansong Xie |
CVPR | 4 |
| 2023 | Procontext: Exploring Progressive Context Transformer for TrackingabstractExisting Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with Progressive Context Encoding Transformer Tracker (ProContEXT), which coherently exploits spatial and temporal contexts to predict object motion trajectories. Specifically, ProContEXT leverages a context-aware self-attention module to encode the spatial and temporal context, refining and updating the multi-scale static and dynamic templates to progressively perform accurately tracking. It explores the complementary between spatial and temporal context, raising a new pathway to multi-context modeling for transformer-based trackers. In addition, ProContEXT revised the token pruning technique to reduce computational complexity. Extensive experiments on popular benchmark datasets such as GOT-10k and TrackingNet demonstrate that the proposed ProContEXT achieves state-of-the-art performance1. Jin-Peng Lan, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li 0007, Bin Luo 0008, Xu Bao 0003, Wangmeng Xiang, Yifeng Geng, Xuansong Xie |
ICASSP | 3 |
| 2023 | Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming PerceptionabstractStreaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns, which restricts their ability to model complex scenes, often leading to poor detection results. To address this limitation, we propose LongShortNet, a novel dual-path network that captures long-term temporal motion and integrates it with short-term spatial semantics for real-time perception. Our proposed LongShortNet is notable as it is the first work to extend long-term temporal modeling to streaming perception, enabling spatiotemporal feature fusion. We evaluate LongShortNet on the challenging Argoverse-HD dataset and demonstrate that it outperforms existing state-of-the-art methods with almost no additional computational cost.1 Chenyang Li 0007, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Yifeng Geng, Jin-Peng Lan, Xuansong Xie |
ICASSP | 3 |
| 2023 | Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance LearningabstractDepth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a persegment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels. Code and models are available at https://github.com/jwh97nn/DeepDPS. Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Bin Luo 0008, Jun-Yan He, Jin-Peng Lan, Yifeng Geng, Xuansong Xie |
ICCV | 6 |
| 2023 | HDFormer: High-order Directed Transformer for 3D Human Pose EstimationabstractHuman pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a novel approach, the High-order Directed Transformer (HDFormer), which leverages high-order bone and joint relationships for improved pose estimation. Specifically, HDFormer incorporates both self-attention and high-order attention to formulate a multi-order attention module. This module facilitates first-order "joint-joint", second-order "bone-joint", and high-order "hyperbone-joint" interactions, effectively addressing issues in complex and occlusion-heavy situations. In addition, modern CNN techniques are integrated into the transformer-based architecture, balancing the trade-off between performance and efficiency. HDFormer significantly outperforms state-of-the-art (SOTA) models on Human3.6M and MPI-INF-3DHP datasets, requiring only 1/10 of the parameters and significantly lower computational costs. Moreover, HDFormer demonstrates broad real-world applicability, enabling real-time, accurate 3D pose estimation. The source code is in https://github.com/hyer/HDFormer. Jun-Yan He, Wangmeng Xiang, Zhi-Qi Cheng, Wei Liu 0015, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
IJCAI | 2 |
| 2023 | DAMO-StreamNet: Optimizing Streaming Perception in Autonomous DrivingabstractIn the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO-StreamNet's main inventions include: (1) a robust neck structure employing deformable convolution, bolstering receptive field and feature alignment capabilities; (2) a dual-branch structure synthesizing short-path semantic features and long-path temporal features, enhancing the accuracy of motion state prediction; (3) logits-level distillation facilitating efficient optimization, which aligns the logits of teacher and student networks in semantic space; and (4) a real-time prediction mechanism that updates the features of support frames with the current frame, providing smooth streaming perception during inference. Our testing shows that DAMO-StreamNet surpasses current state-of-the-art methodologies, achieving 37.8% (normal size (600, 960)) and 43.3% (large size (1200, 1920)) sAP without requiring additional data. This study not only establishes a new standard for real-time perception but also offers valuable insights for future research. The source code is at https://github.com/zhiqic/DAMO-StreamNet. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Wangmeng Xiang, Binghui Chen, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
IJCAI | 1 |
| 2023 | KeyPosS: Plug-and-Play Facial Landmark Detection through GPS-Inspired True-Range MultilaterationabstractIn the realm of facial analysis, accurate landmark detection is crucial for various applications, ranging from face recognition and expression analysis to animation. Conventional heatmap or coordinate regression-based techniques, however, often face challenges in terms of computational burden and quantization errors. To address these issues, we present the KeyPoint Positioning System (KeyPosS) - a groundbreaking facial landmark detection framework that stands out from existing methods. The framework utilizes a fully convolutional network to predict a distance map, which computes the distance between a Point of Interest (POI) and multiple anchor points. These anchor points are ingeniously harnessed to triangulate the POI's position through the True-range Multilateration algorithm. Notably, the plug-and-play nature of KeyPosS enables seamless integration into any decoding stage, ensuring a versatile and adaptable solution. We conducted a thorough evaluation of KeyPosS's performance by benchmarking it against state-of-the-art models on four different datasets. The results show that KeyPosS substantially outperforms leading methods in low-resolution settings while requiring a minimal time overhead.1 The code is available at https://github.com/zhiqic/KeyPosS. Xu Bao 0003, Zhi-Qi Cheng, Jun-Yan He, Wangmeng Xiang, Chenyang Li 0007, Jingdong Sun, Wei Liu 0015, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ACM Multimedia | 3 |
| 2023 | PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose EstimationabstractThe current 3D human pose estimators face challenges in adapting to new datasets due to the scarcity of 2D-3D pose pairs in target domain training sets. We present the Multi-Hypothesis Pose Synthesis Domain Adaptation (PoSynDA) framework to overcome this issue without extensive target domain annotation. Utilizing a diffusion-centric structure, PoSynDA simulates the 3D pose distribution in the target domain, filling the data diversity gap. By incorporating a multi-hypothesis network, it creates diverse pose hypotheses and aligns them with the target domain. Target-specific source augmentation obtains the target domain distribution data from the source domain by decoupling the scale and position parameters. The teacher-student paradigm and low-rank adaptation further refine the process. PoSynDA demonstrates competitive performance on benchmarks, such as Human3.6M, MPI-INF-3DHP, and 3DPW, even comparable with the target-trained MixSTE model. This work paves the way for the practical application of 3D human pose estimation1. The source code is available at https://github.com/hbing-l/PoSynDA. Jun-Yan He, Zhi-Qi Cheng, Wangmeng Xiang, Qize Yang, Wenhao Chai, Gaoang Wang, Xu Bao 0003, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ACM Multimedia | 2 |
| 2022 | Domain-Specific Conditional Jigsaw Adaptation for Enhancing transferability and DiscriminabilityabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a label-rich source domain to a target domain where the label is unavailable. Existing approaches tend to reduce the distribution discrepancy between the source and target domains or assign the pseudo target labels to implement a self-training strategy. However, the transferability or discriminability lackage of the traditional methods results in the limited ability to generalize on the target domain. To remedy this issue, a novel unsupervised domain adaptation framework called Domain-specific Conditional Jigsaw Adaptation Network (DCJAN) is proposed for UDA, which simultaneously encourages the network to extract transferable and discriminative features. To improve the discriminability, a conditional jigsaw module is presented to reconstruct class-aware features of the original images by reconstructing that of corresponding shuffled images. Moreover, in order to enhance the transferability, a domain-specific jigsaw adaptation is proposed to deal with the domain gaps, which utilizes the prior knowledge of jigsaw puzzles to reduce mismatching. It trains conditional jigsaw modules for each domain and updates the shared feature extractor to make the domain-specific conditional jigsaw modules could perform well not only on the corresponding domain but also on the other domain. A consistent conditioning strategy is proposed to ensure the safe training of conditional jigsaw. Experiments conducted on the widely-used Office-31, Office-Home, VisDA-2017, and DomainNet datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods. Qi He 0007, Zhaoquan Yuan, Xiao Wu 0001, Jun-Yan He |
ACM Multimedia | 4 |
| 2022 | SWNet: A Deep Learning Based Approach for Splashed Water Detection on RoadabstractAdverse weather conditions seriously threaten the traffic safety, especially for rainy days with the ponding water on the road surface, which potentially result in vehicle crashes, person injuries and crash fatalities. Automatic splashed water detection based on surveillance videos is an attractive way to effectively prevent the traffic accidents. However, surveillance videos exhibit great variations with lighting changes, illumination conditions and complex backgrounds, which pose great difficulties in automatic recognition. In this paper, a novel deep learning based approach is proposed to detect the splashed water. To the best of our knowledge, this is the first work on this topic based on deep learning. An effective semantic segmentation network, called SWNet, is novelly proposed to extract the potential splashed water regions. An encoder-decoder structure is designed to capture the visual characteristics of splashed water. SWNet achieves high efficiency by reusing pooling indices and adopting the light-weight decoder. With the multi-scale feature fusion structure, SWNet integrates the coarse semantic information and detailed appearance information, which significantly boosts the accuracy and refines the edge segmentation. A weighted cross entropy loss for splashed water is adopted to cope with the unbalanced distribution between splashed water and backgrounds. Moreover, a splashed water attention module is designed to focus on the salient regions of moving vehicles and splashed water, by performing attention mechanism to integrate global contextual information in semantic segmentation. Experiments conducted on a newly collected splashed water dataset demonstrate the effectiveness and efficiency of the proposed approach, which outperforms the state-of-the-art methods. Jian-Jun Qiao, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Qiang Peng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | A novel class restriction loss for unsupervised domain adaptation
Qi He 0007, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He |
Neurocomputing | 4 |
| 2021 | DB-LSTM: Densely-connected Bi-directional LSTM for human action recognition
Jun-Yan He, Xiao Wu 0001, Zhi-Qi Cheng, Zhaoquan Yuan, Yu-Gang Jiang 0001 |
Neurocomputing | 1 |
| 2021 | MGSeg: Multiple Granularity-Based Real-Time Semantic Segmentation NetworkabstractRecent works on semantic segmentation witness significant performance improvement by utilizing global contextual information. In this paper, an efficient multi-granularity based semantic segmentation network (MGSeg) is proposed for real-time semantic segmentation, by modeling the latent relevance between multi-scale geometric details and high-level semantics for fine granularity segmentation. In particular, a light-weight backbone ResNet-18 is first adopted to produce the hierarchical features. Hybrid Attention Feature Aggregation (HAFA) is designed to filter the noisy spatial details of features, acquire the scale-invariance representation, and alleviate the gradient vanishing problem of the early-stage feature learning. After aggregating the learned features, Fine Granularity Refinement (FGR) module is employed to explicitly model the relationship between the multi-level features and categories, generating proper weights for fusion. More importantly, to meet the real-time processing, a series of light-weight strategies and simplified structures are applied to accelerate the efficiency, including light-weight backbone, channel compression, narrow neck structure, and so on. Extensive experiments conducted on benchmark datasets Cityscapes and CamVid demonstrate that the proposed method achieves the state-of-the-art performance, 77.8%@50fps and 72.7%@127fps on Cityscapes and CamVid datasets, respectively, having the capability for real-time applications. Jun-Yan He, Shi-Hua Liang, Xiao Wu 0001, Bo Zhao 0032, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2020 | Learning fashion compatibility across categories with deep multimodal neural networks
Guang-Lu Sun, Jun-Yan He, Xiao Wu 0001, Bo Zhao 0032, Qiang Peng |
Neurocomputing | 2 |
| 2019 | Improving the Learning of Multi-column Convolutional Neural Network for Crowd CountingabstractTremendous variation in the scale of people/head size is a critical problem for crowd counting. To improve the scale invariance of feature representation, recent works extensively employ Convolutional Neural Networks with multi-column structures to handle different scales and resolutions. However, due to the substantial redundant parameters in columns, existing multi-column networks invariably exhibit almost the same scale features in different columns, which severely affects counting accuracy and leads to overfitting. In this paper, we attack this problem by proposing a novel Multicolumn Mutual Learning (McML) strategy. It has two main innovations: 1) A statistical network is incorporated into the multi-column framework to estimate the mutual information between columns, which can approximately indicate the scale correlation between features from different columns. By minimizing the mutual information, each column is guided to learn features with different image scales. 2) We devise a mutual learning scheme that can alternately optimize each column while keeping the other columns fixed on each mini-batch training data. With such asynchronous parameter update process, each column is inclined to learn different feature representation from others, which can efficiently reduce the parameter redundancy and improve generalization ability. More remarkably, McML can be applied to all existing multi-column networks and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that McML can significantly improve the original multi-column networks and outperform the other state-of-the-art approaches. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He, Alex Hauptmann 0001 |
ACM Multimedia | 5 |
| 2019 | BranchGAN: Unsupervised Mutual Image-to-Image Transfer With A Single Encoder and Dual DecodersabstractImage-to-image translation is a fundamental task for a wide range of applications, such as image style transfer, video effect generation, cross-domain retrieval, etc. Due to the limited number of labeled data, complex scenes, abstract semantics and various involved domains, image translation remains a challenging task. Compared to the supervised approaches for image translation that need a large collection of paired images for training, the unsupervised methods can significantly reduce the training cost. In this paper, an unsupervised end-to-end generative adversarial network is proposed, namedBranchGAN, for mutual image-to-image transfer between two domains. A structure with one single encoder and dual decoders is novelly proposed to capture the cross-domain distributions and generate the images in both domains. Three factors, that is, pixel-level overall style, region semantics, and domain distinguishability are comprehensively considered to constrain the training process of the proposed model, corresponding toreconstruction loss,encoding loss, andadversarial loss, respectively. Experiments conducted on three benchmark datasets demonstrate the effectiveness of the proposed method that outperforms the unsupervised state-of-the-art approaches and has the competitive performance as the supervised method. Yi-Fan Zhou, Runhao Jiang, Xiao Wu 0001, Jun-Yan He, Shuang Weng, Qiang Peng |
IEEE Trans. Multim. | 4 |
| 2018 | Hookworm Detection in Wireless Capsule Endoscopy Images With Deep LearningabstractAs one of the most common human helminths, hookworm is a leading cause of maternal and child morbidity, which seriously threatens human health. Recently, wireless capsule endoscopy (WCE) has been applied to automatic hookworm detection. Unfortunately, it remains a challenging task. In recent years, deep convolutional neural network (CNN) has demonstrated impressive performance in various image and video analysis tasks. In this paper, a novel deep hookworm detection framework is proposed for WCE images, which simultaneously models visual appearances and tubular patterns of hookworms. This is the first deep learning framework specifically designed for hookworm detection in WCE images. Two CNN networks, namely edge extraction network and hookworm classification network, are seamlessly integrated in the proposed framework, which avoid the edge feature caching and speed up the classification. Two edge pooling layers are introduced to integrate the tubular regions induced from edge extraction network and the feature maps from hookworm classification network, leading to enhanced feature maps emphasizing the tubular regions. Experiments have been conducted on one of the largest WCE datasets with WCE images, which demonstrate the effectiveness of the proposed hookworm detection framework. It significantly outperforms the state-of-the-art approaches. The high sensitivity and accuracy of the proposed method in detecting hookworms shows its potential for clinical application. Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Qiang Peng, Ramesh Jain 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Sketch Recognition with Deep Visual-Sequential Fusion ModelabstractIn this paper, a deep end-to-end network for sketch recognition, named Deep Visual-Sequential Fusion model (DVSF) is proposed to model the visual and sequential patterns of the strokes. To capture the intermediate states of sketches, a three-way representation learner is first utilized to extract the visual features. These deep features are simultaneously fed into the visual and sequential networks to capture spatial and temporal properties, respectively. More specifically, visual networks are novelly proposed to learn the stroke patterns by stacking the Residual Fully-Connected (R-FC) layers, which integrate ReLU and Tanh activation functions to achieve the sparsity and generalization ability. To learn the patterns of stroke order, sequential networks are constructed by Residual Long Short-Term Memory (R-LSTM) units, which optimize the network architecture by skip connection. Finally, the visual and sequential representations of the sketches are seamlessly integrated with a fusion layer to obtain the final results. Experiments conducted on the benchmark sketch dataset TU-Berlin demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art approaches. Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Bo Zhao 0032, Qiang Peng |
ACM Multimedia | 1 |
| 2016 | Detection of bird nests in overhead catenary system images for high-speed rail
Xiao Wu 0001, Ping Yuan, Qiang Peng, Chong-Wah Ngo, Jun-Yan He |
Pattern Recognit. | 5 |