VLDB 2026 Research / reviewers in the wild / expert
Zhi-Qi Cheng
dblp:188/1193 · also Zhiqi Cheng
· DBLP profile ↗
59ranked-venue papers
10as first author
47since 2021 · last 2026
0000-0002-1720-2085ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 10 first-author · 33 since 2021Artificial intelligence and machine learning · 31 · 4 first-author · 28 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsabstractSign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities.Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage.Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited.In this survey, we present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages.We analyze key challenges such as modality imbalance, annotation granularity, and signer bias, and outline considerations for future dataset design.We also introduce a 24-field Sign-Language Datasheet and release a public GitHub repository 1 to support standardized documentation and reproducible evaluation.Overall, our work provides a unified and practical foundation for developing inclusive, robust, and scalable sign-language technologies in real-world applications. Yiming Ni, Zhi-Qi Cheng |
ACL (1) | 2 |
| 2026 | A survey of robotic manipulation: From bottom-up approaches to end-to-end paradigms with LLMs
Qing Li 0001, Zhijian He, Bowen Zhang 0005, Xianghua Fu, Zhi-Qi Cheng, Yan Yan 0001, Xiaojiang Peng |
Neurocomputing | 8 |
| 2026 | Robust Adaptation of Foundation Models With Black-Box Visual PromptingabstractWith a surge of large-scale pre-trained models, parameter-efficient transfer learning (PETL) of large models has garnered significant attention. While promising, they commonly rely on two optimistic assumptions: 1) full access to the parameters of a PTM and 2) sufficient memory capacity to cache all intermediate activations for gradient computation. However, in most real-world applications, PTMs serve as black-box APIs or proprietary software without full parameter accessibility. Besides, it is hard to meet a large memory requirement for modern PTMs. This work proposes black-box visual prompting (BlackVIP), which efficiently adapts the PTMs without knowledge of their architectures or parameters. BlackVIP has two components: 1) Coordinator and 2) simultaneous perturbation stochastic approximation with gradient correction (SPSA-GC). The Coordinator designs input-dependent visual prompts, which allow the target PTM to adapt in the wild. SPSA-GC efficiently estimates the gradient of PTM to update Coordinator. Besides, we introduce a variant, BlackVIP-SE, which significantly reduces the runtime and computational cost of BlackVIP. Extensive experiments on 19 datasets demonstrate that BlackVIPs enable robust adaptation to diverse domains and tasks with minimal memory requirements. We further provide a theoretical analysis on the generalization of visual prompting methods by presenting their connection to the certified robustness of randomized smoothing, and presenting an empirical support for improved robustness. Changdae Oh, Gyeongdeok Seo, Geunyoung Jung, Zhi-Qi Cheng, Hosik Choi, Jiyoung Jung, Kyungwoo Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | PSGNet: Pure Smoke Image Generation With Gradient and Style LearningabstractThe realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches. Jian-Jun Qiao, Xiao Wu 0001, Zhi-Qi Cheng, Wei Li 0110, Zhaoquan Yuan |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | A Video-grounded Dialogue Dataset and Metric for Event-driven ActivitiesabstractThis paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and more complex video sequences that depict a variety of event-driven activities that require advanced contextual understanding for accurate response generation. The dataset comprises 3,000 dialogues with over 30,000 question-and-answer pairs, derived from 1,000 videos with diverse activity scenarios. VDAct displays a notably challenging characteristic due to its broad spectrum of activity scenarios and wide range of question types. Empirical studies on state-of-the-art vision foundation models highlight their limitations in addressing certain question types on our dataset. Furthermore, VDEval, which integrates dialogue session history and video content summaries extracted from our supplementary Knowledge Graphs to evaluate individual responses, demonstrates a significantly higher correlation with human assessments on the VDAct dataset than existing evaluation metrics that rely solely on the context of single dialogue turns. Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Zhi-Qi Cheng, Ken Fukuda, Teruko Mitamura |
AAAI | 4 |
| 2025 | POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchabstractAchieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs three key contributions: (1) Pseudo-range multilateration is utilized to correct heatmap errors, improving landmark localization accuracy. By integrating multiple anchor points, it reduces the impact of individual heatmap inaccuracies, leading to robust overall positioning. (2) To enhance the pseudo-range accuracy of selected anchor points, a new loss function, named multilateration anchor loss, is proposed. This loss function enhances the accuracy of the distance map, mitigates the risk of local optima, and ensures optimal solutions. (3) A single-step parallel computation algorithm is introduced, boosting computational efficiency and reducing processing time. Extensive evaluations across five benchmark datasets demonstrate that POPoS consistently outperforms existing methods, particularly excelling in low-resolution heatmaps scenarios with minimal computational overhead. These advantages make POPoS as a highly efficient and accurate tool for FLD, with broad applicability in real-world scenarios. Chong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu 0001, Xian-Sheng Hua 0001 |
AAAI | 3 |
| 2025 | StableAnimator: High-Quality Identity-Preserving Human Image AnimationabstractCurrent diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses. Building upon a video diffusion model, StableAnimator contains carefully designed modules for both training and inference striving for identity consistency. In particular, StableAnimator begins by computing image and face embeddings with off-the-shelf extractors, respectively and face embeddings are further refined by interacting with image embeddings using a global content-aware Face Encoder. Then, StableAnimator introduces a novel distribution-aware ID Adapter that prevents interference caused by temporal layers while preserving ID via alignment. During inference, we propose a novel Hamilton-Jacobi-Bellman (HJB) equation-based optimization to further enhance the face quality. We demonstrate that solving the HJB equation can be integrated into the diffusion denoising process, and the resulting solution constrains the denoising path and thus benefits ID preservation. Experiments on multiple benchmarks show the effectiveness of StableAnimator both qualitatively and quantitatively. Shuyuan Tu, Xintong Han, Zhi-Qi Cheng, Qi Dai 0001, Chong Luo 0001, Zuxuan Wu |
CVPR | 4 |
| 2025 | Emphasizing Discriminative Features for Dataset Distillation in Complex ScenariosabstractDataset distillation has demonstrated strong performance on simple datasets like CIFAR, MNIST, and TinyImageNet but struggles to achieve similar results in more complex scenarios. In this paper, we propose EDF (emphasizes the discriminative features), a dataset distillation method that enhances key discriminative regions in synthetic images using Grad-CAM activation maps. Our approach is inspired by a key observation: in simple datasets, high-activation areas typically occupy most of the image, whereas in complex scenarios, the size of these areas is much smaller. Unlike previous methods that treat all pixels equally when synthesizing images, EDF uses Grad-CAM activation maps to enhance high-activation areas. From a supervision perspective, we downplay supervision signals produced by lower trajectory-matching losses, as they contain common patterns. Additionally, to help the DD community better explore complex scenarios, we build the Complex Dataset Distillation (Comp-DD) benchmark by meticulously selecting sixteen subsets, eight easy and eight hard, from ImageNet-1K. In particular, EDF consistently outperforms SOTA results in complex scenarios, such as ImageNet-1K subsets. Hopefully, more researchers will be inspired and encouraged to improve the practicality and efficacy of DD. Our code and benchmark have been made public at NUS-HPC-AI-Lab/EDF. Kai Wang 0036, Zhi-Qi Cheng, Samir Khaki, Ahmad Sajedi, Ramakrishna Vedantam, Konstantinos N. Plataniotis, Alex Hauptmann 0001, Yang You 0001 |
CVPR | 3 |
| 2025 | UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal PromptsabstractEmotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module(EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. To facilitate reproducibility and further research, we have made our code publicly available at https://github.com/KTTRCDL/UMETTS. Zhi-Qi Cheng, Jun-Yan He, Junyao Chen, Xiaomao Fan, Xiaojiang Peng, Alex Hauptmann 0001 |
ICASSP | 2 |
| 2025 | DeformAvatar: Point-Based Human Avatar Re-targeting and RenderingabstractIn this paper, we present the DeformAvatar, a novel architecture for human avatar re-targetting and rendering based on point clouds. Given the multiple views of a person, we first build a point-model-paired human representation containing a raw point cloud and an optimal parametric model. Then, we repurpose several advanced neural point-based rendering and Gaussian Splatting techniques for 3D avatar modeling. Finally, to enhance photorealistic re-targeting of body shapes and poses, we propose a dual point-adaptive (DPA) regularization based on traditional linear blend skinning. Extensive experiments demonstrate that our DeformAvatar framework can synthesize highly realistic novel views in new shape and pose parameters. We also find that 3DGS-based avatar modeling is superior to others in 3D avatar re-targeting. Renyi Zhan, Zhi-Qi Cheng, Junyao Chen, Xiaojiang Peng |
ICASSP | 2 |
| 2025 | MotionFollower: Editing Video Motion via Score-Guided Diffusion
Shuyuan Tu, Qi Dai 0001, Sicheng Xie, Zhi-Qi Cheng, Chong Luo 0001, Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001 |
ICCV | 5 |
| 2025 | MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt SynthesisabstractMetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the creation of customizable WordArt, ranging from semantic enhancements to intricate textural elements. A central feedback mechanism leverages insights from both multimodal models and user evaluations, enabling iterative refinement of design parameters. Through this iterative process, MetaDesigner dynamically adjusts hyperparameters to align with user-defined stylistic and thematic preferences, consistently delivering WordArt that excels in visual quality and contextual resonance. Empirical evaluations underscore the system's versatility and effectiveness across diverse WordArt applications, yielding outputs that are both aesthetically compelling and context-sensitive. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Jingdong Sun, Qi He 0007, Wangmeng Xiang, Jin-Peng Lan, Xianhui Lin, Kang Zhu, Bin Luo 0008, Yifeng Geng, Xuansong Xie, Alex Hauptmann 0001 |
ICLR | 2 |
| 2025 | Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose EstimationabstractAccurate 3D Human Pose Estimation (HPE) in video sequences demands both precision and a robust architectural framework. Building upon the recent success of transformers in computer vision, we introduce the Refined Temporal Pyramidal Compression-and-Amplification (RTPCA) transformer, an approach that tackles a critical issue in current transformer-based methods: the underutilization of intra-and inter-block relations through attention mechanisms. The cornerstone of our approach is the meticulously designed Temporal Pyramidal Compression-and-Amplification (TPCA) module, which ingeniously leverages a temporal pyramid paradigm to significantly enhance multi-scale key and value representations from intra-block attention. Recognizing that focusing solely on individual modules while overlooking their interconnections can limit performance, we introduce the Cross-Layer Refinement (XLR) module. This carefully crafted component is designed to amplify inter-block communication by linking keys and values across adjacent blocks, creating more coherent attention patterns. The seamless integration of TPCA and XLR results in a powerful synergy that facilitates a rich semantic representation through the dynamic interaction of queries, keys, and values. This synergistic approach enables the RTPCA transformer to achieve remarkable performance on leading benchmarks, such as Human3.6M, HumanEva-I, and MPI-INF-3DHP, with only a small computational overhead. We demonstrate the effectiveness of the RTPCA transformer through extensive experiments and comparisons with state-of-the-art methods. The source code is available at https://github.com/hbing-l/RTPCA.git. Zhi-Qi Cheng, Wangmeng Xiang, Jun-Yan He, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ICME | 2 |
| 2025 | SituLM: Leveraging Visual Instruction Tuning and an Augmented SWiG Dataset for Enhanced Grounded Situation RecognitionabstractBridging visual and linguistic modalities is central to advanced scene-understanding tasks such as Grounded Situation Recognition (GSR), which aims to decode complex verb frames and their associated semantic roles in images. In this paper, we offer a comprehensive analysis of the widely used Situation With Grounding (SWiG) dataset, highlighting key issues including skewed role distributions, ambiguous annotations, incomplete bounding boxes, and inconsistent semantic specificity. To address these challenges, we introduce SituLM, a framework that leverages pretrained Large Multimodal Models (LMMs) to jointly model visual cues and language representations. We explore both single-turn and multi-turn variants of SituLM, each capturing distinct inference strategies and incorporating chain-of-thought reasoning. Extensive experiments on SWiG reveal that our approach significantly outperforms existing methods, yielding substantial gains in verb recognition, role prediction, and grounding accuracy. Furthermore, we present an augmented SWiG dataset that accommodates multiple plausible verb frames per image, prioritized by image-text similarity scoring and selectively refined via human verification. This enriched dataset provides a more realistic testbed for evaluating GSR models capable of handling multiple co-occurring actions. By unifying visual and linguistic information in a single instruction-based model, SituLM not only addresses longstanding limitations of SWiG but also paves the way for more robust, context-aware scene interpretation in real-world scenarios.1 Zhi-Qi Cheng |
ICME | 2 |
| 2025 | DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image EditingabstractFashion image editing is a crucial tool for designers to convey their creative ideas by visualizing design concepts interactively. However, current fashion image editing techniques often struggle to accurately identify editing regions and preserve the desired garment texture detail. To address these challenges, we present Detail-Preserved Diffusion Models (DPDEdit), a new multimodal fashion image editing architecture based on latent diffusion models. To precisely locate the editing region, we introduce Grounded-SAM to predict the editing region. To transfer the detail of the given garment texture into the target image, we propose a texture injection and refinement mechanism. This mechanism employs a decoupled cross-attention layer to integrate textual descriptions and texture images, and incorporates an auxiliary U-Net to preserve the high-frequency details of generated garment texture. Additionally, we extend the VITON-HD dataset using a multimodal large language model to generate paired samples with texture images and textual descriptions. Extensive experiments show that our DPDEdit outperforms state-of-the-art methods in terms of image fidelity and coherence with the given multimodal input. Zhi-Qi Cheng, Huizi Xue, Xiaojiang Peng |
ICME | 2 |
| 2025 | A Novel Human Abnormal Posture Detection Method Based on Spatial-Topological Feature Fusion of Skeleton
Yuefeng Ma, Zhi-Qi Cheng, Deheng Liu, Shiying Tang |
MMM (1) | 2 |
| 2025 | ProMQA: Question Answering Dataset for Multimodal Procedural Activity UnderstandingabstractKimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Yuran Wang, Ken Fukuda, Teruko Mitamura. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Ken Fukuda, Teruko Mitamura |
NAACL (Long Papers) | 3 |
| 2025 | MaxSup: Overcoming Representation Collapse in Label SmoothingabstractLabel Smoothing (LS) is widely adopted to reduce overconfidence in neural network predictions and improve generalization. Despite these benefits, recent studies reveal two critical issues with LS. First, LS induces overconfidence in misclassified samples. Second, it compacts feature representations into overly tight clusters, diluting intra-class diversity, although the precise cause of this phenomenon remained elusive. In this paper, we analytically decompose the LS-induced loss, exposing two key terms: (i) a regularization term that dampens overconfidence only when the prediction is correct, and (ii) an error-amplification term that arises under misclassifications. This latter term compels the network to reinforce incorrect predictions with undue certainty, exacerbating representation collapse. To address these shortcomings, we propose Max Suppression (MaxSup), which applies uniform regularization to both correct and incorrect predictions by penalizing the top-1 logit rather than the ground-truth logit. Through extensive feature-space analyses, we show that MaxSup restores intra-class variation and sharpens inter-class boundaries. Experiments on large-scale image classification and multiple downstream tasks confirm that MaxSup is a more robust alternative to LS. Yuxuan Zhou 0004, Zhi-Qi Cheng, Yifei Dong 0002, Mario Fritz, Margret Keuper |
NeurIPS | 3 |
| 2025 | DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming PerceptionabstractThe advancement of autonomous driving systems hinges on the ability to achieve low-latency and high-accuracy perception. To address this critical need, this paper introduces Dynamic Routering Network (DyRoNet), a low-rank enhanced dynamic routing framework designed for streaming perception in autonomous driving systems. DyRoNet integrates a suite of pre-trained branch networks, each meticulously fine-tuned to function under distinct environmental conditions. At its core, the framework offers a speed router module, developed to assess and route input data to the most suitable branch for processing. This approach not only addresses the inherent limitations of conventional models in adapting to diverse driving conditions but also ensures the balance between performance and efficiency. Extensive experimental evaluations demonstrating the adaptability of DyRoNet to diverse branch selection strategies, resulting in significant performance enhancements across different scenarios. This work not only establishes a new benchmark for streaming perception but also provides valuable engineering insights for future work.11Project: https://tastevision.github.io/DyRoNet/ Xiang Huang 0004, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li 0007, Wangmeng Xiang, Baigui Sun |
WACV | 2 |
| 2025 | UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
Zhi-Qi Cheng, Gabriel Moreira, Jiawen Zhu 0003, Jingdong Sun, Bukun Ren, Jun-Yan He, Qi Dai 0001, Xian-Sheng Hua 0001 |
WACV | 2 |
| 2025 | LEAF: Unveiling two sides of the same coin in semi-supervised facial expression recognition
Fan Zhang 0111, Zhi-Qi Cheng, Jian Zhao 0006, Xiaojiang Peng, Xuelong Li 0001 |
Comput. Vis. Image Underst. | 2 |
| 2025 | IVAC-$\mathbf {P^{2}L}$: Leveraging Irregular Repetition Priors for Improving Video Action CountingabstractThe quantification of repetitive actions in videos, a task commonly referred to as Video Action Counting (VAC), is a critical challenge in understanding and analyzing content in sports, fitness, and daily activities. Traditional approaches to VAC have largely overlooked the nuanced irregularities inherent in action repetitions, such as interruptions and variable lengths between cycles. Addressing this gap, our study introduces a novel perspective on VAC, focusing on Irregular Video Action Counting (IVAC), which emphasizes the importance of modeling the irregular repetition priors present in video content. We conceptualize these priors through two key aspects:Inter-cycle ConsistencyandCycle-interval Inconsistency. Inter-cycle Consistency ensures that thespatiotemporalrepresentations across all cycle segments in a videoremainhomogeneous, thereby reflecting the uniformity of actions betweendifferent cycle segments. In contrast, Cycle-interval Inconsistency mandates a clear semantic distinction between the representations of cycle segments and intervals, acknowledging the inherent dissimilarities in content. To effectively encapsulate these priors, we introduce a novel methodology consisting of consistency and inconsistency modules, underpinned by a tailored pull-push loss ($\mathrm {P^{2}~L}$) mechanism. This approach employs a pull loss to enhance the cohesion among cycle segment features and a push loss to distinctly differentiate between cycle and interval segment features. Empirical evaluations on the RepCount dataset illustrate that our IVAC-$\mathrm {P^{2}~L}$model sets a new benchmark in state-of-the-art performance for the VAC task. Moreover, our model demonstrates adaptability and generalization across diverse video content, achieving superior performance on two additional datasets, UCFRep and Countix, without necessitating dataset-specific fine-tuning. These findings not only validate the effectiveness of our approach in addressing the complexities of irregular repetitions in videos but also open new avenues for future research in video understanding and analysis. Zhi-Qi Cheng, Youtian Du, Lei Zhang 0006 |
IEEE Trans. Multim. | 2 |
| 2024 | Music2P: A Multi-Modal AI-Driven Tool for Simplifying Album Cover DesignabstractIn today's music industry, album cover design is as crucial as the music itself, reflecting the artist's vision and brand. However, many AI-driven album cover services require subscriptions or technical expertise, limiting accessibility. To address these challenges, we developed Music2P, an open-source, multi-modal AI-driven tool that streamlines album cover creation, making it efficient, accessible, and cost-effective through Ngrok. Music2P automates the design process using techniques such as Bootstrapping Language Image Pre-training (BLIP), music-to-text conversion (LP-music-caps), image segmentation (LoRA), and album cover and QR code generation (ControlNet). This paper demonstrates the Music2P interface, details our application of these technologies, and outlines future improvements. Our ultimate goal is to provide a tool that empowers musicians and producers, especially those with limited resources or expertise, to create compelling album covers. Joong Ho Choi, Geonyeong Choi, Ji-Eun Han, Wonjin Yang, Zhi-Qi Cheng |
CIKM | 5 |
| 2024 | BlockGCN: Redefine Topology Awareness for Skeleton-Based Action RecognitionabstractGraph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition, leveraging their ability to unravel the complex dynamics of human joint topology through the graph's adjacency matrix. However, an inherent flaw has come to light in these cutting-edge models: they tend to optimize the adjacency matrix jointly with the model weights. This process, while seemingly efficient, causes a gradual decay of bone connectiv-ity data, resulting in a model indifferent to the very topology it sought to represent. To remedy this, we propose a two-fold strategy: (1) We introduce an innovative approach that encodes bone connectivity by harnessing the power of graph distances to describe the physical topology; we further incorporate action-specific topological representation via persistent homology analysis to depict systemic dynamics. This preserves the vital topological nuances often lost in conventional GCNs. (2) Our investigation also reveals the redundancy in existing GCNs for multi-relational modeling, which we address by proposing an efficient refinement to Graph Convolutions (GC) - the BlockGC. This signif-icantly reduces parameters while improving performance beyond original GCNs. Our full model, BlockGCN, es-tablishes new benchmarks in skeleton-based action recognition across all model categories. Its high accuracy and lightweight design, most notably on the large-scale NTU RGB+D 120 dataset, stand as strong validation of the efficacy of BlockGCN. Yuxuan Zhou 0004, Zhi-Qi Cheng, Yan Yan 0001, Qi Dai 0001, Xian-Sheng Hua 0001 |
CVPR | 3 |
| 2024 | ProS: Prompting-to-Simulate Generalized Knowledge for Universal Cross-Domain RetrievalabstractThe goal of Universal Cross-Domain Retrieval (UCDR) is to achieve robust performance in generalized test scenarios, wherein data may belong to strictly unknown do-mains and categories during training. Recently, pre-trained models with prompt tuning have shown strong generalization capabilities and attained noteworthy achievements in various downstream tasks, such as few-shot learning and video-text retrieval. However, applying them directly to UCDR may not be sufficient to handle both domain shift (i.e., adapting to unfamiliar domains) and semantic shift (i.e., transferring to unknown categories). To this end, we propose Prompting-to-Simulate (ProS), the first method to apply prompt tuning for UCDR. ProS employs a two-step process to simulate Content-aware Dynamic Prompts (CaDP) which can impact models to produce generalized features for UCDR. Concretely, in Prompt Units Learning stage, we introduce two Prompt Units to individually capture domain and semantic knowledge in a mask-and-align way. Then, in Context-aware Simulator Learning stage, we train a Content-aware Prompt Simulator under a simulated test scenario to produce the corresponding CaDP. Extensive experiments conducted on three benchmark datasets show that our method achieves new state-of-the-art performance without bringing excessive parameters. Code is available at https://github.com/fangkaipeng/ProS. Kaipeng Fang, Jingkuan Song, Lianli Gao, Pengpeng Zeng, Zhi-Qi Cheng, Xiyao Li, Heng Tao Shen |
CVPR | 5 |
| 2024 | MotionEditor: Editing Video Motion via Content-Aware DiffusionabstractExisting diffusion-based video editing models have made gorgeous advances for editing attributes of a source video over time but struggle to manipulate the motion information while preserving the original protagonist's appearance and background. To address this, we propose MotionEditor, the first diffusion model for video motion editing. MotionEditor incorporates a novel content-aware motion adapter into ControlNet to capture temporal motion correspondence. While ControlNet enables direct generation based on skeleton poses, it encounters challenges when modifying the source motion in the inverted noise due to contradictory signals between the noise (source) and the condition (reference). Our adapter complements Control-Net by involving source content to transfer adapted control signals seamlessly. Further, we build up a two-branch ar-chitecture (a reconstruction branch and an editing branch) with a high-fidelity attention injection mechanism facilitating branch interaction. This mechanism enables the editing branch to query the key and value from the reconstruction branch in a decoupled manner, making the editing branch retain the original background and protagonist appearance. We also propose a skeleton alignment algorithm to address the discrepancies in pose size and position. Experiments demonstrate the promising motion editing ability of MotionEditor, both qualitatively and quantitatively. To the best of our knowledge, MotionEditor is the first to use diffusion models specifically for video motion editing, considering the origin dynamic background and camera movement. Shuyuan Tu, Qi Dai 0001, Zhi-Qi Cheng, Han Hu 0001, Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001 |
CVPR | 3 |
| 2024 | FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioabstractIn this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating vari-ous dynamically audio-consistent talking faces, termed Lis-tening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio. Specifically, it involves two critical challenges: one is to effectively de-couple identity, content, and emotion from entangled au-dio, and the other is to maintain intra-video diversity and inter- video consistency. To tackle the issues, we first dig out the intricate relationships among facial factors and sim-plify the decoupling process, tailoring a Progressive Audio Disentanglement for accurate facial geometry and seman-tics learning, where each stage incorporates a customized training module responsible for a specific factor. Secondly, to achieve visually diverse and audio-synchronized animation solely from input audio within a single model, we intro-duce the Controllable Coherent Frame generation, which involves the flexible integration of three trainable adapters with frozen Latent Diffusion Models (LDMs) to focus on maintaining facial geometry and semantics, as well as texsture and temporal coherence between frames. In this way, we inherit high-quality diverse generation from LDMs while significantly improving their controllability at a low training cost. Extensive experiments demonstrate the flexibility and effectiveness of our method in handling this paradigm. The codes will be released at FaceChain. Chao Xu 0023, Yang Liu 0356, Jiazheng Xing, Weida Wang, Jun Dan, Tianxin Huang, Siyuan Li 0002, Zhi-Qi Cheng, Ying Tai, Baigui Sun |
CVPR | 9 |
| 2024 | DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVsabstractExisting nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable vision system. To address this, we propose a novel architecture called Darkness Clue-Prompted Tracking (DCPT) that achieves robust UAV tracking at night by efficiently learning to generate darkness clue prompts. Without a separate enhancer, DCPT directly encodes anti-dark capabilities into prompts using a darkness clue prompter (DCP). Specifically, DCP iteratively learns emphasizing and undermining projections for darkness clues. It then injects these learned visual prompts into a daytime tracker with fixed parameters across transformer layers. Moreover, a gated feature aggregation mechanism enables adaptive fusion between prompts and between prompts and the base model. Extensive experiments show state-of-the-art performance for DCPT on multiple dark scenario benchmarks. The unified end-to-end learning of enhancement and tracking in DCPT enables a more trainable system. The darkness clue prompting efficiently injects anti-dark knowledge without extra modules. Code is available at https://github.com/bearyi26/DCPT. Jiawen Zhu 0003, Huayi Tang, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Shihao Qiu, Shengming Li, Huchuan Lu |
ICRA | 3 |
| 2024 | A Novel Multi-Pose Person Re-Identification Method Based on Semantic- and Pose-Guided Feature FusionabstractPerson re-identification (ReID) aims to match the query images with images in the gallery. However, ReID traditionally focuses on outdoor scenes and standing pose, neglecting the complexities of different poses and indoor environments. These neglects correspond to some challenges: postural differences and background noise which interfere with feature learning and matching. To address these issues, we propose a novel multi-pose ReID method, Semantic- and Pose-Guided Feature Fusion (SPGFF), which integrates semantic-guided and pose-guided features. Specially, a semantic-guided module is employed to incorporate global contextual semantic information into the feature representation. This global contextual semantic information refers to the comprehensive understanding of the entire image, including the relationships of different pixels. By incorporating this information, even there are significant variations in pose, the model is enabled to focus on the most pertinent parts of the image. Meanwhile, the pose-guided module uses pose feature to cleanly disentangle bodily semantic components and selectively match corresponding body parts. To the best of our knowledge, this is the first work to introduce the concept of multi-pose ReID, we have provided a benchmark to address the issue of a lack of publicly available datasets. We demonstrate the effectiveness of our approach on both public and proprietary datasets, showcasing its potential to significantly improve person re-identification in previously overlooked scenarios. Yuefeng Ma, Deheng Liu, Zhi-Qi Cheng, Shijian Li |
ICTAI | 3 |
| 2024 | Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningabstractAccurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling.
However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset. Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001 |
NeurIPS | 2 |
| 2024 | Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human InteractionsabstractVision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-and-Language Navigation (HA-VLN), extending traditional VLN by incorporating dynamic human activities and relaxing key assumptions. We propose the Human-Aware 3D (HA3D) simulator, which combines dynamic human activities with the Matterport3D dataset, and the Human-Aware Room-to-Room (HA-R2R) dataset, extending R2R with human activity descriptions. To tackle HA-VLN challenges, we present the Expert-Supervised Cross-Modal (VLN-CM) and Non-Expert-Supervised Decision Transformer (VLN-DT) agents, utilizing cross-modal fusion and diverse training strategies for effective navigation in dynamic human environments. A comprehensive evaluation, including metrics considering human activities, and systematic analysis of HA-VLN's unique challenges, underscores the need for further research to enhance HA-VLN agents' real-world robustness and adaptability. Ultimately, this work provides benchmarks and insights for future research on embodied AI and Sim2Real transfer, paving the way for more realistic and applicable VLN systems in human-populated environments. Zhi-Qi Cheng, Yifei Dong 0002, Yuxuan Zhou 0004, Jun-Yan He, Qi Dai 0001, Teruko Mitamura, Alex Hauptmann 0001 |
NeurIPS | 3 |
| 2024 | Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsabstractImproving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation. Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alex Hauptmann 0001, Zhi-Qi Cheng, Kyungwoo Song |
NeurIPS | 8 |
| 2023 | Procontext: Exploring Progressive Context Transformer for TrackingabstractExisting Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with Progressive Context Encoding Transformer Tracker (ProContEXT), which coherently exploits spatial and temporal contexts to predict object motion trajectories. Specifically, ProContEXT leverages a context-aware self-attention module to encode the spatial and temporal context, refining and updating the multi-scale static and dynamic templates to progressively perform accurately tracking. It explores the complementary between spatial and temporal context, raising a new pathway to multi-context modeling for transformer-based trackers. In addition, ProContEXT revised the token pruning technique to reduce computational complexity. Extensive experiments on popular benchmark datasets such as GOT-10k and TrackingNet demonstrate that the proposed ProContEXT achieves state-of-the-art performance1. Jin-Peng Lan, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li 0007, Bin Luo 0008, Xu Bao 0003, Wangmeng Xiang, Yifeng Geng, Xuansong Xie |
ICASSP | 2 |
| 2023 | Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming PerceptionabstractStreaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns, which restricts their ability to model complex scenes, often leading to poor detection results. To address this limitation, we propose LongShortNet, a novel dual-path network that captures long-term temporal motion and integrates it with short-term spatial semantics for real-time perception. Our proposed LongShortNet is notable as it is the first work to extend long-term temporal modeling to streaming perception, enabling spatiotemporal feature fusion. We evaluate LongShortNet on the challenging Argoverse-HD dataset and demonstrate that it outperforms existing state-of-the-art methods with almost no additional computational cost.1 Chenyang Li 0007, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Yifeng Geng, Jin-Peng Lan, Xuansong Xie |
ICASSP | 2 |
| 2023 | ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic RulesabstractCharts are a powerful tool for visually conveying complex data, but their comprehension poses a challenge due to the diverse chart types and intricate components. Existing chart comprehension methods suffer from either heuristic rules or an over-reliance on OCR systems, resulting in suboptimal performance. To address these issues, we present ChartReader, a unified framework that seamlessly integrates chart derendering and comprehension tasks. Our approach includes a transformer-based chart component detection module and an extended pre-trained vision-language model for chart-to-X tasks. By learning the rules of charts automatically from annotated datasets, our approach eliminates the need for manual rule-making, reducing effort and enhancing accuracy. We also introduce a data variable replacement technique and extend the input and position embeddings of the pre-trained model for cross-task training. We evaluate ChartReader on Chart-to-Table, ChartQA, and Chart-to-Text tasks, demonstrating its superiority over existing methods. Our proposed framework can significantly reduce the manual effort involved in chart analysis, providing a step towards a universal chart understanding model. Moreover, our approach offers opportunities for plug-and-play integration with mainstream LLMs such as T5 and TaPas, extending their capability to chart comprehension tasks.1 Zhi-Qi Cheng, Qi Dai 0001, Alex Hauptmann 0001 |
ICCV | 1 |
| 2023 | Implicit Temporal Modeling with Learnable Alignment for Video RecognitionabstractContrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks. However, how to extend CLIP with effective temporal modeling is still an open and crucial problem. Existing factorized or joint spatial-temporal modeling trades off between the efficiency and performance. While modeling temporal information within straight through tube is widely adopted in literature, we find that simple frame alignment already provides enough essence without temporal attention. To this end, in this paper, we proposed a novel Implicit Learnable Alignment (ILA) method, which minimizes the temporal modeling effort while achieving incredibly high performance. Specifically, for a frame pair, an interactive point is predicted in each frame, serving as a mutual information rich region. By enhancing the features around the interactive point, two frames are implicitly aligned. The aligned features are then pooled into a single token, which is leveraged in the subsequent spatial self-attention. Our method allows eliminating the costly or insufficient temporal self-attention in video. Extensive experiments on benchmarks demonstrate the superiority and generality of our module. Particularly, the proposed ILA achieves a top-1 accuracy of 88.7% on Kinetics-400 with much fewer FLOPs compared with Swin-L and ViViT-H. Code is released at https://github.com/Francis-Rings/ILA. Shuyuan Tu, Qi Dai 0001, Zuxuan Wu, Zhi-Qi Cheng, Han Hu 0001, Yu-Gang Jiang 0001 |
ICCV | 4 |
| 2023 | HDFormer: High-order Directed Transformer for 3D Human Pose EstimationabstractHuman pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a novel approach, the High-order Directed Transformer (HDFormer), which leverages high-order bone and joint relationships for improved pose estimation. Specifically, HDFormer incorporates both self-attention and high-order attention to formulate a multi-order attention module. This module facilitates first-order "joint-joint", second-order "bone-joint", and high-order "hyperbone-joint" interactions, effectively addressing issues in complex and occlusion-heavy situations. In addition, modern CNN techniques are integrated into the transformer-based architecture, balancing the trade-off between performance and efficiency. HDFormer significantly outperforms state-of-the-art (SOTA) models on Human3.6M and MPI-INF-3DHP datasets, requiring only 1/10 of the parameters and significantly lower computational costs. Moreover, HDFormer demonstrates broad real-world applicability, enabling real-time, accurate 3D pose estimation. The source code is in https://github.com/hyer/HDFormer. Jun-Yan He, Wangmeng Xiang, Zhi-Qi Cheng, Wei Liu 0015, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
IJCAI | 4 |
| 2023 | DAMO-StreamNet: Optimizing Streaming Perception in Autonomous DrivingabstractIn the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO-StreamNet's main inventions include: (1) a robust neck structure employing deformable convolution, bolstering receptive field and feature alignment capabilities; (2) a dual-branch structure synthesizing short-path semantic features and long-path temporal features, enhancing the accuracy of motion state prediction; (3) logits-level distillation facilitating efficient optimization, which aligns the logits of teacher and student networks in semantic space; and (4) a real-time prediction mechanism that updates the features of support frames with the current frame, providing smooth streaming perception during inference. Our testing shows that DAMO-StreamNet surpasses current state-of-the-art methodologies, achieving 37.8% (normal size (600, 960)) and 43.3% (large size (1200, 1920)) sAP without requiring additional data. This study not only establishes a new standard for real-time perception but also offers valuable insights for future research. The source code is at https://github.com/zhiqic/DAMO-StreamNet. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Wangmeng Xiang, Binghui Chen, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
IJCAI | 2 |
| 2023 | KeyPosS: Plug-and-Play Facial Landmark Detection through GPS-Inspired True-Range MultilaterationabstractIn the realm of facial analysis, accurate landmark detection is crucial for various applications, ranging from face recognition and expression analysis to animation. Conventional heatmap or coordinate regression-based techniques, however, often face challenges in terms of computational burden and quantization errors. To address these issues, we present the KeyPoint Positioning System (KeyPosS) - a groundbreaking facial landmark detection framework that stands out from existing methods. The framework utilizes a fully convolutional network to predict a distance map, which computes the distance between a Point of Interest (POI) and multiple anchor points. These anchor points are ingeniously harnessed to triangulate the POI's position through the True-range Multilateration algorithm. Notably, the plug-and-play nature of KeyPosS enables seamless integration into any decoding stage, ensuring a versatile and adaptable solution. We conducted a thorough evaluation of KeyPosS's performance by benchmarking it against state-of-the-art models on four different datasets. The results show that KeyPosS substantially outperforms leading methods in low-resolution settings while requiring a minimal time overhead.1 The code is available at https://github.com/zhiqic/KeyPosS. Xu Bao 0003, Zhi-Qi Cheng, Jun-Yan He, Wangmeng Xiang, Chenyang Li 0007, Jingdong Sun, Wei Liu 0015, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ACM Multimedia | 2 |
| 2023 | PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose EstimationabstractThe current 3D human pose estimators face challenges in adapting to new datasets due to the scarcity of 2D-3D pose pairs in target domain training sets. We present the Multi-Hypothesis Pose Synthesis Domain Adaptation (PoSynDA) framework to overcome this issue without extensive target domain annotation. Utilizing a diffusion-centric structure, PoSynDA simulates the 3D pose distribution in the target domain, filling the data diversity gap. By incorporating a multi-hypothesis network, it creates diverse pose hypotheses and aligns them with the target domain. Target-specific source augmentation obtains the target domain distribution data from the source domain by decoupling the scale and position parameters. The teacher-student paradigm and low-rank adaptation further refine the process. PoSynDA demonstrates competitive performance on benchmarks, such as Human3.6M, MPI-INF-3DHP, and 3DPW, even comparable with the target-trained MixSTE model. This work paves the way for the practical application of 3D human pose estimation1. The source code is available at https://github.com/hbing-l/PoSynDA. Jun-Yan He, Zhi-Qi Cheng, Wangmeng Xiang, Qize Yang, Wenhao Chai, Gaoang Wang, Xu Bao 0003, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ACM Multimedia | 3 |
| 2023 | Debunking Free Fusion Myth: Online Multi-view Anomaly Detection with Disentangled Product-of-Experts ModelingabstractMulti-view or even multi-modal data is appealing yet challenging for real-world applications. Detecting anomalies in multi-view data is a prominent recent research topic. However, most of the existing methods 1) are only suitable for two views or type-specific anomalies, 2) suffer from the issue of fusion disentanglement, and 3) do not support online detection after model deployment. To address these challenges, our main ideas in this paper are three-fold: multi-view learning, disentangled representation learning, and generative model. To this end, we propose dPoE, a novel multi-view variational autoencoder model that involves (1) a Product-of-Experts (PoE) layer in tackling multi-view data, (2) a Total Correction (TC) discriminator in disentangling view-common and view-specific representations, and (3) a joint loss function in wrapping up all components. In addition, we devise theoretical information bounds to control both view-common and view-specific representations. Extensive experiments on six real-world datasets demonstrate that the proposed dPoE outperforms baselines markedly. Hao Wang 0068, Zhi-Qi Cheng, Jingdong Sun, Xin Yang 0012, Xiao Wu 0001, Hongyang Chen 0001, Yan Yang 0001 |
ACM Multimedia | 2 |
| 2023 | Improving Anomaly Segmentation with Multi-Granularity Cross-Domain AlignmentabstractAnomaly segmentation plays a crucial role in identifying anomalous objects within images, which facilitates the detection of road anomalies for autonomous driving. Although existing methods have shown impressive results in anomaly segmentation using synthetic training data, the domain discrepancies between synthetic training data and real test data are often neglected. To address this issue, Multi-Granularity Cross-Domain Alignment (MGCDA) framework is proposed for anomaly segmentation in complex driving environments. It uniquely combines a new Multi-source Domain Adversarial Training (MDAT) module and a novel Cross-domain Anomaly-aware Contrastive Learning (CACL) method to boost the generality of the model, seamlessly integrating multi-domain data at both scene and sample levels. Multi-source domain adversarial loss and a dynamic label smoothing strategy are integrated into MDAT module to facilitate the acquisition of domain-invariant features at the scene level, through adversarial training across multiple stages. CACL aligns sample-level representations with contrastive loss on cross-domain data, which utilizes an anomaly-aware sampling strategy to efficiently sample hard samples and anchors. The proposed framework has decent properties of parameter-free during the inference stage and is compatible with other anomaly segmentation networks. Experimental conducted on Fishyscapes and RoadAnomaly datasets demonstrate that the proposed framework achieves the state-of-the-art performance. Ji Zhang 0027, Xiao Wu 0001, Zhi-Qi Cheng, Qi He 0007, Wei Li 0110 |
ACM Multimedia | 3 |
| 2022 | Rethinking Spatial Invariance of Convolutional Networks for Object CountingabstractPrevious work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too strict pixel-level spatial invariance would cause overfit noise in the density map generation. In this paper, we try to use locally connected Gaussian kernels to replace the original convolution filter to estimate the spatial position in the density map. The purpose of this is to allow the feature extraction process to potentially stimulate the density map generation process to overcome the annotation noise. Inspired by previous work, we propose a low-rank approximation accompanied with translation invariance to favorably implement the approximation of massive Gaussian convolution. Our work points a new direction for follow-up research, which should investigate how to properly relax the overly strict pixel-level spatial invariance for object counting. We evaluate our methods on 4 mainstream object counting networks (i.e., MCNN, CSRNet, SANet, and ResNet-50). Extensive experiments were conducted on 7 popular benchmarks for 3 applications (i.e., crowd, vehicle, and plant counting). Experimental results show that our methods significantly outperform other state-of-the-art methods and achieve promising learning of the spatial position of objects11Code is at https://github.com/zhiqic/Rethinking-Counting. Zhi-Qi Cheng, Qi Dai 0001, Jingkuan Song, Xiao Wu 0001, Alex Hauptmann 0001 |
CVPR | 1 |
| 2022 | GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementabstractGrounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like'' event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but also predicts all corresponding semantic roles (e.g. agent and goods). Inspired by object detection and image captioning tasks, existing methods typically employ a two-stage framework: 1) detect the activity verb, and then 2) predict semantic roles based on the detected verb. Obviously, this illogical framework constitutes a huge obstacle to semantic understanding. First, pre-detecting verbs solely without semantic roles inevitably fails to distinguish many similar daily activities (e.g., offering and giving, buying and selling). Second, predicting semantic roles in a closed auto-regressive manner can hardly exploit the semantic relations among the verb and roles. To this end, in this paper we propose a novel two-stage framework that focuses on utilizing such bidirectional relations within verbs and roles. In the first stage, instead of pre-detecting the verb, we postpone the detection step and assume a pseudo label, where an intermediate representation for each corresponding semantic role is learned from images. In the second stage, we exploit transformer layers to unearth the potential semantic relations within both verbs and semantic roles. With the help of a set of support images, an alternate learning scheme is designed to simultaneously optimize the results: update the verb using nouns corresponding to the image, and update nouns using verbs from support images. Extensive experimental results on challenging SWiG benchmarks show that our renovated framework outperforms other state-of-the-art methods under various metrics. Zhi-Qi Cheng, Qi Dai 0001, Siyao Li, Teruko Mitamura, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2022 | Real-time Semantic Segmentation with Parallel Multiple Views Feature AugmentationabstractReal-time semantic segmentation is essential for many practical applications, which utilizes attention-based feature aggregation into lightweight structures to improve accuracy and efficiency. However, existing attention-based methods ignore 1) high-level and low-level feature augmentation guided by spatial information, and 2) low-level feature augmentation guided by semantic context, so that feature gaps between multi-level features and noise of low-level spatial details still exist. To address these problems, a new real-time semantic segmentation network, called MvFSeg, is proposed. In MvFSeg, parallel convolution with multiple depths is designed as a context head to generate and integrate multi-view features with larger receptive fields. Moreover, MvFSeg designs multiple views feature augmentation strategies that exploit spatial and semantic guidance for shallow and deep feature augmentation in an inter-layer and intra-layer manner. These strategies eliminate feature gaps between multi-level features, filter out the noise of spatial details, and provide spatial and semantic guidance for multi-level features. By combining multi-view features and augmented features from the lightweight networks with progressive dense aggregation structures, MvFSeg effectively captures invariance at various scales and generates high-quality segmentation results. Experiments conducted on Cityscapes and CamVid benchmark show that MvFSeg outperforms existing state-of-the-art methods. Jian-Jun Qiao, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Ji Zhang 0027 |
ACM Multimedia | 2 |
| 2022 | CrossNet: Boosting Crowd Counting with LocalizationabstractGenerating high-quality density maps is a crucial step in crowd counting. It is obvious that exploiting the head location of the people can naturally highlight the crowded area and eliminate the interference of background noise. However, existing crowd counting methods are still tricky to reasonably use location in density generation. In this paper, a novel location-guided framework named CrossNet is proposed for crowd counting, which integrates location supervision into density maps through dual-branch joint training. First, a new branching network is proposed to localize the potential positions of pedestrians. With the help of supervision induced from the localization branch, Location Enhancement (LE) module is designed to obtain high-quality density maps by positioning foreground regions. Second, Adaptive Density Awareness Attention (ADAA) module is engaged to enhance localization accuracy, which can efficiently use the density of the counting branch to adaptively capture the error-prone dense areas of the location maps. Finally, Density Awareness Localization (DAL) loss is offered to allocate attention to the crowd density levels, which delivers more focus on regions with high densities and less concentration on areas with low densities. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches both in crowd counting and crowd localization. Ji Zhang 0027, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Jian-Jun Qiao |
ACM Multimedia | 2 |
| 2021 | DB-LSTM: Densely-connected Bi-directional LSTM for human action recognition
Jun-Yan He, Xiao Wu 0001, Zhi-Qi Cheng, Zhaoquan Yuan, Yu-Gang Jiang 0001 |
Neurocomputing | 3 |
| 2020 | Stacked Pooling for Boosting Scale Invariance of Crowd CountingabstractIn this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance. Existing deep crowd counting approaches mainly focus on the multi-scale techniques over convolutional layers to capture scale-adaptive features, resulting in high computing costs. In this paper, we propose simple but effective pooling variants, i.e., multi-kernel pooling and stacked pooling, to take place of the vanilla pooling layers in convolutional neural networks (CNNs) for boosting the scale invariance. Our proposed pooling modules do not introduce extra parameters and can be easily implemented in practice. Empirical studies on two benchmark crowd counting datasets show that the proposed pooling modules beat the vanilla pooling layer in most experimental cases. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ICASSP | 3 |
| 2020 | Generating Person Images with Appearance-aware Pose StylizerabstractGeneration of high-quality person images is challenging, due to the sophisticated entanglements among image factors, e.g., appearance, pose, foreground, background, local details, global structures, etc. In this paper, we present a novel end-to-end framework to generate realistic person images based on given person poses and appearances. The core of our framework is a novel generator called Appearance-aware Pose Stylizer (APS) which generates human images by coupling the target pose with the conditioned person appearance progressively. The framework is highly flexible and controllable by effectively decoupling various complex person image factors in the encoding phase, followed by re-coupling them in the decoding phase. In addition, we present a new normalization method named adaptive patch normalization, which enables region-specific normalization and shows a good performance when adopted in person image generation model. Experiments on two benchmark datasets show that our method is capable of generating visually appealing and realistic-looking results using arbitrary image and pose inputs. Siyu Huang, Haoyi Xiong, Zhi-Qi Cheng, Qingzhong Wang, Xingran Zhou, Bihan Wen, Jun Huan, Dejing Dou |
IJCAI | 3 |
| 2019 | Learning Spatial Awareness to Improve Crowd CountingabstractThe aim of crowd counting is to estimate the number of people in images by leveraging the annotation of center positions for pedestrians' heads. Promising progresses have been made with the prevalence of deep Convolutional Neural Networks. Existing methods widely employ the Euclidean distance (i.e., L2loss) to optimize the model, which, however, has two main drawbacks: (1) the loss has difficulty in learning the spatial awareness (i.e., the position of head) since it struggles to retain the high-frequency variation in the density map, and (2) the loss is highly sensitive to various noises in crowd counting, such as the zeromean noise, head size changes, and occlusions. Although the Maximum Excess over SubArrays (MESA) loss has been previously proposed by [16] to address the above issues by finding the rectangular subregion whose predicted density map has the maximum difference from the ground truth, it cannot be solved by gradient descent, thus can hardly be integrated into the deep learning framework. In this paper, we present a novel architecture called SPatial Awareness Network (SPANet) to incorporate spatial context for crowd counting. The Maximum Excess over Pixels (MEP) loss is proposed to achieve this by finding the pixel-level subregion with high discrepancy to the ground truth. To this end, we devise a weakly supervised learning scheme to generate such region with a multi-branch architecture. The proposed framework can be integrated into existing deep crowd counting methods and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that our method can significantly improve the performance of baselines. More remarkably, our approach outperforms the state-of-the-art methods on all benchmark datasets. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Alex Hauptmann 0001 |
ICCV | 1 |
| 2019 | Improving the Learning of Multi-column Convolutional Neural Network for Crowd CountingabstractTremendous variation in the scale of people/head size is a critical problem for crowd counting. To improve the scale invariance of feature representation, recent works extensively employ Convolutional Neural Networks with multi-column structures to handle different scales and resolutions. However, due to the substantial redundant parameters in columns, existing multi-column networks invariably exhibit almost the same scale features in different columns, which severely affects counting accuracy and leads to overfitting. In this paper, we attack this problem by proposing a novel Multicolumn Mutual Learning (McML) strategy. It has two main innovations: 1) A statistical network is incorporated into the multi-column framework to estimate the mutual information between columns, which can approximately indicate the scale correlation between features from different columns. By minimizing the mutual information, each column is guided to learn features with different image scales. 2) We devise a mutual learning scheme that can alternately optimize each column while keeping the other columns fixed on each mini-batch training data. With such asynchronous parameter update process, each column is inclined to learn different feature representation from others, which can efficiently reduce the parameter redundancy and improve generalization ability. More remarkably, McML can be applied to all existing multi-column networks and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that McML can significantly improve the original multi-column networks and outperform the other state-of-the-art approaches. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2018 | Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model SearchabstractAs attribute leaning brings mid-level semantic properties for objects, it can benefit many traditional learning problems in multimedia and computer vision communities. When facing the huge number of attributes, it is extremely challenging to automatically design a generalizable neural network for other attribute learning tasks. Even for a specific attribute domain, the exploration of the neural network architecture is always optimized by a combination of heuristics and grid search, from which there is a large space of possible choices to be searched. In this paper, Generalizable Attribute Learning Model (GALM) is proposed to automatically design the neural networks for generalizable attribute learning. The main novelty of GALM is that it fully exploits the Multi-Task Learning and Reinforcement Learning to speed up the search procedure. With the help of parameter sharing, GALM is able to transfer the pre-searched architecture to different attribute domains. In experiments, we comprehensively evaluate GALM on 251 attributes from three domains: animals, objects, and scenes. Extensive experimental results demonstrate that GALM significantly outperforms the state-of-the-art attribute learning approaches and previous neural architecture search methods on two generalizable attribute learning scenarios. Zhi-Qi Cheng, Xiao Wu 0001, Siyu Huang, Jun-Xiu Li, Alex Hauptmann 0001, Qiang Peng |
ACM Multimedia | 1 |
| 2018 | GNAS: A Greedy Neural Architecture Search Method for Multi-Attribute LearningabstractA key problem in deep multi-attribute learning is to effectively discover the inter-attribute correlation structures. Typically, the conventional deep multi-attribute learning approaches follow the pipeline of manually designing the network architectures based on task-specific expertise prior knowledge and careful network tunings, leading to the inflexibility for various complicated scenarios in practice. Motivated by addressing this problem, we propose an efficient greedy neural architecture search approach (GNAS) to automatically discover the optimal tree-like deep architecture for multi-attribute learning. In a greedy manner, GNAS divides the optimization of global architecture into the optimizations of individual connections step by step. By iteratively updating the local architectures, the global tree-like architecture gets converged where the bottom layers are shared across relevant attributes and the branches in top layers more encode attribute-specific features. Experiments on three benchmark multi-attribute datasets show the effectiveness and compactness of neural architectures derived by GNAS, and also demonstrate the efficiency of GNAS in searching neural architectures. Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2018 | Multi-View Image Generation from a Single-ViewabstractHow to generate multi-view images with realistic-looking appearance from only a single view input is a challenging problem. In this paper, we attack this problem by proposing a novel image generation model termed VariGANs, which combines the merits of the variational inference and the Generative Adversarial Networks (GANs). It generates the target image in a coarse-to-fine manner instead of a single pass which suffers from severe artifacts. It first performs variational inference to model global appearance of the object (e.g., shape and color) and produces coarse images of different views. Conditioned on the generated coarse images, it then performs adversarial learning to fill details consistent with the input and generate the fine images. Extensive experiments conducted on two clothing datasets, MVC and DeepFashion, have demonstrated that the generated images with the proposed VariGANs are more plausible than those generated by existing approaches, which provide more consistent global appearance as well as richer and sharper details. Bo Zhao 0032, Xiao Wu 0001, Zhi-Qi Cheng, Hao Liu 0003, Zequn Jie, Jiashi Feng |
ACM Multimedia | 3 |
| 2018 | Personalized clothing recommendation combining user social circle and fashion style consistency
Guang-Lu Sun, Zhi-Qi Cheng, Xiao Wu 0001, Qiang Peng |
Multim. Tools Appl. | 2 |
| 2017 | Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images
Zhi-Qi Cheng, Xiao Wu 0001, Yang Liu 0155, Xian-Sheng Hua 0001 |
CVPR | 1 |
| 2017 | On the Selection of Anchors and Targets for Video HyperlinkingabstractA problem not well understood in video hyperlinking is what qualifies a fragment as an anchor or target. Ideally, anchors provide good starting points for navigation, and targets supplement anchors with additional details while not distracting users with irrelevant, false and redundant information. The problem is not trivial for intertwining relationship between data characteristics and user expectation. Imagine that in a large dataset, there are clusters of fragments spreading over the feature space. The nature of each cluster can be described by its size (implying popularity) and structure (implying complexity). A principle way of hyperlinking can be carried out by picking centers of clusters as anchors and from there reach out to targets within or outside of clusters with consideration of neighborhood complexity. The question is which fragments should be selected either as anchors or targets, in one way to reflect the rich content of a dataset, and meanwhile to minimize the risk of frustrating user experience. This paper provides some insights to this question from the perspective of hubness and local intrinsic dimensionality, which are two statistical properties in assessing the popularity and complexity of data space. Based these properties, two novel algorithms are proposed for low-risk automatic selection of anchors and targets. Zhi-Qi Cheng, Hao Zhang 0047, Xiao Wu 0001, Chong-Wah Ngo |
ICMR | 1 |
| 2017 | Video eCommerce++: Toward Large Scale Online Video AdvertisingabstractThe prevalence of online videos provides an opportunity for e-commerce companies to recommend their products in videos. In this paper, we propose an online video advertising system named Video eCommerce ++, to exhibit appropriate product ads to particular users at proper time stamps of videos, which takes into account video semantics, user shopping preference, and viewing behavior feedback. First, an incremental co-relation regression (ICRR) model is novelly proposed to construct the semantic association between videos and products. To meet the requirement of online advertising, ICRR is implemented in an incremental way to reduce the time complexity. User preference diffusion (UPD) is induced under the framework of heterogeneous information network to construct user-product association from two different e-commerce platforms, Tmall and MagicBox, which alleviates the problems of data sparsity and cold start. A video scene importance model (VSIM) is proposed to model the scene importance by utilizing the user viewing behavior, so that ads can be embedded at the most attractive positions in the video stream. To combine the outputs of ICRR, UPD, and VSIM, a unified distributed heterogeneous relation matrix factorization (D-HRMF) is applied for online video advertising, which is efficiently conducted in parallel to address the real-time update problem, so that the whole system can be performed in real time. Extensive experiments conducted on a variety of online videos from Tmall MagicBox demonstrate that Video eCommerce++ significantly outperforms the state-of-the-art advertising methods, and can handle large-scale data in real time. Zhi-Qi Cheng, Xiao Wu 0001, Yang Liu 0155, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Video eCommerce: Towards Online Video AdvertisingabstractThe prevalence of online videos provides an opportunity for e-commerce companies to exhibit their product ads in videos by recommendation. In this paper, we propose an advertising system named Video eCommerce to exhibit appropriate product ads to particular users at proper time stamps of videos, which takes into account video semantics, user shopping preference and viewing behavior feedback by a two-level strategy. At the first level, Co-Relation Regression (CRR) model is novelly proposed to construct the semantic association between keyframes and products. Heterogeneous information network (HIN) is adopted to build the user shopping preference from two different e-commerce platforms, Tmall and MagicBox, which alleviates the problems of data sparsity and cold start. In addition, Video Scene Importance Model (VSIM) utilizes the viewing behavior of users to embed ads at the most attractive position within the video stream. At the second level, taking the results of CRR, HIN and VSIM as the input, Heterogeneous Relation Matrix Factorization (HRMF) is applied for product advertising. Extensive evaluation on a variety of online videos from Tmall MagicBox demonstrates that Video eCommerce achieves promising performance, which significantly outperforms the state-of-the-art advertising methods. Zhi-Qi Cheng, Yang Liu 0155, Xiao Wu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 1 |