EDBT 2026 Demo / reviewers in the wild / expert
Zhaoyang Zeng
dblp:207/1887
· DBLP profile ↗
34ranked-venue papers
7as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | T-Rex2++: Toward Generic Object Perception via Text-Visual Prompt SynergyabstractWe present T-Rex2++, a unified and highly practical framework for generic open-set object perception, encompassing both object detection and instance segmentation. Previous methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex object representation due to data scarcity and descriptive limitations. Conversely, visual prompts excel in depicting novel objects through concrete visual examples, but fall short in conveying the abstract concept of objects as effectively as text prompts. Recognizing these complementary strengths, we introduce a text-visual synergy mechanism that aligns both modalities within a single feature space via contrastive learning. Crucially, T-Rex2++ advances beyond the passive perception paradigm of its predecessor by introducing a novel Universal Prompt. This learnable component models generic objectness, empowering the system to autonomously discover and localize arbitrary objects without any user-provided cues, thereby closing the loop between human-guided interaction and fully automatic perception. Furthermore, we extend the synergy verification to the pixel level by integrating a zero-shot instance segmentation module, demonstrating that our contrastive alignment generalizes robustly to fine-grained masks. Comprehensive experiments demonstrate that T-Rex2++ exhibits strong zero-shot object perception capabilities across a wide spectrum of scenarios, validating T-Rex2++ as a versatile foundation for generic object perception. Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Latency Optimization in Hybrid Memory System for GNNsabstractGraph Neural Networks (GNNs) require high-capacity, low-latency memory systems to process large graphs. A hierarchical hybrid memory architecture combining high-capacity Non-Volatile Memory (NVM) and low-latency DRAM offers a promising solution. However, the inherent sparsity of graph data results in poor locality for GNN memory requests, leading to low DRAM cache hit rates and numerous misses, which significantly impairs the hybrid memory system’s performance. A critical issue is that DRAM misses in serial access mode incur substantial latency. While parallel access mode can mitigate this for misses, it introduces long-tail latency and wastes bandwidth for DRAM hits. In this paper, we focus on addressing these issues from two aspects: increasing the cache hit rate and decreasing the miss latency. We mainly propose two predictors: a future data access predictor that enables accurate prefetching to DRAM, thereby improving cache hit rates, and a data location predictor that determines whether data resides in DRAM or NVM, optimizing the choice between serial and parallel access modes to reduce miss latency. By integrating these predictors, we achieve efficient data access in both DRAM and NVM. Our experiments show a 49.5% reduction in memory delay and a 38.1% increase in memory bandwidth utilization compared to baseline. Zhaoyang Zeng, Yujuan Tan, Wei Chen 0101, Zhuoxin Bai, Ao Ren, Duo Liu 0002, Xianzhang Chen |
IEEE Trans. Computers | 1 |
| 2025 | Cocache: An Accurate and Low-Overhead Dynamic Caching Method for GNNs
Zhaoyang Zeng, Yujuan Tan, Zhuoxin Bai, Kan Zhong, Duo Liu 0002, Ao Ren |
Euro-Par (2) | 1 |
| 2025 | Referring to Any PersonabstractHumans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referring to any person, holds substantial practical value. However, we find that existing models generally fail to achieve real-world usability, and current benchmarks are limited by their focus on one-to-one referring, that hinder progress in this area. In this work, we revisit this task from three critical perspectives: task definition, dataset design, and model architecture. We first identify five aspects of referable entities and three distinctive characteristics of this task. Next, we introduce HumanRef, a novel dataset designed to tackle these challenges and better reflect real-world applications. From a model design perspective, we integrate a multimodal large language model with an object detection framework, constructing a robust referring model named RexSeek. Experimental results reveal that state-of-the-art models, which perform well on commonly used benchmarks like RefCOCO/+/g, struggle with HumanRef due to their inability to detect multiple individuals. In contrast, RexSeek not only excels in human referring but also generalizes effectively to common object referring, making it broadly applicable across various perception tasks. Code is available at https://github.com/IDEA-Research/RexSeek Zhaoyang Zeng, Tianhe Ren, Yuda Xiong |
ICCV | 3 |
| 2025 | SMPV: Social Media Prediction for Videos
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2025 | GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren |
J. Syst. Archit. | 3 |
| 2025 | TAPTR3D: Decoupled 3D Point Tracking Boosts 2D and Further Enhances 3D Tracking Accuracy
Hongyang Li 0003, Jinyuan Qu, Zhaoyang Zeng, Lei Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | LAShards: Low-Overhead and Self-Adaptive MRC Construction for Non-Stack AlgorithmsabstractShared cache systems have become increasingly crucial, especially in cloud services, where the Miss Ratio Curve (MRC) is a widely used tool for evaluating cache performance. The MRC depicts the relationship between the cache miss ratio and cache size, indicating how cache performance trends with varying cache sizes. Recent advancements have enabled efficient MRC construction for stack replacement policies. For non-stack policies, miniature simulation downsizes the actual cache size and data stream through spatially hashed sampling, providing a general method for MRC construction. However, this approach still faces significant challenges. Firstly, constructing an MRC requires numerous mini-caches to obtain miss ratios, consuming significant cache resources, leading to tremendous memory and computing overhead. Secondly, it cannot adapt to the dynamic I/O workloads, resulting in less precise MRC.To address these issues, we propose LAShards, a low-overhead and self-adaptive MRC construction method for non-stack replacement policies. The key idea behind LAShards is to exploit the locality and burstiness in access patterns. It can statically reduce memory usage and dynamically adapt to workloads. Compared to previous works, LAShards can save up to 20× of memory resources, and increase throughput by up to 10×. Sanle Zhao, Yujuan Tan, Zhaoyang Zeng, Jing Yu 0026, Zhuoxin Bai, Ao Ren, Xianzhang Chen, Duo Liu 0002 |
IEEE Trans. Computers | 3 |
| 2025 | CMCache: An Adaptive Cross-Level Data Placement Method for Multilevel CacheabstractMultilevel cache systems enhance I/O performance by optimizing data placement across various cache levels from a global perspective. However, existing methods often struggle to place data at the optimal cache level promptly due to their reliance on historical access patterns and inflexible placement strategies. These methods face two main challenges: 1) for already cached data with sufficient access history, existing approaches only optimize movement between adjacent cache levels, potentially delaying data arrival at its globally optimal cache level and leading to unnecessary bandwidth consumption and increased latency and 2) for newly entered data without access history, current methods cannot accurately predict their future hotness and simply place them at a fixed cache level (i.e., first or final level), overlooking future accesses of new data and potentially resulting in high cache miss rates or cache pollution. To address these issues, we propose CMCache, an adaptive cross-level data placement method for multilevel cache. CMCache applies distinct placement strategies for cached and new data to reach the optimal level timely, considering their different characteristics. It also logically divides cache space into two sections to manage cached and new data separately, dynamically adjusting section sizes based on access patterns. This approach significantly improves data placement efficiency, achieving up to an 89% reduction in miss rates and a 79% decrease in average response times compared to existing methods. Zhaoyang Zeng, Yujuan Tan, Zhulin Ma, Sanle Zhao, Duo Liu 0002, Xianzhang Chen, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001 |
ECCV (33) | 3 |
| 2024 | TAPTR: Tracking Any Point with Transformers as Detection
Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Lei Zhang 0001 |
ECCV (16) | 4 |
| 2024 | Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Hao Zhang 0097, Chunyuan Li, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ECCV (47) | 2 |
| 2024 | Reproducibility Companion Paper: Stable Diffusion for Content-Style Disentanglement in Art AnalysisabstractIn this companion paper, we provide the artifacts of the GOYA model for disentangling content and style in art paintings, as presented at ICMR2023. The scripts are written in Python. Yankun Wu, Yuta Nakashima, Noa Garcia, Sheng Li 0010, Zhaoyang Zeng |
ICMR | 5 |
| 2024 | SMP Challenge Summary: Social Media Prediction Challenge
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2024 | TAPTRv2: Attention-based Position Update Improves Tracking Any PointabstractIn this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-like algorithms. TAPTRv2 improves TAPTR by addressing a critical issue regarding its reliance on cost-volume, which contaminates the point query’s content feature and negatively impacts both visibility prediction and cost-volume computation. In TAPTRv2, we propose a novel attention-based position update (APU) operation and use key-aware deformable attention to realize. For each query, this operation uses key-aware attention weights to combine their corresponding deformable sampling positions to predict a new query position. This design is based on the observation that local attention is essentially the same as cost-volume, both of which are computed by dot-production between a query and its surrounding features. By introducing this new operation, TAPTRv2 not only removes the extra burden of cost-volume computation, but also leads to a substantial performance improvement. TAPTRv2 surpasses TAPTR and achieves state-of-the-art performance on many challenging datasets, demonstrating the effectiveness of our approach. Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Feng Li 0040, Tianhe Ren, Lei Zhang 0006 |
NeurIPS | 4 |
| 2024 | BGS: Accelerate GNN training on multiple GPUs
Yujuan Tan, Zhuoxin Bai, Duo Liu 0002, Zhaoyang Zeng, Yan Gan, Ao Ren, Xianzhang Chen, Kan Zhong |
J. Syst. Archit. | 4 |
| 2024 | LightFS: A Lightweight Host-CSD Coordinated File System Optimizing for Heavy Small File AccessesabstractComputational storage drive (CSD) improves the data processing efficiency by processing the data within the storage. However, existing CSDs rely on the host-centric file systems to manage the data, where the layouts of files are retrieved by the host and sent to the CSD, resulting in additional I/O overhead and reduced processing efficiency, especially in heavy small file accesses. Moreover, the lack of consistency mechanisms poses potential consistency issues. To address these challenges, we propose LightFS, a lightweight host-CSD coordinated file system for the CSD file management. To reduce task offloading overhead, LightFS builds an index file$.ndpmeta$which summarizes the files’ metadata and shares between the host and CSD to enable CSD to retrieve the file layout in storage directly. To ensure consistency, LightFS employs a metadata locker and an update synchronizer. The metadata locker leverages the out-of-place update feature of the flash to capture a snapshot of the file to be written without any data copy, while the update synchronizer triggers metadata updates by monitoring the addresses of written blocks to ensure that the modified file is successfully written to the CSD. We implement and evaluate LightFS on a real testbed, and the results demonstrate that LightFS achieves$3.66\times $performance improvement on the average in real-world operations. Zhaoyan Shen, Duo Liu 0002, Xianzhang Chen, Kan Zhong, Zhaoyang Zeng, Yujuan Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | DFA3D: 3D Deformable Attention For 2D-to-3D Feature LiftingabstractIn this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either use estimated depth to get pseudo LiDAR features and then splat them to a 3D space, which is a one-pass operation without feature refinement, or ignore depth and lift features by 2D attention mechanisms, which achieve finer semantics while suffering from a depth ambiguity problem. In contrast, our DFA3D-based method first leverages the estimated depth to expand each view’s 2D feature map to 3D and then utilizes DFA3D to aggregate features from the expanded 3D feature maps. With the help of DFA3D, the depth ambiguity problem can be effectively alleviated from the root, and the lifted features can be progressively refined layer by layer, thanks to the Transformerlike architecture. In addition, we propose a mathematically equivalent implementation of DFA3D which can significantly improve its memory efficiency and computational speed. We integrate DFA3D into several methods that use 2D attention-based feature lifting with only a few modifications in code and evaluate on the nuScenes dataset. The experiment results show a consistent improvement of +1.41% mAP on average, and up to +15.1% mAP improvement when high-quality depth information is available, demonstrating the superiority, applicability, and huge potential of DFA3D. The code is available at https://github.com/IDEAResearch/3D-deformable-attention.git. Hongyang Li 0003, Hao Zhang 0097, Zhaoyang Zeng, Shilong Liu 0004, Feng Li 0040, Tianhe Ren, Lei Zhang 0001 |
ICCV | 3 |
| 2023 | Detection Transformer with Stable MatchingabstractThis paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address this problem, we show that the most important design is to use and only use positional metrics (like IOU) to supervise classification scores of positive examples. Under the principle, we propose two simple yet effective modifications by integrating positional metrics to DETR’s classification loss and matching cost, named position-supervised loss and position-modulated cost. We verify our methods on several DETR variants. Our methods show consistent improvements over baselines. By integrating our methods with DINO, we achieve 50.4 and 51.5 AP on the COCO detection benchmark using ResNet-50 backbones under 1× (12 epochs) and 2× (24 epochs) training settings, achieving a new record under the same setting. We achieve 63.8 AP on COCO detection test-dev with a Swin-Large backbone. Our code will be made available at https://github.com/IDEA-Research/Stable-DINO. Shilong Liu 0004, Tianhe Ren, Zhaoyang Zeng, Hao Zhang 0097, Feng Li 0040, Hongyang Li 0003, Jun Huang 0007, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICCV | 4 |
| 2023 | RadarSSD: A Computational Storage for Radar Signal ProcessingabstractRadar signals contain a multitude of small data items with multidimensional characteristics and various types of errors. It is challenging to store and recognize radar signals in real-time. Traditional computer architectures require data to be moved from storage to the host for processing, resulting in a "storage wall" problem. This problem is caused by low storage bandwidth, long I/O stacks, and excessive data transfers, which significantly reduce the efficiency of radar signal recognition. In this paper, we propose RadarSSD address these challenges by utilizing near-data processing (NDP) architecture to recognize radar signals within the solid-state drive (SSD), through which the high overhead of data movements can be avoided. To support efficient data I/O operations, we design a stripe-like data layout for storing radar signals taking advantage of their time sequential feature. We present a task slicing mechanism to reduce I/O blocking from in-storage data processing, and a dedicated interface for providing highly-efficient direct SSD access. We implement RadarSSD in a real computational SSD platform. Extensive experimental results show that RadarSSD can reduce power consumption while improving I/O and recognize performance, with a maximum improvement of 12.4 ×, 11.5 ×, and 4.1 × recognize speed compared to the systems that manage radar signals using MySQL, MongoDB, and Ext4. Xianzhang Chen, Duo Liu 0002, Ao Ren, Zhaoyang Zeng, Yujuan Tan |
ICPP | 5 |
| 2023 | SMP Challenge: An Overview and Analysis of Social Media Prediction ChallengeabstractSocial Media Popularity Prediction (SMPP) is a crucial task that involves automatically predicting future popularity values of online posts, leveraging vast amounts of multimodal data available on social media platforms. Studying and investigating social media popularity becomes central to various online applications and requires novel methods of comprehensive analysis, multimodal comprehension, and accurate prediction. Bo Wu 0018, Peiye Liu, Wen-Huang Cheng, Bei Liu 0001, Zhaoyang Zeng, Jia Wang 0020, Qiushi Huang, Jiebo Luo 0001 |
ACM Multimedia | 5 |
| 2023 | Stay in Grid: Improving Video Captioning via Fully Grid-Level RepresentationabstractVideo captioning is a challenging task of automatically generating natural and meaningful textual descriptions given some context videos. The state-of-the-art methods aggregate the spatial-wise information in the video encoder at the early stage, which has two drawbacks: 1) Early aggregation in the encoder can cause considerable spatial details missing, which may consequently lead to incorrect word choices in the following text encoder. 2) The spatial attention learned in the video encoder may not be compelling enough without text guidance. To solve these problems, we propose a Stay-in-Grid video CAPtioning method SGCAP, which makes full use of the grid-level spatial features and consists of a Bilinear Sequential Attention Encoder (BSAE) and a Cross-modal Sequential Attention Decoder (CSAD). The former explores and retains fully grid-level discriminative representations in the video encoder, while the latter performs the late spatial aggregation in the decoder to attend to the most relevant regions with the supervision of the input words. Experimental results demonstrate the effectiveness of our method on three public datasets, showing its superior performance over multiple state-of-the-art video captioning models. Source codes and the pre-trained models will be made available to the public. Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng, Xiu Li 0001, Luping Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Tencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity EvaluationabstractMulti-modal video similarity evaluation is important for video recommendation systems such as video de-duplication, relevance matching, ranking, and diversity control. However, there still lacks a benchmark dataset that can support supervised training and accurate evaluation. In this paper, we propose the Tencent-MVSE dataset, which is the first benchmark dataset for the multi-modal video similarity evaluation task. The Tencent-MVSE dataset contains video pairs similarity annotations, and diverse metadata including Chinese title, automatic speech recognition (ASR) text, as well as human-annotated categories/tags. We provide a simple baseline with a multi-modal Transformer architecture to perform supervised multi-modal video similarity evaluation. We also explore pre-training strategies to make use of the unpaired data. The whole dataset as well as our baseline will be released to promote the development of the multi-modal video similarity evaluation. The dataset has been released in https://tencent-mvse.github.io/. Zhaoyang Zeng, Yongsheng Luo, Fengyun Rao, Weidong Guo |
CVPR | 1 |
| 2022 | Horae: A Hybrid I/O Request Scheduling Technique for Near-Data Processing-Based SSDabstractNear-data processing (NDP) architecture is promised to break the bottleneck of data movement in many scenarios (e.g., databases and recommendation systems), which limits the efficiency of data processing. Different from traditional SSD, NDP-based SSD not only needs to handle normal I/Os (e.g., read and write), but also needs to handle NDP requests that contain data processing operations. NDP and normal I/O requests share some function units of NDP-based SSD, such as flash chips and embedded processors. However, existing works ignore the resource competition between normal I/Os and NDP requests, which drastically degrades the performance. In this article, we propose a novel scheduling technique called Horae, which can efficiently schedule hybrid NDP-normal I/O requests in NDP-based SSD to improve performance. Horae exploits the critical paths on critical resources to maximize the parallelism of multiple stages of requests. The experimental results on typical workloads show that Horae can significantly improve the performance of hybrid NDP-normal I/O requests over the state-of-the-art scheduling algorithms of NDP-based SSDs. Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Zhaoyang Zeng, Yujuan Tan, Lei Qiao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Seeing Out of the Box: End-to-End Pre-Training for Vision-Language Representation LearningabstractWe study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of image-text pairs. State-of-the-art approaches extract salient image regions and align regions with words step-by-step. As region-based visual features usually represent parts of an image, it is challenging for existing vision-language models to fully understand the semantics from paired natural languages. In this paper, we propose SOHO to "Seeing Out of tHe bOx" that takes a whole image as input, and learns vision-language representation in an end-to-end manner. SOHO does not require bounding box annotations which enables inference 10 times faster than region-based approaches. In particular, SOHO learns to extract comprehensive yet compact image features through a visual dictionary (VD) that facilitates cross-modal understanding. VD is designed to represent consistent visual abstractions of similar semantics. It is updated on-the-fly and utilized in our proposed pre-training task Masked Visual Modeling (MVM). We conduct experiments on four well-established vision-language tasks by following standard VLPT settings. In particular, SOHO achieves absolute gains of 2.0% R@1 score on MSCOCO text retrieval 5k test split, 1.5% accuracy on NLVR2test-P split, 6.7% accuracy on SNLI-VE test split, respectively. Zhicheng Huang 0002, Zhaoyang Zeng, Yupan Huang, Bei Liu 0001, Dongmei Fu, Jianlong Fu |
CVPR | 2 |
| 2021 | Active Contrastive Learning of Audio-Visual Video Representations
Zhaoyang Zeng, Daniel McDuff, Yale Song |
ICLR | 2 |
| 2021 | Multi-modal Representation Learning for Video Advertisement Content StructuringabstractVideo advertisement content structuring aims to segment a given video advertisement and label each segment on various dimensions, such as presentation form, scene, and style. Different from real-life videos, video advertisements contain sufficient and useful multi-modal content like caption and speech, which provides crucial video semantics and would enhance the structuring process. In this paper, we propose a multi-modal encoder to learn multi-modal representation from video advertisements by interacting between video-audio and text. Based on multi-modal representation, we then apply Boundary-Matching Network to generate temporal proposals. To make the proposals more accurate, we refine generated proposals by scene-guided alignment and re-ranking. Finally, we incorporate proposal located embeddings into the introduced multi-modal encoder to capture temporal relationships between local features of each proposal and global features of the whole video for classification. Experimental results show that our method achieves significantly improvement compared with several baselines and Rank 1 on the task of Multi-modal Ads Video Understanding in ACM Multimedia 2021 Grand Challenge. Ablation study further shows that leveraging multi-modal content like caption and speech in video advertisements significantly improve the performance. Daya Guo, Zhaoyang Zeng |
ACM Multimedia | 2 |
| 2021 | Contrastive Learning of Global and Local Video RepresentationsabstractContrastive learning has delivered impressive results for various tasks in the self-supervised regime. However, existing approaches optimize for learning representations specific to downstream scenarios, i.e., global representations suitable for tasks such as classification or local representations for tasks such as detection and localization. While they produce satisfactory results in the intended downstream scenarios, they often fail to generalize to tasks that they were not originally designed for. In this work, we propose to learn video representations that generalize to both the tasks which require global semantic information (e.g., classification) and the tasks that require local fine-grained spatio-temporal information (e.g., localization). We achieve this by optimizing two contrastive objectives that together encourage our model to learn global-local visual information given audio signals. We show that the two objectives mutually improve the generalizability of the learned global-local representations, significantly outperforming their disjointly learned counterparts. We demonstrate our approach on various tasks including action/sound classification, lipreading, deepfake detection, event and sound localization. Zhaoyang Zeng, Daniel McDuff, Yale Song |
NeurIPS | 2 |
| 2021 | Reference-Based Defect Detection NetworkabstractThe defect detection task can be regarded as a realistic scenario of object detection in the computer vision field and it is widely used in the industrial field. Directly applying vanilla object detector to defect detection task can achieve promising results, while there still exists challenging issues that have not been solved. The first issue is the texture shift which means a trained defect detector model will be easily affected by unseen texture, and the second issue is partial visual confusion which indicates that a partial defect box is visually similar with a complete box. To tackle these two problems, we propose a Reference-based Defect Detection Network (RDDN). Specifically, we introduce template reference and context reference to against those two problems, respectively. Template reference can reduce the texture shift from image, feature or region levels, and encourage the detectors to focus more on the defective area as a result. We can use either well-aligned template images or the outputs of a pseudo template generator as template references in this work, and they are jointly trained with detectors by the supervision of normal samples. To solve the partial visual confusion issue, we propose to leverage the carried context information of context reference, which is the concentric bigger box of each region proposal, to perform more accurate region classification and regression. Experiments on two defect detection datasets demonstrate the effectiveness of our proposed approach. Zhaoyang Zeng, Bei Liu 0001, Jianlong Fu, Hongyang Chao |
IEEE Trans. Image Process. | 1 |
| 2020 | Suppressing Mislabeled Data via Grouping and Self-attention
Xiaojiang Peng, Kai Wang 0036, Zhaoyang Zeng, Qing Li 0058, Jianfei Yang 0001, Yu Qiao 0001 |
ECCV (16) | 3 |
| 2020 | Mind the Discriminability: Asymmetric Adversarial Domain Adaptation
Jianfei Yang 0001, Han Zou, Yuxun Zhou, Zhaoyang Zeng, Lihua Xie 0001 |
ECCV (24) | 4 |
| 2019 | WSOD2: Learning Bottom-Up and Top-Down Objectness Distillation for Weakly-Supervised Object DetectionabstractWe study on weakly-supervised object detection (WSOD) which plays a vital role in relieving human involvement from object-level annotations. Predominant works integrate region proposal mechanisms with convolutional neural networks (CNN). Although CNN is proficient in extracting discriminative local features, grand challenges still exist to measure the likelihood of a bounding box containing a complete object (i.e., “objectness”). In this paper, we propose a novel WSOD framework with Objectness Distillation (i.e., WSOD2) by designing a tailored training mechanism for weakly-supervised object detection. Multiple regression targets are specifically determined by jointly considering bottom-up (BU) and top-down (TD) objectness from low-level measurement and CNN confidences with an adaptive linear combination. As bounding box regression can facilitate a region proposal learning to approach its regression target with high objectness during training, deep objectness representation learned from bottom-up evidences can be gradually distilled into CNN by optimization. We explore different adaptive training curves for BU/TD objectness, and show that the proposed WSOD2 can achieve state-of-the-art results. Zhaoyang Zeng, Bei Liu 0001, Jianlong Fu, Hongyang Chao, Lei Zhang 0001 |
ICCV | 1 |
| 2019 | SMP Challenge: An Overview of Social Media Prediction Challenge 2019abstract"SMP Challenge" aims to discover novel prediction tasks for numerous data on social multimedia and seek excellent research teams. Making predictions via social multimedia data (e.g. photos, videos or news) is not only helps us to make better strategic decisions for the future, but also explores advanced predictive learning and analytic methods on various problems and scenarios, such as multimedia recommendation, advertising system, fashion analysis etc. Bo Wu 0018, Wen-Huang Cheng, Peiye Liu, Bei Liu 0001, Zhaoyang Zeng, Jiebo Luo 0001 |
ACM Multimedia | 5 |
| 2017 | Searching Personal Photos on the Phone with Instant Visual Query Suggestion and Joint Text-Image HashingabstractThe ubiquitous mobile devices have led to the unprecedented growing of personal photo collections on the phone. One significant pain point of today's mobile users is instantly finding specific photos of what they want. Existing applications (e.g., Google Photo and OneDrive) have predominantly focused on cloud-based solutions, while leaving the client-side challenges (e.g., query formulation, photo tagging and search, etc.) unsolved. This considerably hinders user experience on the phone. In this paper, we present an innovative personal photo search system on the phone, which enables instant and accurate photo search by visual query suggestion and joint text-image hashing. Specifically, the system is characterized by several distinctive properties: 1) visual query suggestion (VQS) to facilitate the formulation of queries in a joint text-image form, 2) light-weight convolutional and sequential deep neural networks to extract representations for both photos and queries, and 3) joint text-image hashing (with compact binary codes) to facilitate binary image search and VQS. It is worth noting that all the components run on the phone with client optimization by deep learning techniques. We have collected 270 photo albums taken by 30 mobile users (corresponding to 37,000 personal photos) and conducted a series of field studies. We show that our system significantly outperforms the existing client-based solutions by 10 x in terms of search efficiency, and 92.3% precision in terms of search accuracy, leading to a remarkably better user experience of photo discovery on the phone. Zhaoyang Zeng, Jianlong Fu, Hongyang Chao, Tao Mei 0001 |
ACM Multimedia | 1 |