VLDB 2026 Research / reviewers in the wild / expert
Yuntao Chen
dblp:203/8284 · also YunTao Chen
· DBLP profile ↗
42ranked-venue papers
5as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 4 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge-Enhanced Graph Contrastive Learning for RecommendationsabstractGraph contrastive learning (GCL), which captures essential features from augmented graphs to address data sparsity issues, has recently demonstrated promising potential in improving recommendation performance. Most GCL-based recommendation methods learn consistent entity representations from user-item bipartite graphs through structural perturbations. However, these approaches impose an additional computational cost and have been shown to be insensitive to various graph augmentations, resulting in limited improvements in long-tail recommendation scenarios. To address this issue, we propose a novel framework for recommendation,Knowledge-Enhanced graphContrastiveLearning (KECL), which adopts knowledge graph-based embedding augmentation instead of graph enhancement to construct views for GCL. Specifically, we introduce a knowledge aggregation module with a heterogeneous attentive aggregator to capture relation heterogeneity in the knowledge graph. Furthermore, we propose a knowledge-based augmentation GCL model that adds knowledge-aware embeddings to the learned representations for more efficient representation-level augmentation. Extensive experiments on real-world datasets demonstrate that the knowledge-based augmentation approach effectively enhances recommendation performance and shows superiority over state-of-the-art methods. Xiaofeng Wang 0004, Zhengjie Zhang, Guodong Shen, Shuaiming Lai, Yuntao Chen, Daying Quan |
IEEE Trans. Multim. | 6 |
| 2026 | Enhanced fabric defect detection via snake-shaped global-local feature shrinking network
Yuntao Chen, Hao Liu 0060, Junjie Zhuang, Jiuzhen Liang |
Vis. Comput. | 1 |
| 2025 | AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsabstractUser interface understanding with vision-language models (VLMs) has received much attention due to its potential for enhancing software automation.However, existing datasets used to build UI-VLMs either only contain large-scale context-free element annotations or contextualized functional descriptions for elements at a small scale.In this work, we propose the AutoGUI pipeline for automatically annotating UI elements with detailed functionality descriptions at scale.Specifically, we leverage large language models (LLMs) to infer element functionality by comparing UI state changes before and after simulated interactions. To improve annotation quality, we propose LLM-aided rejection and verification, eliminating invalid annotations without human labor.We construct a high-quality AutoGUI-704k dataset using the proposed pipeline, featuring diverse and detailed functionality annotations that are hardly provided by previous datasets.Human evaluation shows that we achieve annotation correctness comparable to a trained human annotator. Extensive experiments show that our dataset remarkably enhances VLM’s UI grounding capabilities and exhibits significant scaling effects. We also show the interesting potential use of our dataset in UI agent tasks. Please view our project at https://autogui-project.github.io/. Jingfan Chen, Jingran Su, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
ACL (1) | 4 |
| 2025 | Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State InterventionabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, but they frequently suffer from hallucination -generating content inconsistent with visual inputs.In this work, we explore a novel perspective on hallucination mitigation by examining the intermediate activations of LVLMs during generation.Our investigation reveals that hallucinated content manifests as distinct, identifiable patterns in the model's hidden state space.Motivated by this finding, we propose Activation Steering Decoding (ASD), a training-free approach that mitigates hallucination through targeted intervention in the model's intermediate activations.ASD operates by first identifying directional patterns of hallucination in the activation space using a small calibration set, then employing a contrast decoding mechanism that computes the difference between positive and negative steering predictions.This approach effectively suppresses hallucination patterns while preserving the model's general capabilities.Extensive experiments demonstrate that our method significantly reduces hallucination across multiple benchmarks while maintaining performance on general visual understanding tasks.Notably, our approach requires no model re-training or architectural modifications, making it readily applicable to existing deployed models. Jingran Su, Jingfan Chen, Yuntao Chen, Li Qing, Zhaoxiang Zhang 0005 |
ACL (1) | 4 |
| 2025 | DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersabstractWorld model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks. Yuntao Chen, Yuqi Wang 0001, Zhaoxiang Zhang 0001 |
ICCV | 1 |
| 2025 | Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyabstractWhile recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or continuous actions constrains adaptability to heterogeneous action spaces. We present Dita, a scalable framework that leverages Transformer architectures to directly denoise continuous action sequences through a unified multimodal diffusion process. Departing from prior methods that condition denoising on fused embeddings via shallow networks, Dita employs in-context conditioning -- enabling fine-grained alignment between denoised actions and raw visual tokens from historical observations. This design explicitly models action deltas and environmental nuances. By scaling the diffusion action denoiser alongside the Transformer's scalability, Dita effectively integrates cross-embodiment datasets across diverse camera perspectives, observation scenes, tasks, and action spaces. Such synergy enhances robustness against various variances and facilitates the successful execution of long-horizon tasks. Evaluations across extensive benchmarks demonstrate state-of-the-art or comparative performance in simulation. Notably, Dita achieves robust real-world adaptation to environmental variances and complex long-horizon tasks through 10-shot finetuning, using only third-person camera inputs. The architecture establishes a versatile, lightweight and open-source baseline for generalist robot policy learning. Project Page: https://robodita.github.io. Zhi Hou, Yuwen Xiong, Haonan Duan 0001, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai, Yuntao Chen |
ICCV | 11 |
| 2025 | UIPro: Unleashing Superior Interaction Capability for GUI AgentsabstractBuilding autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing methods have tried developing GUI agents based on the multi-modal comprehension ability of vision-language models (VLMs). However, the limited scenario, insufficient size, and heterogeneous action spaces hinder the progress of building generalist GUI agents. To resolve these issues, this paper proposes \textbf{UIPro}, a novel generalist GUI agent trained with extensive multi-platform and multi-task GUI interaction data, coupled with a unified action space. We first curate a comprehensive dataset encompassing 20.6 million GUI understanding tasks to pre-train UIPro, granting it a strong GUI grounding capability, which is key to downstream GUI agent tasks. Subsequently, we establish a unified action space to harmonize heterogeneous GUI agent task datasets and produce a merged dataset to foster the action prediction ability of UIPro via continued fine-tuning. Experimental results demonstrate UIPro's superior performance across multiple GUI task benchmarks on various platforms, highlighting the effectiveness of our approach. Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2025 | Enhancing End-to-End Autonomous Driving with Latent World ModelabstractIn autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we develop better scene feature representations to fully leverage sensor data in end-to-end driving? Self-supervised learning methods show great success in learning rich feature representations in NLP and computer vision. Inspired by this, we propose a novel self-supervised learning approach using the LAtent World model (LAW) for end-to-end driving. LAW predicts future latent scene features based on current features and ego trajectories. This self-supervised task can be seamlessly integrated into perception-free and perception-based frameworks, improving scene feature learning while optimizing trajectory prediction. LAW achieves state-of-the-art performance across multiple benchmarks, including real-world open-loop benchmark nuScenes, NAVSIM, and simulator-based closed-loop benchmark CARLA. The code will be released. Yingyan Li, Lue Fan, Jiawei He 0002, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001, Tieniu Tan |
ICLR | 5 |
| 2025 | FreeVS: Generative View Synthesis on Free Driving TrajectoryabstractExisting reconstruction-based novel view synthesis methods for driving scenes focus on synthesizing camera views along the recorded trajectory of the ego vehicle.
Their image rendering performance will severely degrade on viewpoints falling out of the recorded trajectory, where camera rays are untrained.
We propose FreeVS, a novel fully generative approach that can synthesize camera views on free new trajectories in real driving scenes.
To control the generation results to be 3D consistent with the real scenes and accurate in viewpoint pose, we propose the pseudo-image representation of view priors to control the generation process.
Viewpoint translation simulation is applied on pseudo-images to simulate camera movement in each direction.
Once trained, FreeVS can be applied to any validation sequences without reconstruction process and synthesis views on novel trajectories.
Moreover, we propose two new challenging benchmarks tailored to driving scenes, which are novel camera synthesis and novel trajectory synthesis, emphasizing the freedom of viewpoints.
Given that no ground truth images are available on novel trajectories, we also propose to evaluate the consistency of images synthesized on novel trajectories with 3D perception models.
Experiments on the Waymo Open Dataset show that FreeVS has a strong image synthesis performance on both the recorded trajectories and novel trajectories.
The code is released. Project page: https://freevs24.github.io/. Qitai Wang, Lue Fan, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
ICLR | 4 |
| 2025 | DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving has substantially progressed by directly predicting future trajectories from raw perception inputs, which bypasses traditional modular pipelines. However, mainstream methods trained via imitation learning suffer from critical safety limitations, as they fail to distinguish between trajectories that appear human-like but are potentially unsafe. Some recent approaches attempt to address this by regressing multiple rule-driven scores but decoupling supervision from policy optimization, resulting in suboptimal performance. To tackle these challenges, we propose DriveDPO, a Safety Direct Preference Optimization Policy Learning framework. First, we distill a unified policy distribution from human imitation similarity and rule-based safety scores for direct policy optimization. Further, we introduce an iterative Direct Preference Optimization stage formulated as trajectory-level preference alignment. Extensive experiments on the NAVSIM benchmark demonstrate that DriveDPO achieves a new state-of-the-art PDMS of 90.0. Furthermore, qualitative results across diverse challenging scenarios highlight DriveDPO’s ability to produce safer and more reliable driving behaviors. Shuyao Shang, Yuntao Chen, Yuqi Wang 0001, Yingyan Li, Zhaoxiang Zhang 0001 |
NeurIPS | 2 |
| 2025 | Fabric defect detection via Explicit De-Background
Yuntao Chen, Hao Liu 0060, Jiuzhen Liang |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Uncertain Object Representation for Image-Based 3D Object PerceptionabstractDue to the ill-posed nature of locating 3D objects based on image inputs, objects detected by camera-based detectors tend to have considerable uncertainty in their localization. Previous works in camera-based 3D detection and tracking represent each detected object as a single certain 3D bounding box, ignoring their localization uncertainty. We propose the uncertain representation of 3D objects to meet the indeterminacy of localizing objects in images. We model the localization uncertainty of objects during the detection process and represent the location of objects as a probability distribution in 3D space. For camera-based 3D detection, we propose to gather and suppress redundant predictions about an object to form its uncertain representation. For camera-based 3D multiple object tracking, we generalize the cross-frame association metric under the uncertain representation of objects for better-tracking objects with uncertain and unstable localization. As a plug-in module for camera 3D detectors, our proposed method brings a +3.5%/+3.2%/+3.7% NDS boost to BEVDet4D/BEVDet4D-Depth/DD3D on nuScenes validation set and a +4.7% NDS boost to BEVDet4D-Depth on nuScenes test set. With enhanced cross-frame association, our tracking method achieves a 48.2% AMOTA performance and reduces the remaining identity-switch cases to only 300 on nuScenes test set. Qitai Wang, Yuntao Chen, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Graph Anomaly Detection via Multiscale Contrastive Self-Supervised Learning From Local to GlobalabstractGraph anomaly detection is a challenging task in graph data mining, aiming to recognize unconventional patterns within a network. Recently, there has been increasing attention on graph anomaly detection based on contrastive learning due to its high adaptability to the sample imbalance problem. However, most existing work typically focuses on the contrast of local views while neglecting global comparison information, leading to suboptimal performance. To address this issue, we introduce a new multiscale contrastive self-supervised learning framework for graph anomaly detection (GADMCLG). Our approach incorporates local-level contrasts involving node–node and node–subgraph contrast, and global-level subgraph–subgraph contrast. The former mines localized abnormal information, while the latter is intended to capture global anomalous patterns. Specifically, our proposed subgraph–subgraph contrast adopts theh-order neighbor subgraph sampling instead of augmented subgraphs through edge perturbation. This sampling strategy ensures a comprehensive observation of the neighborhood surrounding the target node, thereby mitigating the introduction of extraneous noise and providing interpretability for the detected results. Furthermore, we incorporate a subgraph centralization technique to reduce the bias caused by the absolute position of subgraphs in the attribute space, which enhances the model's ability to identify anomalies at different scales. Extensive experimental results on six real-world datasets demonstrate the effectiveness of our method and its superiority compared with state-of-the-art approaches. Xiaofeng Wang 0004, Shuaiming Lai, Shuailei Zhu, Yuntao Chen, Laishui Lv |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingabstractIn autonomous driving, predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to better plan their actions, enhancing safety and efficiency on the road. To this end, we propose Drive-Wm, the first driving world model compatible with existing end-to-end planning models. Through a joint spatial-temporal modeling facilitated by view factorization, our model generates high-fidelity multiview videos in driving scenes. Building on its powerful generation ability, we showcase the potential of applying the world model for safe driving planning for the first time. Particularly, our Drive-Wm enables driving into multiple futures based on distinct driving maneuvers, and determines the optimal trajectory according to the image-based rewards. Evaluation on real-world driving datasets verifies that our method could generate high-quality, consistent, and controllable multiview videos, opening up possibilities for real-world simulations and safe planning. Yuqi Wang 0001, Jiawei He 0002, Lue Fan, Yuntao Chen, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2024 | Continual Forgetting for Pre-Trained Vision ModelsabstractFor privacy and security concerns, the need to erase un-wanted information from pre-trained vision models is becoming evident nowadays. In real-world scenar-ios, erasure requests originate at any time from both users and model owners. These requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify two key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. To address them, we propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we use LoRA modules to fine-tune the FFN layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. GS-LoRA is effective, parameter-efficient, data-efficient, and easy to implement. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that GS-LoRA manages to forget specific classes with minimal impact on other classes. Codes will be released on https://github.com/bjzhb666/GS-LoRA. Hongbo Zhao 0006, Bolin Ni, Junsong Fan, Yuxi Wang 0001, Yuntao Chen, Gaofeng Meng, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2024 | PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic SegmentationabstractComprehensive modeling of the surrounding 3D world is crucial for the success of autonomous driving. However, existing perception tasks like object detection, road structure segmentation, depth & elevation estimation, and open-set object localization each only focus on a small facet of the holistic 3D scene understanding task. This divide-and-conquer strategy simplifies the algorithm development process but comes at the cost of losing an end-to-end unified solution to the problem. In this work, we address this limitation by studying camera-based 3D panoptic segmentation, aiming to achieve a unified occupancy representation for camera-only 3D scene understanding. To achieve this, we introduce a novel method called PanoOcc, which utilizes voxel queries to aggregate spatiotemporal information from multi-frame and multi-view images in a coarse-to-fine scheme, integrating feature learning and scene representation into a unified occupancy representation. We have conducted extensive ablation studies to validate the effectiveness and efficiency of the proposed method. Our approach achieves new state-of-the-art results for camera-based semantic segmentation and panoptic segmentation on the nuScenes dataset. Furthermore, our method can be easily extended to dense occupancy prediction and has demonstrated promising performance on the Occ3D benchmark. The code will be made available at https://github.com/Robertwyq/PanoOcc. Yuqi Wang 0001, Yuntao Chen, Xingyu Liao, Lue Fan, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2024 | Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision ApplicationsabstractWe introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of its predecessor, DCNv3, with two key enhancements: 1. removing softmax normalization in spatial aggregation to enhance its dynamic property and expressive power and 2. optimizing memory access to minimize redundant operations for speedup. These improvements result in a significantly faster convergence compared to DCNv3 and a substantial increase in processing speed, with DCNv4 achieving more than three times the forward speed. DCNv4 demonstrates exceptional performance across various tasks, including image classification, instance and semantic segmentation, and notably, image generation. When integrated into generative models like U-Net in the latent diffusion model, DCNv4 outperforms its baseline, underscoring its possibility to enhance generative models. In practical applications, replacing DCNv3 with DCNv4 in the InternImage model to create FlashInternImage results in up to 80% speed increase and further performance improvement without further modifications. The advancements in speed and efficiency of DCNv4, combined with its robust performance across diverse vision tasks, show its potential as a foundational building block for future vision models. Yuwen Xiong, Yuntao Chen, Feng Wang 0015, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu 0002, Hongsheng Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai |
CVPR | 3 |
| 2024 | OneTrack: Demystifying the Conflict Between Detection and Tracking in End-to-End 3D Trackers
Qitai Wang, Jiawei He 0002, Yuntao Chen, Zhaoxiang Zhang 0001 |
ECCV (7) | 3 |
| 2024 | Monocular Occupancy Prediction for Scalable Indoor Scenes
Hongxiao Yu, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
ECCV (30) | 3 |
| 2024 | CSOT: Cross-scan Object Transfer for Semi-Supervised LiDAR Object Detection
Jinglin Zhan, Tiejun Liu, RenGang Li, Zhaoxiang Zhang 0001, Yuntao Chen |
ECCV (17) | 5 |
| 2024 | OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map ConstructionabstractIn this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving. Hongbo Zhao 0006, Lue Fan, Yuntao Chen, Yuran Yang, Xiaojuan Jin, Gaofeng Meng, Zhaoxiang Zhang 0001 |
NeurIPS | 3 |
| 2024 | DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World ModelabstractDriving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving datasets. We introduce DrivingDojo, the first dataset tailor-made for training interactive world models with complex driving dynamics. Our dataset features video clips with a complete set of driving maneuvers, diverse multi-agent interplay, and rich open-world driving knowledge, laying a stepping stone for future world model development. We further define an action instruction following (AIF) benchmark for world models and demonstrate the superiority of the proposed dataset for generating action-controlled future predictions. Yuqi Wang 0001, Jiawei He 0002, Qitai Wang, Hengchen Dai, Yuntao Chen, Zhaoxiang Zhang 0001 |
NeurIPS | 6 |
| 2024 | Self-supervised graph autoencoder with redundancy reduction for community detection
Xiaofeng Wang 0004, Guodong Shen, Zengjie Zhang, Shuaiming Lai, Shuailei Zhu, Yuntao Chen, Daying Quan |
Neurocomputing | 6 |
| 2024 | Fully Sparse Fusion for 3D Object DetectionabstractCurrently prevalent multi-modal 3D detection methods rely on dense detectors that usually use dense Bird's-Eye-View (BEV) feature maps. However, the cost of such BEV feature maps is quadratic to the detection range, making it not scalable for long-range detection. Recently, LiDAR-only fully sparse architecture has been gaining attention for its high efficiency in long-range perception. In this paper, we study how to develop a multi-modal fully sparse detector. Specifically, our proposed detector integrates the well-studied 2D instance segmentation into the LiDAR side, which is parallel to the 3D instance segmentation part in the LiDAR-only baseline. The proposed instance-based fusion framework maintains full sparsity while overcoming the constraints associated with the LiDAR-only fully sparse detector. Our framework showcases state-of-the-art performance on the widely used nuScenes dataset, Waymo Open Dataset, and the long-range Argoverse 2 dataset. Notably, the inference speed of our proposed method under the long-range perception setting is 2.7× faster than that of other state-of-the-art multimodal 3D detection methods. Yingyan Li, Lue Fan, Yang Liu 0347, Zehao Huang, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | 3D Video Object Detection with Learnable Object-Centric Global OptimizationabstractWe explore long-term temporal visual correspondence-based optimization for 3D video object detection in this work. Visual correspondence refers to one-to-one mappings for pixels across multiple images. Correspondence-based optimization is the cornerstone for 3D scene reconstruction but is less studied in 3D video object detection, because moving objects violate multi-view geometry constraints and are treated as outliers during scene reconstruction. We address this issue by treating objects as first-class citizens during correspondence-based optimization. In this work, we propose BA-Det, an end-to-end optimizable object detector with object-centric temporal correspondence learning and featuremetric object bundle adjustment. Empirically, we verify the effectiveness and efficiency of BA-Det for multiple baseline 3D detectors under various setups. Our BA-Det achieves SOTA performance on the large-scale Waymo Open Dataset (WOD) with only marginal computation cost. Our code is available at https://github.com/jiaweihe1996/BA-Det. Jiawei He 0002, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2023 | FrustumFormer: Adaptive Instance-aware Resampling for Multi-view 3D DetectionabstractThe transformation of features from 2D perspective space to 3D space is essential to multi-view 3D object detection. Recent approaches mainly focus on the design of view transformation, either pixel-wisely lifting perspective view features into 3D space with estimated depth or grid-wisely constructing BEV features via 3D projection, treating all pixels or grids equally. However, choosing what to transform is also important but has rarely been discussed before. The pixels of a moving car are more informative than the pixels of the sky. To fully utilize the information contained in images, the view transformation should be able to adapt to different image regions according to their contents. In this paper, we propose a novel framework named FrustumFormer, which pays more attention to the features in instance regions via adaptive instance-aware resampling. Specifically, the model obtains instance frustums on the bird's eye view by leveraging image view object proposals. An adaptive occupancy mask within the instance frustum is learned to refine the instance location. Moreover, the temporal frustum intersection could further reduce the localization uncertainty of objects. Comprehensive experiments on the nuScenes dataset demonstrate the effectiveness of FrustumFormer, and we achieve a new state-of-the-art performance on the benchmark. Codes and models will be made available at https://github.com/Robertwyq/Frustum. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2023 | BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective SupervisionabstractWe present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and bet-suits modern image backbones. Existing state-of-the-art BEV detectors are often tied to certain depth pretrained backbones like Vo Vn et, hindering the synergy between booming image backbones and BEV detectors. To address this limitation, we prioritize easing the optimization of BEV detectors by introducing perspective view supervision. To this end, we propose a two-stage BEV detector; where proposals from the perspective head are fed into the bird’ s-eye-view head for final predictions. To evaluate the effectiveness of our model, we conduct extensive ablation studies focusing on the form of supervision and the gener-ality of the proposed detector. The proposed method is ver-ified with a wide spectrum of traditional and modern image backbones and achieves new SoTA results on the large-scale nuScenes dataset. The code shall be released soon. Yuntao Chen, Hao Tian 0006, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang 0001, Gao Huang 0001, Hongyang Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai |
CVPR | 2 |
| 2023 | Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object DetectionabstractThis paper aims for high-performance offline LiDAR-based 3D object detection. We first observe that experienced human annotators annotate objects from a track-centric perspective. They first label objects in a track with clear shapes, and then leverage the temporal coherence to infer the annotations of obscure objects. Drawing inspiration from this, we propose a high-performance offline detector in a track-centric perspective instead of the conventional object-centric perspective. Our method features a bidirectional tracking module and a track-centric learning module. Such design allows our detector to infer and refine a complete track once the object is detected at a certain moment. We refer this characteristic to "onCe detecTed, neveR Lost" and name the proposed system CTRL. Extensive experiments demonstrate the remarkable performance of our method, surpassing the human-level annotating accuracy and previous state-of-the-art methods in the highly competitive Waymo Open Dataset leaderboard without model ensemble. The code is available at https://github.com/tusen-ai/SST. Lue Fan, Yuxue Yang, Yiming Mao 0008, Feng Wang 0015, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2023 | SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsabstractComputer end users have spent billions of hours completing daily tasks like tabular data processing and project timeline scheduling. Most of these tasks are repetitive and error-prone, yet most end users lack the skill to automate these burdensome works. With the advent of large language models (LLMs), directing software with natural language user requests become a reachable goal. In this work, we propose a SheetCopilot agent that takes natural language task and control spreadsheet to fulfill the requirements. We propose a set of atomic actions as an abstraction of spreadsheet software functionalities. We further design a state machine-based task planning framework for LLMs to robustly interact with spreadsheets. We curate a representative dataset containing 221 spreadsheet control tasks and establish a fully automated evaluation pipeline for rigorously benchmarking the ability of LLMs in software control tasks. Our SheetCopilot correctly completes 44.3\% of tasks for a single generation, outperforming the strong code generation baseline by a wide margin. Our project page: https://sheetcopilot.github.io/. Jingran Su, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 3 |
| 2023 | Object Affinity Learning: Towards Annotation-Free Instance SegmentationabstractWe address the problem of annotation-free instance segmentation in the wild, aiming to relieve the expensive cost of manual mask annotations. Existing approaches utilize appearance cues, such as color, edge, and texture information, to generate pseudo masks for instance segmentation. However, due to the ambiguity of defining an object by visual appearance alone, these methods fail to distinguish objects from the background under complex scenes. Beyond visual cues, objects are one-piece in space and move together over time, which indicates that geometry cues, such as spatial continuity and motion consistency, are also exploitable for this problem. To directly utilize geometry cues, we propose an affinity-based paradigm for annotation-free instance segmentation. The new paradigm is called object affinity learning, a proxy task of annotation-free instance segmentation, which aims to tell whether two pixels come from the same object by learning feature representation from geometry cues. During inference, the learned object affinity could be further converted into instance segmentation masks by some graph partition algorithms. The proposed object affinity learning achieves much better instance segmentation performance than existing pseudo-mask-based methods on the large-scale Waymo Open Dataset and KITTI dataset. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | GIFS: Neural Implicit Function for General Shape RepresentationabstractRecent development of neural implicit function has shown tremendous success on high-quality 3D shape re-construction. However, most works divide the space into inside and outside of the shape, which limits their repre-senting power to single-layer and watertight shapes. This limitation leads to tedious data processing (converting non-watertight raw data to watertight) as well as the incapability of representing general object shapes in the real world. In this work, we propose a novel method to represent general shapes including non-watertight shapes and shapes with multi-layer surfaces. We introduce General Implicit Function for 3D Shape (GIFS), which models the relationships between every two points instead of the relationships between points and surfaces. Instead of dividing 3D space into predefined inside-outside regions, GIFS encodes whether two points are separated by any surface. Experiments on ShapeNet show that GIFS outperforms previous state-of-the-art methods in terms of reconstruction quality, rendering efficiency, and visual fidelity. Project page is available at https://jianglongye.com/gifs. Jianglong Ye, Yuntao Chen, Naiyan Wang, Xiaolong Wang 0004 |
CVPR | 2 |
| 2022 | Densely Constrained Depth Estimator for Monocular 3D Object Detection
Yingyan Li, Yuntao Chen, Jiawei He 0002, Zhaoxiang Zhang 0001 |
ECCV (9) | 2 |
| 2022 | 4D Unsupervised Object DiscoveryabstractObject discovery is a core task in computer vision. While fast progresses have been made in supervised object detection, its unsupervised counterpart remains largely unexplored. With the growth of data volume, the expensive cost of annotations is the major limitation hindering further study. Therefore, discovering objects without annotations has great significance. However, this task seems impractical on still-image or point cloud alone due to the lack of discriminative information. Previous studies underlook the crucial temporal information and constraints naturally behind multi-modal inputs. In this paper, we propose 4D unsupervised object discovery, jointly discovering objects from 4D data -- 3D point clouds and 2D RGB images with temporal information. We present the first practical approach for this task by proposing a ClusterNet on 3D point clouds, which is jointly iteratively optimized with a 2D localization network. Extensive experiments on the large-scale Waymo Open Dataset suggest that the localization network and ClusterNet achieve competitive performance on both class-agnostic 2D object detection and 3D instance segmentation, bridging the gap between unsupervised methods and full supervised ones. Codes and models will be made available at https://github.com/Robertwyq/LSMOL. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
NeurIPS | 2 |
| 2022 | From Individual to Whole: Reducing Intra-class Variance by Feature Aggregation
Zhaoxiang Zhang 0001, Chuanchen Luo, Haiping Wu, Yuntao Chen, Naiyan Wang, Chunfeng Song |
Int. J. Comput. Vis. | 4 |
| 2021 | Unsupervised Object Detection With LIDAR CluesabstractDespite the importance of unsupervised object detection, to the best of our knowledge, there is no previous work addressing this problem. One main issue, widely known to the community, is that object boundaries derived only from 2D image appearance are ambiguous and unreliable. To address this, we exploit LiDAR clues to aid unsupervised object detection. By exploiting the 3D scene structure, the issue of localization can be considerably mitigated. We further identify another major issue, seldom noticed by the community, that the long-tailed and open-ended (sub-)category distribution should be accommodated. In this paper, we present the first practical method for unsupervised object detection with the aid of LiDAR clues. In our approach, candidate object segments based on 3D point clouds are firstly generated. Then, an iterative segment labeling process is conducted to assign segment labels and to train a segment labeling network, which is based on features from both 2D images and 3D point clouds. The labeling process is carefully designed so as to mitigate the issue of long-tailed and open-ended distribution. The final segment labels are set as pseudo annotations for object detection network training. Extensive experiments on the large-scale Waymo Open dataset suggest that the derived unsupervised object detection method achieves reasonable accuracy compared with that of strong supervision within the LiDAR visible range. Hao Tian 0006, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang 0001, Xizhou Zhu |
CVPR | 2 |
| 2020 | Research on evolutionary model of urban rail transit vulnerability based on computer simulation
Yuntao Chen |
Neural Comput. Appl. | 2 |
| 2019 | Scale-Aware Trident Networks for Object DetectionabstractScale variation is one of the key challenges in object detection. In this work, we first present a controlled experiment to investigate the effect of receptive fields for scale variation in object detection. Based on the findings from the exploration experiments, we propose a novel Trident Network (TridentNet) aiming to generate scale-specific feature maps with a uniform representational power. We construct a parallel multi-branch architecture in which each branch shares the same transformation parameters but with different receptive fields. Then, we adopt a scale-aware training scheme to specialize each branch by sampling object instances of proper scales for training. As a bonus, a fast approximation version of TridentNet could achieve significant improvements without any additional parameters and computational cost compared with the vanilla detector. On the COCO dataset, our TridentNet with ResNet-101 backbone achieves state-of-the-art single-model results of 48.4 mAP. Codes are available at https://git.io/fj5vR. Yanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 2 |
| 2019 | Spectral Feature Transformation for Person Re-IdentificationabstractWith the surge of deep learning techniques, the field of person re-identification has witnessed rapid progress in recent years. Deep learning based methods focus on learning a discriminative feature space where data points are clustered compactly according to their corresponding identities. Most existing methods process data points individually or only involves a fraction of samples while building a similarity structure. They ignore dense informative connections among samples more or less. The lack of holistic observation eventually leads to inferior performance. To relieve the issue, we propose to formulate the whole data batch as a similarity graph. Inspired by spectral clustering, a novel module termed Spectral Feature Transformation is developed to facilitate the optimization of group-wise similarities. It adds no burden to the inference and can be applied to various scenarios. As a natural extension, we further derive a lightweight re-ranking method named Local Blurring Re-ranking which makes the underlying clustering structure around the probe set more compact. Empirical studies on four public benchmarks show the superiority of the proposed method. Code is available at https://github.com/LuckyDC/SFT_REID. Chuanchen Luo, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 2 |
| 2019 | Sequence Level Semantics Aggregation for Video Object DetectionabstractVideo objection detection (VID) has been a rising research direction in recent years. A central issue of VID is the appearance degradation of video frames caused by fast motion. This problem is essentially ill-posed for a single frame. Therefore, aggregating features from other frames becomes a natural choice. Existing methods rely heavily on optical flow or recurrent neural networks for feature aggregation. However, these methods emphasize more on the temporally nearby frames. In this work, we argue that aggregating features in the full-sequence level will lead to more discriminative and robust features for video object detection. To achieve this goal, we devise a novel Sequence Level Semantics Aggregation (SELSA) module. We further demonstrate the close relationship between the proposed method and the classic spectral clustering method, providing a novel view for understanding the VID problem. We test the proposed method on the ImageNet VID and the EPIC KITCHENS dataset and achieve new state-of-the-art results. Our method does not need complicated postprocessing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. Haiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 2 |
| 2019 | Fuzzy Network Based Framework for Software Maintainability PredictionabstractSoftware metrics based maintainability prediction is leading to development of new sophisticated techniques to construct prediction models. This paper proposes a new software maintainability prediction framework, which bases on Fuzzy Network, a novel exploratory modeling technique. The proposed framework utilizes both the metric data collected from software system and the subjective appraisals from experts. An application example of the framework is shown. In comparison to the Standard Fuzzy System based models, Fuzzy Network based models improves the transparency more than 71.3% and the accuracy more than 11.0%. It is confirmed that Fuzzy Network based framework is more appropriate for constructing SMP model. Alexander E. Gegov, Farzad Arabikhan, Yuntao Chen |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 4 |
| 2019 | SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance RecognitionabstractObject detection and instance recognition play a central role in many AI applications like autonomous driving, video surveillance and medical image analysis. However, training object detection models on large scale datasets remains computationally expensive and time consuming. This paper presents an efficient and open source object detection framework called SimpleDet which enables the training of state-of-the-art detection models on consumer grade hardware at large scale. SimpleDet covers a wide range of models including both high-performance and high-speed ones. SimpleDet is well-optimized for both low precision training and distributed training and achieves 70% higher throughput for the Mask R-CNN detector compared with existing frameworks. Codes, examples and documents of SimpleDet can be found at https://github.com/tusimple/simpledet. Yuntao Chen, Chenxia Han, Yanghao Li, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
J. Mach. Learn. Res. | 1 |
| 2018 | DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities TransferabstractWe have witnessed rapid evolution of deep neural network architecture design in the past years. These latest progresses greatly facilitate the developments in various areas such as computer vision and natural language processing. However, along with the extraordinary performance, these state-of-the-art models also bring in expensive computational cost. Directly deploying these models into applications with real-time requirement is still infeasible. Recently, Hinton et al. have shown that the dark knowledge within a powerful teacher model can significantly help the training of a smaller and faster student network. These knowledge are vastly beneficial to improve the generalization ability of the student model. Inspired by their work, we introduce a new type of knowledge---cross sample similarities for model compression and acceleration. This knowledge can be naturally derived from deep metric learning model. To transfer them, we bring the "learning to rank" technique into deep metric learning formulation. We test our proposed DarkRank method on various metric learning tasks including pedestrian re-identification, image retrieval and image clustering. The results are quite encouraging. Our method can improve over the baseline method by a large margin. Moreover, it is fully compatible with other existing methods. When combined, the performance can be further boosted. Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
AAAI | 1 |