VLDB 2026 Research / reviewers in the wild / expert
Yecheng Wu
dblp:136/7016
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 64% Vision and language · 13% Robot manipulation · 6% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
autoregressive model |
1.7 | 2 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 HART: Efficient Visual Generation with Hybrid Autoregressive Transformer · ICLR 2025 |
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
0.9 | 1 | 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer · ICCV 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive Transformer · ICLR 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
efficient image generation |
0.9 | 1 | 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer · ICCV 2025 |
Machine learning › Generative modeling
image generation |
0.9 | 1 | 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive Transformer · ICLR 2025 |
Machine learning › Generative modeling
image tokenization |
0.9 | 1 | 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer · ICCV 2025 |
Machine learning › Generative modeling › autoregressive model
masked autoregressive generation |
0.9 | 1 | 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer · ICCV 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer · ICCV 2025 |
Machine learning › Generative modeling › image generation
token-based image generation |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
unified image understanding and generation |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Computer vision › Vision and language › vision-language model › vision-language model architecture
unified vision-language model |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models · CVPR 2025 |
Machine learning › Generative modeling
visual generation |
0.9 | 1 | 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive Transformer · ICLR 2025 |
Image and video coding
image tokenization |
0.9 | 1 | 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive Transformer · ICLR 2025 |
Robotics › Motion planning and robot control › motion planning
manipulation planning |
0.3 | 1 | 2025 | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.7autoregressive transformer · 1.7autoencoder · 1.7vision-language model · 0.9vector quantization · 0.9masked autoregressive modeling · 0.9discrete visual tokenization · 0.9autoregressive next-token prediction · 0.9autoregressive image prediction · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsabstractVision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input–output mappings, lacking the intermediate reasoning steps crucial for complex manipulation tasks. As a result, existing VLAs lack temporal planning or reasoning capabilities. In this paper, we introduce a method that incorporates explicit visual chain-of-thought (CoT) reasoning into vision-language-action models (VLAs) by predicting future image frames autoregressively as visual goals before generating a short action sequence to achieve these goals. We introduce CoT-VLA, a state-of-the-art 7B VLA that can understand and generate visual and action tokens. Our experimental results demonstrate that CoT-VLA achieves strong performance, outperforming the state-of-the-art VLA model by 17% in real-world manipulation tasks and 6% in simulation benchmarks. Videos are available at: https://cot-vla.github.io/. Yao Lu 0006, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Song Han 0003, Chelsea Finn, Ankur Handa, Tsung-Yi Lin, Gordon Wetzstein, Ming-Yu Liu 0001, Donglai Xiang |
CVPR | 6 |
| 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid TokenizerabstractWe introduce DC-AR, a novel masked autoregressive (AR) text-to-image generation framework that delivers superior image generation quality with exceptional computational efficiency. Due to the tokenizers' limitations, prior masked AR models have lagged behind diffusion models in terms of quality or efficiency. We overcome this limitation by introducing DC-HT - a deep compression hybrid tokenizer for AR models that achieves a 32x spatial compression ratio while maintaining high reconstruction fidelity and cross-resolution generalization ability. Building upon DC-HT, we extend MaskGIT and create a new hybrid masked autoregressive image generation framework that first produces the structural elements through discrete tokens and then applies refinements via residual tokens. DC-AR achieves state-of-the-art results with a gFID of 5.49 on MJHQ-30K and an overall score of 0.69 on GenEval, while offering 1.5-7.9x higher throughput and 2.0-3.5x lower latency compared to prior leading diffusion and autoregressive models. Yecheng Wu, Junyu Chen 0003, Zhuoyang Zhang, Enze Xie, Junsong Chen, Jinyi Hu, Yao Lu 0006, Song Han 0003, Han Cai |
ICCV | 1 |
| 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive TransformerabstractWe introduce Hybrid Autoregressive Transformer (HART), the first autoregressive (AR) visual generation model capable of directly generating 1024x1024 images, rivaling diffusion models in image generation quality. Existing AR models face limitations due to the poor image reconstruction quality of their discrete tokenizers and the prohibitive training costs associated with generating 1024px images. To address these challenges, we present the hybrid tokenizer, which decomposes the continuous latents from the autoencoder into two components: discrete tokens representing the big picture and continuous tokens representing the residual components that cannot be represented by the discrete tokens. The discrete component is modeled by a scalable-resolution discrete AR model, while the continuous component is learned with a lightweight residual diffusion module with only 37M parameters. Compared with the discrete-only VAR tokenizer, our hybrid approach improves reconstruction FID from 2.11 to 0.30 on MJHQ-30K, leading to a 31% generation FID improvement from 7.85 to 5.38. HART also outperforms state-of-the-art diffusion models in both FID and CLIP score, with 4.5-7.7$\times$ higher throughput and 6.9-13.4$\times$ lower MACs. Our code is open sourced at https://github.com/mit-han-lab/hart. Haotian Tang, Yecheng Wu, Shang Yang, Enze Xie, Junsong Chen, Junyu Chen 0003, Zhuoyang Zhang, Han Cai, Yao Lu 0006, Song Han 0003 |
ICLR | 2 |
| 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationabstractVILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework. Yecheng Wu, Zhuoyang Zhang, Junyu Chen 0003, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi 0001, Song Han 0003, Yao Lu 0006 |
ICLR | 1 |
| 2024 | A least squares-support vector machine for learning solution to multi-physical transient-state field coupled problems
Yecheng Wu, Zhengwei Qu, Guofeng Li |
Eng. Appl. Artif. Intell. | 3 |
| 2016 | Semigradient-Based Cooperative Caching Algorithm for Mobile Social NetworksabstractWireless caching at users' devices in mobile social network is considered to be a promising solution to alleviate backhaul overload in future wireless networks. However, most of the current works propose caching schemes based on heuristic reasoning and intuition with poor performance or high complexity which are impractical due to individual devices' computing capacity restriction. In this paper, we design a cooperative caching scheme aimed at maximizing hit ratio, incorporating probabilistic modeling of mobility and user interests patterns from mobile social networks. Furthermore, we reformulate this optimization problem into a submodular function maximization and propose a semigradient-based cooperative caching scheme, while this scheme's efficiency is shown to significantly outperform the greedy caching by 99.6%. Yecheng Wu, Sha Yao, Yang Yang 0001, Zeming Hu, Cheng-Xiang Wang 0001 |
GLOBECOM | 1 |
| 1997 | Inversion of the Li-Strahler canopy reflectance model for mapping forest structureabstractAs part of a larger forest vegetation mapping process based on Landsat TM and digital terrain data, inversion of the Li-Strahler model provides estimates of tree size and cover for conifer stands. The vegetation maps are intended for use in natural resource management by the US Forest Service. Analysis of extensive field data in the form of "test" stands from four National Forests indicate the following about the Li-Strahler model: (1) the underlying assumptions of independence between tree size and crown shape are valid, (2) the means for tree geometry parameters vary between forest types, (3) estimates of forest cover are reliable, and (4) estimates of tree size are unreliable due to the breakdown in the relationship between image intra-stand variance and tree size. Improvements in estimates of tree size will require additional data beyond a single Landsat TM image, with multidirectional data a promising possibility. Curtis E. Woodcock, John B. Collins, Vida D. Jakabhazy, Xiaowen Li 0001, Scott A. Macomber, Yecheng Wu |
IEEE Trans. Geosci. Remote. Sens. | 6 |