Qichen Fu

dblp:304/2909 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0008-5976-4095ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 31% Video understanding and tracking · 24% Face, body and person analysis · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
egocentric video understanding
1.422025
Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022
Computer vision › Video understanding and tracking
activity prediction
0.912025
Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention · EMNLP 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
eDKM: An Efficient and Accurate Train-Time Weight Clustering for Large Language Models · HPCA 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
eDKM: An Efficient and Accurate Train-Time Weight Clustering for Large Language Models · HPCA 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression › quantization
weight clustering
0.912025
eDKM: An Efficient and Accurate Train-Time Weight Clustering for Large Language Models · HPCA 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.812024
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation · ICML 2024
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.812024
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation · ICML 2024
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.712023
Deformer: Dynamic Fusion Transformer for Robust Hand Pose Estimation · ICCV 2023
Computer vision › Image recognition and object detection › object detection
active object detection
0.612022
Sequential Voting with Relational Box Fields for Active Object Detection · CVPR 2022
Computer vision › Video understanding and tracking
activity understanding
0.612022
Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.612022
Domain Adaptive Hand Keypoint and Pixel Localization in the Wild · ECCV (9) 2022
Computer vision › Video understanding and tracking › egocentric video understanding
first-person activity recognition
0.612022
Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022
Computer vision › Face, body and person analysis › hand analysis
hand keypoint detection
0.612022
Domain Adaptive Hand Keypoint and Pixel Localization in the Wild · ECCV (9) 2022
Computer vision › Face, body and person analysis
human pose estimation
0.612022
Domain Adaptive Hand Keypoint and Pixel Localization in the Wild · ECCV (9) 2022
Computer vision › Image recognition and object detection
object detection
0.612022
Sequential Voting with Relational Box Fields for Active Object Detection · CVPR 2022
Computer vision › 3D vision
3d scene reconstruction
0.312025
Ego4D: Around the World in 3,600 Hours of Egocentric Video · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Information retrieval › ranking › search relevance › relevance modeling
contextual relevance
0.212024
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation · ICML 2024
Information retrieval
retrieval-augmented generation
0.212024
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation · ICML 2024
Computer vision › 3D vision › 3d human reconstruction
hand mesh reconstruction
0.212023
Deformer: Dynamic Fusion Transformer for Robust Hand Pose Estimation · ICCV 2023
Multimedia analysis and retrieval › video dataset
video dataset benchmark
0.212022
Ego4D: Around the World in 3, 000 Hours of Egocentric Video · CVPR 2022

Methods — techniques the papers use, named apart from their topics

superposition prompting · 1.5parallel prompt paths · 1.5uniquification · 0.9speculative decoding · 0.9sharding · 0.9multi-stream attention · 0.9differentiable k-means clustering · 0.9self-attention · 0.7maxMSE loss · 0.7dynamic fusion · 0.7
YearPublicationVenuePosition
2025 Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention
abstract
Nikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason, Antonie Lin, Mohammad Rastegari, Mahyar Najibi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Nikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason, Antonie Lin, Mohammad Rastegari, Mahyar Najibi
EMNLP3
2025 eDKM: An Efficient and Accurate Train-Time Weight Clustering for Large Language Models
abstract
Since Large Language Models or LLMs have demonstrated high-quality performance on many complex language tasks, there is a great interest in bringing these LLMs to mobile devices for faster responses and better privacy protection. However, the size of LLMs (i.e., billions of parameters) requires highly effective compression to fit into storage-limited devices. Among many compression techniques, weight-clustering, a form of non-linear quantization, is one of the leading candidates for LLM compression, and supported by modern smartphones. Yet, its training overhead is prohibitively significant for LLM fine-tuning. Especially, Differentiable KMeans Clustering, or DKM, has shown the state-of-the-art trade-off between compression ratio and accuracy regression, but its large memory complexity makes it nearly impossible to apply to train-time LLM compression. In this letter, we propose a memory-efficient DKM implementation, eDKM powered by novel techniques to reduce the memory footprint of DKM by orders of magnitudes. For a given tensor to be saved on CPU for the backward pass of DKM, we compressed the tensor by applying uniquification and sharding after checking if there is no duplicated tensor previously copied to CPU. Our experimental results demonstrate that eDKM can fine-tune and compress a pretrained LLaMA 7B model from 12.6 GB to $2.5 \mathrm{~GB}(3 \mathrm{~b} /$ weight) with the Alpaca dataset by reducing the train-time memory footprint of a decoder layer by $130 \times$, while delivering good accuracy on broader LLM benchmarks (i.e., $77.7 \%$ for PIQA, $66.1 \%$ for Winograde, and so on).
Minsik Cho, Keivan Alizadeh-Vahid, Qichen Fu, Saurabh Adya, Carlo C. del Mundo, Mohammad Rastegari, Devang Naik, Peter Zatloukal
HPCA3
2025 Ego4D: Around the World in 3,600 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception.
Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.31
2024 Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
abstract
Despite the successes of large language models (LLMs), they exhibit significant drawbacks, particularly when processing long contexts. Their inference cost scales quadratically with respect to sequence length, making it expensive for deployment in some real-world text processing applications, such as retrieval-augmented generation (RAG). Additionally, LLMs also exhibit the "distraction phenomenon", where irrelevant context in the prompt degrades output quality. To address these drawbacks, we propose a novel RAG prompting methodology, superposition prompting, which can be directly applied to pre-trained transformer-based LLMs without the need for fine-tuning. At a high level, superposition prompting allows the LLM to process input documents in parallel prompt paths, discarding paths once they are deemed irrelevant. We demonstrate the capability of our method to simultaneously enhance time efficiency across a variety of question-answering benchmarks using multiple pre-trained LLMs. Furthermore, our technique significantly improves accuracy when the retrieved context is large relative the context the model was trained on. For example, our approach facilitates a $93\times$ reduction in compute time while improving accuracy by $43%$ on the NaturalQuestions-Open dataset with the MPT-7B instruction-tuned model over naive RAG.
Thomas Merth, Qichen Fu, Mohammad Rastegari, Mahyar Najibi
ICML2
2024 FastSR-NeRF: Improving NeRF Efficiency on Consumer Devices with A Simple Super-Resolution Pipeline
abstract
Super-resolution (SR) techniques have recently been proposed to upscale the outputs of neural radiance fields (NeRF) and generate high-quality images with enhanced inference speeds. However, existing NeRF+SR methods increase training overhead by using extra input features, loss functions, and/or expensive training procedures such as knowledge distillation. In this paper, we aim to leverage SR for efficiency gains without costly training or architectural changes. Specifically, we build a simple NeRF+SR pipeline that directly combines existing modules, and we propose a lightweight augmentation technique, random patch sampling, for training. Compared to existing NeRF+SR methods, our pipeline mitigates the SR computing overhead and can be trained up to 23× faster, making it feasible to run on consumer devices such as the Apple MacBook. Experiments show our pipeline can upscale NeRF outputs by 2-4× while maintaining high quality, increasing inference speeds by up to 18× on an NVIDIA V100 GPU and 12.8× on an M1 Pro chip. We conclude that SR can be a simple but effective technique for improving the efficiency of NeRF models for consumer devices.
Chien-Yu Lin, Qichen Fu, Thomas Merth, Karren D. Yang, Anurag Ranjan
WACV2
2023 Deformer: Dynamic Fusion Transformer for Robust Hand Pose Estimation
abstract
Accurately estimating 3D hand pose is crucial for understanding how humans interact with the world. Despite remarkable progress, existing methods often struggle to generate plausible hand poses when the hand is heavily occluded or blurred. In videos, the movements of the hand allow us to observe various parts of the hand that may be occluded or blurred in a single frame. To adaptively leverage the visual clue before and after the occlusion or blurring for robust hand pose estimation, we propose the Deformer: a framework that implicitly reasons about the relationship between hand parts within the same image (spatial dimension) and different timesteps (temporal dimension). We show that a naive application of the transformer self-attention mechanism is not sufficient because motion blur or occlusions in certain frames can lead to heavily distorted hand features and generate imprecise keys and queries. To address this challenge, we incorporate a Dynamic Fusion Module into Deformer, which predicts the deformation of the hand and warps the hand mesh predictions from nearby frames to explicitly support the current frame estimation. Furthermore, we have observed that errors are unevenly distributed across different hand parts, with vertices around fingertips having disproportionately higher errors than those around the palm. We mitigate this issue by introducing a new loss function called maxMSE that automatically adjusts the weight of every vertex to focus the model on critical hand parts. Extensive experiments show that our method significantly outperforms state-of-the-art methods by 10%, and is more robust to occlusions (over 14%).
Qichen Fu, Xingyu Liu 0001, Ran Xu 0001, Juan Carlos Niebles, Kris Makoto Kitani
ICCV1
2022 Sequential Voting with Relational Box Fields for Active Object Detection
abstract
A key component of understanding hand-object interactions is the ability to identify the active object-the object that is being manipulated by the human hand. In order to accurately localize the active object, any method must reason using information encoded by each image pixel, such as whether it belongs to the hand, the object, or the background. To leverage each pixel as evidence to determine the bounding box of the active object, we propose a pixel-wise voting function. Our pixel-wise voting function takes an initial bounding box as input and produces an improved bounding box of the active object as output. The voting function is designed so that each pixel inside of the input bounding box votes for an improved bounding box, and the box with the majority vote is selected as the output. We call the collection of bounding boxes generated inside of the voting function, the Relational Box Field, as it characterizes a field of bounding boxes defined in relationship to the current bounding box. While our voting function is able to improve the bounding box of the active object, one round of voting is typically not enough to accurately localize the active object. Therefore, we repeatedly apply the voting function to sequentially improve the location of the bounding box. However, since it is known that repeatedly applying a one-step predictor (i.e., auto-regressive processing with our voting function) can cause a data distribution shift, we mitigate this issue using reinforcement learning (RL). We adopt standard RL to learn the voting function parameters and show that it provides a meaningful improvement over a standard supervised learning approach. We perform experiments on two large-scale datasets: 100DOH and MECCANO, improving AP50 performance by 8% and 30%, respectively, over the state of the art. The project page with code and visualizations can be found at https://fuqichen1998.github.io/SequentialVotingDet/.
Qichen Fu, Xingyu Liu 0001, Kris Makoto Kitani
CVPR1
2022 Ego4D: Around the World in 3, 000 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
CVPR30
2022 Domain Adaptive Hand Keypoint and Pixel Localization in the Wild
Takehiko Ohkawa, Yu-Jhe Li, Qichen Fu, Ryosuke Furuta, Kris Makoto Kitani, Yoichi Sato 0001
ECCV (9)3