EDBT 2026 Demo / reviewers in the wild / expert
Xiaofeng Mou
dblp:317/3067
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2024
0009-0009-9480-3667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 55% Image recognition and object detection · 20% Motion planning and robot control · 15% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.4 | 2 | 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model Compression · AAAI 2024 ScaleKD: Distilling Scale-Aware Knowledge in Small Object Detector · CVPR 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model Compression · AAAI 2024 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.8 | 1 | 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model Compression · AAAI 2024 |
Robotics › Motion planning and robot control
robot learning |
0.8 | 1 | 2024 | Retrieval-Augmented Embodied Agents · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
0.8 | 1 | 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model Compression · AAAI 2024 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | ScaleKD: Distilling Scale-Aware Knowledge in Small Object Detector · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.7 | 1 | 2023 | ScaleKD: Distilling Scale-Aware Knowledge in Small Object Detector · CVPR 2023 |
Robotics › Motion planning and robot control › robot learning
manipulation task learning |
0.2 | 1 | 2024 | Retrieval-Augmented Embodied Agents · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
self-distillation · 0.8policy retriever · 0.8policy generator · 0.8multimodal input encoding · 0.8early pruning · 0.8knowledge distillation · 0.7cross-attention · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model CompressionabstractNeural network compression techniques, such as knowledge distillation (KD) and network pruning, have received increasing attention. Recent work `Prune, then Distill' reveals that a pruned student-friendly teacher network can benefit the performance of KD. However, the conventional teacher-student pipeline, which entails cumbersome pre-training of the teacher and complicated compression steps, makes pruning with KD less efficient. In addition to compressing models, recent compression techniques also emphasize the aspect of efficiency. Early pruning demands significantly less computational cost in comparison to the conventional pruning methods as it does not require a large pre-trained model. Likewise, a special case of KD, known as self-distillation (SD), is more efficient since it requires no pre-training or student-teacher pair selection. This inspires us to collaborate early pruning with SD for efficient model compression. In this work, we propose the framework named Early Pruning with Self-Distillation (EPSD), which identifies and preserves distillable weights in early pruning for a given SD task. EPSD efficiently combines early pruning and self-distillation in a two-step process, maintaining the pruned network's trainability for compression. Instead of a simple combination of pruning and SD, EPSD enables the pruned network to favor SD by keeping more distillable weights before training to ensure better distillation of the pruned network. We demonstrated that EPSD improves the training of pruned networks, supported by visual and quantitative analyses. Our evaluation covered diverse benchmarks (CIFAR-10/100, Tiny-ImageNet, full ImageNet, CUB-200-2011, and Pascal VOC), with EPSD outperforming advanced pruning and SD techniques. Dong Chen 0044, Ning Liu 0007, Yichen Zhu 0001, Zhengping Che, Rui Ma 0011, Fachao Zhang, Xiaofeng Mou, Jian Tang 0008 |
AAAI | 7 |
| 2024 | Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language ModelsabstractRecent advancements in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks. Current research enhances the reasoning performance of LLMs by sampling multiple reasoning chains and ensembling based on the answer frequency. However, this approach fails in scenarios where the correct answers are in the minority. We identify this as a primary factor constraining the reasoning capabilities of LLMs, a limitation that cannot be resolved solely based on the predicted answers. To address this shortcoming, we introduce a hierarchical reasoning aggregation framework AoR (Aggregation of Reasoning), which selects answers based on the evaluation of reasoning chains. Additionally, AoR incorporates dynamic sampling, adjusting the number of reasoning chains in accordance with the complexity of the task. Experimental results on a series of complex reasoning tasks show that AoR outperforms prominent ensemble methods. Further analysis reveals that AoR not only adapts various LLMs but also achieves a superior performance ceiling when compared to current methods. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Tianxiang Sun, Qinyuan Cheng, Xiaofeng Mou, Xipeng Qiu, Xuanjing Huang 0001 |
LREC/COLING | 10 |
| 2024 | Retrieval-Augmented Embodied AgentsabstractEmbodied agents operating in complex and uncertain environments face considerable challenges. While some advanced agents handle complex manipulation tasks with proficiency, their success often hinges on extensive training data to develop their capabilities. In contrast, humans typically rely on recalling past experiences and analogous situations to solve new problems. Aiming to emulate this human approach in robotics, we introduce the Retrieval-Augmented Embodied Agent (RAEA). This innovative system equips robots with a form of shared memory, significantly enhancing their performance. Our approach integrates a policy retriever, allowing robots to access relevant strategies from an external policy memory bank based on multi-modal inputs. Additionally, a policy generator is employed to assimilate these strategies into the learning process, enabling robots to formulate effective responses to tasks. Extensive testing of RAEA in both simulated and real-world scenarios demonstrates its superior performance over traditional methods, representing a major leap forward in robotic technology. Yichen Zhu 0001, Zhicai Ou, Xiaofeng Mou, Jian Tang 0008 |
CVPR | 3 |
| 2023 | ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorabstractDespite the prominent success of general object detection, the performance and efficiency of Small Object Detection (SOD) are still unsatisfactory. Unlike existing works that struggle to balance the tradeoff between inference speed and SOD performance, in this paper, we propose a novel Scale-aware Knowledge Distillation (ScaleKD), which transfers knowledge of a complex teacher model to a compact student model. We design two novel modules to boost the quality of knowledge transfer in distillation for SOD: 1) a scale-decoupled feature distillation module that disentangled teacher's feature representation into multi-scale embedding that enables explicit feature mimicking of the student model on small objects. 2) a cross-scale assistant to refine the noisy and uninformative bounding boxes prediction student models, which can mislead the student model and impair the efficacy of knowledge distillation. A multi-scale cross-attention layer is established to capture the multi-scale semantic information to improve the student model. We conduct experiments on COCO and VisDrone datasets with diverse types of models, i.e., two-stage and one-stage detectors, to evaluate our proposed method. Our ScaleKD achieves superior performance on general detection performance and obtains spectacular improvement regarding the SOD performance. Yichen Zhu 0001, Qiqi Zhou 0002, Ning Liu 0007, Zhicai Ou, Xiaofeng Mou, Jian Tang 0008 |
CVPR | 6 |
| 2022 | Few Clean Instances Help Denoising Distant SupervisionabstractExisting distantly supervised relation extractors usually rely on noisy data for both model training and evaluation, which may lead to garbage-in-garbage-out systems. To alleviate the problem, we study whether a small clean dataset could help improve the quality of distantly supervised models. We show that besides getting a more convincing evaluation of models, a small clean dataset also helps us to build more robust denoising models. Specifically, we propose a new criterion for clean instance selection based on influence functions. It collects sample-level evidence for recognizing good instances (which is more informative than loss-level evidence). We also propose a teacher-student mechanism for controlling purity of intermediate results when bootstrapping the clean set. The whole approach is model-agnostic and demonstrates strong performances on both denoising real (NYT) and synthetic noisy datasets. Yufang Liu, Ziyin Huang, Changzhi Sun, Man Lan, Yuanbin Wu, Xiaofeng Mou |
COLING | 7 |