EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Lu
dblp:180/5036
· DBLP profile ↗
6ranked-venue papers
3as first author
2since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Robot manipulation · 41% Speech recognition and synthesis · 27% Face, body and person analysis · 14% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › grasping › grasp detection
grasp pose estimation |
0.6 | 1 | 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022 |
Robotics › Robot manipulation › grasping
grasp quality evaluation |
0.6 | 1 | 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022 |
Computer vision › Face, body and person analysis
person re-identification |
0.4 | 1 | 2020 | Deep Credible Metric Learning for Unsupervised Domain Adaptation Person Re-identification · ECCV (8) 2020 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.4 | 1 | 2020 | FPETS: Fully Parallel End-to-End Text-to-Speech System · AAAI 2020 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech |
0.4 | 1 | 2020 | FPETS: Fully Parallel End-to-End Text-to-Speech System · AAAI 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2020 | Deep Credible Metric Learning for Unsupervised Domain Adaptation Person Re-identification · ECCV (8) 2020 |
Robotics › Robot manipulation › grasping › grasp stability
force-closure grasp |
0.2 | 1 | 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detection · ICRA 2022 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | FPETS: Fully Parallel End-to-End Text-to-Speech System · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
multi-resolution network · 0.6joint loss · 0.6u-shape convolutional network · 0.4two-step training · 0.4trainable position encoding · 0.4metric learning · 0.4credible metric learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor ScenesabstractRobotic grasping faces new challenges in human-robot-interaction scenarios. We consider the task that the robot grasps a target object designated by human's language directives. The robot not only needs to locate a target based on vision-and-language information, but also needs to predict the reasonable grasp pose candidate at various views and postures. In this work, we propose a novel interactive grasp policy, named Visual-Lingual-Grasp (VL-Grasp), to grasp the target specified by human language. First, we build a new challenging visual grounding dataset to provide functional training data for robotic interactive perception in indoor environments. Second, we propose a 6- Dof interactive grasp policy combined with visual grounding and 6- Dof grasp pose detection to extend the universality of interactive grasping. Third, we design a grasp pose filter module to enhance the performance of the policy. Experiments demonstrate the effectiveness and extendibility of the VL-Grasp in real world. The VL-Grasp achieves a success rate of 72.5 % in different indoor scenes. The code and dataset is available at https://github.com/luyh20/VL-Grasp. Yuhao Lu, Yixuan Fan, Beixing Deng, Fangfu Liu, Yali Li 0001, Shengjin Wang |
IROS | 1 |
| 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detectionabstract6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric usually generates several discrete levels of grasp confidence scores, which cannot finely distinguish millions of grasp poses and leads to inaccurate prediction results. In this paper, we propose a hybrid physical metric to solve this evaluation insufficiency. First, we define a novel metric is based on the force-closure metric, supplemented by the measurement of the object flatness, gravity and collision. Second, we leverage this hybrid physical metric to generate elaborate confidence scores. Third, to learn the new confidence scores effectively, we design a multi-resolution network called Flatness Gravity Collision GraspNet (FGC-GraspNet). FGC-GraspNet proposes a multi-resolution features learning architecture for multiple tasks and introduces a new joint loss function that enhances the average precision of the grasp detection. The network evaluation and adequate real robot experiments demonstrate the effectiveness of our hybrid physical metric and FGC-GraspNet. Our method achieves 90.5% success rate in real-world cluttered scenes. Our code is available at https://github.com/luyh20IFGC-GraspNet. Yuhao Lu, Beixing Deng, Zhenyu Wang 0005, Peiyuan Zhi, Yali Li 0001, Shengjin Wang |
ICRA | 1 |
| 2020 | FPETS: Fully Parallel End-to-End Text-to-Speech SystemabstractEnd-to-end Text-to-speech (TTS) system can greatly improve the quality of synthesised speech. But it usually suffers form high time latency due to its auto-regressive structure. And the synthesised speech may also suffer from some error modes, e.g. repeated words, mispronunciations, and skipped words. In this paper, we propose a novel non-autoregressive, fully parallel end-to-end TTS system (FPETS). It utilizes a new alignment model and the recently proposed U-shape convolutional structure, UFANS. Different from RNN, UFANS can capture long term information in a fully parallel manner. Trainable position encoding and two-step training strategy are used for learning better alignments. Experimental results show FPETS utilizes the power of parallel computation and reaches a significant speed up of inference compared with state-of-the-art end-to-end TTS systems. More specifically, FPETS is 600X faster than Tacotron2, 50X faster than DCTTS and 10X faster than Deep Voice3. And FPETS can generates audios with equal or better quality and fewer errors comparing with other system. As far as we know, FPETS is the first end-to-end TTS system which is fully parallel. Dabiao Ma, Zhiba Su, Yuhao Lu |
AAAI | 4 |
| 2020 | Deep Credible Metric Learning for Unsupervised Domain Adaptation Person Re-identification
Guangyi Chen 0002, Yuhao Lu, Jiwen Lu, Jie Zhou 0001 |
ECCV (8) | 2 |
| 2019 | UFANS: U-Shaped Fully-Parallel Acoustic Neural Structure for Statistical Parametric Speech Synthesis
Dabiao Ma, Zhiba Su, Yuhao Lu |
PRICAI (3) | 4 |
| 2016 | Maximum saliency bias in binocular fusionabstractSubjective experience at any instant consists of a single (“unitary”), coherent interpretation of sense data rather than a “Bayesian blur” of alternatives. However, computation of Bayes-optimal actions has no role for unitary perception, instead being required to integrate over every possible action-percept pair to maximise expected utility. So what is the role of unitary coherent percepts, and how are they computed? Recent work provided objective evidence for non-Bayes-optimal, unitary coherent, perception and action in humans; and further suggested that the percept selected is not the maximum a posteriori percept but is instead affected by utility. The present study uses a binocular fusion task first to reproduce the same effect in a new domain, and second, to test multiple hypotheses about exactly how utility may affect the percept. After accounting for high experimental noise, it finds that both Bayes optimality (maximise expected utility) and the previously proposed maximum-utility hypothesis are outperformed in fitting the data by a modified maximum-salience hypothesis, using unsigned utility magnitudes in place of signed utilities in the bias function. Yuhao Lu, Tom Stafford 0002, Charles Fox |
Connect. Sci. | 1 |