Three-dimensional environmental understanding is fundamental to computer vision and robotics, enabling applications from autonomous navigation to robotic manipulation. Monocular depth estimation, inferring depth from a single RGB image, has advanced with deep learning. While self-supervised methods reduce reliance on ground-truth data, supervised techniques currently yield superior performance. This thesis introduces novel contributions to enhance self-supervised monocular depth estimation’s robustness, accuracy, and applicability,narrowing the gap with supervised methods.First, we address dynamic objects, which violate the static world assumption in self supervised models. We propose Dyna-DM, a dynamic object-aware approach that identifies and manages moving objects by estimating poses only for dynamic elements, improving depth quality. Second, depth estimation model performance often degrades in adverse conditions absent in standard datasets. We introduce Robust-Depth, a framework using physics-based and generative data augmentations to simulate diverse weather and corruptions. By introducing a pseudo-supervision loss that exploits correspondences between unaugmented and augmented data, our method significantly improves model generalisation. Third, while larger baselines between frames can improve depth accuracy, particularly for stereo correspondence, they are rarely exploited in self-supervised monocular depth estimation. Base Boost Depth uses a curriculum-based strategy to incorporate wider frame separations, enhancing geometric cues and resulting in more precise depth and sharper boundaries.Finally, extending 3D scene understanding to robotic interaction, we introduce GRASP3R, a novel end-to-end method for 7-DoF grasp estimation for parallel grippers from unposed, uncalibrated RGB images. GRASP3R integrates dense 3D scene reconstruction, supervised by CAD representations, with grasp estimation. This eliminates the need for depth sensors at test time, avoiding common issues with transparent or reflective surfaces.Collectively, these contributions advance self-supervised monocular depth estimation, making it more robust to dynamic scenes and adverse conditions while improving accuracy.These techniques extend to real-world robotics, as shown by accurate, versatile grasp estimation from RGB inputs only.
- monocular depth estimation
- self-supervised learning
- robustness
- dynamic scenes
- data augmentation
- robotic grasping
- 3D reconstruction
- pose estimation
- edge definition
- computer vision
Self-Supervised Monocular Depth Estimation and Beyond
Saunders, K. R. (Author). Sept 2025
Student thesis: Doctoral Thesis › Doctor of Philosophy