Пространственно-временное слияние глубинных данных для оценки трёхмерных расстояний
Работая с сайтом, я даю свое согласие на использование файлов cookie. Это необходимо для нормального функционирования сайта, показа целевой рекламы и анализа трафика. Статистика использования сайта обрабатывается системой Яндекс.Метрика
Научный журнал Моделирование, оптимизация и информационные технологииThe scientific journal Modeling, Optimization and Information Technology
Online media
issn 2310-6018

Spatial-temporal fusion of depth data for estimating three-dimensional distances

Yu X. 

UDC 004.932.2:004.89
DOI: 10.26102/2310-6018/2026.59.8.013

  • Abstract
  • List of references
  • About authors

Depth enhancement is commonly evaluated with pixel-wise errors, although robotic vision and three-dimensional metrology ultimately use geometric distances. This study develops and tests a protocol that evaluates enhancement methods by the error of Euclidean distances between back-projected point pairs. The primary experiments use the TUM Freiburg1 desk and xyz sequences. Additional validation includes the dynamic Freiburg3 walking_static sequence, controlled stress tests with moving depth boundaries, long range and elevated noise, three pseudo-reference variants, and a computational-cost benchmark. With temporal half-window N = 10, temporal mean fusion reduces distance RMSE by 24.07 % on desk and 23.66 % on xyz relative to raw depth. Direct metric comparison shows that temporal median has the lowest pixel MAE, whereas temporal mean has the lowest distance RMSE; Spearman rank correlations are 0.70 and 0.90. Temporal mean remains ranked first under all three pseudo-reference constructions. On the dynamic sequence, a temporal-stability gate retains 18.92 % of pixels and changes the method ranking, demonstrating that moving regions must be excluded or motion-compensated. For 640 × 480 frames at N = 10, the research CPU implementation requires 56.2 ms for temporal mean, 217.5 ms for temporal median, and 3155.6 ms for spatial-temporal median. The results delimit the practical use of the protocol and the systematic limitations of pseudo-reference evaluation.

1. Khoshelham K., Oude Elberink S. Accuracy and resolution of Kinect depth data for indoor mapping applications. Sensors. 2012;12(2):1437–1454. https://doi.org/10.3390/s120201437

2. Nguyen Ch.V., Izadi Sh., Lovell D. Modeling Kinect sensor noise for improved 3D reconstruction and tracking. In: 2012 Second International Conference on 3D Imaging, Modeling, Processing, Visualization and Transmission, 13–15 October 2012, Zurich, Switzerland. IEEE; 2012. P. 524–530. https://doi.org/10.1109/3DIMPVT.2012.84

3. Sarbolandi H., Lefloch D., Kolb A. Kinect range sensing: structured-light versus time-of-flight Kinect. Computer Vision and Image Understanding. 2015;139:1–20. https://doi.org/10.1016/j.cviu.2015.05.006

4. Tomasi C., Manduchi R. Bilateral filtering for gray and color images. In: Sixth International Conference on Computer Vision, 04–07 January 1998, Bombay, India. IEEE; 1998. P. 839–846. https://doi.org/10.1109/ICCV.1998.710815

5. He K., Sun J., Tang X. Guided image filtering. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2013;35(6):1397–1409. https://doi.org/10.1109/TPAMI.2012.213

6. Lin B.-Sh., Su M.-J., Cheng P.-H., et al. Temporal and spatial denoising of depth maps. Sensors. 2015;15(8):18506–18525. https://doi.org/10.3390/s150818506

7. Or-El R., Rosman G., Wetzler A., et al. RGBD-fusion: real-time high precision depth recovery. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 07–12 June 2015, Boston, MA, USA. IEEE; 2015. P. 5407–5416. https://doi.org/10.1109/CVPR.2015.7299179

8. Ronneberger O., Fischer Ph., Brox Th. U-Net: convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015: 18th International Conference: Part III, 05–09 October 2015, Munich, Germany. Cham: Springer; 2015. P. 234–241. https://doi.org/10.1007/978-3-319-24574-4_28

9. Sterzentsenko V., Saroglou L., Chatzitofis A., et al. Self-supervised deep depth denoising. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 27 October – 02 November 2019, Seoul, South Korea. IEEE; 2019. P. 1242–1251. https://doi.org/10.1109/ICCV.2019.00133

10. Sturm J., Engelhard N., Endres F., et al. A benchmark for the evaluation of RGB-D SLAM systems. In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 07–12 October 2012, Vilamoura-Algarve, Portugal. IEEE; 2012. P. 573–580. https://doi.org/10.1109/IROS.2012.6385773

11. Shabanov A., Krotov I., Chinaev N., et al. Self-supervised depth denoising using lower- and higher-quality RGB-D sensors. In: 2020 International Conference on 3D Vision (3DV), 25–28 November 2020, Fukuoka, Japan. IEEE; 2020. P. 743–752. https://doi.org/10.1109/3DV50981.2020.00084

12. Shivakumar Sh.S., Nguyen T., Miller I.D., et al. DFuseNet: deep fusion of RGB and sparse depth information for image guided dense depth completion. In: 2019 IEEE Intelligent Transportation Systems Conference (ITSC), 27–30 October 2019, Auckland, New Zealand. IEEE; 2019. P. 13–20. https://doi.org/10.1109/ITSC.2019.8917294

13. Pyavchenko A.O., Ilchenko A.V. Method of the static objects spatial localization according to depth sensor data and RGB camera. Izvestiya SFedU. Engineering Sciences. 2018;(1):271–284. (In Russ.). https://doi.org/10.23683/2311-3103-2018-1-271-284

14. Vokhmintcev A.V., Pachganov S.A. The algorithm of simultaneous navigation and mapping for mobile robot based on the iterative algorithm of the nearest pixels and descriptor calculated in a circular moving window. Yugra State University Bulletin. 2018;14(3):49–56. (In Russ.). https://doi.org/10.17816/byusu2018049-56

15. Othman W., Gromov V.S. Research of visual simultaneous localization and mapping-based navigation system for mobile robots. Scientific and Technical Journal of Information Technologies, Mechanics and Optics. 2020;20(3):371–376. (In Russ.). https://doi.org/10.17586/2226-1494-2020-20-3-371-376

16. Wu H., Fu K., Zhao Y., et al. Joint self-supervised and reference-guided learning for depth inpainting. Computational Visual Media. 2022;8(4):597–612. https://doi.org/10.1007/s41095-021-0259-z

17. Yang Q., Cui W., Zheng Zh., et al. Unsupervised depth completion based on RGB image and sparse depth map. In: EITCE 2024: Proceedings of the 2024 8th International Conference on Electronic Information Technology and Computer Engineering, 18–20 October 2024, Haikou, China. New York: ACM; 2024. P. 152–157. https://doi.org/10.1145/3711129.3711157

18. Zhang Y., Guo X., Poggi M., et al. CompletionFormer: depth completion with convolutions and vision transformers. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17–24 June 2023, Vancouver, BC, Canada. IEEE; 2023. P. 18527–18536. https://doi.org/10.1109/CVPR52729.2023.01777

19. Wang Y., Li B., Zhang G., et al. LRRU: long-short range recurrent updating networks for depth completion. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 01–06 October 2023, Paris, France. IEEE; 2023. P. 9388–9398. https://doi.org/10.1109/ICCV51070.2023.00864

20. Tang J., Tian F.-P., An B., et al. Bilateral propagation network for depth completion. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 9763–9772. https://doi.org/10.1109/CVPR52733.2024.00932

21. Park J., Li Y.-J., Kitani K. Flexible depth completion for sparse and varying point densities. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 21540–21550. https://doi.org/10.1109/CVPR52733.2024.02035

22. Yan Zh., Lin Y., Wang K., et al. Tri-perspective view decomposition for geometry-aware depth completion. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 4874–4884. https://doi.org/10.1109/CVPR52733.2024.00466

23. Xiang J., Zhu X., Wang X., et al. DEPTHOR: depth enhancement from a practical light-weight dToF sensor and RGB image. In: 2025 IEEE/CVF International Conference on Computer Vision (ICCV), 19–25 October 2025, Honolulu, HI, USA. IEEE; 2025. P. 6101–6111. https://doi.org/10.1109/ICCV51701.2025.00576

24. Kim H., Wang R., Yao Ch., et al. Dense metric depth completion from sparse direct time-of-flight sensors. In: 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 03–07 June 2026, Denver, CO, USA. IEEE; 2026. P. 36518–36528.

25. Li Y., Lou L., Tang Y., et al. LiteSense: lifting lightweight ToF with RGB for high-resolution metric depth estimation. In: 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 03–07 June 2026, Denver, CO, USA. IEEE; 2026. P. 5783–5792.

Yu Xiaohan

Email: yxhpro@yandex.ru

Peoples' Friendship University of Russia named after Patrice Lumumba
Moscow

Moscow, Russian Federation

Keywords: depth data, three-dimensional measurement, depth map, pseudo-reference, spatial-temporal fusion, geometric accuracy, depth enhancement

For citation: Yu X. Spatial-temporal fusion of depth data for estimating three-dimensional distances. Modeling, Optimization and Information Technology. 2026;14(8). URL: https://moitvivt.ru/ru/journal/article?id=2414 DOI: 10.26102/2310-6018/2026.59.8.013 .

© Yu X. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)
37

Full text in PDF

Скачать JATS XML

Received 10.05.2026

Revised 27.07.2026

Accepted 17.08.2026

Published 31.08.2026