<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" dtd-version="1.3" xml:lang="ru" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="https://metafora.rcsi.science/xsd_files/journal3.xsd">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">moitvivt</journal-id>
      <journal-title-group>
        <journal-title xml:lang="ru">Моделирование, оптимизация и информационные технологии</journal-title>
        <trans-title-group xml:lang="en">
          <trans-title>Modeling, Optimization and Information Technology</trans-title>
        </trans-title-group>
      </journal-title-group>
      <issn pub-type="epub">2310-6018</issn>
      <publisher>
        <publisher-name>Издательство</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.26102/2310-6018/2026.59.8.013</article-id>
      <article-id pub-id-type="custom" custom-type="elpub">2414</article-id>
      <title-group>
        <article-title xml:lang="ru">Пространственно-временное слияние глубинных данных для оценки трёхмерных расстояний</article-title>
        <trans-title-group xml:lang="en">
          <trans-title>Spatial-temporal fusion of depth data for estimating three-dimensional distances</trans-title>
        </trans-title-group>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name-alternatives>
            <name name-style="eastern" xml:lang="ru">
              <surname>Юй</surname>
              <given-names>Сяохань</given-names>
            </name>
            <name name-style="western" xml:lang="en">
              <surname>Yu</surname>
              <given-names>Xiaohan</given-names>
            </name>
          </name-alternatives>
          <email>yxhpro@yandex.ru</email>
          <xref ref-type="aff">aff-1</xref>
        </contrib>
      </contrib-group>
      <aff-alternatives id="aff-1">
        <aff xml:lang="ru">Российский университет дружбы народов имени Патриса Лумумбы</aff>
        <aff xml:lang="en">Peoples' Friendship University of Russia named after Patrice Lumumba Moscow</aff>
      </aff-alternatives>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>01</month>
        <year>2026</year>
      </pub-date>
      <volume>1</volume>
      <issue>1</issue>
      <elocation-id>10.26102/2310-6018/2026.59.8.013</elocation-id>
      <permissions>
        <copyright-statement>Copyright © Авторы, 2026</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This work is licensed under a Creative Commons Attribution 4.0 International License</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://moitvivt.ru/ru/journal/article?id=2414"/>
      <abstract xml:lang="ru">
        <p>Актуальность исследования обусловлена тем, что качество обработки глубинных данных часто оценивают по пиксельным ошибкам, хотя в робототехнике и трехмерной метрологии конечным результатом являются пространственные расстояния. Цель работы – разработать и проверить протокол оценки методов улучшения глубины по ошибке евклидовых расстояний между восстановленными точками. Основные эксперименты выполнены на последовательностях TUM Freiburg1 desk и xyz; дополнительная валидация включает динамическую последовательность Freiburg3 walking_static, контролируемые сценарии с движением границ, дальними глубинами и повышенным шумом, три варианта псевдоэталона и измерение вычислительных затрат. При полуширине окна N = 10 временное среднее уменьшило RMSE расстояний на 24,07 % для desk и на 23,66 % для xyz по сравнению с необработанной глубиной. Прямое сопоставление метрик показало, что минимальную пиксельную MAE дает временная медиана, тогда как минимальную RMSE расстояний – временное среднее; коэффициенты ранговой корреляции Спирмена равны 0,70 и 0,90. Временное среднее сохранило первое место при трех способах построения псевдоэталона. На динамической последовательности маска временной стабильности сохранила 18,92 % пикселей и изменила ранжирование, что подтверждает необходимость исключения или компенсации движущихся областей. На CPU обработка кадра 640 × 480 при N = 10 заняла 56,2 мс для временного среднего, 217,5 мс для временной медианы и 3155,6 мс для пространственно-временной медианы в исследовательской реализации. Полученные результаты определяют область применимости протокола и ограничения псевдоэталонной оценки.</p>
      </abstract>
      <trans-abstract xml:lang="en">
        <p>Depth enhancement is commonly evaluated with pixel-wise errors, although robotic vision and three-dimensional metrology ultimately use geometric distances. This study develops and tests a protocol that evaluates enhancement methods by the error of Euclidean distances between back-projected point pairs. The primary experiments use the TUM Freiburg1 desk and xyz sequences. Additional validation includes the dynamic Freiburg3 walking_static sequence, controlled stress tests with moving depth boundaries, long range and elevated noise, three pseudo-reference variants, and a computational-cost benchmark. With temporal half-window N = 10, temporal mean fusion reduces distance RMSE by 24.07 % on desk and 23.66 % on xyz relative to raw depth. Direct metric comparison shows that temporal median has the lowest pixel MAE, whereas temporal mean has the lowest distance RMSE; Spearman rank correlations are 0.70 and 0.90. Temporal mean remains ranked first under all three pseudo-reference constructions. On the dynamic sequence, a temporal-stability gate retains 18.92 % of pixels and changes the method ranking, demonstrating that moving regions must be excluded or motion-compensated. For 640 × 480 frames at N = 10, the research CPU implementation requires 56.2 ms for temporal mean, 217.5 ms for temporal median, and 3155.6 ms for spatial-temporal median. The results delimit the practical use of the protocol and the systematic limitations of pseudo-reference evaluation.</p>
      </trans-abstract>
      <kwd-group xml:lang="ru">
        <kwd>глубинные данные</kwd>
        <kwd>трехмерные измерения</kwd>
        <kwd>карта глубины</kwd>
        <kwd>псевдоэталон</kwd>
        <kwd>пространственно-временное слияние</kwd>
        <kwd>геометрическая точность</kwd>
        <kwd>улучшение глубины</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>depth data</kwd>
        <kwd>three-dimensional measurement</kwd>
        <kwd>depth map</kwd>
        <kwd>pseudo-reference</kwd>
        <kwd>spatial-temporal fusion</kwd>
        <kwd>geometric accuracy</kwd>
        <kwd>depth enhancement</kwd>
      </kwd-group>
      <funding-group>
        <funding-statement xml:lang="ru">Исследование выполнено без спонсорской поддержки.</funding-statement>
        <funding-statement xml:lang="en">The study was performed without external funding.</funding-statement>
      </funding-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="cit1">
        <label>1</label>
        <mixed-citation xml:lang="ru">Khoshelham K., Oude Elberink S. Accuracy and resolution of Kinect depth data for indoor mapping applications. Sensors. 2012;12(2):1437–1454. https://doi.org/10.3390/s120201437</mixed-citation>
      </ref>
      <ref id="cit2">
        <label>2</label>
        <mixed-citation xml:lang="ru">Nguyen Ch.V., Izadi Sh., Lovell D. Modeling Kinect sensor noise for improved 3D reconstruction and tracking. In: 2012 Second International Conference on 3D Imaging, Modeling, Processing, Visualization and Transmission, 13–15 October 2012, Zurich, Switzerland. IEEE; 2012. P. 524–530. https://doi.org/10.1109/3DIMPVT.2012.84</mixed-citation>
      </ref>
      <ref id="cit3">
        <label>3</label>
        <mixed-citation xml:lang="ru">Sarbolandi H., Lefloch D., Kolb A. Kinect range sensing: structured-light versus time-of-flight Kinect. Computer Vision and Image Understanding. 2015;139:1–20. https://doi.org/10.1016/j.cviu.2015.05.006</mixed-citation>
      </ref>
      <ref id="cit4">
        <label>4</label>
        <mixed-citation xml:lang="ru">Tomasi C., Manduchi R. Bilateral filtering for gray and color images. In: Sixth International Conference on Computer Vision, 04–07 January 1998, Bombay, India. IEEE; 1998. P. 839–846. https://doi.org/10.1109/ICCV.1998.710815</mixed-citation>
      </ref>
      <ref id="cit5">
        <label>5</label>
        <mixed-citation xml:lang="ru">He K., Sun J., Tang X. Guided image filtering. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2013;35(6):1397–1409. https://doi.org/10.1109/TPAMI.2012.213</mixed-citation>
      </ref>
      <ref id="cit6">
        <label>6</label>
        <mixed-citation xml:lang="ru">Lin B.-Sh., Su M.-J., Cheng P.-H., et al. Temporal and spatial denoising of depth maps. Sensors. 2015;15(8):18506–18525. https://doi.org/10.3390/s150818506</mixed-citation>
      </ref>
      <ref id="cit7">
        <label>7</label>
        <mixed-citation xml:lang="ru">Or-El R., Rosman G., Wetzler A., et al. RGBD-fusion: real-time high precision depth recovery. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 07–12 June 2015, Boston, MA, USA. IEEE; 2015. P. 5407–5416. https://doi.org/10.1109/CVPR.2015.7299179</mixed-citation>
      </ref>
      <ref id="cit8">
        <label>8</label>
        <mixed-citation xml:lang="ru">Ronneberger O., Fischer Ph., Brox Th. U-Net: convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015: 18th International Conference: Part III, 05–09 October 2015, Munich, Germany. Cham: Springer; 2015. P. 234–241. https://doi.org/10.1007/978-3-319-24574-4_28</mixed-citation>
      </ref>
      <ref id="cit9">
        <label>9</label>
        <mixed-citation xml:lang="ru">Sterzentsenko V., Saroglou L., Chatzitofis A., et al. Self-supervised deep depth denoising. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 27 October – 02 November 2019, Seoul, South Korea. IEEE; 2019. P. 1242–1251. https://doi.org/10.1109/ICCV.2019.00133</mixed-citation>
      </ref>
      <ref id="cit10">
        <label>10</label>
        <mixed-citation xml:lang="ru">Sturm J., Engelhard N., Endres F., et al. A benchmark for the evaluation of RGB-D SLAM systems. In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 07–12 October 2012, Vilamoura-Algarve, Portugal. IEEE; 2012. P. 573–580. https://doi.org/10.1109/IROS.2012.6385773</mixed-citation>
      </ref>
      <ref id="cit11">
        <label>11</label>
        <mixed-citation xml:lang="ru">Shabanov A., Krotov I., Chinaev N., et al. Self-supervised depth denoising using lower- and higher-quality RGB-D sensors. In: 2020 International Conference on 3D Vision (3DV), 25–28 November 2020, Fukuoka, Japan. IEEE; 2020. P. 743–752. https://doi.org/10.1109/3DV50981.2020.00084</mixed-citation>
      </ref>
      <ref id="cit12">
        <label>12</label>
        <mixed-citation xml:lang="ru">Shivakumar Sh.S., Nguyen T., Miller I.D., et al. DFuseNet: deep fusion of RGB and sparse depth information for image guided dense depth completion. In: 2019 IEEE Intelligent Transportation Systems Conference (ITSC), 27–30 October 2019, Auckland, New Zealand. IEEE; 2019. P. 13–20. https://doi.org/10.1109/ITSC.2019.8917294</mixed-citation>
      </ref>
      <ref id="cit13">
        <label>13</label>
        <mixed-citation xml:lang="ru">Пьявченко А.О., Ильченко А.В. Метод пространственной локализации статических объектов по данным датчика глубины и RGB-камеры. Известия ЮФУ. Технические науки. 2018;(1):271–284. https://doi.org/10.23683/2311-3103-2018-1-271-284</mixed-citation>
      </ref>
      <ref id="cit14">
        <label>14</label>
        <mixed-citation xml:lang="ru">Вохминцев А.В., Пачганов С.А. Алгоритм одновременной навигации и составления карты мобильным роботом на основе итеративного алгоритма ближайших точек и дескриптора, вычисляемого в круглом скользящем окне. Вестник Югорского государственного университета. 2018;14(3):49–56. https://doi.org/10.17816/byusu2018049-56</mixed-citation>
      </ref>
      <ref id="cit15">
        <label>15</label>
        <mixed-citation xml:lang="ru">Осман В., Громов В.С. Исследование системы навигации для мобильных роботов на основе одновременной локализации и построения карты. Научно-технический вестник информационных технологий, механики и оптики. 2020;20(3):371–376. https://doi.org/10.17586/2226-1494-2020-20-3-371-376</mixed-citation>
      </ref>
      <ref id="cit16">
        <label>16</label>
        <mixed-citation xml:lang="ru">Wu H., Fu K., Zhao Y., et al. Joint self-supervised and reference-guided learning for depth inpainting. Computational Visual Media. 2022;8(4):597–612. https://doi.org/10.1007/s41095-021-0259-z</mixed-citation>
      </ref>
      <ref id="cit17">
        <label>17</label>
        <mixed-citation xml:lang="ru">Yang Q., Cui W., Zheng Zh., et al. Unsupervised depth completion based on RGB image and sparse depth map. In: EITCE 2024: Proceedings of the 2024 8th International Conference on Electronic Information Technology and Computer Engineering, 18–20 October 2024, Haikou, China. New York: ACM; 2024. P. 152–157. https://doi.org/10.1145/3711129.3711157</mixed-citation>
      </ref>
      <ref id="cit18">
        <label>18</label>
        <mixed-citation xml:lang="ru">Zhang Y., Guo X., Poggi M., et al. CompletionFormer: depth completion with convolutions and vision transformers. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17–24 June 2023, Vancouver, BC, Canada. IEEE; 2023. P. 18527–18536. https://doi.org/10.1109/CVPR52729.2023.01777</mixed-citation>
      </ref>
      <ref id="cit19">
        <label>19</label>
        <mixed-citation xml:lang="ru">Wang Y., Li B., Zhang G., et al. LRRU: long-short range recurrent updating networks for depth completion. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 01–06 October 2023, Paris, France. IEEE; 2023. P. 9388–9398. https://doi.org/10.1109/ICCV51070.2023.00864</mixed-citation>
      </ref>
      <ref id="cit20">
        <label>20</label>
        <mixed-citation xml:lang="ru">Tang J., Tian F.-P., An B., et al. Bilateral propagation network for depth completion. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 9763–9772. https://doi.org/10.1109/CVPR52733.2024.00932</mixed-citation>
      </ref>
      <ref id="cit21">
        <label>21</label>
        <mixed-citation xml:lang="ru">Park J., Li Y.-J., Kitani K. Flexible depth completion for sparse and varying point densities. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 21540–21550. https://doi.org/10.1109/CVPR52733.2024.02035</mixed-citation>
      </ref>
      <ref id="cit22">
        <label>22</label>
        <mixed-citation xml:lang="ru">Yan Zh., Lin Y., Wang K., et al. Tri-perspective view decomposition for geometry-aware depth completion. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16–22 June 2024, Seattle, WA, USA. IEEE; 2024. P. 4874–4884. https://doi.org/10.1109/CVPR52733.2024.00466</mixed-citation>
      </ref>
      <ref id="cit23">
        <label>23</label>
        <mixed-citation xml:lang="ru">Xiang J., Zhu X., Wang X., et al. DEPTHOR: depth enhancement from a practical light-weight dToF sensor and RGB image. In: 2025 IEEE/CVF International Conference on Computer Vision (ICCV), 19–25 October 2025, Honolulu, HI, USA. IEEE; 2025. P. 6101–6111. https://doi.org/10.1109/ICCV51701.2025.00576</mixed-citation>
      </ref>
      <ref id="cit24">
        <label>24</label>
        <mixed-citation xml:lang="ru">Kim H., Wang R., Yao Ch., et al. Dense metric depth completion from sparse direct time-of-flight sensors. In: 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 03–07 June 2026, Denver, CO, USA. IEEE; 2026. P. 36518–36528.</mixed-citation>
      </ref>
      <ref id="cit25">
        <label>25</label>
        <mixed-citation xml:lang="ru">Li Y., Lou L., Tang Y., et al. LiteSense: lifting lightweight ToF with RGB for high-resolution metric depth estimation. In: 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 03–07 June 2026, Denver, CO, USA. IEEE; 2026. P. 5783–5792.</mixed-citation>
      </ref>
    </ref-list>
    <fn-group>
      <fn fn-type="conflict">
        <p>The authors declare that there are no conflicts of interest present.</p>
      </fn>
    </fn-group>
  </back>
</article>