PumpireUnified Benchmark for Metric Distance Estimation

Siyu Chen1,*Zehan Wang1,*,†Jiayang Xu1,*Yihan Wu2Jialei Wang1
Junming Chen1Ziang Zhang1Yutong Ying3Zhou Zhao1,‡
1Zhejiang University2Zhejiang University of Technology3Shanghai Jiaotong University

*Equal contribution†Project leader‡Corresponding author

Pumpire pipeline: RGB views and optional depth priors enter image- or video-level models. Predicted depth and associated camera intrinsics are used for back-projection, then the point-pair distance is compared with the physically measured distance.

Overview of the Pumpire evaluation protocol. Pumpire provides a unified evaluation of image- and video-level 3D foundation models under geometry estimation and completion settings. Estimation models recover geometry from RGB inputs, whereas completion models additionally condition on depth priors. Pumpire back-projects depth maps predicted by 3D foundation models with their associated camera intrinsics and quantifies the results with the proposed metrics. Their predicted distance is compared with the physically measured ground truth using the proposed metrics.

We back-project the two annotated pixels into the camera coordinate system and compute their Euclidean distance:

dpred=∥P0−P1∥2,wherePi=D(ui,vi)K−1[ui,vi,1]⊤,i∈{0,1}.(1)d_{\mathrm{pred}} = \left\|\mathbf{P}_0-\mathbf{P}_1\right\|_2, \qquad \text{where}\quad \mathbf{P}_i = D(u_i,v_i)K^{-1}[u_i,v_i,1]^\top, \quad i\in\{0,1\}.\tag{1}

Here, K denotes the camera intrinsics, and the ground-truth distance dgt is obtained through physical measurement.

Let SS denote the number of scenes and MsM_s the number of evaluated frames in scene ss. Since the point pair remains fixed within each scene, dgt(s)d_{\mathrm{gt}}^{(s)} is shared across all frames. We evaluate each scene using Absolute Error (AE), Relative Error (RE), Logarithmic Error (LE), and Threshold Accuracy (δτ\delta_\tau). Define the frame-wise absolute error and distance ratio as es,j=∣dpred(s,j)−dgt(s)∣e_{s,j}=\left|d_{\mathrm{pred}}^{(s,j)}-d_{\mathrm{gt}}^{(s)}\right| and rs,j=dpred(s,j)/dgt(s)r_{s,j}=d_{\mathrm{pred}}^{(s,j)}/d_{\mathrm{gt}}^{(s)}. The scene-level metrics are then

AEs=1Ms∑j=1Mses,j.(2)\mathrm{AE}_s =\frac{1}{M_s}\sum_{j=1}^{M_s}e_{s,j}.\tag{2}
REs=1Ms∑j=1Mses,jdgt(s).(3)\mathrm{RE}_s =\frac{1}{M_s}\sum_{j=1}^{M_s} \frac{e_{s,j}}{d_{\mathrm{gt}}^{(s)}}.\tag{3}
LEs=1Ms∑j=1Ms∣log⁡rs,j∣.(4)\mathrm{LE}_s =\frac{1}{M_s}\sum_{j=1}^{M_s} \left|\log r_{s,j}\right|.\tag{4}
δτ,s=1Ms∑j=1Ms1 ⁣[max⁡ ⁣(rs,j,rs,j−1)<τ].(5)\delta_{\tau,s} =\frac{1}{M_s}\sum_{j=1}^{M_s} \mathbf{1}\!\left[ \max\!\left(r_{s,j},r_{s,j}^{-1}\right)<\tau \right].\tag{5}

The benchmark-level result for each metric q∈{AE,RE,LE,δτ}q\in\{\mathrm{AE},\mathrm{RE},\mathrm{LE},\delta_\tau\} is obtained by averaging its scene-level values:

q=1S∑s=1Sqs.(6)q=\frac{1}{S}\sum_{s=1}^{S}q_s.\tag{6}

To evaluate cross-view consistency, we utilize Relative Standard Deviation (RSD) to measure the normalized variation of predicted distances within each sequence. For sequence ss, RSD is defined as

RSDs=1dˉpred(s)1Ms−1∑j=1Ms(dpred(s,j)−dˉpred(s))2,dˉpred(s)=1Ms∑j=1Msdpred(s,j).(7)\mathrm{RSD}_s = \frac{1}{\bar d_{\mathrm{pred}}^{(s)}} \sqrt{ \frac{1}{M_s-1} \sum_{j=1}^{M_s} \left( d_{\mathrm{pred}}^{(s,j)} -\bar d_{\mathrm{pred}}^{(s)} \right)^2 }, \qquad \bar d_{\mathrm{pred}}^{(s)} = \frac{1}{M_s} \sum_{j=1}^{M_s}d_{\mathrm{pred}}^{(s,j)}.\tag{7}

Abstract

We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors. In contrast to previous approaches that normally evaluate depth and camera intrinsics separately or evaluate point-clouds with geometric similarity metrics, which cannot directly reflect models’ point-to-point distance estimation capability, Pumpire directly assesses point-to-point distances from the reconstructed geometry. To this end, we collect a large-scale and diverse dataset (pumpire-6k) comprising 100 real-world scenes, each annotated with physically measured point-pair distances and containing 64 frames, for a total of 6,400 frames. Building on this dataset, we establish a holistic evaluation protocol that covers both image- and video-level 3D foundation models and enables direct assessment of point-pair distance errors and cross-setting comparison. We conduct extensive experiments across 29 baseline configurations of representative 3D foundation models and provide a comprehensive analysis of the results. By offering this benchmark, we target the more fundamental ability to perceive and estimate physical scale in the reconstructed 3D space, which prior evaluation protocols have largely overlooked.

Motivation

  • Existing benchmarks for depth and camera intrinsics estimation are insufficient to characterize a model's point-pair distance estimation capability. Finding 1 shows that strong performance on depth and camera intrinsics benchmarks—even when considered jointly—does not necessarily translate into accurate metric distance estimation, highlighting the need for direct evaluation of point-pair distances.
  • Existing point-cloud benchmarks are also not designed to directly evaluate metric point-pair distance estimation. Current benchmarks primarily assess geometric fidelity using metrics such as Chamfer Distance, Accuracy, Completion, and F1-Score, or point-wise reconstruction errors, rather than directly evaluating the accuracy of metric distance estimation between arbitrary point pairs. Meanwhile, recent spatial reasoning benchmarks explicitly evaluate metric distance reasoning, but at a coarse, object-level granularity through question answering. Consequently, neither benchmark type directly assesses metric distance estimation at the point level, leaving the ability to recover physical distances between arbitrary point pairs insufficiently evaluated.
  • There is no evaluation framework that enables unified comparisons of metric 3D foundation models across image- and video-level estimation and completion settings. Existing benchmarks are fragmented: some target image-level estimation or completion, while others focus on video-level estimation or completion. This fragmentation limits unified comparison of 3D foundation models across different settings.

Dataset: pumpire-6k

We curate a real-world benchmark dataset comprising 100 diverse scenes and 6,400 frames, annotated with physically measured point-pair distances and spanning indoor and outdoor environments with varying levels of difficulty, enabling comprehensive evaluation of metric 3D foundation models under diverse real-world conditions.

Quantitative Results

Comprehensive Insights

Finding 1

Direct evaluation of point-pair distance estimation.

The coupling between predicted depth and camera intrinsics renders their separate evaluation insufficient for reliably characterizing point-pair distance estimation.

Depth and intrinsics benchmark rankings

Average rank ↓ across 12 depth metrics and 24 intrinsics metrics. Bold: best; underlined: second best.

Average rank on the depth benchmark (12 metrics) and intrinsics benchmark (24 metrics). Lower is better.
BenchmarkDepth ProMoGe-2MetricAnythingUnidepthUnidepthv2
Depth benchmark4.252.832.752.082.83
Intrinsics benchmark2.891.001.503.463.83
Overall rank3.571.922.132.773.33
See detailed resultsHide detailed results
Depth benchmark results for image-level metric geometry estimation models.
NYU-D / KITTI / DIODE / ETH3D depth benchmark results
Model NYU-D KITTI DIODE ETH3D
AbsRel ↓ L1 (m) ↓ δ1.25 ↑ AbsRel ↓ L1 (m) ↓ δ1.25 ↑ AbsRel ↓ L1 (m) ↓ δ1.25 ↑ AbsRel ↓ L1 (m) ↓ δ1.25 ↑
Depth Pro 0.09 0.25 0.93 0.14 2.34 0.83 0.40 4.46 0.41 0.38 3.25 0.33
MoGe-2 0.08 0.21 0.96 0.21 3.59 0.45 0.33 2.62 0.54 0.10 0.62 0.88
MetricAnything 0.10 0.27 0.94 0.09 1.60 0.94 0.34 2.49 0.65 0.11 0.67 0.90
Unidepth 0.06 0.14 0.98 0.05 1.04 0.98 0.27 2.64 0.67 0.58 3.20 0.14
Unidepthv2 0.07 0.18 0.96 0.09 1.58 0.95 0.78 7.07 0.54 0.21 1.23 0.68
Intrinsics benchmark results for image-level metric geometry estimation models.

@1/5/10° refer to AUC@1/5/10°. vFoV, hFoV refer to vertical FoV and horizontal FoV. Methods trained on evaluated datasets are in gray and excluded from the ranking to ensure a fair comparison

LaMAR-2K
LaMAR-2K intrinsics benchmark results
Model vFoV hFoV
@1° @5° @10° @1° @5° @10°
Depth Pro 0.121 0.235 0.377 0.136 0.260 0.447
MoGe-2 0.243 0.526 0.727 0.287 0.581 0.766
MetricAnything 0.224 0.493 0.698 0.262 0.548 0.743
Unidepth 0.051 0.371 0.491 0.015 0.324 0.479
Unidepthv2 0.040 0.123 0.312 0.055 0.193 0.423
MegaDepth-2K
MegaDepth-2K intrinsics benchmark results
Model vFoV hFoV
@1° @5° @10° @1° @5° @10°
Depth Pro 0.196 0.449 0.660 0.176 0.423 0.636
MoGe-2 0.176 0.412 0.618 0.174 0.397 0.593
MetricAnything 0.270 0.537 0.722 0.252 0.514 0.702
Unidepth 0.068 0.162 0.274 0.058 0.149 0.253
Unidepthv2 0.124 0.305 0.494 0.120 0.289 0.475
Stanford2D3D
Stanford2D3D intrinsics benchmark results
Model vFoV hFoV
@1° @5° @10° @1° @5° @10°
Depth Pro 0.058 0.157 0.263 0.050 0.134 0.233
MoGe-2 0.094 0.269 0.487 0.087 0.249 0.454
MetricAnything 0.078 0.194 0.326 0.062 0.164 0.289
Unidepth 0.032 0.085 0.172 0.028 0.077 0.150
Unidepthv2 0.048 0.132 0.247 0.043 0.116 0.224
TartanAir
TartanAir intrinsics benchmark results
Model vFoV hFoV
@1° @5° @10° @1° @5° @10°
Depth Pro 0.556 0.740 0.859 0.549 0.730 0.851
MoGe-2 0.009 0.080 0.314 0.009 0.074 0.303
MetricAnything 0.364 0.702 0.847 0.353 0.688 0.839
Unidepth 0.008 0.016 0.030 0.006 0.015 0.031
Unidepthv2 0.000 0.004 0.093 0.000 0.004 0.083

Image-level geometry estimation using ground-truth intrinsics

Normal subset. Parentheses show changes relative to the corresponding image-level estimation results: green for improvements and red for degradations. AE is in meters; δ is in %.

Normal-subset results using ground-truth intrinsics. Parentheses show changes relative to the main image-level estimation table. AE in meters; delta accuracies in percent.
Model AE ↓ RE ↓ LE ↓ RSD ↓ δ1.05 ↑ δ1.10 ↑ δ1.25 ↑ Rank ↓
Metric3D1.319 (+0.139)1.895 (+0.200)0.893 (+0.065)0.340 (-0.007)5.16 (-2.86)9.89 (-4.79)23.42 (-4.40)8.00
Depth Pro0.144 (+0.009)0.210 (+0.014)0.240 (+0.006)0.175 (+0.021)16.53 (-1.21)31.68 (-1.30)62.77 (-2.53)5.14
Unidepth0.205 (+0.046)0.287 (+0.067)0.246 (+0.039)0.121 (-0.013)14.68 (-5.32)27.22 (-11.16)52.55 (-15.89)6.86
MoGe-20.172 (-0.004)0.229 (-0.013)0.208 (+0.004)0.095 (+0.009)17.00 (-6.40)31.50 (-9.43)62.29 (-0.79)4.86
MoGe-30.185 (-0.023)0.253 (-0.039)0.220 (-0.016)0.093 (+0.009)17.19 (-1.70)33.68 (-1.81)63.34 (+3.45)4.43
MetricAnything0.150 (+0.005)0.195 (+0.001)0.183 (+0.007)0.098 (+0.009)20.19 (-4.81)37.34 (-3.94)66.86 (-0.10)3.29
Unidepthv20.121 (+0.022)0.165 (+0.029)0.153 (+0.016)0.098 (+0.001)23.58 (-2.43)43.01 (-5.47)75.41 (-4.20)2.29
Metric3D v20.095 (+0.022)0.138 (+0.032)0.126 (+0.022)0.067 (-0.001)26.09 (-4.91)47.94 (-8.06)83.98 (-6.44)1.00

Finding 2

Sensor depth versus learned 3D foundation models.

Raw sensor depth is highly reliable on benign surfaces, but under challenging sensing conditions learned 3D foundation models are far more robust and surpass it.

Sensor comparison: raw RealSense D435 versus image-level estimation and completion models
Normal-scene δ1.05 and challenging-scene RE. Dashed lines mark the raw sensor baseline.

Finding 3

Distinct robustness patterns of estimation and completion.

Completion and estimation models exhibit distinct robustness patterns: the former attains higher peak accuracy but degrades catastrophically on challenging scenes, the latter keeps errors bounded throughout.

Robustness comparison of estimation and completion, including both plots and their legends
Challenging-to-normal RE ratios (left) and multi-view accuracy across scene difficulty (right).

Finding 4

Sparsely-sampled depth prior for image-level estimation models.

Geometry completion requires jointly addressing input sparsity and sensor noise; consequently, denser depth priors do not necessarily yield better performance.

Image-level completion accuracy versus sparse depth input density
Image-level completion with 100, 1,000, 10,000, or all sensor-depth points.

Finding 5

Offering models with different numbers of depth priors.

Additional depth priors generally improve video-level completion, but the benefit depends on the model's ability to integrate metric information across views.

Video-level completion relative error versus the depth-prior view ratio
Video-level completion as the fraction of views with depth priors increases.

Finding 6

Comparison across model settings.

Multi-view estimation improves consistency, depth-conditioned completion improves both accuracy and consistency, and combining both performs best overall.

Relative error versus relative standard deviation for four representative model settings
Representative settings on normal scenes. Lower RE and RSD are better.

Qualitative Results

Citation

@misc{chen2026pumpireunifiedbenchmarkmetric,
      title={Pumpire: Unified Benchmark for Metric Distance Estimation},
      author={Siyu Chen and Zehan Wang and Jiayang Xu and Yihan Wu and Jialei Wang and Junming Chen and Ziang Zhang and Yutong Ying and Zhou Zhao},
      year={2026},
      eprint={2610.12423},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.12423},
}