Method: SAM-FP (depth-aware ranking)

User zwdoescode
Publication
Implementation Python pipeline built on PyTorch (SAM3) and TensorRT (FoundationPose), run on Ubuntu 24.04, CUDA 13.2
Views Single
Test image modalities RGB-D
Description

Method: SAM3 + FoundationPose (single-view RGB-D).

A training-free pipeline combining two published, off-the-shelf models. No network was trained or fine-tuned; all weights are used as released, and the same models handle all objects without per-object training.

Codebase: https://github.com/nvidia-isaac/foundationpose_perception_pipeline

Pipeline (per BOP19 target image):

  1. Input — a single target RGB-D view. The provided sensor depth is used directly; no additional views or depth reconstruction are used.
  2. Detection and segmentation (SAM3) — text-prompted, open-vocabulary detection on the target RGB image, with a confidence threshold of 0.10.
  3. 6D pose estimation (FoundationPose) — model-based registration using the target RGB image, sensor depth, and provided CAD models. FoundationPose uses 128 pose hypotheses and five refinement iterations. All pose proposals are submitted without a score cutoff; the score column contains FoundationPose’s reranking score.

Differences with respect to the linked publications:

1 SAM3 and FoundationPose are composed into an integration pipeline. 2. The individual networks are otherwise unmodified from their released versions.

Training images and rendering: None. No training or offline rendering was performed. During inference, FoundationPose renders CAD-model hypotheses online for pose refinement and scoring. The provided models_eval CAD meshes are used directly.

Networks: One set of networks is used for all objects—a single SAM3 model and a single FoundationPose model. No per-object models are used.

Computer specifications 1x NVIDIA RTX PRO 6000 Blackwell Server Edition

Public submissions

Date Submission name Dataset
2026-10-06 23:36 SAM-FP (depth-aware ranking) HB
2026-10-06 23:37 SAM-FP (depth-aware ranking) LM-O