depth-anything-3-p150
ByteDance-Seed Depth Anything 3 (DA3Metric-Large) port on one Tenstorrent Blackhole p150a. Weights: depth-anything/DA3Metric-Large · Paper: arXiv:2511.10647 · Upstream code: ByteDance-Seed/Depth-Anything-3 · Port: changh95/tt-Depth-Anything-3
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. The whole forward runs on the chip. All numbers on this card were measured in this configuration.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/depth-anything-3-p150 --exclude "image/*" --local-dir depth-anything-3-p150 && cd depth-anything-3-p150
pip install -e code/ # adds only the package; ttnn and torch come from your environment
pip install -e "code/[fast]" # optional: faster PNG decode / encode (the same pixels)
pip install -e "code/[server,test]" # optional: the HTTP server and the tests
from tt_depth_anything_3 import DepthAnything3
with DepthAnything3.from_pretrained(device_id=0) as model: # weights go to your HF cache
out = model("media/source_1.png") # path (relative to the repo root), PIL image, numpy / torch array, or a list
print(out.depth.shape, out.depth.dtype) # (518, 518) float32, metres
print(out.stats()["median"]) # 4.45
out.save("depth16.png") # 16-bit PNG, value / 1000 = metres
out.colorize().save("depth_color.jpg") # turbo view, near = red
- The
--exclude "image/*"option skips the 3 GB container image.git clonealso works, but it downloads that image. - The relative paths in the snippet (
media/source_1.pngand the output files) start at the root of the model repository. Run the snippet from there, or use absolute paths. from_pretraineddownloadsdepth-anything/DA3Metric-Largeto your HF cache, compiles the kernels and captures the metal traces. The first load compiles for approximately 1 minute. A load with a warm kernel cache takes 7-15 s.from_pretrained()warms up the common input sizes (the 4 KITTI sizes and 518x518) and images larger than 4 MP. Thus the first call is as fast as the next calls (a load takes approximately 9 s). To prepare other sizes, usemodel.warmup(size=(H, W))orwarmup_variants={"sizes": [...]}. The console is quiet. For progress and logs, setverbose=True.- Load the model one time and call it many times. A warm call takes approximately 15 ms on the device, plus the image decode.
- The
withblock releases the traces and closes the device at the end. Withoutwith, callmodel.close().
Input (images) |
A file path, encoded PNG/JPEG bytes, a PIL.Image, a numpy array or a torch tensor (HxWx3, 3xHxW or HxW; uint8 0..255, or float 0..1). A list (or a 4-D array) gives a list of outputs. Each side must be 8 px or more. |
| Call options | model(images, resize_to_original=False, max_depth_m=None). resize_to_original=True resizes the depth to the input (H, W). max_depth_m=80 clamps far pixels. |
| Load options | from_pretrained(model_id="depth-anything/DA3Metric-Large", device_id=0, device=None, weights_path=None, revision=None, dispatch="auto"). device accepts a ttnn device that you opened with open_device(...); close() does not close it. |
Output (DepthOutput) |
depth float32 metres, (518, 518) or the input (H, W); original_size, timing_ms; stats() (min, p05, median, p95, max); to_uint16(), to_png16(), colorize(); save(path) for .png (16-bit), .npy, .npz, .jpg. |
- Throughput:
model(["a.png", "b.png"])andmodel.iter_predict(frames)decode the next images on host threads while the chip runs. The chip runs one image at a time (batch 1). - The depth is bit-identical to the HTTP server's
npzoutput.out.save("x.png")gives the same bytes as the server's 16-bit PNG. - One process holds one model on one chip. To use more chips, start one process for each chip.
- Use the
code/of this repo for the Python API. The container image is older and does not containtt_depth_anything_3. - Full reference (all arguments, input forms, speeds and tests):
PYTHON.md. Runnable example:python examples/quickstart.py [image]writes its outputs toexamples/output/.
Serving (HTTP)
tt-model pull changh95/depth-anything-3-p150 --with-weights
tt-model serve changh95/depth-anything-3-p150 # or with tt-cli: tt serve changh95/depth-anything-3-p150
printf '{"image":"%s"}' "$(base64 -w0 media/source_1.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/depth-anything-3-p150
- Weights
depth-anything/DA3Metric-Largego to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete. POST /predict:image(base64 PNG/JPEG); optionaloutput_format(png16-bit PNG default |npzfloat32 |json),resize_to_original(false),max_depth_m(null),png_scale(1000 = millimetres).POST /predict_png: the body is the raw PNG/JPEG; the response is the 16-bit depth PNG (image/png), with the stats inX-headers; optional?png_scale=. Only the currentcode/has this route. The container image does not have it.GET /health,GET /info.
Response (shortened):
{"depth": {"format": "png16", "data": "...", "scale": 1000.0, "saturation_m": 65.535, "saturated_pixels": 0},
"shape": [518, 518], "original_size": [375, 1242], "model_input_size": [518, 518], "units": "metres",
"stats_m": {"min": 1.64, "p05": 1.76, "median": 4.45, "p95": 20.09, "max": 42.52},
"timing_ms": {"preprocess": ..., "infer": ..., "postprocess": ..., "total": ...}}
depth.datais a base64 16-bit grayscale PNG at 518×518 (shapeis[H, W]);depth_m = uint16 / scale, saturating at65535 / png_scalem.output_format: "npz"returns a lossless float32depth_marray.
Demo
| Input (KITTI Eigen test frames) | Metric depth on p150a (turbo colormap, 1–50 m log-clipped; bright = near) |
|---|---|
![]() |
![]() |
Demo & Performances
Warm, batch 1, media/source_1.png (1242×375) squashed to 518×518, one metal trace per request; bench_da3_stages.py with 30 iterations per run (median / min). Configuration of every row: dispatch on the ETH cores, 1 command queue (CQ), 12×10 compute grid (2026-10-05 build, commit f877beb).
| Metric | Performance |
|---|---|
End-to-end run() (pixel upload + trace + D2H of the 518×518 log-depth) |
14.93 ms median (14.79–15.03, 3 runs) · 14.57 ms min |
| Device trace (patch embed, 24 DINOv2-L blocks, DPT head; replay + sync) | 13.49 ms median (13.47–13.52, 3 runs) |
| Trace back-to-back (20 replays / 20) | 13.84–13.95 ms |
| Profiler span of one traced replay · kernel sum (265 ops) | 11.63–11.79 ms · 11.46–11.50 ms |
| H2D · D2H | 0.28 · 0.46 ms |
Python model() call (tt_depth_anything_3, array input, 30 ms pause before each call) |
14.44 ms median (14.43–14.65, 3 runs) · 14.09 ms min (back-to-back: 15.15–15.67 median · 14.73 min) |
Served /predict, uvicorn, 120 requests |
1 client: total 19.8–20.7 ms (25.7 ms under host load) · infer 13.7–13.8 ms · client p50 24.3–26.2 ms · 37–38 req/s. 4 clients: 49.0–49.6 req/s |
Served /predict_png, uvicorn, 120 requests |
1 client: client p50 24.2–24.9 ms · 37.7–40.1 req/s. 4 clients: 50.4–52.8 req/s |
Served path, served-optimised mode (2026-10-03 build), bench_served.py, 50 iterations |
total 42.0 ms median (41.90–42.17, 2 runs) · infer 14.08–14.23 ms |
| Served path, like-for-like mode (the GPU run's host code; 2026-10-03 build), 50 iterations | total 55.95–56.55 ms median · infer 13.53–13.67 ms |
The measurement hardware is a Blackhole chip in the p150-equivalent configuration: dispatch on the ETH cores, 1 CQ and a 12×10 compute grid. All rows use this configuration, also the rows of the earlier builds. An independent verifier measured two runs of each f877beb row on one chip. The developer measured the third run on a different chip of the same type. The H2D · D2H row comes from the verification of round 121dcaf; later rounds did not change the transfers. The 2026-10-05 change sets ETH dispatch, 1 CQ and the 12×10 grid as the code default on all paths. It does not change the kernels, and the output is bit-identical to the 2026-10-04 build. The 2026-10-04 build gave run() 14.67 ms and device trace 13.17 ms on a third chip; the difference is chip-to-chip and host-load variation. The depth from model() is bit-identical to the server's npz output. In the 2026-10-04 test, the Python API added 0.38 ms to the server's own device call (gate: 0.5 ms). The served-optimised mode is the server default: the chip does the resize and normalisation, and the host uses fast PNG decode and encode. The like-for-like mode uses the same host code as the GPU run. The uvicorn host was shared (load average 11–28), so the served numbers are noisy. Accuracy on KITTI Eigen-652 (semi-dense GT, median scaling, Garg crop): AbsRel 0.0780, δ1 0.9526, RMSE 2.904 m (fp32 reference AbsRel 0.0771); PCC vs the fp32 reference mean 0.999960, min 0.999820. The 2026-10-05 build gives the same AbsRel and PCC values. Details: VERIFICATION_2026-10-03.md.
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on media/source_1.png. The "incl. H2D/D2H" column compares with our run() (14.93 ms). The "forward only" column compares with our device trace (13.49 ms). Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (14.93 ms) | GPU forward only vs ours (13.49 ms) |
|---|---|---|---|
| fp32 strict, eager | 45.3 ms | Blackhole 3.04× faster | 45.1 ms: Blackhole 3.34× faster |
| tf32, eager | 27.2 ms | Blackhole 1.82× faster | 26.9 ms: Blackhole 2.00× faster |
| bf16 autocast, eager | 23.2 ms | Blackhole 1.56× faster | 22.9 ms: Blackhole 1.70× faster |
| fp16 autocast, eager | 22.9 ms | Blackhole 1.53× faster | 22.6 ms: Blackhole 1.68× faster |
tf32 + torch.compile (CUDA graphs) |
25.4 ms | Blackhole 1.70× faster | 25.1 ms: Blackhole 1.86× faster |
fp16 autocast + torch.compile (CUDA graphs) |
13.4 ms | GPU 1.11× faster | 13.0 ms: GPU 1.04× faster |
The Blackhole lead comes from one metal trace that runs the whole network on all 120 compute cores. Custom kernels fuse qkv, proj, fc1 + GELU and the head 3x3 convs. A one-pass SDPA without the row max runs in 23 of 24 layers. Eager GPU at batch 1 is launch-bound; torch.compile with CUDA graphs removes this cost, and then the GPU is 1.04–1.11× faster. The served totals are host-bound on both sides. The GPU served-like bf16 total is 52.84 ms: the served-optimised mode of the 2026-10-03 build is 1.26× faster, and the like-for-like mode is 1.06× slower.
Caveats
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(orTT_FUSED_DISPATCH=worker) is an opt-in for a Galaxy chip only. On a p150, worker dispatch gives an 11×10 grid, and the model cannot use this grid. The default (auto) selects ETH dispatch.- Every image is squashed (no letterbox) to 518×518 and ImageNet-normalised; one image per request, batch 1.
- The whole forward runs on the chip in one metal trace: patch embed, 24 DINOv2-L blocks and the DPT head. Matmuls and head convs use bf16 with HiFi3 and fp32 accumulation. The backbone qkv, proj and fc2 matmuls use HiFi2 with a weight debias. Some optimizations change numerics.
OPT_REPORT.mdlists each change.TT_FUSED_BB_HIFI2=nonerestores HiFi3 in the backbone, andTT_FUSED=0restores the legacy eager path. - The HiFi2 change of 2026-10-04 changed the accuracy a little. KITTI-652 AbsRel stays at 0.0780. The PCC vs the fp32 reference went from mean 0.999965 to 0.999960, and from min 0.999860 to 0.999820. The worst image is 0.000020 above the gate floor (0.99980). The developer selected the debias value (0.0020) on the KITTI-652 gate set. The δ1 and RMSE values come from the 2026-10-03 build.
- 23 of 24 attention layers use the one-pass SDPA with calibrated exp offsets and key shifts. The calibration used a subset of the KITTI-652 images. The 489 images outside that subset also give AbsRel 0.0780. If a row leaves the calibrated window, a guard detects it. The request then runs again with the stock SDPA (about 115 ms).
TT_FUSED_SDPA_NOMAX=0restores the stock SDPA in all layers. - The port implements no sky head:
depth = exp(raw), so far-sky pixels can exceed 80 m; passmax_depth_m(the KITTI eval caps at 80) or usenpz, since the default 16-bit PNG saturates at 65.5 m. - Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150a power was not measured, so no efficiency comparison is made.
Licensing
- Weights: depth-anything/DA3Metric-Large, Apache-2.0 (upstream LICENSE).
- Port and serving code (
code/): Apache-2.0, from changh95/tt-Depth-Anything-3.code/exp/nomax/kernels/holds model-local copies of tt-metal SDPA kernels (Apache-2.0). Only the experiment scripts use them.
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-05 build with the Python API and the ETH-dispatch defaults; see OPT_REPORT.md and PYTHON.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
05a1c6793b218b2c (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:50:35+00:00 by tt-model 0.1.0 |

