How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "jihochoi/DR-MV3D-R-SFT" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "jihochoi/DR-MV3D-R-SFT",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "jihochoi/DR-MV3D-R-SFT" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "jihochoi/DR-MV3D-R-SFT",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Quick Links

DR-MV3D (SFT)

Official checkpoint for Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views (ECCV 2026).

Paper · Project page · Code · SFT + GRPO checkpoint

DR-MV3D answers questions about a 3D scene shown from several viewpoints. This is the supervised fine-tuning stage: Qwen2.5-VL-3B-Instruct trained on map-grounded reasoning traces. For general use prefer the final SFT + GRPO checkpoint.

Usage

The model expects the structured task prompt it was trained with, and produces a cognitive map, per-view ego-centric maps, a reasoning chain and the answer. See the GitHub repository for the prompt format, data preparation, inference and evaluation.

python scripts/run_inference.py \
  --model-path jihochoi/DR-MV3D-R-SFT \
  -i data/prompts/MindCube_tinybench_drmv3d.jsonl \
  -o results/tinybench_responses.jsonl \
  --image-root data --batch-size 8

License

The weights are a fine-tune of Qwen2.5-VL-3B-Instruct and are released under Apache-2.0, the base model's license. The code is MIT-licensed.

Citation

@inproceedings{drmv3d2026,
  title={Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views},
  author={Choi, Jiho and Lee, Seonho and Park, Seojeong and Shim, Hyunjung},
  booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
  year={2026},
  eprint={2606.23557},
  archivePrefix={arXiv},
  primaryClass={cs.CV}
}
Downloads last month
83
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jihochoi/DR-MV3D-R-SFT

Finetuned
(884)
this model

Dataset used to train jihochoi/DR-MV3D-R-SFT

Collection including jihochoi/DR-MV3D-R-SFT

Paper for jihochoi/DR-MV3D-R-SFT