RF-DETR Medium Finetuned on VisDrone-DET

Fine-tuned RF-DETR Medium object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Detection Showcase

VisDrone-DET Detection Demo


Performance

Metric Score (%)
mAP@50 36.82
mAP@50-95 20.14
Precision 64.0
Recall 47.05
F1 Score 54.24
Parameters 33.7M
FLOPs N/A (not published upstream)

Evaluation Protocol

Metrics reported in this model card are computed on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).


VisDrone-DET Model Zoo

Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

Model mAP@50 mAP@50-95 Precision Recall
RF-DETR Medium 36.82 20.14 64.0 47.05
RF-DETR Small 33.25 17.88 62.62 43.51
RF-DETR Nano 25.15 12.77 58.99 35.0

External VisDrone-DET Comparison

The YOLO/RT-DETR rows below were trained and evaluated on VisDrone2019-DET's test split via a separate companion codebase (VisDrone-dataset-python-toolkit), not reproduced inside DetectionBench -- included here purely for context. The RF-DETR rows are this repository's own DetectionBench-trained runs (see the Model Zoo table above).

Model mAP@50 mAP@50-95 Precision Recall
YOLOv9e 40.02 23.73 54.78 42.42
YOLOv11x 38.44 22.6 52.41 41.43
YOLOv26x 38.33 22.48 52.91 41.06
YOLOv11l 37.14 21.85 51.87 40.33
YOLOv10x 37.24 21.81 52.59 39.84
YOLOv26l 37.65 21.75 51.6 40.42
YOLOv9c 37.22 21.73 51.99 39.77
YOLOv8x 36.81 21.52 51.91 39.78
YOLOv26m 36.67 21.22 51.03 39.79
YOLOv10l 35.95 21.09 52.13 38.48
YOLOv11m 36.35 21.02 50.24 39.46
YOLOv9m 36.19 20.95 51.05 39.12
RF-DETR-Medium 36.82 20.14 64.0 47.05
YOLOv8m 34.39 19.95 48.18 38.2
YOLOv9s 33.52 19.26 46.16 37.43
YOLOv11s 32.3 18.47 45.49 35.31
YOLOv8s 31.95 18.24 45.99 35.49
YOLOv26s 32.1 18.06 45.75 35.05
RF-DETR-Small 33.25 17.88 62.62 43.51
YOLOv9t 29.09 16.22 42.57 32.66
YOLOv8n 28.18 15.77 40.86 31.81
YOLOv11n 27.59 15.46 39.58 31.74
YOLOv10n 27.65 15.32 41.02 31.68
YOLOv26n 26.73 14.64 38.6 31.14
RF-DETR-Nano 25.15 12.77 58.99 35.0
rt_detr_l 21.68 9.34 35.76 26.3
Source: https://huggingface.co/collections/dronefreak/visdrone-detection-model-zoo

Per-Class Performance

Class mAP@50 mAP@50-95
pedestrian 28.15 11.15
people 23.02 8.2
bicycle 16.46 6.71
car 73.13 44.06
van 39.72 24.55
truck 46.54 29.03
tricycle 24.14 12.39
awning-tricycle 20.15 10.71
bus 62.35 41.68
motor 34.53 12.93
others 0.0 0.0

Evaluation Visualizations

This model was evaluated with Supervision's detection metrics, which report mAP/Precision/Recall directly but don't produce PR-curve, F1-curve, or confusion-matrix plot images the way Ultralytics' validator does. See the Performance table above for Precision/Recall/F1 and the per-class table above for the full per-class mAP breakdown.


Dataset

This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/Voxel51/VisDrone2019-DET

Classes

  • pedestrian
  • people
  • bicycle
  • car
  • van
  • truck
  • tricycle
  • awning-tricycle
  • bus
  • motor
  • others

Usage

Install Dependencies

pip install rfdetr huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
import rfdetr

weights = hf_hub_download(
    repo_id="dronefreak/visdrone-rfdetr-medium",
    filename="checkpoint_best_total.pth"
)

model = rfdetr.RFDETRMedium(pretrain_weights=weights)

Run Inference

detections = model.predict("image.jpg", threshold=0.25)

Training Configuration

Setting Value
Dataset VisDrone-DET
Framework RF-DETR
Training Toolkit DetectionBench
Epochs (configured max) 500
Epochs (actually trained) 141
Early Stopping Patience 100
Batch Size 10
Resolution 576
Optimizer adamw
Learning Rate 0.0001
Seed 42

Repository Contents

checkpoint_best_total.pth
metrics.csv
config.json
visdrone_rfdetr-medium_showcase.jpg
README.md

Related Resources


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Severe class imbalance: car (42.21%) and pedestrian (23.12%) account for two-thirds of all annotated boxes in the training set, while awning-tricycle (0.95%) and tricycle (1.40%) are rare -- the others class has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP).
  • Extreme small-object density: ~53 annotated boxes per image on average, with roughly 69% of boxes covering under 0.1% of the image area -- consistent with VisDrone's aerial small-object detection challenge (objects captured from significant altitude).
  • The original authors license VisDrone under CC BY-NC-SA 3.0 -- non-commercial research use only (see the dataset's homepage); this applies to any model trained on it, not only the raw images.
  • These RF-DETR checkpoints were trained/evaluated directly through DetectionBench. The YOLO/RT-DETR rows in the External VisDrone Model Zoo comparison below were trained via a separate companion codebase, not reproduced inside DetectionBench -- see that collection for their own training details and caveats.

Citation

If you use this model in your research, please consider citing:

  1. The VisDrone-DET dataset (see below)
  2. The original RF-DETR Medium architecture (see below)
  3. DetectionBench, the training/evaluation framework used to produce this checkpoint
@article{zhu2018vision,
  title={Vision meets drones: A challenge},
  author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
  journal={arXiv preprint arXiv:1804.07437},
  year={2018}
}
@inproceedings{robinson2026rfdetr,
  title     = {RF-DETR: Real-Time Detection Transformer},
  author    = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2511.09554}
}

@article{oquab2023dinov2,
  title={DINOv2: Learning Robust Visual Features without Supervision},
  author={Oquab, Maxime and Darcet, Timoth{\'e}e and Moutakanni, Theo and Vo, Huy and Szafraniec, Marc and Khalidov, Vasil and Fernandez, Pierre and Haziza, Daniel and Massa, Francisco and El-Nouby, Alaaeldin and others},
  journal={arXiv preprint arXiv:2304.07193},
  year={2023}
}
@software{Saksena_DetectionBench_2026,
  author = {Saksena, Saumya Kumaar},
  title = {DetectionBench: Reproducible Benchmarks for Modern Object Detectors on Real-World Datasets},
  url = {https://github.com/dronefreak/DetectionBench},
  year = {2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dronefreak/visdrone-rfdetr-medium

Finetuned
(13)
this model

Dataset used to train dronefreak/visdrone-rfdetr-medium

Collection including dronefreak/visdrone-rfdetr-medium

Papers for dronefreak/visdrone-rfdetr-medium