--- library_name: sentence-transformers pipeline_tag: sentence-similarity tags: - sentence-transformers - feature-extraction - sentence-similarity - visual-document-retrieval - cross-modal-distillation - knowledge-distillation - nanovdr base_model: google-bert/bert-base-uncased language: - en license: apache-2.0 datasets: - openbmb/VisRAG-Ret-Train-Synthetic-data - openbmb/VisRAG-Ret-Train-In-domain-data - vidore/colpali_train_set - llamaindex/vdr-multilingual-train model-index: - name: NanoVDR-Q-BERT-Qwen3VL2B-2048 results: - task: type: retrieval dataset: name: ViDoRe v1 type: vidore/vidore-benchmark-667173f98e70a1c0fa4d metrics: - name: NDCG@5 type: ndcg_at_5 value: 82.1 - task: type: retrieval dataset: name: ViDoRe v2 type: vidore/vidore-benchmark-v2 metrics: - name: NDCG@5 type: ndcg_at_5 value: 62.2 ---
-- - [-ML]`, and the > rule is simply that the teacher and the width have to match. This model is > `Q` (query tower), `BERT` (backbone), distilled from > Qwen3-VL-Embedding-2B into 2048 dimensions. > `-ML` used to be spelled `-Multi`, which read as multi-vector when it meant > multilingual. **BERT-base ablation variant.** For production use, we recommend **[NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML)**. NanoVDR-Q-BERT-Qwen3VL2B-2048 is a 112M-parameter text-only query encoder for visual document retrieval, trained via asymmetric cross-modal distillation from [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B). It uses BERT-base + a 2-layer MLP projector and achieves the highest v2 score among all NanoVDR variants. ### Highlights - **Single-vector retrieval** — queries and documents share the same 2048-dim embedding space as [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B); retrieval is a plain dot product, FAISS-compatible, **4 KB per page** (float16) - **Lightweight on storage** — 454 MB model; doc index costs 64× less than ColPali's multi-vector patches - **Asymmetric setup** — tiny 112M text encoder at query time; large VLM indexes documents offline once ## Results | Model | Params | ViDoRe v1 | ViDoRe v2 | ViDoRe v3 | Avg Retention | |-------|--------|-----------|-----------|-----------|---------------| | Qwen3-VL-Emb (Teacher) | 2.0B | 84.3 | 65.3 | 50.0 | — | | **NanoVDR-Q-BERT-Qwen3VL2B-2048** | **112M** | **82.1** | **62.2** | **44.7** | **94.0%** | | NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML | 69M | 82.2 | 61.9 | 46.5 | 95.1% | NDCG@5 (×100). Retention = Student / Teacher averaged across v1/v2/v3. ## Usage > **Prerequisite:** Documents must be indexed offline using [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) (the teacher model). See the [NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML model page](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML#prerequisites-document-indexing-with-teacher-model) for a complete indexing guide. ```python from sentence_transformers import SentenceTransformer import numpy as np # doc_embeddings: (N, 2048) from teacher indexing (see prerequisite above) model = SentenceTransformer("nanovdr/NanoVDR-Q-BERT-Qwen3VL2B-2048") query_embeddings = model.encode(["What was the revenue growth in Q3?"]) # (1, 2048) scores = query_embeddings @ doc_embeddings.T top_k_indices = np.argsort(scores[0])[-5:][::-1] ``` ## Training Details | | Value | |--|-------| | Architecture | BERT-base (109M) + MLP projector (768 → 768 → 2048, 2.4M) = 112M | | Objective | Pointwise cosine alignment with teacher query embeddings | | Data | 711K query-document pairs | | Epochs / lr | 20 / 2e-4 | | Training cost | ~10.5 GPU-hours (1× H200) | | CPU query latency | 101 ms | ## All NanoVDR Models | Model | Backbone | Params | v1 | v2 | v3 | Retention | |-------|----------|--------|----|----|----| ----------| | **[NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML)** | **DistilBERT** | **69M** | **82.2** | **61.9** | **46.5** | **95.1%** | | [NanoVDR-Q-DistilBERT-Qwen3VL2B-2048](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048) | DistilBERT | 69M | 82.2 | 60.5 | 43.5 | 92.4% | | [NanoVDR-Q-BERT-Qwen3VL2B-2048](https://huggingface.co/nanovdr/NanoVDR-Q-BERT-Qwen3VL2B-2048) | BERT-base | 112M | 82.1 | 62.2 | 44.7 | 94.0% | | [NanoVDR-Q-ModernBERT-Qwen3VL2B-2048](https://huggingface.co/nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048) | ModernBERT | 151M | 82.4 | 61.5 | 44.2 | 93.4% | ### Want a system with no teacher at all? Every model in this table targets **Qwen3-VL-Embedding-2B at 2048 dimensions**, so documents still have to be indexed by that 2B teacher. A second family targets the 8B teacher at 4096 dimensions and includes a **document tower**, so nothing multi-billion runs at indexing time either: | Model | Role | Params | |---|---|---| | [NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML) | query tower | 70M | | [NanoVDR-D-HiRes-Qwen3VL8B-4096](https://huggingface.co/nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096) | document tower | 457M | | [NanoVDR-D-Fast-Qwen3VL8B-4096](https://huggingface.co/nanovdr/NanoVDR-D-Fast-Qwen3VL8B-4096) | document tower, 3x fewer visual tokens | 457M | The two families are **not interchangeable**: 2048-d and 4096-d vectors live in different spaces and will not score against each other. ## Citation ```bibtex @article{nanovdr2026, title={NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval}, author={Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu}, journal={arXiv preprint arXiv:2603.12824}, year={2026} } ``` ## License Apache 2.0