|
Download ADDITIONAL_README.md from fdurant/colbert-xm-for-inference-api: direct link, hf CLI and curl.
- Browser
- Download file 1.12 kB
-
https://huggingface.co/fdurant/colbert-xm-for-inference-api/resolve/main/ADDITIONAL_README.md
- Command line
-
hf download hf://fdurant/colbert-xm-for-inference-api/ADDITIONAL_README.md
-
curl -L -o ADDITIONAL_README.md https://huggingface.co/fdurant/colbert-xm-for-inference-api/resolve/main/ADDITIONAL_README.md
1.12 kB
Multilingual Colbert embeddings as a service
Goal
- Deploy Antoine Louis' colbert-xm as an inference service: text(s) in, vector(s) out
Motivation
- use the service in a broader RAG solution
Steps followed
- Clone the original repo following this procedure
- Add a custom handler script as described here
Local development and testing
Build and start docker container hf_endpoints_emulator
docker-compose up -d --build
This can take a few moments to load, given the size of the model (> 3 GB)!
How to test locally
./embed_single_query.sh
./embed_two_chunks.sh
docker-compose exec hf_endpoints_emulator pytest
Check output
docker-compose logs --follow hf_endpoints_emulator