|
Download README.md from tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 1.39 kB
-
https://huggingface.co/tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4/resolve/main/README.md
- Command line
-
hf download hf://tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4/README.md
-
curl -L -o README.md https://huggingface.co/tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4/resolve/main/README.md
1.39 kB
| license: llama3.1 | |
| base_model: meta-llama/Llama-3.1-8B-Instruct | |
| tags: | |
| - llama | |
| - quantized | |
| - nvidia-modeloptimizer | |
| - NVFP4 | |
| library_name: nvidia-modeloptimizer | |
| # Llama-3.1-8B-Instruct Quantized (ModelOpt NVFP4) | |
| This is a quantized version of [meta-llama/Llama-3.1-8B-Instruct](https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct) using [modelopt](https://github.com/NVIDIA/Model-Optimizer) with NVFP4 weight quantization. | |
| ## Model Details | |
| - **Base Model:** meta-llama/Llama-3.1-8B-Instruct | |
| - **Quantization Method:** modelopt NVFP4 Post-Training Quantization (PTQ) | |
| - **Weight Precision:** NVFP4 | |
| - **Original Size:** ~16 GB (bfloat16) | |
| - **Quantized Size:** ~6 GB (nvfp4) | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| # Load base model structure | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4", | |
| torch_dtype=torch.bfloat16, | |
| low_cpu_mem_usage=True | |
| ) | |
| # Load tokenizer and generate | |
| tokenizer = AutoTokenizer.from_pretrained("tokenlabsdotrun/Llama-3.1-8B-ModelOpt-NVFP4") | |
| inputs = tokenizer("Hello, my name is", return_tensors="pt") | |
| outputs = model.generate(**inputs, max_new_tokens=10) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| ## License | |
| This model inherits the [Llama 3.1 Community License](https://llama.meta.com/llama3_1/license/). | |