Instructions to use bratao/portugueseT5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bratao/portugueseT5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bratao/portugueseT5")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("bratao/portugueseT5") model = AutoModelForSeq2SeqLM.from_pretrained("bratao/portugueseT5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bratao/portugueseT5 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bratao/portugueseT5" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/portugueseT5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bratao/portugueseT5
- SGLang
How to use bratao/portugueseT5 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bratao/portugueseT5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/portugueseT5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bratao/portugueseT5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/portugueseT5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bratao/portugueseT5 with Docker Model Runner:
docker model run hf.co/bratao/portugueseT5
YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other
portugueseT5
portugueseT5 is a Portuguese encoder-decoder research checkpoint trained from
scratch with the architecture and configuration of
google/t5-v1_1-large. The doctoral
thesis describes pre-training on a 20% sample of GigaVerbo.
This repository is a base/intermediate language-model artifact, not a ready-made
OpenIE model and not a backend registered by
portuguese-openie. For OpenIE, use
bratao/PortugueseT5Oie or bratao/PortugueseT5OieAbstractive and their documented
task prompt.
Model details
| Field | Value |
|---|---|
| Public repository | bratao/portugueseT5 |
| Architecture reference | google/t5-v1_1-large |
| Architecture | T5 encoder-decoder, 24 encoder and 24 decoder layers, GEGLU |
| Task status | Portuguese pre-training checkpoint; no task-specific contract |
| Parameters in published configuration | 783,150,080 (approximately 783M) |
| Published weight precision | bfloat16 |
| Approximate repository size | 1.57 GB |
| Audited revision | 96e9ee96be4f6fedcfece958f74e895b19352acf (2026-08-30) |
The thesis rounds the model size to 770M; 783,150,080 is the exact count reported by the public configuration. The repository metadata does not identify a tokenizer or weights from Google as the base model: it records an architecture/configuration reference, while the thesis says the model was pre-trained from scratch.
Direct Transformers use
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_id = "bratao/portugueseT5"
revision = "96e9ee96be4f6fedcfece958f74e895b19352acf"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForSeq2SeqLM.from_pretrained(
model_id,
revision=revision,
dtype="auto",
device_map="auto",
)
# This is only a low-level generation example. The base checkpoint has no
# documented instruction or OpenIE prompt contract.
text = "A UFBA está localizada em Salvador."
inputs = tokenizer(text, return_tensors="pt", truncation=True).to(model.device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=64, do_sample=False)
decoded = tokenizer.decode(output[0], skip_special_tokens=True)
print(decoded) # plain generated text; content is not guaranteed
Output contract
The API returns a decoded string. No specific response to the example above is claimed because the base checkpoint has no published downstream instruction format, and an invented answer would be misleading. Evaluate a prompt and downstream fine-tuning protocol for the intended task before deployment.
Training-data provenance
The thesis states that this checkpoint was pre-trained on 20% of GigaVerbo. The
model repository does not declare a Hugging Face dataset identifier, the exact
sample/revision is not documented in its metadata, and the corpus is not bundled
with this card. The YAML therefore intentionally omits datasets.
Evaluation and status
No intrinsic or downstream evaluation metric is attributable to this exact base
checkpoint in the public repository. Metrics reported for PortugueseT5Oie family
members must not be transferred to this artifact. Treat it as an intermediate
research checkpoint that requires task-specific evaluation.
Requirements and hardware
- Recent Python, PyTorch, Transformers, and Accelerate.
- The published bfloat16 weights occupy about 1.57 GB. Around 4–6 GB of available RAM/VRAM is a practical starting point; actual use depends on input and generation length, runtime, and device.
- GPU execution is recommended for training and large-scale inference but is not required for a small CPU smoke test.
Limitations
- There is no stable instruction, QA, summarization, or OpenIE prompt contract.
- No public evaluation, model-completion statement, or detailed pre-training recipe is included in the repository.
- Generated text can be incorrect, biased, unsafe, or unrelated to the input.
- The model has not been audited for demographic bias or high-impact use.
License
No license is declared in the public model repository as of 2026-08-30. Absence of a license is not permission to copy, modify, or redistribute the weights. Obtain clarification from the author and review the terms of the architecture reference and training data before reuse. This card does not assign a license by inference.
Citation
@phdthesis{cabral2025evolving,
author = {Cabral, Bruno Souza},
title = {Evolving Open Information Extraction for Portuguese employing Language Models},
school = {Universidade Federal da Bahia},
year = {2025}
}
@inproceedings{cabral2022portnoie,
author = {Cabral, Bruno and Souza, Marlo and Claro, Daniela Barreiro},
title = {PortNOIE: A Neural Framework for Open Information Extraction for the Portuguese Language},
booktitle = {Computational Processing of the Portuguese Language (PROPOR 2022)},
year = {2022},
doi = {10.1007/978-3-030-98305-5_23}
}
Project: Portuguese-OpenIE · PortNOIE paper
- Downloads last month
- 223