VietNerm - Giấy ra viện NER Model
PhoBERT-based Named Entity Recognition model for Vietnamese Giấy ra viện documents.
⚠️ DISCLAIMER: SYNTHETIC / MOCKUP DATA
Model này được train hoàn toàn trên dữ liệu giả lập (synthetic/mockup data), KHÔNG sử dụng dữ liệu cá nhân thật.
- Tất cả dữ liệu training được sinh tự động bằng hệ thống template + generator
- Không sử dụng giấy tờ thật, thông tin cá nhân thật, hoặc dữ liệu thu thập từ người dùng
- Số định danh (ID, CCCD...) được sinh ngẫu nhiên, thiết kế để không trùng với dữ liệu thật
- Dữ liệu có inject nhiễu OCR (noise) để giả lập điều kiện thực tế
- Mục đích: nghiên cứu AI, Document AI, OCR/NER pipeline
- Không được sử dụng để giả mạo giấy tờ, tạo giấy tờ giả, lừa đảo hoặc gian lận
Model Description
This model is fine-tuned from vinai/phobert-base for token-level NER on Vietnamese administrative/medical documents. It extracts structured fields from OCR text output.
- Base model: vinai/phobert-base
- Task: Token Classification (NER)
- Language: Vietnamese (vi)
- Document type: Giấy ra viện
- Number of labels: 29
- Training data: Synthetic/Mockup (not real personal data)
Labels
B-admission_dateB-bhxh_codeB-departmentB-dept_mgmtB-diagnosisB-discharge_dateB-hospital_nameB-medical_codeB-notesB-patient_addressB-patient_dobB-patient_ethnicityB-patient_genderB-patient_nameB-patient_occupationB-treatment_methodI-admission_dateI-bhxh_codeI-departmentI-dept_mgmtI-diagnosisI-discharge_dateI-hospital_nameI-notesI-patient_addressI-patient_nameI-patient_occupationI-treatment_method
Usage
With VietNerm SDK
from vietnerm import VietNerm
ner = VietNerm(doc_type="giay_ra_vien", model_path="phatdatpq/phobert-giay_ra_vien-ner")
result = ner.extract("your document text here")
print(result)
With Transformers
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("phatdatpq/phobert-giay_ra_vien-ner")
model = AutoModelForTokenClassification.from_pretrained("phatdatpq/phobert-giay_ra_vien-ner")
text = "your document text here"
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=-1)
Training
- Dataset: Synthetically generated (mockup data) with OCR noise simulation
- Data source: Auto-generated from Jinja2 templates + random generators (no real personal data)
- Framework: HuggingFace Transformers + Trainer API
- Optimizer: AdamW (lr=2e-5)
- Epochs: 5-7 (with early stopping)
Ethical Use
This model is built for research and development purposes only:
- ✅ AI/NLP research
- ✅ Document AI development
- ✅ OCR/NER pipeline prototyping
- ✅ Educational purposes
- ❌ Forging documents
- ❌ Creating fake identity papers
- ❌ Fraud or deception
About VietNerm
VietNerm is a Document AI Factory for Vietnamese documents. It provides a complete pipeline from template-based synthetic data generation to model training and deployment.
- Repository: Devhub-Solutions/VietNerm
- Training dataset: ngocthanhdoan/vietnerm-giay_ra_vien-dataset
- SDK:
pip install vietnerm - License: MIT — Copyright (c) 2026 Devhub Solutions
- Downloads last month
- 8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for phatdatpq/phobert-giay_ra_vien-ner
Base model
vinai/phobert-base