The base model ships with 8 coarse PII categories and English-only training. This
model trades that for a
6.75× more granular vocabulary
spanning identity,
contact, address, financial, vehicle, digital, and crypto labels — all evaluated
across 16 languages.
Family at a glance.
Same architecture, three runtimes:
PyTorch (this repo)
— CPU + CUDA, anywhere transformers runs.
OpenMed gives you
extract_pii()
/
deidentify()
with built-in BIOES Viterbi
decoding, span refinement, and a Faker-backed obfuscation engine. Same call
on every host — Apple Silicon picks up MLX automatically; everywhere else uses
this PyTorch checkpoint.
pip install -U "openmed[hf]"
from openmed import extract_pii, deidentify
text = (
"Patient Sarah Johnson (DOB 03/15/1985), MRN 4872910, ""phone 415-555-0123, email [email protected]."
)
# Extract grouped entity spans
result = extract_pii(text, model_name="OpenMed/privacy-filter-multilingual")
for ent in result.entities:
print(f"{ent.label:30s}{ent.text!r} conf={ent.confidence:.2f}")
# De-identify with any of the supported methods
masked = deidentify(text, method="mask", model_name="OpenMed/privacy-filter-multilingual")
removed = deidentify(text, method="remove", model_name="OpenMed/privacy-filter-multilingual")
hashed = deidentify(text, method="hash", model_name="OpenMed/privacy-filter-multilingual")
# Faker-backed locale-aware obfuscation, deterministic with consistent=True+seed
fake = deidentify(
text,
method="replace",
model_name="OpenMed/privacy-filter-multilingual",
consistent=True,
seed=42,
)
print(fake.deidentified_text)
OpenMed/privacy-filter-multilingual-mlx*
model names also work in the same
extract_pii()
/
deidentify()
calls — on a non-Apple-Silicon host they
automatically fall back to
this PyTorch checkpoint
with a one-time warning.
So you can ship MLX names in code and still run on Linux/Windows.
The OpenMed wrapper passes
trust_remote_code=True
for you, runs the model's
own BIOES Viterbi decoder, and skips OpenMed's regex smart-merging (the model
already produces clean spans).
The output space is
O
plus
B-
,
I-
,
E-
,
S-
for each of the 54 categories
(4 × 54 + 1 = 217). The
id2label
mapping is shipped with the model.
Limitations & intended use
Multilingual but uneven.
Strongest on languages with rich PII training
data (German, Spanish, French, Italian, Hindi, Telugu, English). CJK languages
(Japanese, Korean, Chinese) and some morphologically-marked low-resource
languages remain the main bottleneck on the current training mix.
Synthetic training data.
The AI4Privacy datasets are template-synthesized;
real clinical notes, legal documents, and web text may show different
surface forms. For high-stakes deployments, collect a domain-specific eval
set and re-calibrate thresholds.
Not a substitute for legal compliance review.
Use alongside a governance
layer (human review, deterministic regex pre-filters, etc.).
Not a clinical PHI model.
Healthcare-specific PHI and clinical entity
training is planned as a separate branch.
Head initialization
:
opf
's default "copy-from-matching-base" head init.
Of the 217 new BIOES classes, the few with exact base-vocabulary matches
(
O
,
B/I/E/S-account_name
, etc.) were copied directly; the rest were copied
from semantically-adjacent coarse rows and fine-tuned end-to-end.
Router
: base model has 128 MoE experts per layer with top-4 routing.
Routers were kept trainable during full fine-tuning; no collapse was observed.
Credits & Acknowledgements
This model wouldn't exist without two open-source releases — sincere thanks
to both teams:
OpenAI
for
open-sourcing the Privacy Filter
(architecture, modeling code, and
opf
training/eval CLI). Everything in
this repo is a fine-tune on top of that release.
privacy-filter-multilingual huggingface.co is an AI model on huggingface.co that provides privacy-filter-multilingual's model effect (), which can be used instantly with this OpenMed privacy-filter-multilingual model. huggingface.co supports a free trial of the privacy-filter-multilingual model, and also provides paid use of the privacy-filter-multilingual. Support call privacy-filter-multilingual model through api, including Node.js, Python, http.
privacy-filter-multilingual huggingface.co is an online trial and call api platform, which integrates privacy-filter-multilingual's modeling effects, including api services, and provides a free online trial of privacy-filter-multilingual, you can try privacy-filter-multilingual online for free by clicking the link below.
OpenMed privacy-filter-multilingual online free url in huggingface.co:
privacy-filter-multilingual is an open source model from GitHub that offers a free installation service, and any user can find privacy-filter-multilingual on GitHub to install. At the same time, huggingface.co provides the effect of privacy-filter-multilingual install, users can directly use privacy-filter-multilingual installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
privacy-filter-multilingual install url in huggingface.co: