Rizzo-AI-Academy/rizzo-pii
Local-first privacy guard: anonymize your documents before sharing with LLMs.
About Rizzo-AI-Academy/rizzo-pii
Rizzo-AI-Academy/rizzo-pii is an open-source project on GitHub, mainly written in Python. Local-first privacy guard: anonymize your documents before sharing with LLMs. It currently holds 1,094 stars and 80 forks with 77 open issues, and was last pushed on 2026-09-25 (repository created 2026-06-28).
Project Overview
Git Homed tracks it on the Today's Trending board, currently at rank #88 with 26 new stars today.
GitHub Repository Details
README
rizzo-pii
Local, reversible PII anonymization for Italian legal text
_Use frontier models without giving up your data._
📄 Read the full technical report (PDF) — model, dataset, method and experiments in detail
🪟 Windows installer · 🍎 macOS (Apple Silicon) · 🐧 Linux AppImage — all available now
rizzo-pii:0.3B is a lightweight, CPU-friendly, Italian-first token-classification
model (≈0.3B parameters, mmBERT / ModernBERT
backbone) that detects 22 categories of personal data — including the Italian-legal
identifiers (codice fiscale, partita IVA, dati catastali) that no other open model
covers — and drives a fully reversible anonymization workflow:
🔒 anonymize locally → 🏷️ placeholder + reversible local dictionary → ☁️ frontier LLM → 🔓 restore locally
It is built for law firms, accountants, notaries and anyone bound by the GDPR who wants to keep using ChatGPT / Claude / Gemini on sensitive documents without ever sending the real data out.
|
| ≈0.3B | ~0.5 GB | 22 | 0.989 | |:---:|:---:|:---:|:---:| | parameters (mmBERT-base) | RAM footprint, CPU | PII categories | micro-F1 (real IT validation) | The hedgehog mascot does one job: it grabs your document, blacks out every identifier, and stays inside the EU while doing it. |
---
The problem: convenience is leaking your data
People summarize contracts, draft replies and ask legal questions simply by pasting the document in. It is fast and useful — and it quietly moves enormous amounts of personal and confidential data off the user's device. Names, addresses, tax codes, IBANs, health details, unsigned-contract clauses: all of it crosses the network to servers the user does not control, where it may be logged, cached, retained or exposed in a breach. For a law firm or a hospital this is not hypothetical; under the GDPR it can be a direct compliance failure.
The intuitive fix is to stop sending data out and run an open model locally — but a frontier-grade open model is large and expensive to serve (€9,000–€10,000 of hardware), and the small models that fit a normal laptop are not in the same league on the hard tasks (legal reasoning, dense contracts, long official documents) — exactly where Italian professionals need the most help.
The trade-off we actually want: keep the frontier model and *remove the data from the equation*. Anonymize the document locally on a CPU, send only placeholders to the cloud, and restore the real values locally from the answer. The sensitive content never leaves the machine.
---
rizzo-pii in one picture
The workflow has three local steps and one remote step. Locally, rizzo-pii tags every span of
personal data and replaces each one with a stable, type-aware placeholder
([FULLNAME_1], [IBAN_1], [CF_1]), recording the mapping placeholder → real value in a
dictionary that stays on disk. Identical values share the same placeholder, so the frontier
model still sees a coherent text and can reason about it. The anonymized text is sent to
ChatGPT / Claude / Gemini; when the answer comes back, a local pass swaps the placeholders for the
true values. The cloud provider never receives a single real name, code or number.
Everything except the frontier query happens on the user's CPU; only placeholder text crosses the boundary, and the answer is re-identified locally.
---
Why this is different: privacy that is actually private
This is not "yet another PII detector". It is an architecture for using powerful models without surrendering data, built so the privacy guarantee is structural rather than a promise:
- The data never leaves the device. Detection and re-identification run locally on a CPU.
- GDPR by design. The workflow implements data minimization (Art. 5) almost literally:
- Aligned with the EU AI Act. Keeping personal data under local control and out of
- Accessible to everyone. The model is ≈0.3B parameters and runs on a CPU in well under 1 GB
- Reversible, not destructive. Classic redaction throws information away. rizzo-pii
How it compares
| Property | rizzo-pii:0.3B | OpenAI Privacy Filter | MS Presidio | |---|---|---|---| | Type | Dense encoder (mmBERT / ModernBERT) | Sparse MoE encoder | NER + rules pipeline | | Parameters | ≈0.3B dense (all active) | 1.5B total / ≈50M active | spaCy model + rules | | Memory to load | 0.5–1.2 GB | 1.5B params resident | Varies (spaCy) | | Runs on | CPU, under 1 GB RAM | On-device | CPU | | Categories | 22 (incl. IT-legal) | 8 generic | Configurable; EN defaults | | Italian CF / PIVA / catasto | Yes | No | Not by default | | Primary language | Italian (+7 more) | English | English | | Checksum validation | Yes (IBAN/CF/PIVA/card) | No | Some recognizers | | Reversible mapping | Yes (local dict) | Masking | Anonymization |
The differentiators are Italian-legal coverage, a smaller memory footprint, and a checksum-backed safety net (mod-97 for IBAN, Luhn for cards, the official CF/PIVA algorithms) that the larger generic models do not provide.
Concrete example. Take *"Il Sig. Mario Rossi, C.F. RSSMRA85H12F205Y, P.IVA 12345678903, è
titolare dell'immobile al Foglio 12, particella 345, sub. 6."* rizzo-pii tagsFULLNAME,CF,
PIVAandCATASTOand rewrites it as *"Il Sig. [FULLNAME_1], C.F. [CF_1], P.IVA [PIVA_1], è
titolare dell'immobile al [CATASTO_1]."* A generic English-first model has no label for the
fiscal code, the VAT number or the cadastral reference — the three most sensitive identifiers in
the sentence — and would leave them in the clear.
---
The taxonomy: 22 tags, and why
rizzo-pii predicts 22 entity types in BIO format (a B-/I- label per tag, plus O). The
raw datasets are left untouched; the mapping to these 22 tags is applied at load time through a
single TAG_MAP in train_pii.py, so the taxonomy can be changed in one place without
re-annotating anything. Details in docs/TASSONOMIA_TAG.md.
| Tag | Meaning | Example | Source |
|---|---|---|---|
| FULLNAME | Person name (incl. legal roles: judge, lawyer, parties, witness) | Mario Rossi | real+synth |
| AGE | Age | 45 anni | real |
| GENDER | Sex / gender | Femmina | real |
| DATE | Calendar date | 12/06/1985 | real+synth |
| TIME | Time of day | ore 15:30 | real |
| STREET | Street / square | Via Garibaldi | real+synth |
| BUILDINGNUM | Street number | 24 | real+synth |
| ZIPCODE | Postal code (CAP) | 00185 | real+synth |
| CITY | City | Milano | real+synth |
| PROVINCE | Province abbreviation | MI | synth |
| EMAIL | E-mail (incl. PEC) | [email protected] | real+synth |
| TELEPHONENUM | Phone number | +39 333 1234567 | real+synth |
| CF | Codice fiscale (personal tax code) | RSSMRA85H12F205Y | synth |
| PIVA | Partita IVA (VAT number) | 12345678903 | real+synth |
| ID_DOC | ID / passport / licence / social number | CA12345AB | real+synth |
| IBAN | IBAN / bank account | IT60X05428… | synth |
| CREDITCARDNUMBER | Credit-card number | 4111 1111 1111 1111 | real |
| AMOUNT | Money amount | € 12.500,00 | synth |
| TARGA | Vehicle plate | AB 123 CD | synth |
| ORG | Private company / firm / bank | Edilnord S.r.l. | synth |
| DOCID | Act identifier (RG, protocol, repertory, ruling) | 1234/2024 | synth |
| CATASTO | Cadastral data (sheet, parcel, sub.) | Foglio 12, part. 345 | synth |
Two design decisions stand out. First, legal roles collapse into FULLNAME — whether
"Mario Rossi" is the judge, the lawyer or a witness is not a property of the string; the role, if
needed, is recovered downstream as metadata. Second, raw types that mean the same thing are
merged: names + surnames → FULLNAME; PEC → EMAIL; TAXNUM → PIVA; ID card / passport /
licence / social number → ID_DOC; account number → IBAN. Honorifics (Dott., Avv.) and the
name of the court itself are dropped to O because they are not identifiers to mask.
The five Italian-legal tags (CF, PIVA, CATASTO, DOCID, PROVINCE) are the reason
rizzo-pii exists: they do not appear as labeled data in any public corpus, so they are created
through synthesis with mathematically valid checksums.
The app adds a 23rd tag, URL, handled by the regex net alone — the model is not trained on it.
It matches http(s)://…, www.… and bare domains on a closed TLD list; the closed list is what
keeps Italian legalese (p.iva, n.ro, S.r.l.) from being read as a domain, at the price of
letting an exotic TLD through.
---
Dataset & training
The model is fine-tuned on a multilingual pool of ≈745k labeled rows (Italian reinforced to
~45%) assembled from four sources — real (Ai4Privacy, DeepMount) and synthetic — all remapped to the
22 tags at load time. The synthetic part follows the "LLM author, code labeler" principle: an
LLM writes only Italian legal prose with placeholders ({SLOT}) and our code injects the real
values (CF/PIVA/IBAN with valid checksums), so BIO labels are exact, identifiers are valid by
construction, and no real personal data is ever produced by the LLM. The backbone is
mmBERT-base (ModernBERT architecture, native 8192-token context) and training runs in a single
epoch on one 16 GB consumer GPU.
📄 The full dataset composition, the synthesis method, the training recipe and the
experiments are described in the technical report. See also
docs/DATASET.md and CLAUDE.md for the operational details.
---
Results
Training was a single epoch (~26.6k steps over 744,912 rows). The loss fell from ≈6.2 to below 0.1 within the first few hundred steps, then settled into a clean, monotone, low regime; the validation loss decreased monotonically and was still falling when the epoch ended. Final training loss ≈0.003 and validation loss ≈0.006 are both very low and very close, so the model is not over-fitting — a second epoch would very likely push it lower still.
Left: training loss (smoothed zoom) — fast drop, then low and stable. Right: validation loss — monotone, still decreasing when training stopped.
On the 7,000-row held-out real Italian benchmark (validation_real.jsonl):
| 0.987 | 0.990 | 0.989 | 0.998 | |:---:|:---:|:---:|:---:| | micro precision | micro recall | micro F1 | token accuracy |
The unweighted per-tag mean (macro-F1) across all 22 tags is 0.987, and every one of the five Italian-legal identifiers scores a perfect 1.000.
Per-tag precision / recall / F1 (v1.2.0)
| Tag | Sup. | P | R | F1 | | Tag | Sup. | P | R | F1 | |---|---:|---:|---:|---:|---|---|---:|---:|---:|---:| | FULLNAME | 4390 | .989 | .990 | .990 | | GENDER | 472 | 1.00 | 1.00 | 1.00 | | CATASTO | 1200 | 1.00 | 1.00 | 1.00 | | PROVINCE | 400 | 1.00 | 1.00 | 1.00 | | CITY | 953 | .961 | .963 | .962 | | DOCID | 400 | 1.00 | 1.00 | 1.00 | | DATE | 922 | 1.00 | 1.00 | 1.00 | | CF | 400 | 1.00 | 1.00 | 1.00 | | TELEPHONENUM | 874 | 1.00 | 1.00 | 1.00 | | AGE | 385 | .979 | .977 | .978 | | ID_DOC | 800 | 1.00 | 1.00 | 1.00 | | ZIPCODE | 299 | .938 | .967 | .952 | | EMAIL | 748 | .999 | .999 | .999 | | IBAN | 278 | .996 | .996 | .996 | | TIME | 637 | .991 | .992 | .991 | | CREDITCARD | 257 | .919 | .973 | .945 | | STREET | 617 | .951 | .969 | .960 | | AMOUNT | 146 | 1.00 | .993 | .997 | | BUILDINGNUM | 594 | .969 | .958 | .964 | | ORG | 145 | .967 | 1.00 | .983 | | PIVA | 514 | 1.00 | 1.00 | 1.00 | | TARGA | 43 | 1.00 | 1.00 | 1.00 |
All five Italian-legal identifiers (CF, PIVA, CATASTO, DOCID, PROVINCE) score a perfect
1.000, as do ID_DOC, DATE, TELEPHONENUM, GENDER and TARGA. The remaining soft spots are
the open, high-variability classes (ZIPCODE, CREDITCARDNUMBER, STREET, CITY) and ORG,
which is exactly where a larger, better-balanced dataset would help.
---
Deployment: it runs on a normal computer
The released checkpoint is ~1.2 GB on disk in fp32 and runs comfortably on a CPU: quantized, its memory footprint is on the order of 0.5 GB, with no GPU required. That is the whole point — the privacy layer is cheap enough to run on the laptop the user already owns.
In production the neural model is never used alone. It is paired with a deterministic regex + checksum network for the structured identifiers (EMAIL, phone, IBAN, CF, PIVA, credit card, amount, plate), where IBAN/PIVA/card must pass their checksum (mod-97 / Luhn) to be accepted, and a valid checksum overrides the model. This eliminates the classic failure mode of a neural tagger fragmenting a long code, and gives mathematically certain detection for exactly the identifiers whose leakage is most damaging. The app adds the reversible layer (stable placeholders, downloadable local dictionary, a "restore" tab tolerant to markdown/format drift), chunking with overlap for long PDFs, and a colored per-tag UI.
| To use the model (inference) | To retrain the model |
|---|---|
| Any 64-bit CPU (no GPU) | A single 16 GB GPU is enough |
| 0.5–1.2 GB RAM for the model | Reference run: RTX 5060 Ti, ~2 h |
| Windows (installer) / Linux / macOS | PyTorch cu128 for Blackwell |
| Fully offline; no API key | ~745k rows, regenerable from scripts |
The desktop app Rizzo PII (Tauri) launches the Python/Flask backend as a bundled CPU "sidecar"; a CPU-only PyTorch build keeps it fully offline on Windows (WebView2), macOS and Linux. Packaging instructions in docs/BUILD.md.
⬇️ Download. Grab the ready-to-use build from the
Releases page — no Python or
setup required: a Windows installer (double-click), a macOS .dmg (Apple Silicon /
arm64 — signed & notarized by Apple, just open it), and a Linux AppImage (chmod +x then
run) are all available now.
---
Inference quickstart
Just want to anonymize documents? You do not need the dataset, the training scripts or a GPU — only the model. Three ways, all of them fully local. (Regenerating the data and retraining is a different flow: Quickstart, further down.)
Run with Docker
Nothing to install but Docker; the image carries the CPU dependencies and the model.
git clone https://github.com/Rizzo-AI-Academy/rizzo-pii
cd rizzo-pii
docker build -t rizzo-pii . # ~10 min, 2.65 GB image
docker run -d --name rizzo-pii \
-p 127.0.0.1:5005:5005 rizzo-pii # -> http://127.0.0.1:5005
docker logs -f rizzo-pii # startup / gunicorn logs
docker rm -f rizzo-pii # stop and remove
The image is self-contained: the CPU-only PyTorch stack and the model
(rizzoaiacademy/rizzo-pii-0.3B) are baked
in at build time, so docker build needs the network but docker run never does —
HF_HUB_OFFLINE=1 makes sure of it. Same deal as the desktop app: the documents never leave the
machine. Inside the container the server is gunicorn with 1 worker (the model is loaded once,
~1.2 GB) and 4 threads; /health answers 503 until the model is ready, which is what the
HEALTHCHECK polls — the first request is served when docker ps says healthy.
Publish the port as 127.0.0.1:5005:5005, not 5005:5005, unless you actually mean to expose the
service to your LAN — inside the container the bind is 0.0.0.0 on purpose, and the network
boundary is Docker's job.
Useful knobs (all optional):
# a different port, tags left in the clear, irreversible anonymization
docker run -d -p 127.0.0.1:8080:8080 -e PII_PORT=8080 \
-e PII_EXCLUDE_TAGS=AGE,GENDER -e PII_MAPPING=0 rizzo-pii
The image is CPU-only, which is the intended deployment (see the table above); torch and
transformersare pinned to the versions it was verified with. A GPU build would need thecu128
wheels and the NVIDIA container runtime. Note that Dockerfile.linux is a different thing: it is
the reproducible build environment for the Linux .deb/AppImage bundles, described in
docs/BUILD.md.
Run the web app from source
Same server, without Docker — this is what the desktop app runs internally.
git clone https://github.com/Rizzo-AI-Academy/rizzo-pii
cd rizzo-pii
python -m venv .venv && source .venv/bin/activate # Windows: .\.venv\Scripts\Activate.ps1
pip install -r requirements.txt # CPU torch is fine, no GPU needed
the model (~1.2 GB) into the folder the app looks for. The folder name carries the
version: app.py pins it in APP_MODEL_VERSION, so revision and directory must match.
hf download rizzoaiacademy/rizzo-pii-0.3B --revision v1.5.0 \
--local-dir models/rizzo-pii-0.3B-v1.5.0
python src/app/app.py # -> http://127.0.0.1:5005
python src/app/app.py --port 8080 --exclude-tags AGE,GENDER --no-mapping # same knobs as above
One document, from the CLI
No server at all: entities and anonymized text printed to stdout.
PII_MODEL_DIR=models/rizzo-pii-0.3B-v1.5.0 \
python src/training/test_pii.py "Mi chiamo Mario Rossi, IBAN IT60X0542811101000000123456"
Call it over HTTP
Whichever way you started it, the process is a plain HTTP service — see
The local HTTP API below for /analyze, /pdf, /preview and the rest.
curl localhost:5005/health # 200 = model loaded and ready
curl -X POST localhost:5005/analyze -H 'Content-Type: application/json' \
-d '{"text": "Mario Rossi, CF RSSMRA85M01H501Q"}'
---
Quickstart
Prerequisites and critical environment constraints (Blackwell GPU, torch cu128, etc.) are in
CLAUDE.md. The dataset/raw/ sources are downloaded from Hugging Face (see the
hf download commands in CLAUDE.md). All scripts force UTF-8 and resolve their paths from
__file__, so they run from any working directory.
💡 Just want to use the app? You don't need any of this — download the ready-to-use build
(Windows / macOS / Linux) from the
Releases page, or run the web app
with Docker / from source: Inference quickstart. The steps
below are for developers who want to regenerate the data and retrain the model.
0) Install
git clone https://github.com/Rizzo-AI-Academy/rizzo-pii
cd rizzo-pii
python -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -r requirements.txt # NVIDIA Blackwell? install torch cu128 first — see requirements.txt
copy .env.example .env # optional: add W&B / Gemini keys
1) Generate the data
python src/data_pipeline/llm_template_bank.py --per-type 5 --append # (opt.) legal templates via Gemini
python src/data_pipeline/generate_synthetic_pii.py -n 200000 --out dataset/synthetic/synthetic_pii_it_200k.jsonl
python src/data_pipeline/augment_real_pii.py -n 40000 --out dataset/synthetic/synthetic_pii_it_realaug.jsonl
python src/data_pipeline/prepare_deepmount.py # requires HF login
python src/data_pipeline/build_validation.py # real validation (7k, it)
python src/data_pipeline/build_subset.py # 10k/5k subsets for smoke tests
2) Train
# fast smoke test / tuning on the subset (~3 min) -> experiments/subset_smoke/
python src/training/train_pii.py --type subset
full run on the whole dataset -> models/rizzo-pii-0.3B-v{VERSION}/ + experiments/full_run_v{VERSION}/
python src/training/train_pii.py --type full
python src/training/train_pii.py --type full --version 1.2.0 # or an explicit version
3) Use the model
python src/training/test_pii.py "Mi chiamo Mario Rossi, IBAN IT60X0542811101000000123456"
python src/app/app.py # http://127.0.0.1:5005 (paste text or upload a PDF)
The web app assigns every PII a reversible ID ([FULLNAME_1], [IBAN_1]…) plus a local
dictionary, pairing the model with the regex/checksum net. You copy the anonymized text into an LLM
and restore the real values from the response. Input: pasted text, a PDF, or a
.md / .txt file.
The local HTTP API
The same process is a plain HTTP service, so you can use it as a sidecar in a fail-closed pipeline:
```bash curl localhost:5005/health