Back to BlogGDPR & Compliance

NAIH Hungary: AI Governance & DPA Rules

NAIH requires DPIAs for all AI systems processing personal data. Hungarian NER accuracy is 67% — well below the EU 82% average.

March 6, 20268 minute read
Hungary NAIHAI GDPR complianceTAJ-szám detectionCentral Europe DPAHungarian data protection

NAIH Hungary: AI Governance and DPA Rules

Hungary's data body is NAIH — Nemzeti Adatvédelmi és Információszabadság Hatóság. The authority has issued the most detailed AI guidance of any Central European DPA. In 2024 it issued 38 enforcement decisions. It also published rules requiring a DPIA for every AI system that handles personal data. These rules go further than the GDPR baseline.

NAIH's AI Enforcement Rules

Most EU DPAs publish broad AI guidance. Hungary's DPA went further. Its 2024 guidance is operationally specific.

DPIAs required for all AI systems: Every AI system that touches personal data needs a DPIA first. The regulator requires this before deployment. This applies even when the processing is not "high-risk" under GDPR Article 35. That is stricter than the GDPR's own risk-based approach.

What a NAIH DPIA must include:

  • A technical description of the AI model's data inputs and outputs
  • Evidence that training data was anonymized or had a valid legal basis
  • An assessment of algorithmic discrimination risk
  • A human review step for automated decisions
  • A retention and deletion schedule for AI-processed data

Annual review: The authority requires DPIAs to be updated each year. This applies when an AI system is retrained or significantly changed.

Hungary handled over 890,000 GDPR data requests in 2024. That is a large volume for a country of 10 million. It signals active rights use and real pressure on compliance teams.

The NER Accuracy Gap

The authority's 2024 review tested NER models on Hungarian text. They scored only 67% accuracy. The EU average is 82%. That 15-point gap has real compliance costs.

Hungarian is an agglutinative language. It builds words through many suffixes. Names, addresses, and IDs in Hungarian look very different from data in English or German. Tools trained on those languages miss a large share of personal data in Hungarian. See our multilingual PII detection guide for how this gap affects GDPR compliance across languages.

The regulator found that generic NLP tools miss the TAJ-szám in 61% of documents. Format variation and no checksum support are the main causes.

Hungarian National Identifiers

Teams processing documents in Hungary must detect these ID types accurately. See our EU national tax ID detection guide for full EU coverage context.

TAJ-szám (Társadalombiztosítási Azonosító Jel): A 9-digit social security number. It appears in health, benefit, and pension records. Validation uses a weighted checksum set by the Social Insurance authority.

Adóazonosító jel: A 10-digit personal tax ID. The format is an 8-digit core plus 2 check digits. It appears in payroll, tax filings, and employment contracts.

Személyi igazolvány number: The national ID card number. Format and check digit rules follow the issuing authority.

Útlevél szám: The passport number. Format and check digit also follow rules set by the issuing authority.

The Ügyfélkapu Context

Hungary runs most public services through one platform — Ügyfélkapu (Client Gateway). Over 4 million citizens use it for tax, benefits, healthcare, and licensing. Private firms connect to Ügyfélkapu for payroll, benefits, or identity checks. Those firms process the same identifiers in a regulated context.

The authority has found that these firms often use international PII tools. Most of those tools lack support for the identifiers above. That leads to missed data and direct compliance risk.

EU AI Act Overlap

Hungary was early to fold AI Act rules into DPA guidance. The regulator's stance is clear.

High-risk AI systems are listed in AI Act Annex III. These cover jobs, credit scoring, and essential services. They require both AI Act conformity assessment and a NAIH DPIA.

General-purpose AI models that process data of people in Hungary also need a NAIH DPIA. This applies even when the model is not listed as high-risk under the AI Act.

For teams deploying AI in Hungary, the core checklist has three items. Complete a NAIH DPIA before launch. Verify that your NER tool covers the entities above in Hungarian text. Confirm TAJ-szám and adóazonosító jel detection with checksum validation.

When This Approach Has Limits

Pairing a NAIH DPIA with Hungarian-aware NER and checksum-validated identifier detection is the correct shape for compliance here, but limits remain worth stating plainly.

Hungarian is where generic detection breaks down. The authority's own review measured NER at 67 percent on Hungarian text against an 82 percent EU average, and found generic tools miss the TAJ-szám in 61 percent of documents. Hungarian is agglutinative: names, addresses, and inflected forms attach stacks of suffixes that Latin-trained models were never built to segment. Multilingual detection accuracy varies sharply by language, and Hungarian sits well below the mean. The 9-digit TAJ-szám and 10-digit adóazonosító jel each need their own checksum logic plus held-out testing on real Hungarian documents; without that, the residual false-negative rate is the number that actually governs your exposure.

A DPIA is governance, not a detection guarantee. NAIH requires a DPIA for every AI system touching personal data, including a technical description, evidence training data was anonymized, and an annual review. A completed DPIA documents intent and process; it does not prove your pipeline caught every TAJ-szám in a given corpus. The regulator audits the organization's whole posture, including the human review step for automated decisions, not the existence of one tool. Detection quality and documented governance are separate obligations that each have to hold on their own.

Removing identifiers may leave pseudonymized data in scope. Strip the TAJ-szám and adóazonosító jel and quasi-identifiers can still re-identify people, particularly through Ügyfélkapu-linked records where employment, benefit, and health context cluster around the same individuals. Pseudonymized data remains within GDPR scope, and Hungary's 890,000-plus annual rights requests mean re-identification risk is something individuals actively test. Verify the output resists linkage rather than assuming direct-identifier removal was sufficient.

Sources

Limitations / When this doesn't apply

Hungarian is where generic detection breaks down. NAIH's own review measured NER at 67% on Hungarian against an 82% EU average, and found generic tools miss the TAJ-szám in 61% of documents. Hungarian is agglutinative — names and inflected forms stack suffixes Latin-trained models were never built to segment. The 9-digit TAJ-szám and 10-digit adóazonosító jel each need their own checksum logic plus held-out testing on real Hungarian documents; without that, the residual false-negative rate governs your exposure.

A DPIA is governance, not a detection guarantee. NAIH requires a DPIA for every AI system touching personal data, with a technical description, evidence training data was anonymized, and an annual review — none of which proves your pipeline caught every TAJ-szám. The regulator audits the whole posture, including the human-review step for automated decisions.

Removing the TAJ-szám and adóazonosító jel may still leave pseudonymized data in scope, since quasi-identifiers re-identify people, particularly across Ügyfélkapu-linked employment, benefit, and health records. This is educational guidance on evolving Hungarian guidance, not legal advice or a substitute for counsel.

Ready to protect your data?

Start anonymizing PII with 285+ entity types across 48 languages.

About this page

We update this page when our platform or the law changes.

Read our founder note for how we work.

Each change shows up in the timestamp at the top.

We follow these rules

  • GDPR (EU 2016/679).
  • ISO/IEC 27001:2022.
  • NIS2 (EU 2022/2555).
  • HIPAA safe harbor under 45 CFR § 164.514(b)(2).

Our promise

We do not sell your data.

We do not train models on your text.

We store your files in Germany.

You can delete your account at any time.

You own your work.

Where we run

Our company HQ is in Saarbrücken, Germany. Our servers run in Hetzner's Falkenstein datacenter.

Hetzner holds ISO 27001 certification.

All data stays in the EU.

Backups run every day.

Need help?

Email support@anonym.legal.

We reply within one business day.

How we test

We run a full check suite on every release.

Each surface gets its own sweep script and report.

Human reviewers spot-check the output each week.

We track recall and precision on a labelled set.

Bad runs block the deploy.

What we never do

  • We never sell your information to third parties.
  • We never train models on what you upload.
  • We never keep your work after you delete it.
  • We never share keys with any outside firm.
  • We never run ads inside the product.

Plans in plain words

We sell credits, not seats.

One credit covers one short job.

Long jobs use a few credits each.

You can top up at any time.

Unused credits roll over each month.

Read the plans page for current rates.

Who built this

A small team of engineers and lawyers built this.

We ship from Europe and work in the open.

Our founder note spells out why we started.

Where to start

How the parts fit

A browser add-on cleans text inside Chrome.

A Word plug-in handles drafts in Office.

A small desktop tool works on whole folders.

An agent protocol link feeds large models safely.

All four share one core engine and one rule set.

Words from our team

We started this work after a lunch about cookies.

One friend kept getting odd ads on her phone.

We asked why a court file leaked through a draft.

We sketched the first build on a napkin that week.

By month three we had a tiny demo for a friend.

She used it on her first case the next day.

Common questions we hear

Can the tool read scanned PDFs? Yes, with OCR.

Does it work on long files? Yes, in small chunks.

Can I roll my own rule set? Yes, save it as a preset.

Does it run offline? The desktop build runs offline.

Do you keep my files? No, the cloud build wipes after each run.

Will it learn from my work? No, we never train on inputs.

A short tour of the workflow

Upload a file or paste a snippet of prose.

Pick the entities you want gone from the draft.

Choose a method: replace, mask, hash, encrypt, or redact.

Press run and watch the side panel show each hit.

Skim the result and tweak any rule that misfired.

Save the cleaned file or send it to a teammate.