Skip to content
Literature · Future Internet EN LLM context A cited markdown file you can paste into your AI assistant (ChatGPT, Claude, a RAG or project knowledge base) to ground it in this document. Contains: this document’s text, its sections with their topics, and the full text of every law provision it applies. Everything links back to its source on overview.legal — legal information, not advice.

A Hybrid Approach to the Automatic Detection of Personal Data in Latvian-Language Texts

Henrihs Gorskis, Jūlija Strebko, Jurijs Korņijenko, Vitaly Zabiniako et al. — Future Internet

Future Internet
DOI

How it connects

Full text

In the modern world, hybrid and fully remote work formats are becoming increasingly widespread. The volume of digital communication continues to grow, and the need to exchange documents and information through email, corporate messengers (such as Microsoft Teams), and collaborative workspaces is rising. Personal data is often involved in these exchanges, which increases the risk of unintentional disclosure. To comply with GDPR requirements and ensure information security, it is necessary to implement methods for the automatic detection of personal data. This task is particularly relevant for low-resource languages, such as Latvian, for which existing tools often operate with limited accuracy and efficiency. This work proposes a hybrid approach to the automatic detection of personal data in Latvian texts from Microsoft Teams messages, emails, and documents, combining a transformer-based NER model with rule-based detection of structured identifiers. The approach builds on a multilingual NER model, supplements it with Latvian-specific rules and structured-identifier detectors, and demonstrates the potential of adapted solutions to improve the accuracy and robustness of personal-data detection in real-world workflows.

Similar Content