A Hybrid Approach to the Automatic Detection of Personal Data in Latvian-Language Texts
Henrihs Gorskis, Jūlija Strebko, Jurijs Korņijenko, Vitaly Zabiniako et al. — Future Internet
How it connects
Related across sources
Full text
In the modern world, hybrid and fully remote work formats are becoming increasingly widespread. The volume of digital communication continues to grow, and the need to exchange documents and information through email, corporate messengers (such as Microsoft Teams), and collaborative workspaces is rising. Personal data is often involved in these exchanges, which increases the risk of unintentional disclosure. To comply with GDPR requirements and ensure information security, it is necessary to implement methods for the automatic detection of personal data. This task is particularly relevant for low-resource languages, such as Latvian, for which existing tools often operate with limited accuracy and efficiency. This work proposes a hybrid approach to the automatic detection of personal data in Latvian texts from Microsoft Teams messages, emails, and documents, combining a transformer-based NER model with rule-based detection of structured identifiers. The approach builds on a multilingual NER model, supplements it with Latvian-specific rules and structured-identifier detectors, and demonstrates the potential of adapted solutions to improve the accuracy and robustness of personal-data detection in real-world workflows.