The Construction of Corporate Texts
The construction of structured corpora is the foundation of Arabic processing, and it reflects the real-world language of the public sector. It includes:
- Civic data
- Anonymised service data
- Data on citizen inquiries
- Data on policy analysis
- Data on regulations
These data sets are characterised by dialects, spelling variations, service terminology, and non-standardised service phrases. Annotation protocols establish boundaries for named entities, as well as intent, sequence, and classification in the semantic fields of citizen-government interaction