Voice agents attract one persistent misconception: that if you do not store recordings, there is nothing to worry about. In practice the transcript, the call metadata and the audio are each potentially personal data. Not recording does not remove your obligations; it only removes your evidence.
What follows is a practical order of work under KVKK. It is not legal advice, but it should shorten your first meeting with the compliance team.
Which category does voice data fall into?
A call recording is personal data as a rule. If you extract a template from the physical characteristics of a voice in order to verify who is speaking, you are processing biometric data, which brings a stricter regime. When a single call contains health, religious, union or criminal-record details, the weight of the file changes entirely.
- Audio recording: personal data, requiring retention limits and access control
- Transcript: carries every piece of personal data spoken on the call
- Voice template or embedding: can fall under biometric data
- Emotion and tone analysis: treated as behavioural data
- Call metadata: number, time, duration and routing history
That classification drives design. Verifying identity by voice template is technically possible, but do it only where the conditions for biometric data are genuinely met. For most operations a one-time code or a knowledge question is enough, and it removes a large amount of compliance load.
Notice and explicit consent are separate steps
The short announcement at the start of a call is notice: it states that the call is recorded, why, and for how long the recording is kept. Explicit consent does not replace it and must be obtained separately. Folding the two into one sentence weakens both comprehension and your evidence.
Notice informs. Explicit consent authorises. Neither replaces the other.
You also need to keep the consent record itself: who agreed, to which text, on what date, through which channel. In an audit that record is the first document anyone asks for. Verbal consent given mid-call should be timestamped and mapped to the transcript.
Retention and deletion
- Retention follows the purpose, not habit
- When the period ends, the recording is deleted or anonymised automatically
- A deletion request covers every copy, including backups
- Access is role-based and every access is logged
The most common gap is a retention period that exists on paper but not in the system. Do not leave deletion to somebody remembering to do it. Implement it as a rule and test that it fires once a quarter.
Anonymisation is harder than it sounds. Separating a speaker from an audio file completely is difficult; even after names, addresses and numbers are stripped from the transcript, the voice itself may still identify someone. Treat anonymisation as a specialist task, not as a substitute for deletion.
Cross-border transfer and vendor selection
If your agent runs in the cloud, audio and transcripts are most likely processed on servers abroad. The contract should be explicit about the transfer mechanism, which data sits in which countries, and who the sub-processors are.
Three questions for any vendor: where do you process and store the data? Do we set the retention period? Within how many days do you close a deletion request, and with what evidence? If you cannot get written answers, your compliance file has a hole in it.
In short: classify voice data on day one, keep notice and consent apart, enforce retention in the system rather than in a policy document, and document the transfer chain. Those four steps cover both an audit and your customers' questions.


