AASIST
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
The last open version came out in Oct 2021. The family has not been updated for a long time: the model still works, but do not expect fixes or new sizes.
- Developer
- NAVER Clova AI Research and EURECOM, South Korea
- First release
- Oct 2021
- Latest release
- Oct 2021
- Sizes
- weight files of 0.4 and 1.3 MB
- License
- Commercial use allowedMIT
- Running
- On your own serverAlso runs without a GPU
- Industries
- Security, Finance, Customer support, Public sector
What it does
- Voice check during phone authentication
- Filtering replays and synthesis in a voice menu
- A baseline when comparing voice detectors
- Fine-tuning for your own communication channel
Where it is used
Hardware requirements
Versions
- AASIST-L (облегчённая)
- AASIST
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can AASIST be used in a commercial project?
Yes. License: MIT. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does AASIST need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does AASIST support Russian?
Language does not matter for this model: it does not work with text.
Where can I download AASIST and what does it cost?
The AASIST weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
DetailsDeepfake detectionAntiDeepfake (NII)National Institute of Informatics, Yamagishi Lab · JapanNon-commercial onlySeven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
DetailsVoice: speakers and soundWeSpeakerWeNet community · ChinaCommercial use allowedA set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
DetailsSource: github.com/clovaai/aasist. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


