OccCANINE
Turns a free-form occupation description into a standard HISCO code in 13 languages. Built for historical archives, but also useful for cleaning up job title reference lists. A human makes the decision about a candidate; automatic screening without review must not be used.
The last open version came out in Apr 2024. The family has not been updated for a long time: the model still works, but do not expect fixes or new sizes.
- Developer
- University of Southern Denmark, Denmark
- First release
- Feb 2024
- Latest release
- Apr 2024
- Sizes
- based on CANINE-s, size not stated on the model card
- License
- Commercial use allowedApache 2.0
- Russian
- Not supported
- Running
- On your own serverAlso runs without a GPU
- Industries
- HR, Science and research, Public sector
What it does
- Mapping mixed occupation names onto a single code
- Processing archives of HR and statistical data
- Preparing data for reporting
Where it is used
Hardware requirements
Versions
- OccCANINE
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can OccCANINE be used in a commercial project?
Yes. License: Apache 2.0. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does OccCANINE need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does OccCANINE support Russian?
No. The developer documentation lists the supported languages and Russian is not among them.
Where can I download OccCANINE and what does it cost?
The OccCANINE weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
A research line of models for labour market texts: trained on job postings and the European ESCO occupation taxonomy, they pull skills and requirements out of vacancies. A human makes the decision about a candidate; automatic screening without review must not be used.
DetailsSearch and RAGCareerBERTThe authors of the CareerBERT paper, German universities · GermanyCommercial use with conditionsA German-language model that matches a resume to occupations from the European ESCO reference list and suggests suitable directions. A human makes the decision about a candidate; automatic screening without review must not be used.
DetailsText analysisXLM-RoBERTaMeta · USACommercial use allowedA classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.
DetailsSource: huggingface.co/Christianvedel/OccCANINE. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


