Acoustic monitoring2019–2026
Google · USA
A tiny model covering 521 sound events, including alarms, breaking glass and screams: it fits on a microcontroller and runs without a GPU. A microphone can capture people voices, which is personal data, so check the procedure with a lawyer.
- Fast sound labelling right on the device
- Detecting alarm sounds on a site
- Picking interesting fragments out of a continuous recording
- Sizes
- 3.7M
- Hardware
- from: Laptop
Acoustic monitoring2021–2026
MLCommons · USA
A reference autoencoder for finding anomalies in machine sound: it learns only from recordings of a healthy unit and flags deviations. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- Listening to a machine tool, pump or conveyor
- Flagging deviations from a unit usual noise
- Training on your own recordings of healthy equipment
- Sizes
- a tiny fully connected autoencoder sized for microcontrollers
- Hardware
- from: Laptop
Acoustic monitoring2023–2026
Cornell Lab of Ornithology and Chemnitz University of Technology · USA and Germany
Recognition of more than 6000 bird species by voice, the baseline tool for acoustic monitoring of an area. A microphone on site also records people voices, which is personal data, so check the procedure with a lawyer.
- Long-term monitoring of the sound background of a site
- Assessing biodiversity for an environmental review
- Selecting events from round-the-clock recordings
- Sizes
- a compact EfficientNet-B0 based model
- Hardware
- from: Laptop
Acoustic monitoring2019–2026
University of Surrey · UK
The classic set of convolutional networks that label sound across the 527 AudioSet categories, from machinery noise and alarms to breaking glass and screams. The system does not make a diagnosis, it gives you a reason to check the unit before it fails. A microphone can also capture people voices, which is personal data, so check the procedure with a lawyer.
- Labelling what is happening in a microphone recording
- Detecting alarms, breaking glass, screams
- A base for your own model tuned to one shop floor
- Sizes
- from 5.9M (CNN6) to 81.9M (CNN14)
- Hardware
- from: Laptop
Acoustic monitoring2023–2026
Xiaomi · China
Compact sound-labelling models from 5.5M to 86M, with ONNX and INT8 builds that run on an ordinary CPU and on a board next to the equipment. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- Labelling sound on modest hardware and on site
- Detecting alarm sounds and abnormal noise
- Fine-tuning for your own set of equipment sounds
- Sizes
- 5.5M - 86M
- Hardware
- from: Laptop
Acoustic monitoring2024–2026
Xiaomi · China
A general-purpose audio encoder trained on 272 thousand hours of speech, music and noise: it gives features on top of which you train your own classifier of abnormal sounds. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- Audio features for your own anomaly model
- Finding similar fragments in a recording archive
- Fine-tuning for the sounds of one production line
- Sizes
- 86M - 1.2B
- Hardware
- from: Laptop
Acoustic monitoring2025
Carnegie Mellon University and co-authors (the ESPnet project) · USA
A fully open reproduction of BEATs: code, training recipes and weights are all published. Teams pick it when they need a transparent base for their own sound model. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- A transparent base for your own audio model
- Fine-tuning on your own equipment recordings
- Labelling sound events
- Sizes
- base and large versions, exact sizes are listed on the model cards
- Hardware
- from: Laptop
Acoustic monitoring2025
Google DeepMind · USA
A model for acoustic monitoring of nature: it recognises about 15,000 species and provides features for your own tasks, and the license allows commercial use. A microphone on site also records people voices, which is personal data, so check the procedure with a lawyer.
- Monitoring the sound background of a site or water area
- Environmental audio features for your own model
- Selecting events from round-the-clock recordings
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Acoustic monitoring2024–2025
Shanghai Jiao Tong University and Peng Cheng Laboratory · China
A self-supervised sound understanding model that is markedly cheaper to train on your own data than its predecessors. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- Training your own audio model on modest hardware
- Audio features for spotting abnormal operating modes
- Labelling sound events on a site
- Sizes
- 90M (base) and 309M (large)
- Hardware
- from: Laptop
Acoustic monitoring2024–2025
MIT CSAIL and MIT-IBM Watson AI Lab · USA
A sound understanding model built on state spaces instead of a transformer: it handles recordings hours long, which suits continuous listening to a production line. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- Processing long continuous recordings
- Finding a rare event across a multi-hour shift
- Labelling sound events on a site
- Sizes
- 30M (small) and 49M (medium)
- Hardware
- from: Laptop
Acoustic monitoring2024–2025
Google · USA
A model for health sounds: coughing, breathing, throat clearing. It provides features on top of which a researcher builds their own model and draws no conclusions about illness itself. It does not replace a doctor; decisions are made by a specialist. Voice recordings are personal data, so check the procedure with a lawyer.
- Features of cough and breathing sounds for your own model
- Research projects in health acoustics
- Selecting cough and breathing fragments from a recording
- Sizes
- a ViT-Large class model, trained on more than 300 million two-second clips
- Hardware
- from: Laptop
Acoustic monitoringNot maintained2022
Microsoft · USA
A self-supervised sound understanding model that became the base for many current systems: teams take it as a starting point and fine-tune it on their own equipment sounds. The system does not make a diagnosis, it gives you a reason to check the unit before it fails.
- A base for your own sound understanding model
- Audio features for spotting abnormal operating modes
- Labelling sound events on a site
- Sizes
- a ViT-base class model, around 90M parameters
- Hardware
- from: Laptop