TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
Image generationRU2022–2026
Sber (Kandinsky Lab) · Russia
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
- Images from Russian-language descriptions
- Short promo videos from text or a photo
- Instruction-based image editing
- Sizes
- 2B – 19B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Speech to textRUGGUF2026
Alibaba (Qwen) · China
Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
- Transcribing calls and meetings
- Video subtitles
- Multilingual recognition
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
AvatarsGGUF2025–2026
Meituan · China
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
- News or course presenter videos
- Promo videos with a talking character
- Singing and voice-over
- Sizes
- based on LongCat-Video 13.6B
- Hardware
- from: 1 GPU
MedicineOllama2023–2026
EPFL · Switzerland
Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.
- Answering staff questions based on clinical guidelines
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 2B – 70B
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
Text to speechRU2025–2026
Supertone · South Korea
Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.
- Voicing voice bot replies on an ordinary server
- Voiceover in offline and mobile apps
- Reading texts and notifications aloud
- Sizes
- about 99M
- Hardware
- from: Laptop
Avatars2024–2026
Fudan University · China
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
- Presenter video from a photo and audio
- Long training videos
- Live avatar
- Sizes
- about 1B – 5B
- Hardware
- from: 1 GPU
3D2025–2026
Tencent · China
Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.
- 3D scenes for games and virtual tours
- Backgrounds and environments for video production
- Draft locations for simulations
- Sizes
- set of several models
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
Math and reasoningGGUF2025–2026
Princeton University · USA
Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.
- Formal verification of mathematical workings
- Verifying code correctness
- Training
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Text to speechRU2026
k2-fsa (Next-gen Kaldi) · China
Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.
- Voiceover in rare languages
- Voice cloning from a sample
- Research and prototypes of multilingual voiceover
- Sizes
- 0.6B
- Hardware
- from: Laptop
Video2025–2026
Skywork AI (Kunlun) · China
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
- Game world prototypes without an engine
- Interactive demos and simulations
- Generating data to train agents
- Sizes
- 1.8B – 17B
- Hardware
- from: 1 GPU
Avatars2026
SII-GAIR and Sand.ai · China
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
- Sizes
- 15B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2026
LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France
A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.
- Checking the logic of reasoning and workings
- Step-by-step explanations of solutions
- Training and olympiad preparation
- Sizes
- 4B
- Hardware
- from: Laptop
Image generation2026
Zhipu AI (Z.ai) · China
A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.
- Posters and banners with text
- Infographics
- Illustrations for presentations
- Sizes
- 9B + 7B
- Hardware
- from: 1 GPU
Avatars2024–2026
Ant Group · China
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
- Presenter video from a photo and voice
- Avatar with gestures for presentations
- Voiced characters
- Sizes
- up to 1.3B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Alibaba (Quark) · China
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
- Live avatar for customer dialogue
- Endless broadcasts with a presenter
- Interactive characters
- Sizes
- 14B
- Hardware
- from: 1 GPU
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
Voice assistantsGGUF2026
NVIDIA · USA
A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.
- A voice assistant with a set role
- A conversation simulator for staff training
- Voice interfaces without delay
- Sizes
- 7B
- Hardware
- from: 1 GPU
Speech to textRU2025
Meta · USA
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
- Transcription in rare and local languages
- Digitizing oral archives
- Subtitles in many languages
- Sizes
- 300M – 7B
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
3D2025
Tencent Hunyuan · China
Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.
- Character animation from a text description
- Draft animation for games and videos
- Motion library for avatars
- Sizes
- 0.46B – 1B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek's maths models. The first 7B version introduced the GRPO training method; the 685B V2 writes and checks its own olympiad-level proofs.
- Calculations and formula checks
- Checking mathematical workings in reports
- Working through problems step by step
- Sizes
- 7B – 685B
- Hardware
- from: Laptop
Text to speechRU2022–2025
Silero · Russia
Lightweight Russian speech synthesis that runs on a regular CPU. Version v5 added CIS languages and languages of Russia's peoples: Tatar, Bashkir, Yakut, Kazakh and others.
- Voicing voice bot replies
- Reading texts in Russian
- Voices in the languages of Russia's peoples
- Sizes
- tens of megabytes
- Hardware
- from: Laptop
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Deepfake detection2023–2025
IBM Research and The Chinese University of Hong Kong · USA
An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.
- Checking texts that may have been rewritten after generation
- First-pass filtering in a newsroom or admissions office
- Comparison against simpler detectors
- Sizes
- about 355M (RoBERTa-large)
- Hardware
- from: Laptop
Avatars2025
Meituan · China
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
- Video dubbing with matched facial expressions
- Dialogue between two characters from audio
- Long videos with a presenter
- Sizes
- 14B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Moonshot AI and Project Numina · China, France
Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.
- Formal verification of mathematical workings
- Translating a problem from plain language into Lean
- Training and olympiad preparation
- Sizes
- 0.6B – 72B
- Hardware
- from: Laptop
TextRU2024–2025
Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia
Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.
- Russian-language assistant on your own server
- Answers based on company documents (RAG) in Russian
- Analysis and summaries of long Russian texts
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Text to speechRU2025
ESpeech (independent group of Russian-speaking developers) · Russia
Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.
- Voicing videos and audiobooks in Russian
- Cloning a narrator's voice from a sample
- Voice for a bot or assistant in Russian
- Sizes
- about 340M
- Hardware
- from: Laptop
Video2025
Tencent Hunyuan · China
Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.
- Interactive video prototypes of game locations
- Camera walkthrough videos of a scene
- Level demos for pitches
- Sizes
- based on HunyuanVideo
- Hardware
- from: 1 GPU
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoning2025
NVIDIA · USA
NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Working through programming problems
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Robotics2025
Hugging Face · USA
A small robot control model that runs on a regular laptop. Trained on open data from the LeRobot community, suited to low-cost robot arms.
- Controlling a low-cost robot arm
- Quick robotization pilots and demos
- Training staff and students
- Sizes
- 450M
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
RVC-Boss and community · China
Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.
- Voicing texts with a specific narrator's voice
- Voice for a bot or assistant
- Dubbing training videos
- Sizes
- under 1B
- Hardware
- from: Laptop
Avatars2025
ByteDance · China
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
- Sizes
- requires 8–18 GB of VRAM
- Hardware
- from: Laptop
AvatarsGGUF2025
Tencent · China
Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.
- Presenter video from a photo and audio
- Scenes with several speakers
- Cartoon characters
- Sizes
- about 13B
- Hardware
- from: 1 GPU
Fact-checking and judgesRU2024–2025
NVIDIA · USA
Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data to fine-tune your own model
- Scoring assistant answers in Russian and other languages
- Sizes
- 49B, 70B and 340B
- Hardware
- from: Cluster
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.
- Formal verification of mathematical workings
- Verifying algorithm correctness
- Training and olympiad preparation
- Sizes
- 7B – 671B
- Hardware
- from: Laptop
TextRUGGUF2023–2025
Ilya Gusev (IlyaGusev) · Russia
The best-known Russian community fine-tune: open models (Llama, Mistral, Gemma, YandexGPT) trained to act as a Russian-speaking assistant. A convenient starting point for a Russian chatbot on your own server.
- Russian-language chat assistant
- Answers based on the company knowledge base
- Drafts of emails, descriptions and posts in Russian
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of contracts and internal policies
- Sizes
- 32B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Stanford University · USA
A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.
- Calculations and formula checks
- Working through complex problems step by step
- Training
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Qihoo 360 · China
Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Deepfake detection2025
Desklib · India
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
- Checking submitted articles and reports
- Filtering templated reviews and applications
- First-pass check of student work
- Sizes
- 0.4B (DeBERTa-v3-large)
- Hardware
- from: Laptop
Math and reasoningGGUF2025
NovaSky (Sky Computing Lab, Berkeley) · USA
A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
TextGGUF2025
HUMAIN (formerly SDAIA) · Saudi Arabia
A Saudi model for Arabic and English, trained from scratch. One 7B version is openly available.
- An Arabic-language assistant
- Answering questions about documents
- Writing and editing texts in Arabic
- Sizes
- 7B
- Hardware
- from: Laptop
Search and RAG2025
The authors of the CareerBERT paper, German universities · Germany
A German-language model that matches a resume to occupations from the European ESCO reference list and suggests suitable directions. A human makes the decision about a candidate; automatic screening without review must not be used.
- Suggesting occupations that fit the candidate experience
- Matching resumes against job descriptions
- Hints on internal career moves
- Sizes
- 110M
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.
- Calculations and formula checks
- Checking calculations in estimates and reports
- Working through problems step by step
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Search and RAG2024
TechWolf · Belgium
Finds mentions of skills in a vacancy or resume text and maps them to the company skill reference list. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting skills from a job description
- Matching candidate skills against requirements
- Building a competence map across departments
- Sizes
- 109M
- Hardware
- from: Laptop
Video2024
Zhipu AI (Z.ai) and Tsinghua University · China
A 2–5B video model that runs on a single gaming GPU. A popular base for research and add-ons.
- Short clips from text
- Animating images
- Video fine-tuning experiments
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
CodeOllama2024
INF Technology · China
Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.
- Code generation and completion
- Training your own coding model from an open recipe
- A programming assistant on low-end hardware
- Sizes
- 1.5B – 8B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
TextRUNot maintained2023–2024
Sber (ai-forever) · Russia
Sber's Russian text-to-text model, successor to ruT5 (2021). Small and fast: fine-tuned for summarizing, paraphrasing and fixing errors in Russian text; ready-made SAGE spell-checking versions exist.
- Fixing spelling mistakes and typos in Russian text
- Short summaries and paraphrasing
- Normalizing requests and inquiries before processing
- Sizes
- 95M – 1.7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
KAIST · South Korea
A research try-on model from CVPR 2024, one of the first built on Stable Diffusion. Now mostly used as a comparison baseline.
- Pilot of upper-body garment try-on
- Comparing quality of different try-on models
- Training your own try-on on the open code
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
CodeOllamaNot maintained2023–2024
Meta · USA
A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.
- Code autocompletion and explanation
- Generating Python scripts
- Base model for fine-tuning on your own stack
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2023
Resemble AI · USA
A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.
- Restoring old and phone recordings
- Cleaning a voice before voiceover and cloning
- Improving audio in videos and podcasts
- Sizes
- under 1B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
FinanceGGUFNot maintained2023
AdaptLLM · not disclosed
Finance-tuned versions of Llama 2: reading industry texts, reports and questions about terminology. Not investment advice: decisions are made by a specialist.
- Reading financial news and reports
- Answering questions about financial terminology
- A base for fine-tuning to your own financial task
- Sizes
- 7B и 13B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023
Hello-SimpleAI · China
One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.
- First-pass check of student work
- Filtering templated applications and reviews
- Flagging suspicious texts for manual review
- Sizes
- about 125M (RoBERTa-base)
- Hardware
- from: Laptop
RerankersNot maintained2023
NetEase Youdao · China
An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.
- Search across a knowledge base and reference materials
- Reordering retrieved passages
- Picking answers for a support chatbot
- Sizes
- about 280M
- Hardware
- from: Laptop
Text to speechNot maintained2023
Columbia University · USA
A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.
- Voicing texts in English
- A base for fine-tuning your own voice
- Voice service prototypes
- Sizes
- about 150M
- Hardware
- from: Laptop
Voice assistantsRUNot maintained2023
Meta · USA
Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.
- Speech-to-speech translation
- Translating and transcribing recordings
- Streaming translation
- Sizes
- 281M – 2.3B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Microsoft Research · USA
Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.
- Research on reasoning methods
- Comparison with modern small models
- Training specialists
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
Text analysisNot maintained2022–2023
Mike Zhang, Rob van der Goot, Barbara Plank (IT University of Copenhagen and LMU Munich) · Denmark
A research line of models for labour market texts: trained on job postings and the European ESCO occupation taxonomy, they pull skills and requirements out of vacancies. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting skills and requirements from vacancy text
- Mapping skills to the single ESCO reference list
- Classifying vacancies and job titles
- Sizes
- 110M – 560M
- Hardware
- from: Laptop
Text to speechRUNot maintained2023
Coqui · Germany
A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.
- Voice cloning from a sample
- Multilingual voiceover
- Research and prototypes
- Sizes
- about 470M
- Hardware
- from: Laptop
Documents and OCRNot maintained2023
Meta · USA
An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.
- Converting scientific papers from PDF into text with formulas
- Digitising technical documentation
- Sizes
- 250M – 350M
- Hardware
- from: Laptop
TextRUGGUFNot maintained2022–2023
Sber (ai-forever) · Russia
Sber's multilingual model covering 61 languages, including languages of the peoples of Russia and the CIS. Separate fine-tunes exist for Buryat, Yakut, Tatar, Bashkir, Kazakh and others, rare for open models.
- Texts in languages of the peoples of Russia and the CIS
- Base for fine-tuning on a less common language
- Drafts and templates in several languages
- Sizes
- 1.3B – 13B
- Hardware
- from: Laptop
AvatarsNot maintained2023
Xi'an Jiaotong University and Tencent AI Lab · China
An older lightweight talking-head model: one photo plus audio becomes a video. Runs on weak hardware, but quality is noticeably below newer models.
- Talking photo for greetings
- Simple voiced avatars
- Sizes
- under 1B
- Hardware
- from: Laptop
Weather and climateNot maintained2023
Huawei Cloud · China
One of the first weather neural networks, published in Nature and added to ECMWF charts. The weights are open for research only; commercial use is prohibited.
- Research weather forecasts a week ahead
- Comparison with other weather models on your own data
- Training courses on weather neural networks
- Sizes
- 4 models of ~1.1 GB each (1, 3, 6 and 24-hour steps)
- Hardware
- from: Laptop
3DNot maintained2022–2023
OpenAI · USA
Early open OpenAI models that create a 3D object from text or an image in seconds. Quality is basic, but they are fast and easy to run.
- Rough 3D mock-ups from a description
- Quick object prototypes for games and AR
- Training and research pilots in 3D
- Sizes
- 40M – 1B
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2022–2023
Meta · USA
A translator for 200 languages, including rare and minor ones. Russian is supported. Strong language coverage, but the license prohibits commercial use.
- Translating texts between 200 languages
- Translating into rare languages where no other models exist
- Comparing quality when choosing a translator
- Sizes
- 600M to 3.3B (plus 54B MoE)
- Hardware
- from: Laptop
TextNot maintained2022–2023
EleutherAI · USA
Fully open models from the non-profit lab EleutherAI: GPT-NeoX-20B and the Pythia series with published intermediate training checkpoints.
- Base model for fine-tuning
- Research into model behavior
- Simple text generation and completion
- Sizes
- 70M – 20B
- Hardware
- from: Laptop
TextNot maintained2022
Meta · USA
An early open Meta series matching GPT-3 in size. Outdated; useful for research and comparison.
- Research experiments
- Training specialists
- Comparison with modern models
- Sizes
- 125M – 66B (175B on request)
- Hardware
- from: Laptop
TextGGUFNot maintained2022
BigScience (Hugging Face and the community) · France
One of the first large open models, trained by a community of hundreds of researchers in 46 languages. Today it is interesting mostly as a historical milestone.
- Text generation and translation in many languages
- Experiments and team training
- Base model for fine-tuning on a narrow task
- Sizes
- 560M – 176B
- Hardware
- from: Laptop
Documents and OCRNot maintained2021–2022
Microsoft · USA
Recognizes a single line of text, including handwriting. The official weights are English only, but the model is often fine-tuned for other languages; there are community Russian versions.
- Recognizing handwritten lines in questionnaires and forms
- Recognizing printed lines after text detection on the page
- A base for fine-tuning to your own handwriting or font
- Sizes
- 62M – 608M
- Hardware
- from: Laptop
AvatarsNot maintained2020
IIIT Hyderabad · India
The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.
- Quick dubbing tests
- Comparison with newer lip-sync models
- Educational and research projects
- Sizes
- small model, 96-pixel face
- Hardware
- from: Laptop