Florence-2
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
The last open version came out in Jun 2024. The family has not been updated for a long time: the model still works, but do not expect fixes or new sizes.
- Developer
- Microsoft, USA
- First release
- Jun 2024
- Latest release
- Jun 2024
- Sizes
- 0.23B – 0.77B
- License
- Commercial use allowedMIT
- Russian
- Not supported
- Ready-made builds
- MLX (Apple)
- Running
- On your own serverAlso runs without a GPU
- Industries
- Manufacturing and logistics, Media and production, Retail and marketplaces
What it does
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Labelling data for training
Where it is used
Hardware requirements
Versions
- Florence-2 base / large
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems. Quantization compresses a model so it takes less video memory and runs on more modest hardware. Answers change slightly, so quality is checked on your own examples.
Frequently asked questions
Can Florence-2 be used in a commercial project?
Yes. License: MIT. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does Florence-2 need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Florence-2 support Russian?
No. The model card lists its languages and Russian is not among them.
Where can I download Florence-2 and what does it cost?
The Florence-2 weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Comparisons
Similar models
Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.
DetailsImage + textMoondreamMoondream (M87 Labs) · USACommercial use with conditionsA small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
DetailsComputer visionGrounding DINO / Rex-OmniIDEA Research · ChinaCommercial use with conditionsFinds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.
DetailsSource: huggingface.co/microsoft/Florence-2-large. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


