Hi all,
I would like to share an open-source project that integrates Microsoft Florence-2 with NVIDIA DeepStream.
Florence-2 is a versatile vision foundation model capable of performing a wide range of tasks through a unified prompt-based interface. This project demonstrates how these capabilities can be leveraged in video analytics applications powered by DeepStream.
Features
- Image and scene captioning
- OCR and text extraction
- Phrase grounding
- Dense region captioning
- Open-vocabulary object detection
- Prompt-based visual understanding
GitHub Repository: