Florence-2 on NVIDIA DeepStream: Vision-Language AI for Real-Time Video Analytics

Hi all,

I would like to share an open-source project that integrates Microsoft Florence-2 with NVIDIA DeepStream.

Florence-2 is a versatile vision foundation model capable of performing a wide range of tasks through a unified prompt-based interface. This project demonstrates how these capabilities can be leveraged in video analytics applications powered by DeepStream.

Features

  • Image and scene captioning
  • OCR and text extraction
  • Phrase grounding
  • Dense region captioning
  • Open-vocabulary object detection
  • Prompt-based visual understanding

GitHub Repository:

Thank you for sharing with the community!