AI for Media Product Update: New Real-Time AI Models and Infrastructure Microservices

We are excited to share the latest product update for NVIDIA AI for Media, designed to enhance real-time live video pipelines and post-production workflows.

Here are the major product and feature updates now available:

  • New SMPTE ST 2110 NIM Microservices: We have introduced a lineup of SMPTE ST 2110-compliant microservices to simplify deployment directly into live video pipelines, including:

    • NVIDIA LipSync SMPTE ST 2110 NIM

    • NVIDIA Active Speaker Detection SMPTE ST 2110 NIM

    • NVIDIA Studio Voice SMPTE ST 2110 NIM

    • NVIDIA Video Super Resolution SMPTE ST 2110 NIM

  • Multilingual LipSync – Private Access: Expanded language-optimized models now support French, German, and Spanish for realistic lip synchronization during translation workflows.

  • Enhanced Active Speaker Detection (ASD) – General Availability: Now features cross-video speaker identity correlation, allowing accurate tracking across multi-camera and multi-microphone broadcast environments.

  • Synthetic Video Detector (SVD) – Private Access: A new NIM model optimized to predict the percentage probability that a video was generated by AI diffusion models for use in media integrity pipelines. The model achieves 92% accuracy on uncompressed video and processes frames in as little as 22 milliseconds.

You can explore and test these new features today: https://developer.nvidia.com/ai-for-media