# Calculating utilization (core load) of Tensor core and cuda core seperately

**URL:** <https://forums.developer.nvidia.com/t/calculating-utilization-core-load-of-tensor-core-and-cuda-core-seperately/165390>\
**Category:** CUPTI – CUDA Profiler Tools Interface\
**Created:** [January 7, 2021, 10:14am UTC](https://forums.developer.nvidia.com/t/calculating-utilization-core-load-of-tensor-core-and-cuda-core-seperately/165390 "2021-01-07T10:14:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![VivekM](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@VivekM](https://forums.developer.nvidia.com/u/VivekM)\
**Post date:** [January 7, 2021, 10:14am UTC](https://forums.developer.nvidia.com/t/calculating-utilization-core-load-of-tensor-core-and-cuda-core-seperately/165390/1 "2021-01-07T10:14:00Z")

</div>

Hi,  
Architecture: Turing  
DL: using TensorRT  
I know using the tensor\_precision\_fu\_utilization and tensor\_int\_fu\_utilization the tensor core utilization can be found for each kernel scaling 0 to 10. Is there a convenient way to find out the total utilization of tensor cores lets say over 1 second ? without actually going into each kernel utilization? I want to use the CUPTI APIs to basically Segregate the tensor core and CUDA core utilization complete deep learning network over a period of time frame. we are using right now nvmlDeviceGetUtilizationRates() nvml library function to know the GPU utilization but I think this API returns the total GPU core load and the bifurcation between tensor and cuda cores is not in it.

Thanks

---

<div class="post-metadata">

**Author:** ![mjain](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/mjain/32/14043_2.png) [@mjain](https://forums.developer.nvidia.com/u/mjain)\
**Post date:** [January 15, 2021, 12:44pm UTC](https://forums.developer.nvidia.com/t/calculating-utilization-core-load-of-tensor-core-and-cuda-core-seperately/165390/2 "2021-01-15T12:44:49Z")

</div>

Hi Vivek,

If I understand the use case, you want to capture the tensor core usage data without serializing the kernels in the application, is that correct? I think tool like DCGM (Data Center GPU Manager) is better suited for this use case as it can provide a set of metrics at the device-level with low performance overhead in a continuous manner. I assume “Tensor Activity” is the metric you are interested in. More details can be found at [Welcome — NVIDIA DCGM Documentation latest documentation](https://docs.nvidia.com/datacenter/dcgm/latest/dcgm-user-guide/feature-overview.html#profiling)
