Drive PX 2 Caffe performance

san_kim · February 27, 2018, 9:36am

Hi all

I used a caffe-based Single shot multibox detector(SSD) algorithm.

When we made the inference, we only got 5 fps based on the HD images.

Under the same conditions, Jetson TX2 gets 8.5 fps.

My environment)

vi .bashrc

export PYTHONPATH=/home/nvidia/ssd-caffe/python
export PATH=/usr/local/cuda-8.0/bin/:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda-8.0/targets/aarch64-linux/lib:$LD_LIBRARY_PATH

nvidia@nvidia:~$ nvcc --version

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2016 NVIDIA Corporation
Built on Mon_Mar_20_17:07:33_CDT_2017
Cuda compilation tools, release 8.0, V8.0.72

./deviceQuery

CUDA Device Query (Runtime API) version (CUDART static linking)

Detected 2 CUDA Capable device(s)

Device 0: “GP106”
CUDA Driver Version / Runtime Version 8.0 / 8.0
CUDA Capability Major/Minor version number: 6.1
Total amount of global memory: 3840 MBytes (4026466304 bytes)
( 9) Multiprocessors, (128) CUDA Cores/MP: 1152 CUDA Cores
GPU Max Clock rate: 1290 MHz (1.29 GHz)
Memory Clock rate: 3003 Mhz
Memory Bus Width: 128-bit
L2 Cache Size: 1048576 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(131072), 2D=(131072, 65536), 3D=(16384, 16384, 16384)
Maximum Layered 1D Texture Size, (num) layers 1D=(32768), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(32768, 32768), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 65536
Warp size: 32
Maximum number of threads per multiprocessor: 2048
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 2 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: No
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Disabled
Device supports Unified Addressing (UVA): Yes
Device PCI Domain ID / Bus ID / location ID: 0 / 4 / 0
Compute Mode:
< Default (multiple host threads can use ::cudaSetDevice() with device simultaneously) >

Device 1: “GP10B”
CUDA Driver Version / Runtime Version 8.0 / 8.0
CUDA Capability Major/Minor version number: 6.2
Total amount of global memory: 6660 MBytes (6983639040 bytes)
( 2) Multiprocessors, (128) CUDA Cores/MP: 256 CUDA Cores
GPU Max Clock rate: 1275 MHz (1.27 GHz)
Memory Clock rate: 1600 Mhz
Memory Bus Width: 128-bit
L2 Cache Size: 524288 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(131072), 2D=(131072, 65536), 3D=(16384, 16384, 16384)
Maximum Layered 1D Texture Size, (num) layers 1D=(32768), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(32768, 32768), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 32768
Warp size: 32
Maximum number of threads per multiprocessor: 2048
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 1 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: Yes
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Disabled
Device supports Unified Addressing (UVA): Yes
Device PCI Domain ID / Bus ID / location ID: 0 / 0 / 0
Compute Mode:
< Default (multiple host threads can use ::cudaSetDevice() with device simultaneously) >

Peer access from GP106 (GPU0) → GP10B (GPU1) : No
Peer access from GP10B (GPU1) → GP106 (GPU0) : No

deviceQuery, CUDA Driver = CUDART, CUDA Driver Version = 8.0, CUDA Runtime Version = 8.0, NumDevs = 2, Device0 = GP106, Device1 = GP10B
Result = PASS

Thanks

SivaRamaKrishnaNV · March 1, 2018, 12:02pm

Dear san_kim,
By default it runs on dGPU. Please set env variable CUDA_VISIBLE_DEVICES= 1 to get the timings on iGPU

san_kim · March 5, 2018, 9:48am

I got
GP106: 5 fps
GP10B: 3 fps

I try this
ssd-caffe/examples/ssd/ssd_pascal_video.py

Do I need to modify the code for multi thread inference?

Topic		Replies	Views
Speed of SSD on TX1. Jetson TX1	7	1082	October 18, 2021
TX2: caffe model that runs slower on nvcaffe GPU than on OpenCV CPU Jetson TX2	3	681	October 18, 2021
Maximize PX2 performance by setting max frequency to CPU, GPU General	7	1555	November 15, 2017
Jetson TX2 and notebook(965m) performance problem Jetson TX2	2	479	October 18, 2021
jetson tx2 not using gpu for my the opencv caffe-model Jetson TX2	7	992	October 18, 2019
How to run Caffe on GPU (TX2) Jetson TX2	3	1326	July 5, 2018
can the nvidia TensorRT accelerate SSD(single shot detector)? Jetson TX2	22	9059	October 18, 2021
Performance analysis for TX2 Jetson TX2	2	845	October 18, 2021
Object Detection Performance Jetson Tx2 slower than expected Jetson TX2	22	14887	October 18, 2021
Running with GPU is slower than CPU on TX2 Jetson TX2	2	777	October 18, 2021

Drive PX 2 Caffe performance

Related topics