It depends on how big your frame buffers are and what you’re doing with the data after the downloadPixels() in that loop.
1.) First you would need to determine what the bottleneck actually is.
Are you limited by how long the ray tracing takes, or are you limited by the device-to-host data transfer of the pixels and whatever you do with them afterwards inside that loop. Nsight Systems can determine that.
If you’re unable to do anything asynchronously with the pixels data while the next image gets rendered on the GPU, you’re going to be limited by device-to-host data transfers.
2.) If the fbsize isn’t changing per camera, you can move these lines before the loop.
std::vector<uint32_t> pixels(fbSize.x*fbSize.y);
renderer->resize(fbSize);
3.) I assume the renderer->setCamera(camera); updates the single camera information inside the launch parameters.
You could instead upload the information of that whole camera array into a CUDA device buffer and put that device pointer and number of cameras into the launch parameters outside the loop, and change your render() call to take a camera index as argument.
4.) If the fbsize is comparably small, like 256x256 or even smaller, and your ray tracing algorithm is rather fast, high-end GPUs could not be fully saturated by that launch dimension.
If you placed many of these cameras as tiles into a single bigger framebuffer, e.g. for 256x256 per camera image, you could render 100 cameras at once into a 2560x2560 buffer in a single optixLaunch instead.
That would later require some more logic to split these again to get individual images on the host, but that could happen on the CPU on the copied pixels while the GPU renderer works on the next cameras already.
5.) Or similarly to the tiles approach but with more asynchronous operations, when allocating room for multiple camera output images on the device and on the host, you could render multiple cameras with (asynchronous) optixLaunch calls and asynchronous memory copies into the same CUDA stream.
For that each optixLaunch would need to know which camera it renders and how many optixLaunch calls are done per render() call to be able to calculate its output buffer pointer.
I think when inserting CUDA events after the asynchronous copies, the CPU could also be triggered to start working on already finished pixel data.
So it really depends on what workload you’re talking about.
If this is doing something like rendering 4K images and saving them to disk, then you’re data transfer limited anyway and that could be alleviated a little by double buffering the rendering and pixel copies maybe.