Dear All,
This issue is related to the later development in this topic:
https://forums.developer.nvidia.com/t/error-in-trying-to-use-stdpar-in-nvc/351502
The topic was about I was writing a code in nvc++ using feature like cartesian_product which was unsuccessful as it required system with HMM which I did not have on my laptop.
Recently I updated my driver so HMM is now available on the system of my laptop.
So I can tried to use these nvc++ features now. And new issues arise that I describe here.
Here is the current version, which is still illustrated by the example of solving 2D Laplace’s eqt with relaxation:
#include <iostream>
#include <vector>
#include <ranges>
#include <execution>
#include <fstream>
#include <algorithm>
#include <mdspan>
#include <chrono>
auto main() -> int {
int size1 = 200;
int size2 = 200; // array size= size1*size2
int niter=1000; // test perf 1000000
std::vector<double> A(size1*size2, 0.0);
auto A_v = std::mdspan (A.data(), size1, size2); //create view
// Some initialization code
// Set one side of the rectangular grid in A to unity
for (int i = 0; i < size1; ++i) {
A_v(i,0) = 1.0; // Set to unity on one side
}
std::vector<double> B(A);
auto B_v = std::mdspan (B.data(), size1, size2); //create view
//view for using for_each
auto v = std::ranges::views::cartesian_product(
std::ranges::views::iota(1, size1 - 1),
std::ranges::views::iota(1, size2 - 1));
auto beg = std::chrono::high_resolution_clock::now();
//iteration
for (int i=0; i<niter; i++) {
std::for_each(std::execution::par, std::begin(v), std::end(v),
[=](auto idx) {
auto [i, j] = idx;
B_v(i,j) = 0.25*(A_v(i-1,j) + A_v(i+1,j) + A_v(i,j-1) + A_v(i,j+1));
});
std::swap(A_v,B_v);
}
auto end = std::chrono::high_resolution_clock::now();
auto duration = std::chrono::duration_cast<std::chrono::milliseconds>(end - beg);
// Displaying the elapsed time
std::cout << "Elapsed Time for iteration part: " << duration.count()<< " ms" <<std::endl;
// Open file for writing
std::ofstream outputFile("output.txt");
if (!outputFile.is_open()) {
std::cerr << "Error: Unable to open output file!" << std::endl;
return 1;
}
// Write contents of A to file
for (int i = 0; i < size1; ++i) {
for (int j = 0; j < size2; ++j) {
outputFile << i << " " << j << " " << A_v(i,j) << std::endl;
}
outputFile << std::endl; // Skip a line after each row
}
// Close file
outputFile.close();
// Confirmation message
std::cout << "Output file 'output.txt' generated successfully." << std::endl;
}
I compiled with nvc++ -std=c++23 -stdpar=gpu -gpu=cc89
It runs normally, giving correct results. However, for niter=10^6, the iteration part is 40% slower
than not using cartesian_product but just use
auto flat_range = std::views::iota(0, (size1 - 2) * (size2 - 2));
together with
int i = idx / (size2 - 2) + 1;
int j = idx % (size2 - 2) + 1;
inside the lambda, suggested by Mat in the previous topic.
Also, there are unknown H2D and D2H transfer in the profiling, which were not present for the case not using cartesian_product. I attached the zoomed timeline here.
The red parts were H2D and D2H transfer with the migration cause: Page fault.
What are the causes for the performance decrease? And why is page fault introduced, or what it is,
why is it leading to data transfer between H and D, and how to avoid them?
Thanks.
Yours,
huccpp
