# Stdpar unknown H2D D2H transfer with nvc++

**URL:** <https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824>\
**Category:** nvc, nvc++ and nvfortran\
**Created:** [September 20, 2026, 12:00pm UTC](https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824 "2026-09-20T12:00:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![huccpp](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@huccpp](https://forums.developer.nvidia.com/u/huccpp)\
**Post date:** [September 20, 2026, 12:00pm UTC](https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824/1 "2026-09-20T12:00:45Z")

</div>

Dear All,

This issue is related to the later development in this topic:

[https://forums.developer.nvidia.com/t/error-in-trying-to-use-stdpar-in-nvc/351502](https://forums.developer.nvidia.com/t/error-in-trying-to-use-stdpar-in-nvc/351502)

The topic was about I was writing a code in nvc++ using feature like cartesian\_product which was unsuccessful as it required system with HMM which I did not have on my laptop.

Recently I updated my driver so HMM is now available on the system of my laptop.

So I can tried to use these nvc++ features now. And new issues arise that I describe here.

Here is the current version, which is still illustrated by the example of solving 2D Laplace’s eqt with relaxation:

```cpp
#include <iostream>
#include <vector>
#include <ranges>
#include <execution>
#include <fstream>
#include <algorithm>
#include <mdspan>
#include <chrono>

auto main() -> int {
int size1 = 200;
int size2 = 200; // array size= size1*size2
int niter=1000; // test perf 1000000
std::vector<double> A(size1*size2, 0.0);
auto A_v = std::mdspan (A.data(), size1, size2); //create view
// Some initialization code
// Set one side of the rectangular grid in A to unity
    for (int i = 0; i < size1; ++i) {
        A_v(i,0) = 1.0; // Set to unity on one side
    }
std::vector<double> B(A);
auto B_v = std::mdspan (B.data(), size1, size2); //create view

//view for using for_each
auto v = std::ranges::views::cartesian_product(
std::ranges::views::iota(1, size1 - 1),
std::ranges::views::iota(1, size2 - 1));

auto beg = std::chrono::high_resolution_clock::now();
//iteration
for (int i=0; i<niter; i++) {
std::for_each(std::execution::par, std::begin(v), std::end(v),
[=](auto idx) {
auto [i, j] = idx;
B_v(i,j) = 0.25*(A_v(i-1,j) + A_v(i+1,j) + A_v(i,j-1) + A_v(i,j+1));
});
std::swap(A_v,B_v);

}
auto end = std::chrono::high_resolution_clock::now();
auto duration = std::chrono::duration_cast<std::chrono::milliseconds>(end - beg);
// Displaying the elapsed time
std::cout << "Elapsed Time for iteration part: " << duration.count()<< " ms" <<std::endl;

    // Open file for writing
    std::ofstream outputFile("output.txt");
    if (!outputFile.is_open()) {
        std::cerr << "Error: Unable to open output file!" << std::endl;
        return 1;
    }

    // Write contents of A to file
    for (int i = 0; i < size1; ++i) {
        for (int j = 0; j < size2; ++j) {
            outputFile << i << " " << j << " " << A_v(i,j) << std::endl;
        }
        outputFile << std::endl; // Skip a line after each row
    }

    // Close file
    outputFile.close();

    // Confirmation message
    std::cout << "Output file 'output.txt' generated successfully." << std::endl;

}

```

I compiled with nvc++ -std=c++23 -stdpar=gpu -gpu=cc89

It runs normally, giving correct results. However, for niter=10^6, the iteration part is 40% slower

than not using cartesian\_product but just use

auto flat\_range = std::views::iota(0, (size1 - 2) \* (size2 - 2));

together with

```auto
            int i = idx / (size2 - 2) + 1;
            int j = idx % (size2 - 2) + 1;

```

inside the lambda, suggested by Mat in the previous topic.

Also, there are unknown H2D and D2H transfer in the profiling, which were not present for the case not using cartesian\_product. I attached the zoomed timeline here.

 ![test_lp](https://global.discourse-cdn.com/nvidia/original/4X/9/e/4/9e4aba57ca0bfdede746df16eade96ba81ed8f72.png)

The red parts were H2D and D2H transfer with the migration cause: Page fault.

What are the causes for the performance decrease? And why is page fault introduced, or what it is,

why is it leading to data transfer between H and D, and how to avoid them?

Thanks.

Yours,

huccpp

---

<div class="post-metadata">

**Author:** ![scamp1](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/scamp1/32/312607_2.png) [@scamp1](https://forums.developer.nvidia.com/u/scamp1)\
**Post date:** [September 22, 2026, 5:00pm UTC](https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824/2 "2026-09-22T17:00:59Z")

</div>

Hi huccpp,

I believe this behavior is consistent with how std::ranges::cartesian\_product\_view iterators are implemented, combined with demand paging under HMM.

A Cartesian-product iterator is not generally self-contained. In common standard-library implementations, it contains:

- The current iterator for each component range.
- A pointer back to the parent cartesian\_product\_view."

When the random-access iterator is advanced, it reads the component sizes from the parent view and converts the linear offset into Cartesian coordinates. For example, libstdc++’s implementation stores \_M\_parent and accesses \_M\_parent-\>\_M\_bases while advancing the iterator. It also performs division and remainder operations during this conversion.

In this program, v is an automatic variable stored on the CPU stack:

` auto v = std::ranges::views::cartesian_product(...);`

The iterator passed to the GPU therefore contains a pointer into the CPU stack. Previously this caused an illegal address because the GPU could not access that memory. HMM makes the stack address accessible, so the program now runs, but accessibility does not mean that the memory is permanently resident on the GPU.

In this mode of data handling, HMM uses CPU and GPU page faults to place system-allocated pages where they are needed. When the GPU dereferences the iterator’s parent pointer, the stack page may migrate from host memory to device memory. When the CPU subsequently evaluates v.begin() and v.end() for the next iteration—or accesses another variable sharing that stack page—the page may migrate back. Nsight Systems reports those migrations as H2D and D2H transfers with “Page fault” as the cause. This is normal demand-paging behavior, not an error. NVIDIA’s HMM description ([Simplifying GPU Application Development with Heterogeneous Memory Management | NVIDIA Technical Blog](https://developer.nvidia.com/blog/simplifying-gpu-application-development-with-heterogeneous-memory-management/)) explains this mechanism.

The flat iota\_view does not have this problem. Its iterator is essentially just an integer, so it can be copied completely into the kernel arguments and does not need a pointer back to the host-side view.

There is also a second performance difference. The Cartesian iterator performs a generic multidimensional offset calculation, commonly using ptrdiff\_t, which is 64-bit. The flat version performs one 32-bit division and remainder directly in the lambda:

```auto
  int i = idx / (size2 - 2) + 1;
  int j = idx % (size2 - 2) + 1;

```

Thus the 40% difference may be a combination of:

1. HMM page migration involving the parent view or its stack page.
2. More expensive generic iterator arithmetic, potentially including 64-bit division/remainder.
3. One million synchronous std::for\_each calls and kernel launches. Each iteration is relatively small, so fixed overhead is significant.

I believe the best solution for this particular problem is to retain the flat index range. It is compact, avoids the parent-view pointer, and gives the compiler a simpler kernel, something like:

```auto
  const int nx = size1 - 2;
  const int ny = size2 - 2;
  auto flat_range = std::views::iota(0, nx * ny);

  for (int iter = 0; iter < niter; ++iter) {
      std::for_each(std::execution::par,
                    flat_range.begin(), flat_range.end(),
                    [=](int idx) {
          int i = idx / ny + 1;
          int j = idx % ny + 1;

          B_v(i,j) = 0.25 * (A_v(i-1,j) + A_v(i+1,j)
                           + A_v(i,j-1) + A_v(i,j+1));
      });

      std::swap(A_v, B_v);
  }

```

One-time H2D migration when the arrays are first used by the GPU, and D2H migration when the CPU writes the result file, are expected. Repeated page-sized migrations between individual iterations are the important ones to look for. Comparing kernel duration separately from Unified Memory migration time in Nsight Systems should show how much of the slowdown comes from iterator arithmetic versus page migration.

Let me know how that works out for you!

Cheers,

Seth.

---

<div class="post-metadata">

**Author:** ![huccpp](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@huccpp](https://forums.developer.nvidia.com/u/huccpp)\
**Post date:** [September 26, 2026, 11:44pm UTC](https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824/3 "2026-09-26T23:44:46Z")

</div>

Hello Seth,

Thanks for your reply and useful advice.

Yes, I profiled the flat\_range case as well and also notice the kernel run time is far shorter in addition to the absence of the H2D D2H transfer.

Then, what is the advantage of using the Cartesian iterator? Or what is the motivation for its development? I was thinking it could perform better than, or at least close to manual division in the lambda, so that users don’t need to explicitly code the division, but it’s seemed not the case, at least in this particular example.

Thanks.

huccpp

---

<div class="post-metadata">

**Author:** ![scamp1](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/scamp1/32/312607_2.png) [@scamp1](https://forums.developer.nvidia.com/u/scamp1)\
**Post date:** [September 27, 2026, 8:27pm UTC](https://forums.developer.nvidia.com/t/stdpar-unknown-h2d-d2h-transfer-with-nvc/383824/4 "2026-09-27T20:27:47Z")

</div>

Hi huccpp,

The main reason to use std::views::cartesian\_product is to express “visit every combination of these ranges” without writing nested loops or converting a flat index back to coordinates. That can make code easier to reuse when the number of dimensions, bounds, or input ranges change. However, it’s a general C++23 range feature, rather than a GPU optimization. If you were to have more work to be done on the GPU in between iterator calculations, it might increase the optimization a lot because you’d be spending large amounts of your time on the GPU before having page migrations. However, we’d likely still be able to find alternatives that are faster when memory location is maintained- though perhaps in the right problem where you spend lots of time on the gpu, that difference may be small.

Often times when converting code to run efficiently on GPUs, changing the core approach to maintain memory location or optimize GPU performance is an important detail. Sometimes the best CPU code and the best GPU code look different.

Cheers,

Seth.
