Improving Performance When Creating Many References

Isaac Sim Version

5.0.0

Operating System

Windows 11

GPU Information

  • Model: 5090
  • Driver Version:

Topic Description

Hi everyone,

I’m currently developing an extension that automatically generates a warehouse based on a floor plan. Part of this setup includes a pallet storage system designed to hold approximately 8,000 pallets.

While creating the storage structure itself is quite fast, populating it by adding each pallet as a reference takes a very long time. Currently, I’m creating an individual reference for each pallet, and the process seems to be a bottleneck.

Here’s a relevant excerpt of my code:


def position_elements_in_racks(elements_path, element_matrix, layout_element):
    stage = get_context().get_stage()
    # Load root palette asset
    palett_root_prim = load_element(layout_element['pallet']['asset_path'], layout_element['pallet']['asset_name'])
    print(f"  📊 Processing {len(element_matrix)} elements...")

    for i, row in enumerate(element_matrix):
        if i % 100 == 0 and i > 0:
            print(f"    Processed {i}/{len(element_matrix)} elements")
        rack_id = int(row[0])
        palett_id = int(row[1])
        object_type = int(row[2])
        rack_rot = row[3]
        rack_pos = row[4:7]
        element_pos = row[7:10]
        element_position = element_pos + rack_pos
        element_prim_path = f"{elements_path}/element_{rack_id}_{palett_id}"
        new_prim = stage.DefinePrim(element_prim_path)
        new_prim.GetReferences().AddInternalReference(palett_root_prim.GetPrimPath())
        position_prim(new_prim, element_position, (0, 0, rack_rot-90))

The data is stored in a numpy array, where each entry contains IDs, position coordinates, and a type identifier. The type field enables me to place different pallet variants, not just a single type everywhere.

My assumption is that sequentially adding all these prim references is what makes this process so slow.
Is there a way to significantly speed this up, perhaps with some kind of parallelization or batch operation? Has anyone dealt with a similar situation or has optimization tips?

Any suggestions or advice would be greatly appreciated!

Best regards,

Axel

Hi @axel.goedrich,

The bottleneck is not the Python loop itself but USD change notifications. Each call to
GetReferences().AddInternalReference() individually notifies the Kit scene-graph, triggering
a viewport update and a Hydra re-sync. With 8,000 prims that is 8,000 round-trips through the
notification system. Parallelization is not available (the USD stage is not thread-safe for
authoring), but batching is – and the speedup is dramatic.

Fix 1: Wrap the loop in Sdf.ChangeBlock

Sdf.ChangeBlock defers all change notifications until the block exits. The entire loop
fires exactly one consolidated notification at the end instead of 8,000.

from pxr import Sdf

def position_elements_in_racks(elements_path, element_matrix, layout_element):
    stage = get_context().get_stage()
    palett_root_prim = load_element(
        layout_element['pallet']['asset_path'],
        layout_element['pallet']['asset_name']
    )
    palett_root_path = palett_root_prim.GetPrimPath()

    with Sdf.ChangeBlock():  # <-- only change needed
        for row in element_matrix:
            rack_id     = int(row[0])
            palett_id   = int(row[1])
            rack_rot    = row[3]
            element_pos = row[7:10] + row[4:7]
            element_prim_path = f"{elements_path}/element_{rack_id}_{palett_id}"
            new_prim = stage.DefinePrim(element_prim_path)
            new_prim.GetReferences().AddInternalReference(palett_root_path)
            position_prim(new_prim, element_pos, (0, 0, rack_rot - 90))

This is the same technique Isaac Sim uses internally for bulk stage edits (for example,
in the material-deduplication and geometry-routing rules in isaacsim.asset.transformer.rules).
See Sdf.ChangeBlock in the
OpenUSD API docs for details.

Fix 2: Temp-stage + Sdf.CopySpec (zero intermediate notifications)

For even better results, build all prims on a temporary in-memory stage first, then
copy the entire subtree to the live stage in one atomic operation. This is the approach
used by Isaac Sim’s Warehouse Creator extension when generating large warehouse layouts
(see Warehouse Creator):

from pxr import Gf, Sdf, Usd, UsdGeom

def position_elements_in_racks(elements_path, element_matrix, layout_element):
    ctx = get_context()
    live_stage = ctx.get_stage()
    palett_root_prim = load_element(
        layout_element['pallet']['asset_path'],
        layout_element['pallet']['asset_name']
    )
    palett_root_path = palett_root_prim.GetPrimPath()

    # Build on an in-memory stage -- no live notifications whatsoever
    temp_stage = Usd.Stage.CreateInMemory()

    for row in element_matrix:
        rack_id     = int(row[0])
        palett_id   = int(row[1])
        rack_rot    = row[3]
        element_pos = row[7:10] + row[4:7]
        element_prim_path = f"{elements_path}/element_{rack_id}_{palett_id}"
        new_prim = temp_stage.DefinePrim(element_prim_path, "Xform")
        new_prim.GetReferences().AddInternalReference(palett_root_path)
        xf = UsdGeom.Xformable(new_prim)
        xf.AddTranslateOp().Set(Gf.Vec3d(*element_pos))
        xf.AddRotateXYZOp().Set(Gf.Vec3f(0.0, 0.0, rack_rot - 90))

    # Atomic write to live stage: one notification for the whole subtree
    elements_path_sdf = Sdf.Path(elements_path)
    with Sdf.ChangeBlock():
        if live_stage.GetPrimAtPath(elements_path_sdf):
            live_stage.RemovePrim(elements_path_sdf)
        Sdf.CopySpec(
            temp_stage.GetRootLayer(), elements_path_sdf,
            live_stage.GetRootLayer(), elements_path_sdf,
        )

Fix 3: SetInstanceable(True) for identical pallet variants

For pallet variants that are visually identical (no per-pallet material overrides or
child-prim visibility differences), mark the prim instanceable:

new_prim.SetInstanceable(True)

This tells Kit’s renderer to treat all instanceable prims referencing the same asset as
a single hardware instance group, cutting GPU draw calls significantly. You cannot use
this for pallets that need individual material swaps or descendant overrides – in those
cases, leave instanceable=False (the default).

If pallets do not need independent per-instance rigid-body simulation

If each pallet does not need to move, collide, or be simulated independently as its own
rigid body, UsdGeom.PointInstancer
is the most efficient option. It stores all 8,000 positions as a single array and renders
with hardware instancing. As the Isaac Sim scripting docs
note, PointInstancer does support physics interaction – instances share the prototype’s
collision and physics properties – but each instance is not a separate simulated body.

from pxr import UsdGeom, Gf, Vt

point_instancer = UsdGeom.PointInstancer(
    stage.DefinePrim(f"{elements_path}/pallets", "PointInstancer")
)
point_instancer.CreatePrototypesRel().SetTargets([palett_root_prim.GetPath()])
point_instancer.CreateProtoIndicesAttr().Set([0] * len(element_matrix))
point_instancer.CreatePositionsAttr().Set(
    Vt.Vec3fArray([Gf.Vec3f(*(row[7:10] + row[4:7])) for row in element_matrix])
)

If each pallet must be an independent rigid body (separate mass, contacts, forces), stick
with Fix 1 or Fix 2 instead.


Summary:

Approach Effort Best for
Sdf.ChangeBlock around loop 1 line All cases; drop-in replacement
Temp stage + Sdf.CopySpec Moderate refactor Largest scenes; eliminates all intermediate updates
SetInstanceable(True) 1 line (supplement) Visually identical variants; cuts GPU draw calls
UsdGeom.PointInstancer Refactor Shared physics or visual-only; no independent per-pallet rigid bodies

Since this topic is stale, we are closing
it now. If you are still working on this or ran into any issues with the approaches above,
please open a new topic
and link back to this one for context.