Perplexity Computer on DGX Spark — first impressions + memory/telemetry numbers for local-first folks

Posting measured observations after installing Perplexity Computer v26.9.0 (arm64) on a
DGX Spark (GB10, 128 GB unified), since I didn’t find community notes on its memory footprint and what leaves the machine. Everything captured below is measured with ss, tshark, nvidia-smi, and the app bundle itself. Summary below pulled from Claude Code CLI.

What’s good

  • The local model server is a standard OpenAI-compatible endpoint at http://127.0.0.1:45793/v1`` (llama-server style), bound to loopback (not 0.0.0.0). So it’s composable — you can point your own tooling at it — and not exposed to the LAN.
  • Cloud escalation is gated per-action. Every step that would leave the box (web search,
    adviser call, GitHub PR, etc.) throws an explicit Allow/Deny with a plain-language “what leaves.”
  • On-device inference runs fully on Spark; no per-token cost for local steps.

Memory reality (biggest disappointment)

  • Install pulls PPLX 27B (Qwen-27B post-trained, 4-bit, ~27.6 GB download). PPLX is
    selectable as the orchestrator; Nemotron 3.5 Lightning (~19 GB) is “coming soon.”
  • A 27B 4-bit model is ~15–20 GB of weights. But the loaded vLLM engine reserved ~82 GB
    (nvidia-smi: 83,984 MiB on one engine core) on the 128 GB unified box.
  • That’s vLLM’s default gpu-memory-utilization ≈ 0.9 KV-cache preallocation, not model size. On unified memory (CPU+GPU shared) it pushed system RAM to ~121 GB, into swap, and the GPU allocator hit NVRM ... NV_ERR_NO_MEMORY — any other GPU app (I tried launching Firefox) crashes while the model is loaded. There is no user-facing knob to lower the reservation (user-settings.json only picks the model).
  • “Close window” ≠ quit. Closing left ~12 Electron processes + the model resident. Right-click
    the local model → Stop to release the ~82 GB; the app minimizes to background rather than exit.

Takeaway:
Perplexity Computer behaves like the DGX Spark is its dedicated inference server and expects to own the box — may be great on a headless node, rough as a desktop app if you are on the Spark through NVIDIA OS. Plan for one-model-owns-the-machine reserving memory for orchestration.

Telemetry (heavier than “local-first” implies)

Static scan of /opt/Perplexity shows an embedded analytics stack: Sentry, Datadog, Segment, Heap, Amplitude, New Relic. Observed on the wire (tshark, SNI/DNS):

  • Persistent HTTPS to Datadog intakes (browser-intake-datadoghq.com, http-intake.logs.us5.datadoghq.com).
  • ~195 Datadog “span” batches generated locally in a single ~2–3 hr session (~/.config/Perplexity/spans/).
  • At idle (no task, no apps, browser closed) the Perplexity Computer app holds ~8 persistent connections to its backend (Cloudflare-fronted) + Google/GCP. Toggling Remote Access off dropped ~4 of them (the GCP relay).
  • The crash reporter is mandatory — the app FATALs on launch if chrome_crashpad_handler can’t spawn.

None of this is required for local inference to run — it’s vendor-side observability/product
analytics. Blocking it has no functional downside to on-device work but it’s hard-set as mandatory out telemetry for Perplexity Computer to work.

One caveat worth knowing: the Privacy Gate has a gap

The pre-escalation Privacy Gate checks plain-text and code files only. Per its own disclaimer, PDFs, Office docs, images, audio, and video are sent without a PII check — i.e., exactly the formats sensitive documents usually are. For confidential-doc workflows, keep them local and don’t approve escalation.

Practical tips

  • Reclaim memory: right-click model → Stop; don’t rely on closing the window.
  • See what it talks to: ss -tunp | grep -i perplex and
    sudo tshark -i any -Y 'tls.handshake.type==1' -T fields -e tls.handshake.extensions_server_name.
  • Silence the behavioral telemetry (null-route both IPv4 and IPv6 in /etc/hosts, since the app connects over v6): browser-intake-datadoghq.com, http-intake.logs.us5.datadoghq.com,
    o<id>.ingest.us.sentry.io, api.segment.io, api.heapanalytics.com, api.amplitude.com,bam.nr-data.net. Leave *.perplexity.ai (that’s the app’s functional backend).
  • Heads-up: blocking the intakes makes the local span batches pile up (they can’t upload), so
    sweep ~/.config/Perplexity/spans/ on a cron.

Initial Verdict

Promising, but unless the app behavior changes, budget for a dedicated node — the model appears to claim the whole box — and go in knowing the “local-first” label covers inference, not the telemetry plane. Unless you choose to share tracking, terminate the behavioral analytics; it maps usage for the vendor, not performance for you.

Side Note:
Perplexity Computer offers access to your Enterprise/Pro subscription so your local activity can be accessed online if you toggle remote on, you can also allow for web search and escalation; so you can choose local-only without remote access, web search or escalation, or local-accessible where you can access your local files while allowing cloud escalation to Perplexity frontier models, web search/ deep research, remote access from your Perplexity account.

These modes govern the app’s intentional egress. The observability/telemetry plane and the idle backend connection persist regardless — so treat the mode selector as convenience, not a guarantee. True local-only comes from enforcing it at the network (see the /etc/hosts block above), not from the app dropdown.

Will share more impressions after running a few local orchestration tasks.

Tested: DGX Spark GB10 / 128 GB / DGX OS (aarch64), Perplexity Computer 26.9.0, Perplexity Enterprise Max.

I am going to test it against a custom inference model I have already running in my Spark. Another option is running PPLX Qwen model as if it were a regular model and point perplexity to that model. That way you can control memory and KV cache allocation.

Thanks, I don’t have Perplexity myself but know a few people who do. I did watch the Nvidia/Perpexity youtube, I’m more of an “under the hood” type of person so it’s probably not really my thing. Nice to read some real experience other than the marketing talk ;)

Thanks for testing it against your custom inference and sharing with the community! PPLX grabs memory for it’s app to run, so you may see minimal difference since you’d still launch PPLX and point to PPLX’s model. Some of the consumed memory was released when app behavioral telemetry was disabled. I didn’t check the exact breakdown of how much was consumed by outbound telemetry, vs.running PPLX app on NIVIDIA OS vs. headless.