DeepSeek v4 Flash (Aiden Recipe from Reddit) - 1M token session operational, Cuda 12.1 tailored for DGX Spark GB10

Thank you!


Use ./run-recipe.sh deepseek-v4-flash -d --no-ray to run the containers indaemon mode or detached mode

➜  spark-vllm-docker git:(main) ✗ ./run-recipe.sh deepseek-v4-flash -d --no-ray                                                                                                                                                                                 git:(main|…2⚑1 
Recipe: DeepSeek-V4-Flash
  vLLM serving deepseek-ai/DeepSeek-V4-Flash on a DGX Spark cluster

Using cluster nodes from .env: 192.168.177.11, 192.168.177.12

=== Launching ===
Container: vllm-node
Cluster: 2 nodes

Loading configuration from .env file...
Loaded .env variables: DOTENV_CLUSTER_NODES DOTENV_ETH_IF DOTENV_IB_IF DOTENV_LOCAL_IP 
Using launch script: /tmp/tmp1163vpa7.sh
Head Node: 192.168.177.11
Worker Nodes: 192.168.177.12
Container Name: vllm_node
Image Name: vllm-node
Action: exec
Checking SSH connectivity to worker nodes...
  SSH to 192.168.177.12: OK
Starting Head Node on 192.168.177.11...
44bd94ebec2e93a444bc6a20601ad19254631e9d4045684b5b4d578771286dc1
Starting Worker Node on 192.168.177.12...
eee57dc9af4e316af48caa5494f7b34e071d4949a1f02ed7c6340d075725497c
Copying launch script to head node (192.168.177.11)...
Successfully copied 3.07kB to vllm_node:/workspace/exec-script.sh
Copying launch script to worker 192.168.177.12...
vllm_node_script_PYtTsr.sh                                                                                                                                                                                                                       100% 1097   708.8KB/s   00:00    
Executing command: /workspace/exec-script.sh
Launching worker (rank 1) on 192.168.177.12...
Executing command on head node (rank 0): /workspace/exec-script.sh
Command dispatched in background (Daemon mode). Container: vllm_node