# Question on Inference Performance Results of Qwen3 235B A22B on 2× DGX Spark

**URL:** <https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053>\
**Category:** DGX Spark / GB10\
**Tags:** cuda\
**Created:** [December 18, 2025, 6:30am UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053 "2025-12-18T06:30:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Turtle7777](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/turtle7777/32/457652_2.png) [@Turtle7777](https://forums.developer.nvidia.com/u/Turtle7777)\
**Post date:** [December 18, 2025, 6:30am UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/1 "2025-12-18T06:30:48Z")

</div>

Hi NV member:  
I tested the inference performance of Qwen3 235B A22B using two DGX SPARK systems. By following the steps below, I obtained the following results. Could you please let me know whether these numbers look reasonable?

Thank you.

- Test log:

[multi-node\_test\_log.txt](https://forums.developer.nvidia.com/uploads/short-url/2Fd9CQlO96lLpihfE9PMUFU1FLQ.txt) (24.6 KB)

- Test Result 4 times:

 ![image](https://global.discourse-cdn.com/nvidia/original/4X/0/6/5/065a2bce122fbcf7f18679318eecd3b382eb70ea.png)

- Test ways:

> _ **Instructions to run Qwen3 235B A22B on 2xDGX Spark** _
> 
> # Set permissions for trtllm-mn-entrypoint.sh file attached in the folder
> 
> chmod +x ./trtllm-mn-entrypoint.sh
> 
> # Run this command on both Spark nodes to start the TensorRT-LLM containers with proper networking and GPU access
> 
> docker run --name trtllm --rm -d   
> –gpus all --network host --ipc=host   
> –ulimit memlock=-1 --ulimit stack=67108864   
> -e UCX\_NET\_DEVICES=enp1s0f0np0,enp1s0f1np1   
> -e NCCL\_SOCKET\_IFNAME=enp1s0f0np0,enp1s0f1np1   
> -e OMPI\_MCA\_btl\_tcp\_if\_include=enp1s0f0np0,enp1s0f1np1   
> -e OMPI\_ALLOW\_RUN\_AS\_ROOT=1   
> -e OMPI\_ALLOW\_RUN\_AS\_ROOT\_CONFIRM=1   
> -v $HOME/.cache/huggingface/:/root/.cache/huggingface/   
> -v ./trtllm-mn-entrypoint.sh:/opt/trtllm-mn-entrypoint.sh   
> -v ~/.ssh:/tmp/.ssh:ro   
> –entrypoint /opt/trtllm-mn-entrypoint.sh   
> [nvcr.io/nvidia/tensorrt-llm/release:1.0.0rc3](http://nvcr.io/nvidia/tensorrt-llm/release:1.0.0rc3)
> 
> # Run the inference from one of the two DGX Spark systems
> 
> docker exec trtllm bash -c ‘cat \< /tmp/extra-llm-api-config.yml  
> print\_iter\_log: false  
> kv\_cache\_config:  
> dtype: “fp8”  
> free\_gpu\_memory\_fraction: 0.9  
> cuda\_graph\_config:  
> enable\_padding: true  
> EOF’
> 
> # Initiate the LLM benchmarking for Qwen3 235B from one of the two DGX Spark systems. Specify huggingface token for downloading the model.
> 
> export HF\_TOKEN=
> 
> docker exec   
> -e ISL=128 -e OSL=128   
> -e MODEL=“nvidia/Qwen3-235B-A22B-FP4”   
> -e HF\_TOKEN=$HF\_TOKEN   
> -it trtllm bash -c ’  
> mpirun -x HF\_TOKEN=$HF\_TOKEN -np 2 -H 192.168.1.10:1,192.168.1.11:1 bash -c “huggingface-cli download $MODEL” &&   
> mpirun -x HF\_TOKEN=$HF\_TOKEN -np 2 -H 192.168.1.10:1,192.168.1.11:1 bash -c “python benchmarks/cpp/prepare\_dataset.py --tokenizer=$MODEL --stdout token-norm-dist --num-requests=1 --input-mean=$ISL --output-mean=$OSL --input-stdev=0 --output-stdev=0 \> /tmp/dataset.txt” &&   
> mpirun -x HF\_TOKEN=$HF\_TOKEN -np 2 -H 192.168.1.10:1,192.168.1.11:1 trtllm-llmapi-launch trtllm-bench -m $MODEL throughput   
> –tp 2   
> –dataset /tmp/dataset.txt   
> –backend pytorch   
> –max\_num\_tokens 4096   
> –concurrency 1   
> –max\_batch\_size 4   
> –extra\_llm\_api\_options /tmp/extra-llm-api-config.yml   
> –streaming’

---

<div class="post-metadata">

**Author:** ![NVES](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/nves/32/14043_2.png) [@NVES](https://forums.developer.nvidia.com/u/NVES)\
**Post date:** [December 18, 2025, 4:00pm UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/2 "2025-12-18T16:00:18Z")

</div>

Hi, I encourage the community to weigh in on performance expectations, as we don’t comment beyond the workloads already published here [How NVIDIA DGX Spark’s Performance Enables Intensive AI Tasks | NVIDIA Technical Blog](https://developer.nvidia.com/blog/how-nvidia-dgx-sparks-performance-enables-intensive-ai-tasks/)

---

<div class="post-metadata">

**Author:** ![eugr](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/eugr/32/449615_2.png) [@eugr](https://forums.developer.nvidia.com/u/eugr)\
**Post date:** [December 18, 2025, 6:17pm UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/3 "2025-12-18T18:17:36Z")

</div>

The inference speeds seems to be slow. I get ~25 tokens/sec on my two node cluster using QuantTrio/Qwen3-VL-235B-A22B-Instruct-AWQ in vLLM.

A few things to keep in mind:

- TRTLLM container you are using is outdated, there is a newer version available.
- NVFP4 support on DGX Spark is still lacking, as of today at least, you’ll get noticeably better performance using AWQ quants (that actually have slightly better accuracy as they are activation aware and keep activation weights at 16 bits).
- vLLM is the best way to run LLMs on Spark. You can either use NVIDIA’s 25.11-py3 container, or one of the community builds here if you want the latest vllm features not supported in 0.11.2.

---

<div class="post-metadata">

**Author:** ![Turtle7777](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/turtle7777/32/457652_2.png) [@Turtle7777](https://forums.developer.nvidia.com/u/Turtle7777)\
**Post date:** [December 19, 2025, 3:45am UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/4 "2025-12-19T03:45:41Z")

</div>

Hi @eugr  
As you mentioned, 15 tokens/s is indeed too slow, which is why I wanted to clarify whether the data was correct. After re-validating using the approach you suggested, the results now reach over 20 tokens/s, which meets expectations.

Thank you for your suggestions and explanations.

- TRTLLM container you are using is outdated, there is a newer version available.  
[Turtle7777] I verified two DGX Spark systems by following NVIDIA’s TensorRT-LLM SOP. As you suggested, I removed the TensorRT-LLM container and downloaded it again, but the version is still TensorRT-LLM version: 1.0.0rc3. After re-running the tests, the performance results now match what I previously saw online, i.e., above 20 tokens/s.  
I am not sure whether this improvement is related to the kernel version change from 6.14.0-1013-nvidia to 6.14.0-1015-nvidia. Please refer to the data and log files below.

- Environments:  
========================  
EC: 2.75.3.3  
SOC FW Version: 3.0.4  
PD0 FW1: 5.0  
PD1 FW1: 5.0  
GOP Driver Version: 9000AE0  
DGX SPARK OS: 7.3.1 2025-11-12-09-12-21  
Kernel: 6.14.0-1015-nvidia  
SSD: 4TB Gen4 Phison  
=========================  
[20251219\_multinode\_test\_log.txt](https://forums.developer.nvidia.com/uploads/short-url/q1QGmylXjw2VOvb5Bv1zSFktH6A.txt) (27.2 KB)  
 ![image](https://global.discourse-cdn.com/nvidia/original/4X/0/3/8/03802c30c13c68dcc4cb9560cb55571cf1cd12ce.png)

- NVFP4 support on DGX Spark is still lacking, as of today at least, you’ll get noticeably better performance using AWQ quants (that actually have slightly better accuracy as they are activation aware and keep activation weights at 16 bits).
- vLLM is the best way to run LLMs on Spark. You can either use NVIDIA’s 25.11-py3 container, or one of the community builds here if you want the latest vllm features not supported in 0.11.2.  
[Turtle7777] I will further follow your suggestions and verify the performance of multiple DGX Spark systems using vLLM and AWQ quants, to see whether the performance still meets expectations.

---

<div class="post-metadata">

**Author:** ![vgoklani](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/vgoklani/32/11912_2.png) [@vgoklani](https://forums.developer.nvidia.com/u/vgoklani)\
**Post date:** [December 19, 2025, 5:09am UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/5 "2025-12-19T05:09:47Z")

</div>

The images are published here:

> **[TensorRT-LLM Release | NVIDIA NGC](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags?version=1.2.0rc5)**
>
> TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.

This is the latest version: `nvcr.io/nvidia/tensorrt-llm/release:1.2.0rc5`

---

<div class="post-metadata">

**Author:** ![Turtle7777](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/turtle7777/32/457652_2.png) [@Turtle7777](https://forums.developer.nvidia.com/u/Turtle7777)\
**Post date:** [December 19, 2025, 7:32am UTC](https://forums.developer.nvidia.com/t/question-on-inference-performance-results-of-qwen3-235b-a22b-on-2x-dgx-spark/355053/6 "2025-12-19T07:32:22Z")

</div>

Hi @vgoklani :  
After switching to the latest container, [nvcr.io/nvidia/tensorrt-llm/release:1.2.0rc5](http://nvcr.io/nvidia/tensorrt-llm/release:1.2.0rc5), I was able to run the tests, but I noticed the following related errors and am not sure whether they affect the results. I saw the message “triton is not supported on current platform, roll back to CPU” as well as another backtrace log shown below.

Does this error mean that the test is falling back to using the CPU? However, the measured performance is still 21.65 tokens/s. Is there anything that still needs to be changed in the command or configuration to obtain more accurate results?

Thank you.

[20251219\_twonode\_test\_log\_1.2.0rc5.txt](https://forums.developer.nvidia.com/uploads/short-url/5CN80N5R1jCbAKIfnw71afWitmA.txt) (61.1 KB)

 ![image](https://global.discourse-cdn.com/nvidia/original/4X/f/b/0/fb080cded617375b09dc01a478dec77b94dd60e6.png)

> W1219 07:19:56.237000 1900 torch/utils/cpp\_extension.py:2422] If this is not desired, please set os.environ[‘TORCH\_CUDA\_ARCH\_LIST’] to specific architectures.  
> /tmp/tmpvif\_h2p3/cuda\_utils.c:1:10: fatal error: cuda.h: No such file or directory  
> 1 | #include “cuda.h”  
> | ^ ~~~~~~~  
> compilation terminated.  
> /usr/local/lib/python3.12/dist-packages/tensorrt\_llm/\_torch/modules/fla/utils.py:216: UserWarning: Triton is not supported on current platform, roll back to CPU.  
> warnings.warn(  
> /tmp/tmprc0xglrj/cuda\_utils.c:1:10: fatal error: cuda.h: No such file or directory  
> 1 | #include “cuda.h”  
> | ^ ~~~~~~~  
> compilation terminated.  
> …  
> [12/19/2025-07:25:49] [TRT-LLM] [RANK 0] [W] [Autotuner] Failed when profiling runner=\<tensorrt\_llm.\_torch.custom\_ops.torch\_custom\_ops.MoERunner object at 0xf3836e0c6330\>, tactic=6, shapes=[torch.Size([1, 2048]), torch.Size([128, 1536, 256]), torch.Size([0]), torch.Size([128, 4096, 48]), torch.Size([0])]. Error: [TensorRT-LLM][ERROR] Assertion failed: Failed to initialize cutlass TMA WS grouped gemm. Error: Error Internal (tensorrt\_llm/kernels/cutlass\_kernels/cutlass\_instantiations/gemm\_grouped/120/cutlass\_kernel\_file\_gemm\_grouped\_sm120\_M128\_BS\_group2.generated.cu:39)  
> 1 0xf387a2495ec8 tensorrt\_llm::common::throwRuntimeError(char const\*, int, char const\*) + 120  
> 2 0xf387a50fe114 /usr/local/lib/python3.12/dist-packages/tensorrt\_llm/libs/libth\_common.so(+0x328e114) [0xf387a50fe114]  
> 3 0xf387a50fe4e4 void tensorrt\_llm::kernels::cutlass\_kernels\_oss::tma\_warp\_specialized\_generic\_moe\_gemm\_kernelLauncher\<cutlass::arch::Sm120, \_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, void, tensorrt\_llm::cutlass\_extensions::EpilogueOpDefault, (tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput::EpilogueFusion)3, cute::tuple\<cute::C\<128\>, cute::C\<256\>, cute::C\<128\> \>, cute::tuple\<cute::C\<1\>, cute::C\<1\>, cute::C\<1\> \>, false, false, false, false\>(tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput, int, int, CUstream\_st\*, int\*, unsigned long\*, cute::tuple\<int, int, cute::C\<1\> \>, cute::tuple\<int, int, cute::C\<1\> \>) + 84  
> 4 0xf387a37cd488 void tensorrt\_llm::kernels::cutlass\_kernels\_oss::dispatchMoeGemmSelectClusterShapeTmaWarpSpecialized\<cutlass::arch::Sm120, \_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, tensorrt\_llm::cutlass\_extensions::EpilogueOpDefault, (tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput::EpilogueFusion)3, cute::tuple\<cute::C\<128\>, cute::C\<256\>, cute::C\<128\> \> \>(tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput, int, tensorrt\_llm::cutlass\_extensions::CutlassGemmConfig, int, CUstream\_st\*, int\*, unsigned long\*) + 216  
> 5 0xf387a37ce1b8 void tensorrt\_llm::kernels::cutlass\_kernels\_oss::dispatchMoeGemmSelectTileShapeTmaWarpSpecialized\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, tensorrt\_llm::cutlass\_extensions::EpilogueOpDefault, (tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput::EpilogueFusion)3\>(tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput, int, tensorrt\_llm::cutlass\_extensions::CutlassGemmConfig, int, CUstream\_st\*, int\*, unsigned long\*) + 1000  
> 6 0xf387a37b143c void tensorrt\_llm::kernels::cutlass\_kernels::MoeGemmRunner\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_bfloat16\>::dispatchToArch\<tensorrt\_llm::cutlass\_extensions::EpilogueOpDefault\>(tensorrt\_llm::kernels::cutlass\_kernels::GroupedGemmInput\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_bfloat16\>, tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput) + 236  
> 7 0xf387a37b1f10 tensorrt\_llm::kernels::cutlass\_kernels::MoeGemmRunner\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_bfloat16\>::moeGemmBiasAct(tensorrt\_llm::kernels::cutlass\_kernels::GroupedGemmInput\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_bfloat16\>, tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput) + 272  
> 8 0xf387a3797ce8 tensorrt\_llm::kernels::cutlass\_kernels::CutlassMoeFCRunner\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, void\>::gemm2(tensorrt\_llm::kernels::cutlass\_kernels::MoeGemmRunner\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_bfloat16\>&, tensorrt\_llm::kernels::fp8\_blockscale\_gemm::CutlassFp8BlockScaleGemmRunnerInterface\*, \_\_nv\_fp4\_e2m1 const\*, void\*, \_\_nv\_bfloat16\*, long const\*, tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput, \_\_nv\_fp4\_e2m1 const\*, \_\_nv\_bfloat16 const\*, \_\_nv\_bfloat16 const\*, float const\*, unsigned char const\*, tensorrt\_llm::kernels::cutlass\_kernels::QuantParams, float const\*, float const\*, int const\*, int const\*, int const\*, long const\*, long, long, long, long, long, long, int, long, float const\*\*, bool, void\*, CUstream\_st\*, tensorrt\_llm::kernels::cutlass\_kernels::MOEParallelismConfig, bool, tensorrt\_llm::cutlass\_extensions::CutlassGemmConfig, bool, int\*, int\*) + 744  
> 9 0xf387a3799478 tensorrt\_llm::kernels::cutlass\_kernels::CutlassMoeFCRunner\<\_\_nv\_fp4\_e2m1, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, \_\_nv\_fp4\_e2m1, \_\_nv\_bfloat16, void\>::gemm2(void const\*, void\*, void\*, long const\*, tensorrt\_llm::kernels::cutlass\_kernels::TmaWarpSpecializedGroupedGemmInput, void const\*, void const\*, void const\*, float const\*, unsigned char const\*, tensorrt\_llm::kernels::cutlass\_kernels::QuantParams, float const\*, float const\*, int const\*, int const\*, int const\*, long const\*, long, long, long, long, long, long, int, long, float const\*\*, bool, void\*, bool, CUstream\_st\*, tensorrt\_llm::kernels::cutlass\_kernels::MOEParallelismConfig, bool, tensorrt\_llm::cutlass\_extensions::CutlassGemmConfig, bool, int\*, int\*) + 408  
> 10 0xf387a36c7c8c tensorrt\_llm::kernels::cutlass\_kernels::GemmProfilerBackend::runProfiler(int, tensorrt\_llm::cutlass\_extensions::CutlassGemmConfig const&, char\*, void const\*, CUstream\_st\* const&) + 2776  
> 11 0xf387a2894cec torch\_ext::FusedMoeRunner::runGemmProfile(at::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, long, long, long, long, long, long, long, bool, bool, long, long, bool, long, long) + 476  
> 12 0xf387a28a4de0 std::_Function\_handler\<void (std::vector\<c10::IValue, std::allocatorc10::IValue \>&), torch::class_\<torch\_ext::FusedMoeRunner\>::defineMethod\<torch::detail::WrapMethod\<void (torch\_ext::FusedMoeRunner::_)(at::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, long, long, long, long, long, long, long, bool, bool, long, long, bool, long, long)\> \>(std::\_\_cxx11::basic\_string\<char, std::char\_traits, std::allocator \>, torch::detail::WrapMethod\<void (torch\_ext::FusedMoeRunner::_)(at::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, at::Tensor const&, std::optionalat::Tensor const&, long, long, long, long, long, long, long, bool, bool, long, long, bool, long, long)\>, std::\_\_cxx11::basic\_string\<char, std::char\_traits, std::allocator \>, std::initializer\_listtorch::arg)::{lambda(std::vector\<c10::IValue, std::allocatorc10::IValue \>&)#1}\>::\_M\_invoke(std::\_Any\_data const&, std::vector\<c10::IValue, std::allocatorc10::IValue \>&) + 576  
> 13 0xf38907ecaea0 /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch\_python.so(+0xceaea0) [0xf38907ecaea0]  
> 14 0xf38907ecb650 /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch\_python.so(+0xceb650) [0xf38907ecb650]  
> 15 0xf38907fac478 /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch\_python.so(+0xdcc478) [0xf38907fac478]  
> 16 0xf38907faca70 /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch\_python.so(+0xdcca70) [0xf38907faca70]  
> 17 0xf3890778eb40 /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch\_python.so(+0x5aeb40) [0xf3890778eb40]  
> 18 0x503454 python3() [0x503454]  
> 19 0x4c2d1c \_PyObject\_MakeTpCall + 124  
> 20 0x4c6f88 python3() [0x4c6f88]  
> 21 0x528ab4 python3() [0x528ab4]  
> 22 0x4c2d1c \_PyObject\_MakeTpCall + 124  
> 23 0x563824 \_PyEval\_EvalFrameDefault + 2208  
> 24 0x4c6ee8 python3() [0x4c6ee8]  
> 25 0x4c5278 PyObject\_Call + 280  
> 26 0x566dd4 \_PyEval\_EvalFrameDefault + 15952  
> 27 0x4c48b4 \_PyObject\_Call\_Prepend + 436  
> 28 0x528970 python3() [0x528970]  
> 29 0x4c51cc PyObject\_Call + 108  
> 30 0x566dd4 \_PyEval\_EvalFrameDefault + 15952  
> 31 0x4c6ee8 python3() [0x4c6ee8]  
> 32 0x4c5278 PyObject\_Call + 280  
> 33 0x566dd4 \_PyEval\_EvalFrameDefault + 15952  
> 34 0x4c6ee8 python3() [0x4c6ee8]  
> 35 0x4c5278 PyObject\_Call + 280  
> 36 0x566dd4 \_PyEval\_EvalFrameDefault + 15952  
> …
