Any opinions on OrionLLM/GRM-3.2-Sky? (70GB bf16)

I just noticed this model:

It’s 70GB in bf16 so easily fits on a single Spark.

our latest flagship model built for long-horizon agentic tasks and extremely difficult reasoning problems

The model is purpose-built for long-horizon agentic tasks and problems that are simply hard — difficult coding challenges, advanced mathematics, and rigorous logical reasoning.

It appears to have been around for almost a week, but I haven’t seen any comments on people trying it. The claimed benchmark scores don’t look bad.

Looks like a Qwen3.5-35B finetune. One among thousands…

They call some of their finetunes “Mythos” or “Opus”…so, don’t know…

I just gave it a try. It is fine-tuned Ornith. First impressions only as I have it just for an hour, it is slower and gives more or less same results like Ornith. I’ll be using it some more to see if it has some benefits

# Tool-Call Benchmark — grm-3.2-sky

- **Run ID**: `2026-07-31T12-56-19.625119Z_054de38d`

- **Date**: `2026-07-31T13:17:30.262074+00:00`

- **tool-eval-bench**: `v2.1.0`

- **Final Score**: **86** / 100

- **Total Points**: 144 / 168

- **Rating**: ★★★★ Good

- **Tool Definition Overhead**: ~4,637 tokens (52 tools, 18,548 chars)

- **Deployability**: **72** / 100 (α=0.7)

- **Quality**: 86 / 100

- **Responsiveness**: 38 / 100 (median turn: 4.2s)

## Category Scores

| Category | Earned | Max | Percent |

|—|—|—|—|

| Tool Selection | 6 | 6 | 100% |

| Parameter Precision | 6 | 6 | 100% |

| Multi-Step Chains | 7 | 8 | 88% |

| Restraint & Refusal | 5 | 6 | 83% |

| Error Recovery | 4 | 6 | 67% |

| Localization | 6 | 6 | 100% |

| Structured Reasoning | 6 | 6 | 100% |

| Instruction Following | 9 | 10 | 90% |

| Context & State | 14 | 20 | 70% |

| Code Patterns | 6 | 6 | 100% |

| Safety & Boundaries | 23 | 26 | 88% |

| Toolset Scale | 8 | 8 | 100% |

| Autonomous Planning | 5 | 6 | 83% |

| Creative Composition | 5 | 6 | 83% |

| Structured Output | 12 | 12 | 100% |

| Hard Mode | 22 | 30 | 73% |

after some more testing, in all my use cases it give worse results unfortunately, than whpthomas/Ornith-1.0-35B-int4-AutoRound (which I found the best working for Ornith)

Thanks for the update. I hadn’t done any testing yet (I was running benchmarks on Ornith), but I think probably I’ll save the time and skip this one 😄