I just gave it a try. It is fine-tuned Ornith. First impressions only as I have it just for an hour, it is slower and gives more or less same results like Ornith. I’ll be using it some more to see if it has some benefits
# Tool-Call Benchmark — grm-3.2-sky
- **Run ID**: `2026-07-31T12-56-19.625119Z_054de38d`
- **Date**: `2026-07-31T13:17:30.262074+00:00`
- **tool-eval-bench**: `v2.1.0`
- **Final Score**: **86** / 100
- **Total Points**: 144 / 168
- **Rating**: ★★★★ Good
- **Tool Definition Overhead**: ~4,637 tokens (52 tools, 18,548 chars)
- **Deployability**: **72** / 100 (α=0.7)
- **Quality**: 86 / 100
- **Responsiveness**: 38 / 100 (median turn: 4.2s)
## Category Scores
| Category | Earned | Max | Percent |
|—|—|—|—|
| Tool Selection | 6 | 6 | 100% |
| Parameter Precision | 6 | 6 | 100% |
| Multi-Step Chains | 7 | 8 | 88% |
| Restraint & Refusal | 5 | 6 | 83% |
| Error Recovery | 4 | 6 | 67% |
| Localization | 6 | 6 | 100% |
| Structured Reasoning | 6 | 6 | 100% |
| Instruction Following | 9 | 10 | 90% |
| Context & State | 14 | 20 | 70% |
| Code Patterns | 6 | 6 | 100% |
| Safety & Boundaries | 23 | 26 | 88% |
| Toolset Scale | 8 | 8 | 100% |
| Autonomous Planning | 5 | 6 | 83% |
| Creative Composition | 5 | 6 | 83% |
| Structured Output | 12 | 12 | 100% |
| Hard Mode | 22 | 30 | 73% |