An interactive explorer over a real eval-harness snapshot: rubric-graded scores across reasoning, code repair, extraction, long context, tool use, and a domain task where a fine-tuned 8B beats the frontier models at 4% of the cost.
Heatmap below — swipe sideways to see every model →
| Task | GPT-4o | Claude | Gemini | Llama | Mixtral | Mashdun |
|---|---|---|---|---|---|---|
| Reasoning | 88 | 92 | 90 | 71 | 62 | 58 |
| Code | 91 | 94 | 86 | 74 | 60 | 41 |
| Extraction | 93 | 90 | 94 | 68 | 57 | 63 |
| Long context | 84 | 91 | 93 | 55 | 48 | 45 |
| Tool use | 92 | 93 | 88 | 70 | 52 | 66 |
| Domain | 78 | 82 | 76 | 64 | 58 | 95 |
Constraint-satisfaction scheduling puzzle: 7 staff, 12 shifts, overlapping rules. Graded on whether the final schedule violates zero constraints.
Worked the constraints symbolically before assigning shifts; 5/5 valid schedules.
4/5 valid; one run double-booked a split shift.
Valid on relaxed rules, missed the rest-period constraint twice.
Domain fine-tune trades general reasoning for concierge tasks — expected.
The fine-tuned 8B wins its domain at ~4% of frontier cost — the point of custom models: not beating GPT-4o everywhere, beating it where your product lives.