Mashdun
AI GrowthWorkCapabilitiesIntegrationsProcessLabsMarketplaceBlogAbout
Get in Touch
Mashdun

Full-stack web developer, AI engineer & growth marketer. Building production-grade apps and intelligent solutions.

Navigation

  • Portfolio
  • Capabilities
  • Process
  • Labs
  • About

Resources

  • Blog / Notes
  • Marketplace
  • Contact
  • RSS Feed

Legal

  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Do Not Sell or Share
  • Delete my data

© 2026 Mashdun. All rights reserved.

Built with Next.js, Tailwind CSS & Prisma

    All Labs
    AI
    Evals
    Fine-tuning
    GPT-4o
    Claude
    Gemini

    LLM Benchmark Arena

    An interactive explorer over a real eval-harness snapshot: rubric-graded scores across reasoning, code repair, extraction, long context, tool use, and a domain task where a fine-tuned 8B beats the frontier models at 4% of the cost.

    LLM Benchmark Arena

    Eval Harness
    6 Models
    Rubric Grading
    Fine-tuning

    Overall (avg across 6 hard tasks)

    Claude Sonnet 4.5Anthropic
    90
    Anthropic
    GPT-4oOpenAI
    88
    OpenAI
    Gemini 2.5 ProGoogle
    88
    Google
    Llama 3.3 70BMeta (open)
    67
    Meta (open)
    Mashdun Concierge-8BCustom fine-tune
    61
    Custom fine-tune
    Mixtral 8×22BMistral (open)
    56
    Mistral (open)

    Score matrix — pick a task to see model verdicts

    Heatmap below — swipe sideways to see every model →

    TaskGPT-4oClaudeGeminiLlamaMixtralMashdun
    Reasoning889290716258
    Code919486746041
    Extraction939094685763
    Long context849193554845
    Tool use929388705266
    Domain788276645895

    Multi-step reasoning

    Constraint-satisfaction scheduling puzzle: 7 staff, 12 shifts, overlapping rules. Graded on whether the final schedule violates zero constraints.

    Claude Sonnet 4.5pass

    Worked the constraints symbolically before assigning shifts; 5/5 valid schedules.

    92
    GPT-4opass

    4/5 valid; one run double-booked a split shift.

    88
    Llama 3.3 70Bpartial

    Valid on relaxed rules, missed the rest-period constraint twice.

    71
    Mashdun Concierge-8Bfail

    Domain fine-tune trades general reasoning for concierge tasks — expected.

    58

    How it works

    1. Fixed prompt set, 6 task families, 5 runs per model per task
    2. Rubric-graded by a judge model, spot-checked by hand
    3. Scores averaged; this page explores the snapshot
    4. Same harness we run before picking a model for client work

    Cost & speed

    GPT-4o3.1s48/1k
    Claude Sonnet 4.52.8s42/1k
    Gemini 2.5 Pro3.4s39/1k
    Llama 3.3 70B2.2s9/1k
    Mixtral 8×22B1.9s7/1k
    Mashdun Concierge-8B0.9s2/1k

    The fine-tuned 8B wins its domain at ~4% of frontier cost — the point of custom models: not beating GPT-4o everywhere, beating it where your product lives.

    Want your product benchmarked?