vibehacker
Discuss
Omar Hassan
21 days ago

mistral-small ran my whole eval suite for $1.80

Mistral
Le Chat and frontier open-weight models from Mistral AI

Tried swapping Claude Sonnet out of our nightly eval harness for mistral-small-latest on the API.

120 cases, mostly JSON schema + a few tool-call traces. Finished in ~18 min. Claude was ~$11 last week for the same run. This one landed at $1.80.

Quality dip on the multi-step tool cases though — case 47 invents a cancel_invoice endpoint we never had. Still keeping it for the cheap regression pass and only burning Sonnet on the hard suite.

Anyone else running dual-model evals like this or am I overcomplicating it?

5 comments

Join the discussion

Log in to comment.

  • Renee

    we do the same split. small model for smoke, big model for the golden set.

    watch the schema-strict mode though — mistral-small sometimes drops required fields even when you pass response_format. we had to add a jsonschema validate step before scoring or the pass rate looks fake-good.

    • Pixel

      yeah the schema-strict thing bit us too. got ZodError: Required at "customer_id" on like 12/40 smoke cases even with response_format json_object.

      we just validate with zod before scoring now. otherwise the $1.80 suite lies to you. did you pin a dated mistral revision or still on -latest?

  • Cedar Syntax

    $1.80 is cute until case 47 ships to prod because someone trusted the cheap suite.

    we pin mistral-small only on unit-ish prompts now. anything with tools stays on sonnet. learned that after it "fixed" our zod schema by deleting half the fields.

    • Eden Chen

      same. we run mistral-small for the 90 "does the json even parse" cases and keep sonnet on the 30 tool-trace ones.

      case 47 vibes — our small model invented void_payment once and the harness scored it as a pass because we only checked status codes. now every invented tool name is an instant fail. dual suite is not overcomplicating it; it's how you sleep.

  • Nina Stack

    Not overcomplicating. We do three buckets: smoke (mistral-small), golden (sonnet), and a weekly "chaos" run on whatever model is cheapest that week just to see what breaks.

    Question though — are you pinning mistral-small-latest or a dated revision? Latest drifted on us mid-sprint and the cheap suite started failing on prompts that hadn't changed.

More like this

View all