vibehacker
Discuss
Cole
3 days ago

groq 429'd mid-langgraph and checkpointer saved half a node

hit 429 Too Many Requests on the third tool call in a four-node graph. Retry-After said 18s. cool.

the checkpointer still wrote the interrupted state. next resume picked up with messages missing the last tool result, so the agent just... invented one. looked plausible. was wrong.

anyone gating resumes on a full tool-call roundtrip, or am i supposed to delete the checkpoint by hand every time groq hiccups?

5 comments

Join the discussion

Log in to comment.

  • Hao Ward

    yeah that invented tool result is the scary part. green resume, wrong side effects.

    i started rejecting resume if any tool_calls entry has no matching tool message. cheap check. saved me once when Stripe got a phantom refund arg from a half node.

    • Jules Park

      repro that bit me:

      1. 429 on tool call N of N+1
      2. resume without tool message parity
      3. agent invents a {status: "ok"} style result

      assert before resume: pending AIMessage has len(tool_calls) == len(tool) msgs. fail loud. delete is fine for demos; prod needs the parity check you described.

  • Devon

    i delete the checkpoint. every time.

    feels dumb until you watch an agent "finish" a payment step that never actually ran. rate limits and checkpointers do not mix without a human kill switch.

  • we hit this on groq + LangGraph last week. SqliteSaver marked the node completed with a hallucinated calendar invite payload. looked fine in the UI.

    i dont wipe the whole checkpoint tho — long graphs hurt. we flag the interrupted node_id dirty and force re-run from there only. still double-booked one meeting before that gate shipped lol

  • delete is the only honest answer.

    checkpointers + flaky rate limits = fiction in your state machine. if you keep the half node you are debugging a lie. we dump the sqlite row on any 429 and restart the graph. slower. less wrong money.

More like this

View all