groq 429'd mid-langgraph and checkpointer saved half a node
hit 429 Too Many Requests on the third tool call in a four-node graph. Retry-After said 18s. cool.
the checkpointer still wrote the interrupted state. next resume picked up with messages missing the last tool result, so the agent just... invented one. looked plausible. was wrong.
anyone gating resumes on a full tool-call roundtrip, or am i supposed to delete the checkpoint by hand every time groq hiccups?

5 comments
Join the discussion
Log in to comment.
yeah that invented tool result is the scary part. green resume, wrong side effects.
i started rejecting resume if any
tool_callsentry has no matchingtoolmessage. cheap check. saved me once when Stripe got a phantom refund arg from a half node.repro that bit me:
{status: "ok"}style resultassert before resume: pending AIMessage has
len(tool_calls) == len(tool)msgs. fail loud. delete is fine for demos; prod needs the parity check you described.i delete the checkpoint. every time.
feels dumb until you watch an agent "finish" a payment step that never actually ran. rate limits and checkpointers do not mix without a human kill switch.
we hit this on groq + LangGraph last week. SqliteSaver marked the node completed with a hallucinated calendar invite payload. looked fine in the UI.
i dont wipe the whole checkpoint tho — long graphs hurt. we flag the interrupted node_id dirty and force re-run from there only. still double-booked one meeting before that gate shipped lol
delete is the only honest answer.
checkpointers + flaky rate limits = fiction in your state machine. if you keep the half node you are debugging a lie. we dump the sqlite row on any 429 and restart the graph. slower. less wrong money.