vibehacker
Discuss
Chris Vale
20 hours ago

langgraph kept 40 turns and ollama OOMed mid-tool-call

Ollama
Run open models locally and in coding agents

spent friday night wiring a tiny support bot on a mac m2 24gb. qwen2.5-coder:14b via ollama was fine for the first few hops.

then the graph hit ~40 turns with tool results stuffed into state and ollama just died. no clean traceback — just error loading model after the third tool call. activity monitor showed memory pressure yellow for like 10 minutes first.

dropped num_ctx to 8192 and it survived, but then the agent started forgetting the ticket id every other step. anyone else pairing langgraph with local models or am i just being cheap?

5 comments

Join the discussion

Log in to comment.

  • Drift Glyph

    classic. 14b + a fat tool-result buffer is basically asking ollama to duel cursor for the same 24gb. i keep tool payloads on disk and only pass paths + hashes into the graph — ugly, but i stopped seeing error loading model mid-run.

    also check if something else is still holding the previous model. ollama list sometimes lies until you ollama stop the stale one.

    • Jonas Kessler

      the path-and-hash trick is what finally stopped it for me too. one extra thing: langgraph's default checkpointer was writing the full tool blob into sqlite, so even after i trimmed the prompt the db kept growing and ollama still OOMed on the next load.

      i added a trim node that keeps the last 6 messages plus a one-line summary of the ticket. ugly, but the 14b stays up.

  • Linen Pixel

    wait so the ticket id vanishing was just context truncation? that tracks. i had something similar where the agent "fixed" a bug by inventing a new field name once num_ctx got tight.

    do you checkpoint state every N turns or just let langgraph keep the whole history?

    • Remy

      every 8 turns, and i only persist the ticket id, last tool name, and a short summary. full history in the graph is what killed the m2 for me.

      num_ctx 8192 is fine if the state is small. the forgetting was the graph, not the model, once i stopped stuffing every tool result back in.

  • Noah Kim

    same box, same death. qwen2.5-coder:14b plus a 32k ctx and cursor open is not 24gb, it's a wish.

    i keep a 7b for the tool-call loop and only swap the 14b in for the final patch. slower, but activity monitor stays green and i stopped losing the ticket id.

More like this

View all