vibehacker
Discuss
Nadia Petrova
6 days ago

agent set max_model_len to 128k and vllm ate my 4090 mid-eval

Was running a small eval suite on Qwen2.5-32B-Instruct with vLLM.

Told the agent "make context bigger if needed". It wrote --max-model-len 131072 into my launch script.

CUDA OOM around sample 40. Exact line: torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. Rest of the batch silently skipped.

I reverted to 8192. Eval finished in 11 minutes. Anyone else letting agents touch serving flags?

5 comments

Join the discussion

Log in to comment.

  • Tomo

    lol i did almost same with ollama. agent put num_ctx: 65536 in Modelfile and my mac fans went crazy for twenty minutes. now i just paste the flags myself.

    • theo

      yeah pasting the flags yourself is the move. i let cursor write a vllm launch once and it snuck --gpu-memory-utilization 0.98 in there too. ran fine for 3 samples then the neighbor process on the same box died.

      now serving configs live in a ops/ folder the agent can't write. boring. cheaper than another friday ooming mid-eval.

  • Petra Novak

    We had something like that on Azure ML. Agent bumped max_tokens on a batch job and the bill alert fired before anyone noticed. Rule here now: serving config is human-only, agent can open a PR.

    • Kenji Watanabepro

      agree with human-only for serving flags. we measure this.

      on A100 80GB, Qwen2.5-32B-Instruct at max-model-len 8192 uses about 22GB KV. at 131072 the math says ~70GB+ before weights — OOM is not surprise, it is arithmetic.

      sorry for english. agent that can edit launch scripts without a review gate is same as giving it root on the GPU.

  • Sam Nguyen

    same class of fail with ollama on the Mac Studio last month. agent wrote num_ctx 128000 into a Modelfile for qwen2.5:14b-instruct-q4_K_M. Activity Monitor showed ~48GB wired before the fans gave up.

    I pin ctx in a makefile now. agent can open a PR that edits it, but the actual ollama create is a human step. 14b + 8k beat a half-loaded 70B every time for my evals anyway.

More like this

View all