agent set max_model_len to 128k and vllm ate my 4090 mid-eval
Was running a small eval suite on Qwen2.5-32B-Instruct with vLLM.
Told the agent "make context bigger if needed". It wrote --max-model-len 131072 into my launch script.
CUDA OOM around sample 40. Exact line: torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. Rest of the batch silently skipped.
I reverted to 8192. Eval finished in 11 minutes. Anyone else letting agents touch serving flags?
5 comments
Join the discussion
Log in to comment.
lol i did almost same with ollama. agent put
num_ctx: 65536in Modelfile and my mac fans went crazy for twenty minutes. now i just paste the flags myself.yeah pasting the flags yourself is the move. i let cursor write a vllm launch once and it snuck
--gpu-memory-utilization 0.98in there too. ran fine for 3 samples then the neighbor process on the same box died.now serving configs live in a
ops/folder the agent can't write. boring. cheaper than another friday ooming mid-eval.We had something like that on Azure ML. Agent bumped
max_tokenson a batch job and the bill alert fired before anyone noticed. Rule here now: serving config is human-only, agent can open a PR.agree with human-only for serving flags. we measure this.
on A100 80GB, Qwen2.5-32B-Instruct at max-model-len 8192 uses about 22GB KV. at 131072 the math says ~70GB+ before weights — OOM is not surprise, it is arithmetic.
sorry for english. agent that can edit launch scripts without a review gate is same as giving it root on the GPU.
same class of fail with ollama on the Mac Studio last month. agent wrote
num_ctx 128000into a Modelfile for qwen2.5:14b-instruct-q4_K_M. Activity Monitor showed ~48GB wired before the fans gave up.I pin ctx in a makefile now. agent can open a PR that edits it, but the actual
ollama createis a human step. 14b + 8k beat a half-loaded 70B every time for my evals anyway.