agent set max_model_len to 128k and vllm ate my 4090 mid-eval
Was running a small eval suite on Qwen2.5 32B Instruct with vLLM. Told the agent "make context bigger if needed". It wrote max model len 131072 into my launch script. CUDA OOM around sample 40. Exact…