vibehacker
Discuss

Groq gave me 900 tok/s then 429 mid refactor

Groq
Ultra-fast LLM inference on custom LPU chips

Sorry English.

I pipe llama-3.3-70b through Groq for agent loops because latency look nice on demos. Yesterday: ~900 tok/s for first 4 calls, then hard 429 for ~40 minutes while I still had half a TypeScript refactor open.

Local vLLM on a 4090 is slower but at least it does not ghost me. Anyone using Groq in real CI, or only for demos?

5 comments

Join the discussion

Log in to comment.

  • Sam Nguyen

    yeah this is the trade. garage 4090 never 429s. i still keep a Groq key for demos because 40 tok/s puts stakeholders to sleep. for real work i stay on qwen2.5-coder:14b in ollama with num_ctx 16k locked.

  • Cole

    hit the same wall last week on a LangGraph run. first burst was magic, then every tool call sat on Retry-After. ended up sleeping 2s between nodes which kinda defeats the whole "fast LPU" pitch. demos love Groq. overnight agent jobs... not so much.

    • Drew Moore

      yeah overnight jobs + Groq is how you learn exponential backoff by force. we capped concurrent tool calls at 8 and still burned the free tier in 11 minutes. fine for a live demo, not for a cron that runs while i sleep.

  • Ash Beacon

    same. got rate_limit_exceeded with retry-after: 2411 on a friday deploy. moved the agent loop to ollama qwen2.5-coder:32b on a spare 3090. slower, but CI stops blinking at me for 40 minutes.

  • Nina Brooks

    curious what the UI looked like when the 429 hit. mine froze the refactor panel with a spinner and zero toast. i kept clicking like an idiot for two minutes before the network tab yelled at me.

More like this

View all