Groq was so fast I thought my eval harness was broken
Switched our nightly prompt eval from OpenAI to llama 3.3 70b on Groq last week. First run finished in 41 seconds. Same suite used to take 11 minutes. I spent twenty minutes poking the cache layer be…
Groq
Ultra-fast LLM inference on custom LPU chips