Show: internal docs bot stack

LangChain + Mem0 + vLLM behind Open WebUI. Sounds like alphabet soup but it runs our internal docs bot.
Happy to share the boring parts (evals, chunking, failure modes) if useful.

LangChain + Mem0 + vLLM behind Open WebUI. Sounds like alphabet soup but it runs our internal docs bot.
Happy to share the boring parts (evals, chunking, failure modes) if useful.
ran one of Quandora's open research workflows on three midcap 10 Ks this week. got a clean markdown brief, then a table where two PE cells were empty and one was from FY2022. no error toast. just qui…

Official now: https://www.cnbc.com/2026/09/03/nvidia agrees to buy hugging face for almost 13 billion ai expansion.html $12.9 billion. Second biggest Nvidia deal ever after the Groq assets. Delangue…
Running Qwen + Llama locally on a 4090. Open WebUI + Ollama is fine, but I'm curious who switched to Jan or LM Studio and why. Also: anyone got a sane setup for sharing a local model across a small t…
3 comments
Join the discussion
Log in to comment.
Would love the chunking notes. That's where we keep bleeding.
Mem0 vs rolling your own memory table — which did you pick and why?
vLLM throughput numbers would help. We're deciding between that and Ollama.