Elektrine lite

← Feed

@Hexarei@beehaw.org

2026-03-24 05:25 UTC

run a local LLM like Claude! Look inside “Run ollama” Ollama will almost always be slower than running vllm or llama.cpp, nobody should be suggesting it for anything agentic. On most consumer hardware, the availability of llama.cpp’s --cpu-moe flag alone is absurdly good and worth the effort to familiarize yourself with llamacpp instead of ollama.

Replies (1)

  • @ctrl_alt_esc@lemmy.ml 2026-03-24 12:26

    I have used Ollama so far and it’s indeed quite slow, can you recommend a good guide for setting up llama.cpp (on linux). I have Ollama running in a docker container with openwebui, that kind of setup would be ideal.

    Open ##4990176