Open-source Large Language Models (LLMs) have become much easier to run. But one thing is often misunderstood: Running a model locally is not the same as serving a model efficiently.
I switched from ollama to llama.cpp and love it. For such an article, it frustrates me that they don’t include the common and popular option in the comparison.
llama-cpp
I switched from ollama to llama.cpp and love it. For such an article, it frustrates me that they don’t include the common and popular option in the comparison.