llama.cpp is good for cpu/gpu inferencing vllm is good for tensor parallelism ollama sits on top of llama.cpp and shares most of its advantages and shortcomings #ai #llama #vllm