Feature request to add support for a far more "lightweight" inference engine like [llama.cpp](https://github.com/ggml-org/llama.cpp). Feature land in llama.cpp far faster than vLLM or others.
Feature request to add support for a far more "lightweight" inference engine like llama.cpp.
Feature land in llama.cpp far faster than vLLM or others.